License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmtrd8j1z05ryo107d2m09ulg
Release Date: 9/7/2026
Format: WEBM, TSV
Size: 982.50 MB
Share
This dataset features video recordings capturing the lip movements of 6 native speakers of Bahasa Indonesia, representing diverse regional dialect groups. To provide comprehensive visual speech data, the spoken vocabulary spans three distinct linguistic domains: everyday casual speech patterns, socioenvironmental discussions, and formal literary texts to reflect complex syntactic structures.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
The dataset allows the commercialization of models or applications developed using the data, provided that the raw biometric data itself is not redistributed, restricted, or placed behind proprietary access barriers. Users are encouraged to maintain openness and transparency regarding the use of the dataset. Users are highly appreciated to provide a brief introduction or explanation on the dataset access request page describing their explicit intended use. This practice supports responsible data usage, improves transparency, and helps the dataset creators understand how the resource contributes to research, development, and innovation.
Forbidden Usage
This dataset is subject to strict restrictions to protect the privacy, identity, and security of the individuals represented. Users are prohibited from attempting to re-identify any speaker or extract personal identity information from the data. The dataset must not be used to train, develop, or deploy systems for voice cloning, synthetic voice generation, deepfake creation, facial spoofing, or any technology intended to imitate the voice, facial structure, facial expressions, or other unique characteristics of specific individuals. Furthermore, the use of facial data for surveillance, unauthorized biometric identification, or tracking purposes is strictly forbidden. These limitations are established to prevent identity misuse and ensure that the dataset is used responsibly for ethical research and development purposes.
Ethical Review
The texts are owned by the Bahasa Bahana community. The dataset creator has permission to adapt the texts and make derivations for the lip-record dataset. The dataset was created using MoLiÈRe, an Audio-Visual Speech (Lip-reading) Dataset Creation Tool by Mozilla Data Collective https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/moliere .
Intended Use
This dataset is intended primarily for research, data science development, and educational purposes, with a focus on Indonesian language studies, automated lip-reading (visual speech recognition), accessibility technologies for individuals with hearing impairments, and multimodal machine learning model development. The dataset is specifically designed to support the development and training of Indonesian-based lip-reading AI systems by providing resources for advancing visual speech recognition technologies. Its intended applications emphasize ethical research, AI improvement, and the creation of inclusive technologies while maintaining responsible data usage practices.
Indonesian language in formal and informal expressions, representing daily use language in abroad contexts.
There are 6 speakers consisting of males and females with different social backgrounds, ages between 20 and 30, and representing different dialects. The majority of the contributors are bilingual in Indonesian and Javanese, but a few of them had L1 in Indonesia. Each folder name informs L1 of each contributor and the dialect (if they have), for example Jawa_Multidialect and Jawa_Lumajang Dialect, etc.
Corpus texts of this dataset are owned by Bahasa Bahana and this dataset creator has permission to create lip-reading dataset.
General domains.
6 hours.
Public open access is permitted with proper attribution and citation of the dataset source. Recommended citation:
Rosyid Hidayatul Fadilah, Yunus Aji Lumintang, Masyhuri Farhan, Muqsit Harjuno. (2026). Bahasa Indonesia Lip-Reading Dataset [Data set]. Mozilla Data Collective. URL [link dataset].
id, video_filename, audio_filename, sentence
Latin alphabet (A–Z), Arabic numerals (0–9).