License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmtk5j4s2004unv07serrr1n8
Release Date: 9/2/2026
Format: WEBM, TSV
Size: 981.20 MB
Share
This dataset contains video recordings of speakers reading Indonesian words and sentences. The recordings show the speakers’ lip movements while speaking. The dataset is designed for Indonesian lip-reading research and development. It can be used to develop and evaluate lip-reading and visual speech recognition systems. This dataset may also support further research on visual speech processing.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
The dataset allows the commercialization of models or applications developed using the data, provided that the raw biometric data itself is not redistributed, restricted, or placed behind proprietary access barriers. Users are encouraged to maintain openness and transparency regarding the use of the dataset. Users are highly appreciated to provide a brief introduction or explanation on the dataset access request page describing their explicit intended use. This practice supports responsible data usage, improves transparency, and helps the dataset creators understand how the resource contributes to research, development, and innovation.
Forbidden Usage
This dataset is subject to strict restrictions to protect the privacy, identity, and security of the individuals represented. Users are prohibited from attempting to re-identify any speaker or extract personal identity information from the data. The dataset must not be used to train, develop, or deploy systems for voice cloning, synthetic voice generation, deepfake creation, facial spoofing, or any technology intended to imitate the voice, facial structure, facial expressions, or other unique characteristics of specific individuals. Furthermore, the use of facial data for surveillance, unauthorized biometric identification, or tracking purposes is strictly forbidden. These limitations are established to prevent identity misuse and ensure that the dataset is used responsibly for ethical research and development purposes.
Ethical Review
The texts are owned by the Bahasa Bahana community and the dataset creator has permission to adapt the texts for lip-record dataset. The dataset was created using MoLiÈRe, an Audio-Visual Speech (Lip-reading) Dataset Creation Tool by Mozilla Data Collective https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/moliere .
Intended Use
The dataset allows the commercialization of models or applications developed using the data, provided that the raw biometric data itself is not redistributed, restricted, or placed behind proprietary access barriers. Users are encouraged to maintain openness and transparency regarding the use of the dataset. Users are highly appreciated to provide a brief introduction or explanation on the dataset access request page describing their explicit intended use. This practice supports responsible data usage, improves transparency, and helps the dataset creators understand how the resource contributes to research, development, and innovation.
Indonesian language in formal and informal expressions, representing daily use language in abroad contexts.
There are 6 speakers consisting of males and females with different social backgrounds, ages between 20 and 30, and representing different dialects. The majority of the contributors are bilingual in Indonesian and Javanese, but a few of them had L1 in Indonesia. Each folder name informs L1 of each contributor and the dialect (if they have), for example Jawa_Jombang Dialect and Jawa_Lumajang Dialect, etc.
Corpus texts of this dataset are owned by Bahasa Bahana and this dataset creator has permission to create lip-reading dataset.
The dataset covers various domains, including daily life, literature, environment, and health.
6 hours.
Public open access is permitted with proper attribution and citation of the dataset source.
Alifa Ardla Farahdila. (2026). Indonesian-Multidialect Lip-Record Speech Corpus [data set]. Mozilla Data Collective. URL [link dataset].
id, video_filename, audio_filename, sentence
Latin alphabet (A–Z), Arabic numerals (0–9).