License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmu1ieso2001bo107h8377g5v
Release Date: 9/14/2026
Format: WEBM, TSV
Size: 940.42 MB
This dataset is a multimodal visual speech resource featuring synchronized video and audio recordings of six native speakers representing diverse regional Indonesian and javanese dialects. The vocabulary encompasses casual speech, socioenvironmental discussions, and formal literary texts to support research in visual speech recognition and multimodal processing.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
The dataset allows the commercialization of models or applications developed using the data, provided that the raw biometric data itself is not redistributed, restricted, or placed behind proprietary access barriers. Users are encouraged to maintain openness and transparency regarding the use of the dataset. Users are highly appreciated to provide a brief introduction or explanation on the dataset access request page describing their explicit intended use. This practice supports responsible data usage, improves transparency, and helps the dataset creators understand how the resource contributes to research, development, and innovation.
Forbidden Usage
This dataset is subject to strict restrictions to protect the privacy, identity, and security of the individuals represented. Users are prohibited from attempting to re-identify any speaker or extract personal identity information from the data. The dataset must not be used to train, develop, or deploy systems for voice cloning, synthetic voice generation, deepfake creation, facial spoofing, or any technology intended to imitate the voice, facial structure, facial expressions, or other unique characteristics of specific individuals. Furthermore, the use of facial data for surveillance, unauthorized biometric identification, or tracking purposes is strictly forbidden. These limitations are established to prevent identity misuse and ensure that the dataset is used responsibly for ethical research and development purposes.
Ethical Review
The texts are owned by the Bahasa Bahana community. The dataset creator has permission to adapt the texts and make derivations for the lip-record dataset. The dataset was created using MoLiÈRe, an Audio-Visual Speech (Lip-reading) Dataset Creation Tool by Mozilla Data Collective https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/moliere.
Intended Use
The dataset allows the commercialization of models or applications developed using the data, provided that the raw biometric data itself is not redistributed, restricted, or placed behind proprietary access barriers. Users are encouraged to maintain openness and transparency regarding the use of the dataset. Users are highly appreciated to provide a brief introduction or explanation on the dataset access request page describing their explicit intended use. This practice supports responsible data usage, improves transparency, and helps the dataset creators understand how the resource contributes to research, development, and innovation.
Indonesian language in formal and informal expressions, representing daily use language in abroad contexts.
There are 6 speakers consisting of males and females with different social backgrounds, ages between 20 and 30, and representing different dialects. The majority of the contributors are bilingual in Indonesian and Javanese, but a few of them had L1 in Indonesia. Each folder name informs L1 of each contributor and the dialect (if they have), for example Jawa_Jombang Dialect and Jawa_Lumajang Dialect, etc.
Corpus texts of this dataset are owned by Bahasa Bahana and this dataset creator has permission to create lip-reading dataset.
Diverse subject domains
6 hours.
Public open access is permitted with proper attribution and citation of the dataset source. Recommended citation:
Muqsit Harjuno, Rosyid Hidayatul Fadilah, Yunus Aji Lumintang, Masyhuri Farhan, Alifa Ardla Farahdila, Yustri Agung Prastiyono. (2026). Indonesian Lip Reading and Visual Speech Dataset [Data set]. Mozilla Data Collective. URL [link dataset].
id, video_filename, audio_filename, sentence
Latin alphabet (A–Z), Arabic numerals (0–9).