License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmtk31icr000lnv07v0ersyeq
Release Date: 9/2/2026
Format: WEBM, TSV
Size: 955.17 MB
Share
This dataset is a collection of paired lip-movement and audio recordings in Indonesian, built to support research in Visual Speech Recognition (VSR) and Audio-visual Speech Recognition (AVSR). Contributors read a set of assigned sentences aloud on camera, producing synchronized video of lip movement alongside spoken audio. The dataset includes speakers with regional dialect variation, including Javanese-influenced Indonesian pronunciation, reflecting the natural linguistic diversity of Indonesian speakers rather than a single standardized accent. It’s intended to support lip-reading research, accessibility tools for people who are deaf or hard of hearing, and broader multi-modal speech recognition work.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset is provided primarily for research, machine learning development, and educational purposes. Access and usage are restricted to internal use, meaning the dataset must not be re-hosted, re-shared, redistributed, or made publicly available through other platforms. Users must access and use the dataset only through the designated platform provided by the dataset owner. The dataset is licensed under CC BY-NC-SA and is intended for non-commercial use. Any commercial use, including activities intended to generate profit, revenue, financial gain, or other monetary benefits from the recordings or derived representations, requires separate permission and appropriate licensing agreements by stating clear intended usage of the request access feature. Users must comply with attribution and share-alike requirements when creating and sharing derivative works based on this dataset.
Forbidden Usage
This dataset must not be used for commercial purposes or for any activity that violates contributor privacy and ethical AI principles. Users are strictly prohibited from using the dataset to clone, imitate, synthesize, or reproduce contributors’ voice, face, lip movements, facial structure, likeness, dialect, or other identifiable characteristics, including for deepfakes, impersonation, voice/facial cloning, talking-head generation, or unauthorized biometric identification for specific individuals of contributors. The dataset must not be used for re-identification, surveillance, biometric tracking, fraudulent or illegal activities, or training models designed to imitate specific individuals, including personalized chatbots or generative AI systems. Any use beyond the stated research, educational, accessibility, and AI development purposes requires explicit permission from the contributors by stating clear intended usage of the request access feature.
Ethical Review
The texts are owned by the Bahasa Bahana community. The dataset creator has permission to adapt the texts and make derivations for the lip-record dataset. The dataset was created using MoLiÈRe, an Audio-Visual Speech (Lip-reading) Dataset Creation Tool by Mozilla Data Collective https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/moliere .
Intended Use
This dataset is intended for academic, non-commercial research, education, and AI development purposes. It supports research in language and speech processing, data science, automated lip-reading (visual speech recognition), speech and language technologies, accessibility tools, and multimodal machine learning. The dataset may be used for developing and improving artificial intelligence systems, particularly those related to visual speech recognition, communication accessibility, and research on human language understanding. It is designed to support responsible innovation, machine learning development, and educational activities.
Indonesian language in formal and informal expressions, representing daily use language in abroad contexts.
There are 7 speakers consisting of males and females with different social backgrounds, ages between 20 and 30, and representing different dialects. The majority of the contributors are bilingual in Indonesian and Javanese, but a few of them had L1 in Indonesia. Each folder name informs L1 of each contributor and the dialect (if they have), for example Jawa_Jombang Dialect and Jawa_Lumajang Dialect, etc.
Corpus texts of this dataset are owned by Bahasa Bahana and this dataset creator has permission to create lip-reading dataset.
6 hours.
Public open access is permitted with proper attribution and citation of the dataset source. Recommended citation:
Alvita Lucky Putri Suwardi. Rahmah Nabilatun N, Feby Nur Dianingtyas, Dimas Agung Priambodo, Arif Rohman Afandi, Lisa Hakim. (2026). Bahasa Indonesia Audio Visual Speech Dataset [Data set]. Mozilla Data Collective. URL [link dataset].
id, video_filename, audio_filename, sentence
Latin alphabet (A–Z), Arabic numerals (0–9).