License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmu2mq96s00ugo1072rf2p6e3
Release Date: 9/15/2026
Format: WEBM, TSV
Size: 943.86 MB
The Indonesian Speech Motion Dataset (Lip-Record) is a collection of high-resolution video recordings focusing on the spatial and dynamic aspects of lip movements during the pronunciation of various Indonesian vocabulary items. This dataset was created to address resource limitations in research regarding visual speech recognition and audio-visual speech processing. Each sample features vocabulary recorded in a structured manner, enabling machine learning and deep learning models to accurately learn phoneme formation patterns and mouth articulation. Lip-Record is suitable for use by researchers, developers, and academics in building assistive technology applications, extracting visual articulation features, and developing multimodal AI models specifically for the Indonesian language.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
Restrictions/Special Constraints
This dataset is provided primarily for research, machine learning development, and educational purposes. Access and usage are restricted to internal use, meaning the dataset must not be re-hosted, re-shared, redistributed, or made publicly available through other platforms. Users must access and use the dataset only through the designated platform provided by the dataset owner. The dataset is licensed under CC BY-NC-SA and is intended for non-commercial use. Any commercial use, including activities intended to generate profit, revenue, financial gain, or other monetary benefits from the recordings or derived representations, requires separate permission and appropriate licensing agreements by stating clear intended usage of the request access feature. Users must comply with attribution and share-alike requirements when creating and sharing derivative works based on this dataset.
Forbidden Usage
This dataset must not be used for commercial purposes or for any activity that violates contributor privacy and ethical AI principles. Users are strictly prohibited from using the dataset to clone, imitate, synthesize, or reproduce contributors’ voice, face, lip movements, facial structure, likeness, dialect, or other identifiable characteristics, including for deepfakes, impersonation, voice/facial cloning, talking-head generation, or unauthorized biometric identification for specific individuals of contributors. The dataset must not be used for re-identification, surveillance, biometric tracking, fraudulent or illegal activities, or training models designed to imitate specific individuals, including personalized chatbots or generative AI systems. Any use beyond the stated research, educational, accessibility, and AI development purposes requires explicit permission from the contributors by stating clear intended usage of the request access feature.
Ethical Review
The texts are owned by the Bahasa Bahana community. The dataset creator has permission to adapt the texts and make derivations for the lip-record dataset. The dataset was created using MoLiÈRe, an Audio-Visual Speech (Lip-reading) Dataset Creation Tool by Mozilla Data Collective https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/moliere .
Intended Use
This dataset is intended for academic, non-commercial research, education, and AI development purposes. It supports research in language and speech processing, data science, automated lip-reading (visual speech recognition), speech and language technologies, accessibility tools, and multimodal machine learning. The dataset may be used for developing and improving artificial intelligence systems, particularly those related to visual speech recognition, communication accessibility, and research on human language understanding. It is designed to support responsible innovation, machine learning development, and educational activities.
Indonesian language in formal and informal expressions, representing daily use language in abroad contexts.
There are 6 speakers consisting of males and females with different social backgrounds, ages between 20 and 30, and representing different dialects. The majority of the contributors are bilingual in Indonesian and Javanese, but a few of them had L1 in Indonesia. Each folder name informs L1 of each contributor and the dialect (if they have), for example Jawa_Jombang Dialect and Jawa_Lumajang Dialect, etc.
Corpus texts of this dataset are owned by Bahasa Bahana and this dataset creator has permission to create lip-reading dataset.
General domain, including daily social interaction, health and medical terms, environmental context, and conversational Indonesian.
6 hours.
Public open access is permitted with proper attribution and citation of the dataset source.
Dimas Agung Priambodo, Feby Nur Dianingtyas, Rahmah Nabilatun N, Arif Rohman Afandi, Lisa Hakim, Alvita Lucky Putri Suwardi. (2026). Indonesian Speech Motion Dataset [Dataset]. Mozilla Data Collective. URL [link dataset].
id, video_filename, audio_filename, sentence
Latin alphabet (A–Z), Arabic numerals (0–9).