License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmtk59j4a0043nv07toiemii6
Release Date: 9/2/2026
Format: WEBM, TSV
Size: 1.22 GB
Share
INALVS (Indonesian Lip-reading & Visual Speech Dataset) is a dataset of Indonesian visual speech recordings created to support research, education, and development in lip-reading, visual speech recognition, and artificial intelligence. The dataset contains recordings of spoken Indonesian language with corresponding visual information of the speaker’s lip and facial movements. It is intended for non-commercial academic and research purposes, including machine learning, visual speech recognition, accessibility technology, and multimodal speech research.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset is provided primarily for research, machine learning development, and educational purposes. Access and usage are restricted to internal use, meaning the dataset must not be re-hosted, re-shared, redistributed, or made publicly available through other platforms. Users must access and use the dataset only through the designated platform provided by the dataset owner. The dataset is licensed under CC BY-NC-SA and is intended for non-commercial use. Any commercial use, including activities intended to generate profit, revenue, financial gain, or other monetary benefits from the recordings or derived representations, requires separate permission and appropriate licensing agreements by stating clear intended usage of the request access feature. Users must comply with attribution and share-alike requirements when creating and sharing derivative works based on this dataset.
Forbidden Usage
This dataset must not be used for commercial purposes or for any activity that violates contributor privacy and ethical AI principles. Users are strictly prohibited from using the dataset to clone, imitate, synthesize, or reproduce contributors’ voice, face, lip movements, facial structure, likeness, dialect, or other identifiable characteristics, including for deepfakes, impersonation, voice/facial cloning, talking-head generation, or unauthorized biometric identification for specific individuals of contributors. The dataset must not be used for re-identification, surveillance, biometric tracking, fraudulent or illegal activities, or training models designed to imitate specific individuals, including personalized chatbots or generative AI systems. Any use beyond the stated research, educational, accessibility, and AI development purposes requires explicit permission from the contributors by stating clear intended usage of the request access feature.
Ethical Review
The texts are owned by the Bahasa Bahana community. The dataset creator has permission to adapt the texts and make derivations for the lip-record dataset. The dataset was created using MoLiÈRe, an Audio-Visual Speech (Lip-reading) Dataset Creation Tool by Mozilla Data Collective https://mdc-dataset-toolbox-ifuhj.ondigitalocean.app/app/moliere .
Intended Use
This dataset is intended for academic, non-commercial research, education, and AI development purposes. It supports research in language and speech processing, data science, automated lip-reading (visual speech recognition), speech and language technologies, accessibility tools, and multimodal machine learning. The dataset may be used for developing and improving artificial intelligence systems, particularly those related to visual speech recognition, communication accessibility, and research on human language understanding. It is designed to support responsible innovation, machine learning development, and educational activities.
Indonesian language in formal and informal expressions, representing daily use language in abroad contexts.
There are 7 speakers consisting of males and females with different social backgrounds, ages between 20 and 30, and representing different dialects. The majority of the contributors are bilingual in Indonesian and Javanese, but a few of them had L1 in Indonesia. Each folder name informs L1 of each contributor and the dialect (if they have), for example Jawa_Multidialect and Jawa_Lumajang Dialect, etc.
Corpus texts of this dataset are owned by Bahasa Bahana and this dataset creator has permission to create lip-reading dataset.
General topic domains.
6 hours.
Public open access is permitted with proper attribution and citation of the dataset source. Recommended citation:
Arif Rohman Afandi, Feby Nur Dianingtyas, Rahmah Nabilatun N, Dimas Agung Priambodo, Lisa Hakim, Alvita Lucky Putri Suwardi. (2026). INALVS (Indonesian Lip-reading & Visual Speech Dataset) [Data set]. Mozilla Data Collective. URL [link dataset].
id, video_filename, audio_filename, sentence
Latin alphabet (A–Z), Arabic numerals (0–9).