License:
CC-BY-NC-SA-4.0
Steward:
CommunityDataset ID:
cmumq51mj0aacny078eindv36
Release Date: 9/29/2026
Format: OPUS, TSV
Size: 81.38 MB
The ASR Javanese-Lumajang Dialect dataset contains a 10-hour audio corpus of spontaneous speech record (ASR) from native Javanese speakers in Lumajang Regency, East Java Province, Indonesia. The recording is captured Javanese in informal settings which reflect authentic daily vocabulary and local dialect featuring natural code-switching with Indonesian.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
This dataset is open to the public and is not exclusively granted to any party. This dataset is intended for linguistic research and NLP (AI training) related to Javanese-Lumajang dialect and Indonesian in switching code context. All academic or research purposes are required to properly cite this dataset. For derivative products or modification, users must explain explicitly the modification and respect privacy, copy right, and law. Users must comply with the dataset license terms and provide appropriate attribution for research uses, article publications, and products. If you wish to use it for other purposes, please contact the administrator by submitting an access request with a clear statement of your intended use.
Forbidden Usage
You must not attempt to identify or re-identify any individual speaker. You must not use this dataset for malicious, deceptive, or harmful purposes according to laws and social norms. You must not use this dataset to generate misleading, hate-speech, discrimination, bullying, and fraudulent audio content. Re-distribute the dataset into an individual or organization is not allowed. Accessing this dataset is only via Mozilla Data Collective and comply with the terms and conditions. Any use that is intended to commit war and violates privacy rights, human rights (animate and inanimate), or applicable laws is strictly prohibited. This dataset may not be used by companies, individuals, or affiliates with any known human rights violations or war crimes. Unattributed use of the dataset and causing disadvantages to contributors is prohibited.
Ethical Review
All participants were informed and gave consent for the creation of this dataset. Participants used their own devices to record their voices. Audio recordings were submitted via WhatsApp community groups. The audio recordings were compiled by group administrators and then transcribed by participants. Finally, the collection of audio recordings and their transcriptions was compiled into a comprehensive dataset.
Intended Use
This dataset is made for development of AI training and linguistics research on phonology, dialectology, and sociolinguistics of the Javanese-Lumajang dialect within code-switching to Indonesian from East Java Province, Indonesia.
Commercial derivative products require contacting the admin by submitting access request. In addition, the dataset owner accepts compensation in any form either for commercial or non-commercial use of this dataset.
The recordings capture natural and spontaneous speech in the Javanese-Lumajang dialect, including code-mixing and code-switching with Indonesian. The data were produced by adult speakers aged 17 to 25 years from diverse social backgrounds.
This dataset was created collectively with the native speakers of the community and the data collection was managed by the owner of this dataset.
This dataset covers a broad range of everyday topics, including daily activities, family life, work, community, travel, education, and health.
Approximately 10 hours.
This dataset is open to the public with proper source attribution or citation.
Yustri Agung Prastiyono & Masyhuri Farhan. (2026). ASR Javanese-Lumajang Dialect [Data set]. Mozilla Data Collective. URL [dataset link].
audio_filename, speaker_number, proposed transcription, post-edited transcription.
This dataset consists of audio compilation and TSV, unfortunately code-switching and code-mixing is not annotated yet.
,,,Mama ya gak tau lali ngelingno aku ben sarapan, jaga kesehatan, karo ati-ati lek misale aku dolan utawa metu-metu ngono,,,
This sentence contains code-mixing in Indonesian (e.g. “jaga kesehatan”)
,,,Sing paling tak kagumi maneh iku, sifate sing pantang menyerah,,
This sentence contains code-mixing in Indonesian (e.g. “pantang menyerah”)
,,,Tapi semangate Kartini iki tetep urip ning dunia pendidikan Indonesia nganti saiki,,,
This sentence contains code-mixing in Indonesian (e.g. “dunia pendidikan Indonesia”)
,,,Pemikirane mulai nyebar lewat surat-surat sing ditulise, nganti akeh tokoh seng sadar nek gagasan iki penting gawe masa depan bangsa,,,
This sentence contains code-mixing in Indonesian (e.g. “masa depan bangsa”)
The dataset uses the latin alphabet with spelling that follows everyday language usage. Latin alphabet (A–Z), Arabic numerals (0–9).