License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmu1ax60f03zjnx081lh28hty
Release Date: 9/14/2026
Format: WEBM, TSV
Size: 266.65 MB
The dataset is based on the Jepara dialect of the Javanese language spoken in Jepara Regency, Central Java Province, Indonesia. The corpus represents the natural linguistic characteristics of Jepara Javanese, which reflects the local identity and communication patterns of communities in the Jepara region. Compared with other Javanese varieties, the Jepara dialect has distinctive vocabulary choices, pronunciation patterns, and expressions influenced by local cultural contexts and daily interactions. The dataset captures authentic language usage across different conversational situations, including personal experiences, social relationships, cultural practices, and contemporary topics. Through the collection of speech recordings and corresponding text transcriptions, this dataset provides a valuable resource for speech technology development, linguistic research, and preservation of regional Javanese language varieties.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
https://spdx.org/licenses/CC-BY-SA-4.0.htmlRestrictions/Special Constraints
This dataset may be used for public benefit, including research, educational activities, and commercial applications with proper attribution. Any use of the dataset requires appropriate citation and acknowledgment of the dataset owner.
Forbidden Usage
Redistributing or republishing this dataset is prohibited. Any attempt to identify individual speakers or clone specific speakers' voices is not allowed.
Ethical Review
This dataset was created by writing texts in the Jepara dialect of Javanese, representing the language variety spoken by communities in Jepara City, Central Java Province, Indonesia. The files were read and recorded by native speakers through the hosting platform https://sabre-2.onrender.com/. The collected audio recordings were compiled into a comprehensive dataset for linguistic research, speech technology development, and Sundanese language preservation.
Intended Use
This dataset is intended to support the preservation, documentation, and dissemination of the Javanese language, particularly the Jepara dialect spoken in Jepara Regency, Central Java Province, Indonesia. It is developed to facilitate linguistic studies, educational activities, cultural documentation, and the development of speech technology applications such as text-to-speech (TTS) systems.
This dataset uses the Jepara dialect of the Javanese language with possible Indonesian and English code-mixing in several utterances.
Native speakers of the Jepara Javanese dialect, aged between 20-30 years old, and spoken in Jepara Regency, Central Java Province, Indonesia.
Created by the owner of the dataset, considered as a linguist and native speaker.
General domain, including culture, local traditions, tourism, family, education, social interaction, technological development, national issues, and everyday activities.
5 hours
Sholikhah, Nikmatus. (2026). TTS - Jepara Javanese Speech Corpus (JJSC) [Data set]. Mozilla Data Collective. URL [link dataset].
Audio file name, text
“Salah siji destinasi sing paling pengen tak kunjungi ning masa depan yaiku negara Makkah”
“Aku lan akeh mahasiswa liyane kudu cepet adaptasi karo cara belajar sing anyar”
“Pengalaman pertama teka ning konser musik dadi salah siji momen sing paling berkesan ning uripku”
“Pak Ali dikenal minangka guru sing sabar, ramah, lan cara ngajare gampang dipahami”
“Miturutku, diwenehi nasihat nalika lagi sedhih utawa mental down iku rasane nyenengake lan nguatake”
Latin alphabet (A–Z), Arabic numerals (0–9)