License:
CC-BY-SA-4.0
Steward:
CommunityDataset ID:
cmu1b03dq03zrnx08kumwzt3i
Release Date: 9/14/2026
Format: WEBM, TSV
Size: 282.63 MB
The dataset is based on the Bogor dialect of the Sundanese language, West Java Province, Indonesia. Unlike the East Priangan Sundanese variety, which strongly maintains the “undak usuk basa” (Sundanese speech levels) system with distinct informal (kasar), neutral (loma), and polite (lemes) forms, the Bogor Sundanese dialect tends to employ a more flexible and informal communication style. It reflects a natural conversational pattern in which speakers place less emphasis on hierarchical language distinctions and more on social familiarity, directness, and everyday interaction. The dataset captures authentic language use across various contexts, including personal experiences, family, culture, education, social interaction, technological changes, and contemporary issues, thereby representing the linguistic characteristics and sociocultural identity of Sundanese speakers in the Bogor region.
Licensing
Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0)
Restrictions/Special Constraints
This dataset may be used for public benefit, including commercial applications. Any use of the dataset requires prior notification and permission from the dataset owner.
Forbidden Usage
Modifying this dataset without the owner’s permission is prohibited.
Ethical Review
This dataset was created by writing texts in the Bogor dialect of Sundanese, representing the language variety spoken by communities in Bogor City and Bogor Regency, West Java Province, Indonesia. The files were read and recorded by native speakers through the hosting platform https://sabre-2.onrender.com/. The collected audio recordings were compiled into a comprehensive dataset for linguistic research, speech technology development, and Sundanese language preservation.
Intended Use
This dataset is intended to support the preservation, documentation, and dissemination of the Sundanese language, particularly the Bogor Sundanese dialect spoken in Bogor City and Bogor Regency, West Java Province, Indonesia. It is developed to facilitate educational activities, linguistic studies, and cultural research by providing a representation of the local Sundanese language variety.
This dataset uses the Bogor dialect of the Sundanese language with Indonesian and English code-mixing.
Native of Sundanese speaker, aged between 20 and 30 years, situated in Bogor city and Bogor regency, West Java Province, Indonesia.
Created by the owner of the dataset, considered as a linguist and native speaker.
General domain, including culture, tourism travel, family, education, social interaction, technological development, national issues, and everyday activities.
5 hours
Rahmawati. (2026). BOSSCO : Bogor Sundanese Speech Corpus [Data set]. Mozilla Data Collective. URL [link dataset].
Audio file name, text
“Masyarakat di ditu meuni taat jasa ka adat anu geus diturunkeun ti moyang.”
“Aya hiji bangunan anu dibangun maké témbok, teu lila bangunanna diurugkeun.”
“Taun kamari, abdi jeung kaluarga ulin ka Yogyakarta.”
“Bapa jeung ema meuni reuseupeun kana suasana Yogyakarta.”
“Bapa jeung ema teu biasa barang dahar anu amis sedangkeun kadaharan di Yogyakarta lolobana rasana amis.”
Latin alphabet (A–Z), Arabic numerals (0–9).