Release Date: 9/17/2026
Format: MP3
Size: 462.06 MB
A collection of read speech recordings in Vietnamese (Việt).
Restrictions/Special Constraints
None provided.
Forbidden Usage
It is forbidden to attempt to determine the identity of speakers in the Common Voice datasets. It is forbidden to re-host or re-share this dataset.
Intended Use
This dataset is intended to be used for training and evaluating automatic speech recognition (ASR) models. It may also be used for applications relating to computer-aided language learning (CALL) and language or heritage revitalisation.
vi)This datasheet is for cv-corpus-27.0-2026-09-11 of the Mozilla Common Voice Scripted Speech dataset for Vietnamese [Việt - vi]. The dataset contains 20605 clips representing 22.97 hours of recorded speech (7.81 hours validated) from 419 speakers, recorded from a text corpus of 22,058 sentences.
The official language spoken by Vietnamese people.
There are three main variations of the language: Hà Nội, Huế and Saigon.
| Code | Variant | Clips | Speakers |
|---|---|---|---|
| vi-huett | Huế | 1,561 (7.6%) | 6 (1.4%) |
| vi-saigon | Sài Gòn | 404 (2.0%) | 8 (1.9%) |
| vi-hanoi | Hà Nội | 351 (1.7%) | 43 (10.3%) |
| Code | Accent | Clips | Speakers |
|---|---|---|---|
| - | 1,582 (7.7%) | 18 (4.3%) |
The dataset includes the following self-declared age and gender distributions. A coverage summary is shown below each table.
Self-declared gender information. The table shows clip and speaker counts with percentages. Speakers who did not declare a gender are listed as Unspecified. A dash (-) indicates zero.
| Code | Gender | Clips | Speakers |
|---|---|---|---|
| male_masculine | Male, masculine | 9,506 (46.1%) | 78 (18.6%) |
| female_feminine | Female, feminine | 4,114 (20.0%) | 26 (6.2%) |
| transgender | Transgender | - | - |
| non-binary | Non-binary | - | - |
| do_not_wish_to_say | Prefer not to say | 5 (0.0%) | 1 (0.2%) |
| - | Unspecified | 6,980 (33.9%) | 347 (82.8%) |
Gender declared: 13,625 of 20,605 clips (66.1%), 72 of 419 speakers (17.2%)
Self-declared age information. The table shows clip and speaker counts with percentages. Speakers who did not declare an age are listed as Unspecified. A dash (-) indicates zero.
| Code | Age | Clips | Speakers |
|---|---|---|---|
| teens | Teens | 3,909 (19.0%) | 12 (2.9%) |
| twenties | Twenties | 4,411 (21.4%) | 114 (27.2%) |
| thirties | Thirties | 890 (4.3%) | 26 (6.2%) |
| fourties | Fourties | 357 (1.7%) | 5 (1.2%) |
| fifties | Fifties | 5 (0.0%) | 1 (0.2%) |
| sixties | Sixties | 5,044 (24.5%) | 1 (0.2%) |
| seventies | Seventies | 840 (4.1%) | 3 (0.7%) |
| eighties | Eighties | - | - |
| nineties | Nineties | - | - |
| - | Unspecified | 5,149 (25.0%) | 290 (69.2%) |
Age declared: 15,456 of 20,605 clips (75.0%), 129 of 419 speakers (30.8%)
Clip buckets
| Bucket | Clips |
|---|---|
| Validated | 7,013 (34.0%) |
| Invalidated | 505 (2.5%) |
| Other | 13,087 (63.5%) |
Training splits
| Split | Clips |
|---|---|
| Train | 1,848 (26.4%) |
| Dev | 1,515 (21.6%) |
| Test | 1,593 (22.7%) |
Training split coverage: 4,956 of 7,013 validated clips (70.7%)
The dataset contains 7013 validated, 505 invalidated, and 13087 unresolved clips. The average clip duration is 4.013 seconds.
"This text corpus is for the Huế variation only. The source is the Dictionary of the Huế Variation (Từ Điển Tiếng Huế) by Dr. Bùi Minh Đức. The author has granted a permission to use the dictionary. There are about 10,700 sentences posted on Common Voice."
Validated sentences: 16,663
| Category | Count |
|---|---|
| Unvalidated sentences | 5,395 |
| Pending sentences | 5,292 |
| Rejected sentences | 103 |
| Reported sentences | 199 |
The corpus contains 22,058 sentences: 16,663 validated and 5,395 unvalidated (5,292 pending review, 103 rejected), with 199 reported for review.
"The writing system consists of the 26 letters of the English alphabet minus f, j, w, and z. ̣ (These letters are, however, found in foreign loanwords.) and seven modified letters using diacritics: đ, ă, â, ê, ô, ơ, and ư. "
Please refer to Alphabet and Character Frequency: Vietnamese (Việt) (https://www.sttmedia.com/characterfrequency-vietnamese#alphabet)
There follows a randomly selected sample of five sentences from the corpus.
Anh chị cho ăn bánh bèo đến ngã ngửa
Nuôi nhiều người trong nhà mà toàn là thứ vô tích sự
Ngó bộ mặt giận nên mần thinh mần thấu
Ra đường không chịu đi cho duyên dáng mà cứ ưng đi ắc dơ
Nhảy đầm sinh nhảy đì, sẽ sinh lắm chuyện
Dictionary of the Huế Variation (Từ Điển Tiếng Huế) by Dr. Bùi Minh Đức.
| Source | Sentences |
|---|---|
| Từ Điển Tiếng Huế by Bùi Minh Đức, 2001, First Edition | 10,751 (64.5%) |
| sentence-collector | 5,710 (34.3%) |
| Other | 202 (1.2%) |
General
| Code | Domain | Clips | Speakers |
|---|---|---|---|
| general | General | 2,128 (10.3%) | 79 (18.9%) |
| agriculture_food | Agriculture and Food | - | - |
| automotive_transport | Automotive and Transport | - | - |
| finance | Finance | - | - |
| service_retail | Service and Retail | - | - |
| healthcare | Healthcare | - | - |
| history_law_government | History, Law and Government | - | - |
| media_entertainment | Media and Entertainment | - | - |
| nature_environment | Nature and Environment | - | - |
| news_current_affairs | News and Current Affairs | - | - |
| technology_robotics | Technology and Robotics | - | - |
| language_fundamentals | Language Fundamentals | - | - |
The dictionary is copied page by page in the PDF format.
Use the online OCR to text conversion service Convertio to convert the files to a text format.
Process the text files using Python to extract example sentences.
Manually proof read the result to correct OCR typos and remove extra sentences.
Due to Common Voice limits, only sentences with 5 or more words and less than 15 words are submitted.
Each row of a tsv file represents a single audio clip, and contains the following information:
client_id - hashed UUID of a given user
path - relative path of the audio file
sentence - the sentence to be read aloud
sentence_id - unique identifier for the sentence
sentence_domain - domain classification(s) of the sentence
up_votes - number of people who said audio matches the text
down_votes - number of people who said audio does not match text
age - age of the speaker1
gender - gender of the speaker1
accents - accents of the speaker1
variant - variant of the language1
locale - locale code of the language
segment - if sentence belongs to a custom dataset segment, it will be listed here
validated_sentences.tsvThe validated_sentences.tsv file contains one row per validated sentence in the text corpus:
sentence_id - unique identifier for the sentence
sentence - the sentence text
variant - the variant of the language
sentence_domain - the domain(s) the sentence belongs to
source - the source the sentence was collected from
is_used - whether the sentence is still in circulation for recording
clips_count - number of clips recorded for this sentence
unvalidated_sentences.tsvThe unvalidated_sentences.tsv file contains one row per unvalidated sentence in the text corpus:
sentence_id - unique identifier for the sentence
sentence - the sentence text
variant - the variant of the language
sentence_domain - the domain(s) the sentence belongs to
source - the source the sentence was collected from
up_votes - number of upvotes the sentence received
down_votes - number of downvotes the sentence received
status - current status of the sentence (pending or rejected)
Presentation in Vietnamese at https://cvtienghue.github.io/presentation
https://cvtienghue.github.io/presentation
Chung V Le , Phung Xung Thi Le
This dataset is released under the Creative Commons Zero (CC-0) licence. By downloading this data you agree to not determine the identity of speakers in the dataset.
For a full list of age, gender, and accent options, see the demographics spec. These will only be reported if the speaker opted in to provide that information. ↩ ↩2 ↩3 ↩4