Release Date: 10/2/2026
Format: JSONL
Size: 391.72 KB
# MALOBA — Lingala OCR transcriptions, harmonised text (v1.1) Private dataset of Congo Digital Services (CDS), MALOBA project, with UNDP Republic of Congo. Text only (no images): reference transcriptions of handwritten and printed Lingala documents, annotated in Label Studio, harmonised to the MALOBA spelling convention (Unicode look-alikes fixed, tone marks removed, majority ɔ/ɛ spelling). Monolingual Lingala, intended as source text for back-translation (NLLB-Maloba). - `ocr.jsonl`: 1,282 unique lines (duplicates and unrevised pre-annotations removed), with `texte_brut` (original) and `modifications` (each replacement: step, before, after). - `ocr_pages.jsonl`: 104 pages rebuilt in reading order. - `ocr_nllb.jsonl` / `.parquet`: 535 units for NLLB (67 pages + 468 lines, 22,498 words); exclusions listed in `ocr_nllb_exclusions.jsonl` (lines under 3 words, 2 fragments close to the OCR test set). - The 40-line OCR test set is NOT included and must stay out of any training. - Known limits: two French loanwords lost their accent through the rule (`Thème` → `Theme` ×6, `bénéfice` → `benefice` ×1); 61 accented forms absent from the lexicon are kept as written; a few rare letters (ɑ, ı, ł, š, õ, ĩ) have no rule. Details in `rapport_finalisation_ocr.md`. - No annotator e-mail, name or pseudonym.
Licensing
Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)
https://licensingafricandatasets.com/nwulite-obodo-licenseRestrictions/Special Constraints
Private dataset: access limited to the CDS community on Mozilla Data Collective. Use under the NOODL 1.0 licence.
Forbidden Usage
Mixing with the MALOBA OCR test set (not included) for training.
Ethical Review
These datasets contain no personal data. The texts come from documents provided by SIL Congo and from prompts written for the project. They were transcribed, translated or annotated by Lingala-speaking annotators recruited by CDS, then checked by a lead linguist. Handwritten documents were anonymised before annotation. The identities of annotators and validators are not published.
Intended Use
Monolingual Lingala source text for back-translation and language modelling (NLLB-Maloba).
Collection. Source images come from Lingala documents held by SIL Congo: 26 handwritten documents, photographed and anonymised before annotation, and printed documents converted to images on the project's collection platform (Partkole). Annotators drew a rectangle or polygon around each text region in Label Studio and transcribed it exactly, labelling it as printed or handwritten. The project's 117 annotation tasks produced 1,747 text regions.
From annotations to this dataset. The regions were cleaned before use:
empty or incomplete boxes were removed (18); machine pre-annotations never corrected by a person were removed (263); duplicate texts were removed (143).
This leaves 1,323 unique transcriptions validated by a person. This dataset is a sample built from them to target the two Lingala open vowels:
164 lines containing ɔ or ɛ and 164 ordinary lines (328 in total), split 85/15 before any oversampling into 278 training and 50 test lines; the training lines containing ɔ or ɛ were then duplicated once, giving 417 training rows.
The test split is not oversampled. Duplicated training rows are intentional and should not be counted as independent examples.
Format. Each row is a multimodal chat example (ChatML): the user message holds the image of one text region (a crop, not a full page) and the instruction to transcribe the text exactly; the assistant message holds the reference transcription. Transcriptions keep the original spelling, including ɔ, ɛ, Ɔ and Ɛ.
Quality. Three independent Lingala-speaking validators reviewed all 50 test items and 41 training items. On the test split, 84 % of reference transcriptions were judged correct and 96 % at least partly correct. On the training sample, the figures are 80.5 % and 90.2 %. Inter-rater agreement (Krippendorff's α = 0.53) is below the project's 0.60 threshold, so these rates are indicative.
Reference result. Qwen2-VL-2B fine-tuned on this dataset (Congo-digital-service/qwen-vl-lingala-qlora-vf) reaches 88.7 % character accuracy on the test split (CER 11.3 %, micro-averaged, normalised), against 60.0 % for the base model.
Known limitations.
The test split is small (50 items) and carries no source-document identifier, so results cannot be broken down by document. The collection is dominated by educational, agricultural and radio-script content from a single institutional source. Spelling follows each annotator; it is not harmonised in this dataset.
Related versions. qwen-vl-lingala-dataset-augmented (884 / 50 items) applies image degradations to 140 distinct training texts. It was not used to train the reference model. An earlier intermediate repository (1,480 / 82 rows) has been withdrawn.