Release Date: 9/2/2026
Format: TXT
Size: 2.76 MB
Share
The Associated Press Pakistan 1.8 Million Chinese Corpus is a Chinese-language news dataset comprising 8,867 articles (approximately 1.8 million words) published by the Associated Press of Pakistan on its Chinese-language section at app.com.pk, covering the period 2022–2026. APP retains full and exclusive ownership of the data, with MirasAI serving solely as documentation and hosting facilitator access is restricted and approval-based via the Mozilla Data Collective, for non-commercial research use only, with no sale, sublicensing, or commercial redistribution permitted, and APP may request takedown within a 48-hour compliance window. The dataset is intended for non-commercial NLP research, including Chinese language modeling and Pakistan–China media studies, and all users are required to credit the Associated Press of Pakistan as the sole and primary source.
Licensing
Creative Commons Attribution Non Commercial Share Alike 4.0 International (CC-BY-NC-SA-4.0)
https://spdx.org/licenses/CC-BY-NC-SA-4.0.htmlRestrictions/Special Constraints
Access is restricted and approval-based for non-commercial research, including Chinese language modeling and Pakistan–China media studies. The Associated Press of Pakistan retains full ownership, while MirasAI acts solely as a documentation and hosting facilitator. Sale, sublicensing, and commercial redistribution are prohibited.
Forbidden Usage
Commercial use, sale, sublicensing, redistribution, and commercial derivative works are prohibited. Use is limited to non-commercial NLP research, including Chinese language modeling and Pakistan–China media studies. MirasAI LLC must not be represented as the data owner; the Associated Press of Pakistan retains full ownership.
Ethical Review
This corpus is uploaded and hosted with the express permission of the Associated Press of Pakistan. Associated Press of Pakistan retains full and exclusive ownership of the data. MirasAI acts solely as documentation and hosting facilitator. Access on the Mozilla Data Collective is restricted and approval-based, for non-commercial research use only.
Intended Use
This dataset is intended for non-commercial NLP research, including Chinese language modeling, cross-lingual research, and Pakistan–China media studies. All users must acknowledge the Associated Press of Pakistan as the primary source.
The Associated Press Pakistan (APP) 1.8 Million Chinese Corpus is a Chinese-language news corpus comprising articles published through the Chinese-language section of the Associated Press of Pakistan, Pakistan’s national news agency. Established in 1947, the Associated Press of Pakistan serves as the country’s premier news service, providing news coverage across multiple languages, including English, Urdu, Chinese, Arabic, Sindhi, Saraiki, Pashto, Balochi, and Brahvi.
This dataset specifically contains Chinese-language text and represents news coverage related to Pakistan–China relations, bilateral cooperation, diplomacy, politics, economics, culture, and other areas of shared interest. With approximately 1.8 million words, the corpus provides a valuable resource for studying contemporary Chinese-language media produced by a Pakistani national news agency.
Chinese (Mandarin) is a Sino-Tibetan language and the most widely spoken language in the world, serving as the official language of China and one of the official languages of Pakistan's key regional partnerships. It is widely used in education, media, literature, business, and everyday communication across the Chinese-speaking world. Chinese has a rich linguistic and literary heritage spanning thousands of years and plays an important role in cross-lingual and machine translation research, particularly in the context of Pakistan–China media and economic ties.
| Field | Value |
|---|---|
| Articles | 8,867 |
| Word count | ~1.8 million |
| Coverage | 2022–2026 |
| Language | Chinese (zho) |
| Source | app.com.pk (Chinese section) |
Articles were collected from Associated Press Pakistan's publicly published Chinese-language section on app.com.pk using an automated web crawler developed by MirasAI. Each article's HTML was parsed to extract the headline, body text, publication date, and source URL; articles were deduplicated, cleaned of boilerplate, and normalized to plain text in UTF-8. Crawling was rate-limited and conducted with the knowledge and permission of Associated Press Pakistan.
Chinese.tar
└── Chinese/
├── 2022.txt
├── 2023.txt
├── 2024.txt
├── 2025.txt
└── 2026.txt
巴通社伊斯兰堡5月17日电 周二中国驻巴基斯坦大使馆发言人表示,巴基斯坦的所有孔子学院都在运作,并未关闭。
中国大使馆发言人在回答()提问时表示,巴基斯坦各孔子学院和孔子课堂的所有教学活动都将由中巴教师和中方合作大学通过线上或线下方式开展。
4月26日,卡拉奇大学孔子学院外发生自杀式袭击,造成至少四人死亡,其中包括三名中国教师,四位受伤。
中国大使馆发言人表示,经与巴基斯坦有关部门协商,部分中国教师为过暑假已回国,将根据要求,适时返回巴基斯坦。
同时表示,为满足巴基斯坦学生学习汉语需求,中方计划提供优质教学资源。
This corpus is uploaded and hosted with the express permission of the Associated Press of Pakistan. The Associated Press of Pakistan retains full and exclusive ownership of the data; MirasAI acts solely as documentation and hosting facilitator. Access on the Mozilla Data Collective is restricted and approval-based, for non-commercial research use only. No sale, sublicensing, or commercial redistribution is permitted.
Intended for non-commercial NLP research, including Chinese language modeling and Pakistan–China media studies. All users must credit the Associated Press of Pakistan as the sole and primary source.