Task: N/A
Release Date: 8/28/2026
Format: MP4, SRT, CSV
Size: 130.40 GB
Share
This dataset provides 20 hours of premium, rights-cleared vintage Hindi scripted film content sourced from established production catalogues. Each title is delivered as a 1080p+ MP4 file accompanied by a structured metadata sheet covering title, production year, synopsis, genre, bitrate, resolution, cast, director, country of origin, and tags. English subtitles are included across all titles. The collection spans a broad range of narrative genres and is purpose-built for vision-language model training, scene understanding, and multimodal fine-tuning workflows.
Pricing details
Your purchase is the license to the raw data. Once purchased, you're responsible for storing and using this dataset.
Your purchase includes a license to the data, paid directly to the dataset vendor, and a 5% ($100.00) platform fee paid to MDC.
Licensing
CONTENT X LABS DATA LICENCE AGREEMENT 1.0
https://pebble-mouse-502.notion.site/mdc-data-licence-agreement-1-0Restrictions/Special Constraints
Solely for purpose of training artificial intelligence learning models. No encouragement of violence, fraud or illegal activities in accordance to Hollywood practices and guidelines.
Forbidden Usage
No cloning or reproduction of any measure of content and specifically no cloning of actors' faces, voices and music in the output, etc... Shall not be used for illegal purposes according to local law including commiting cyber-crimes or acts of violence.
Ethical Review
All content within this dataset is supplied by rights-holding content owners who have explicitly granted licensing permissions for AI training purposes under direct contractual agreement. Providers are fully informed of the intended use of their content at the point of licensing, and no content is included without express grant of rights covering AI and machine learning applications. Content owners retain ownership of their intellectual property throughout, and no personally identifiable information pertaining to individuals featured in the content is collected or distributed as part of this dataset.
Intended Use
The collection spans a broad range of narrative genres and is purpose-built for vision-language model training, scene understanding, and multimodal fine-tuning workflows.
Rights-cleared Hindi vintage scripted narrative video (~20 hours) with metadata and English subtitles.
Source catalogue is ContentC. Each title includes MP4 video plus sidecar subtitles and per-title metadata.
Packaged from the ContentC / Allrites rights-cleared catalogue for submission to Mozilla Data Collective.
Commercial catalogue titles already cleared for this distribution. No new crowdsourced recording was commissioned for this package.
Movies: 9
Series episodes: 0
Subtitle files: 9
Total runtime: 20.59 hours (74137.9 sec)
Release years: 1967–1992
Declared locale: hi
Typical runtime: 1:59:12–2:46:13 (mean 2:17:18)
Video bitrate: 15.0–15.0 Mbps (mean 15.0 Mbps)
Hindi (9)
English (9)
India (9)
Drama (6)
Romance (3)
Family (2)
Fantasy (1)
Musical (1)
Thriller (1)
NR (9)
Container: MP4
Resolutions:
1920x1080 (9)
Frame rate:
25.0 (9)
Aspect ratio:
16:9 (9)
Audio:
stereo (9)
Each content.csv row also includes bitrate, synopsis, director, cast, and tags.
Titles were selected from a rights-cleared scripted-narrative catalogue (Tier-1 / ContentC). Metadata (title, year, rating, synopsis, language, country, genres, credits, and media technical fields) comes from the production catalogue database. Subtitle files are the catalogue sidecar captions/subtitles stored with each video, not ASR generated for this package.
No additional annotation campaign was run. Ratings, tags, and synopses are catalogue values and may be incomplete or inconsistent across titles.
README.md — package overview and directory layout
datasheet.md — this technical datasheet
metadata/content.csv — one row per movie
metadata/subtitle.csv — subtitle rows joined to video via Filename
metadata/schema.json — column names and types
media/movies/ contains one .mp4 per title (path in content.csv Filename)
subtitles/movies/ contains matching subtitle files (path in subtitle.csv)
Join path: content.csv.Filename = video path = subtitle.csv.Filename.
Subtitle file location is subtitle.csv.Subtitle Path.
import pandas as pd
from pathlib import Path
root = Path("hindi-vintage-scripted-narrative-rights-cleared-20-hours")
c pd.read_csv(root / "metadata" / "content.csv")
subs = pd.read_csv(root / "metadata" / "subtitle.csv")
df = content.merge(subs, on="Filename", how="left", suffixes=("", "_subtitle"))
# video path: root / row.Filename
# subtitle path: root / row["Subtitle Path"]
Kasak (1992) — 1920x1080, 2:22:24
Bade Ghar Ki Beti (1989) — 1920x1080, 2:15:47
Santosh (1989) — 1920x1080, 2:07:55
Maa Beti (1987) — 1920x1080, 2:25:32
Chhoti Si Baat (1975) — 1920x1080, 1:59:12
Khote Sikkay (1974) — 1920x1080, 2:05:39
Andaz (1971) — 1920x1080, 2:22:39
Bramha Vishnu Mahesh (1971) — 1920x1080, 2:10:17
Full list:
| Title | Year | Rating | Runtime | Resolution | Country |
|---|---|---|---|---|---|
| Kasak | 1992 | NR | 2:22:24 | 1920x1080 | India |
| Bade Ghar Ki Beti | 1989 | NR | 2:15:47 | 1920x1080 | India |
| Santosh | 1989 | NR | 2:07:55 | 1920x1080 | India |
| Maa Beti | 1987 | NR | 2:25:32 | 1920x1080 | India |
| Chhoti Si Baat | 1975 | NR | 1:59:12 | 1920x1080 | India |
| Khote Sikkay | 1974 | NR | 2:05:39 | 1920x1080 | India |
| Andaz | 1971 | NR | 2:22:39 | 1920x1080 | India |
| Bramha Vishnu Mahesh | 1971 | NR | 2:10:17 | 1920x1080 | India |
| Hamraaz | 1967 | NR | 2:46:13 | 1920x1080 | India |
Intended: ML research / evaluation on rights-cleared narrative video and subtitles.
Not intended: treating this corpus as openly licensed / public-domain video; re-hosting titles outside Mozilla Data Collective and the catalogue license; using tags or synopses as gold-standard labels without review.
Some catalogue tags are noisy or thematically mismatched to the title.
A few country, director, or cast fields are empty.
Subtitles may be captions rather than verbatim transcripts and may not cover on-screen text or songs.
Sample packages (if marked above) contain a middle clip per title, not the full feature.
Content includes commercially rated narrative films and may contain mature themes consistent with the listed ratings.
Custom license — rights-cleared catalogue titles. Redistribution is limited to Mozilla Data Collective terms and any additional conditions on the listing.
This snapshot is generated from the catalogue export used to build the tarball. Updates require a new package build rather than in-place mutation of files already uploaded to Mozilla Data Collective.
Generated: 2026-08-27 11:43 UTC