Datasets:
- π At a glance
- π§ How it's built
- π Quick start
- π Languages
- π§± Schema
- π₯ Loading guide
- 1. π€
datasets(recommended) - 2. Streaming (no full download)
- 3. One language, or a language pair
- 4. Load only the columns you need (faster, smaller)
- 5. Pandas straight from the Hub
- 6. Polars
- 7. DuckDB (SQL over the shards)
- 8. PyArrow
- 9. Download a specific file
- 10. Read the JSONL line by line
- 11. Snapshot the whole repo
- 12. CLI
- 13. Git LFS
- 14. Wide β long (translation-pair) format
- 15. Filter, sample, and split
- 16. Export to CSV
- 1. π€
- ποΈ Repository files
- β
Suggested use cases
- β οΈ Limitations
- π Access
- π Citation
- π€ Maintainer
π BharatSetu 22-Indic 1M
One English sentence β 22 Indic languages. Nearly a million rows. One wide table.
Quick start Β· Languages Β· Schema Β· Loading guide Β· Files Β· Limitations
π At a glance
| Rows | 984,764 |
| Source language | English (en) |
| Target languages | 22 Indic languages |
| Layout | Wide: one row per English sentence, one column per target language |
| Translation model family | sarvamai/sarvam-translate |
Viewer / load_dataset files |
data/train-*.parquet |
| Preserved artifacts | JSONL, XLSX, single-file Parquet, merge manifest, metadata |
π§ How it's built
flowchart LR
A[English source corpus<br/>CSV with provenance] --> B[sarvam-translate<br/>22 target languages]
B --> C[Merge and packaging job]
C --> D[data/ Parquet shards<br/>Viewer + load_dataset]
C --> E[jsonl/ full export]
C --> F[full/ XLSX workbook]
C --> G[parquet/ single file]
C --> H[metadata.json + manifest]
π Quick start
from datasets import load_dataset
ds = load_dataset("Omarrran/BharatSetu-22Indic-1M", split="train")
print(ds)
print(ds[0]["english"], "->", ds[0]["Hindi"])
π Languages
| # | Language | Code | # | Language | Code |
|---|---|---|---|---|---|
| 1 | Assamese | as |
12 | Manipuri | mni |
| 2 | Bengali | bn |
13 | Marathi | mr |
| 3 | Bodo | brx |
14 | Nepali | ne |
| 4 | Dogri | doi |
15 | Odia | or |
| 5 | Gujarati | gu |
16 | Punjabi | pa |
| 6 | Hindi | hi |
17 | Sanskrit | sa |
| 7 | Kannada | kn |
18 | Santali | sat |
| 8 | Kashmiri | ks |
19 | Sindhi | sd |
| 9 | Konkani | kok |
20 | Tamil | ta |
| 10 | Maithili | mai |
21 | Telugu | te |
| 11 | Malayalam | ml |
22 | Urdu | ur |
Column names in the dataset use the full language name (e.g. Kashmiri), not the code.
π§± Schema
| Column | Type | Description |
|---|---|---|
id |
int |
Row id aligned to the original source corpus |
english |
string |
Source English text |
source |
string |
Provenance field preserved from the input CSV |
Assamese β¦ Urdu |
string |
Machine translation into each of the 22 target languages |
Full column list (25 columns)
id, english, source, Assamese, Bengali, Bodo, Dogri, Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Odia, Punjabi, Sanskrit, Santali, Sindhi, Tamil, Telugu, Urdu
Example row (shape only):
{
"id": 0,
"english": "β¦",
"source": "β¦",
"Assamese": "β¦",
"Bengali": "β¦",
"β¦": "β¦",
"Urdu": "β¦"
}
π₯ Loading guide
1. π€ datasets (recommended)
from datasets import load_dataset
ds = load_dataset("Omarrran/BharatSetu-22Indic-1M", split="train")
print(ds.column_names)
2. Streaming (no full download)
from datasets import load_dataset
stream = load_dataset("Omarrran/BharatSetu-22Indic-1M", split="train", streaming=True)
for i, row in enumerate(stream):
print(row["english"], "|", row["Tamil"])
if i == 4:
break
3. One language, or a language pair
# English -> Kashmiri
ks = ds.select_columns(["id", "source", "english", "Kashmiri"])
print(ks[0])
# Hindi <-> Urdu pair
hi_ur = ds.select_columns(["id", "english", "Hindi", "Urdu"])
4. Load only the columns you need (faster, smaller)
from datasets import load_dataset
ds = load_dataset(
"Omarrran/BharatSetu-22Indic-1M",
split="train",
columns=["id", "english", "Kashmiri"],
)
5. Pandas straight from the Hub
import pandas as pd
df = pd.read_parquet(
"hf://datasets/Omarrran/BharatSetu-22Indic-1M/parquet/english_to_22_indic_all_languages.parquet",
columns=["id", "english", "Hindi"],
)
df.head()
6. Polars
import polars as pl
df = pl.read_parquet(
"hf://datasets/Omarrran/BharatSetu-22Indic-1M/data/train-*.parquet",
columns=["id", "english", "Bengali"],
)
print(df.head())
7. DuckDB (SQL over the shards)
import duckdb
duckdb.sql("""
SELECT id, english, Sanskrit
FROM 'hf://datasets/Omarrran/BharatSetu-22Indic-1M/data/train-*.parquet'
LIMIT 5
""").show()
8. PyArrow
import pyarrow.parquet as pq
from huggingface_hub import hf_hub_download
path = hf_hub_download(
repo_id="Omarrran/BharatSetu-22Indic-1M",
repo_type="dataset",
filename="parquet/english_to_22_indic_all_languages.parquet",
)
table = pq.read_table(path, columns=["id", "english", "Marathi"])
print(table.num_rows)
9. Download a specific file
from huggingface_hub import hf_hub_download
jsonl_path = hf_hub_download(
repo_id="Omarrran/BharatSetu-22Indic-1M",
repo_type="dataset",
filename="jsonl/english_to_22_indic_all_languages.jsonl",
)
10. Read the JSONL line by line
import json
with open(jsonl_path, encoding="utf-8") as f:
for line in f:
row = json.loads(line)
print(row["english"], "|", row["Gujarati"])
break
11. Snapshot the whole repo
from huggingface_hub import snapshot_download
local_dir = snapshot_download(
repo_id="Omarrran/BharatSetu-22Indic-1M",
repo_type="dataset",
local_dir="./BharatSetu-22Indic-1M",
)
Only the Viewer shards:
snapshot_download(
repo_id="Omarrran/BharatSetu-22Indic-1M",
repo_type="dataset",
allow_patterns=["data/*.parquet", "metadata.json"],
local_dir="./bharatsetu-data",
)
12. CLI
huggingface-cli download Omarrran/BharatSetu-22Indic-1M \
--repo-type dataset \
--include "data/*.parquet" \
--local-dir ./bharatsetu-data
13. Git LFS
git lfs install
git clone https://huggingface.co/datasets/Omarrran/BharatSetu-22Indic-1M
14. Wide β long (translation-pair) format
Handy for training NMT models on (source, target, language) triples.
TARGET_LANGUAGES = [
"Assamese", "Bengali", "Bodo", "Dogri", "Gujarati", "Hindi", "Kannada",
"Kashmiri", "Konkani", "Maithili", "Malayalam", "Manipuri", "Marathi",
"Nepali", "Odia", "Punjabi", "Sanskrit", "Santali", "Sindhi", "Tamil",
"Telugu", "Urdu",
]
def to_long(dataset, languages=TARGET_LANGUAGES):
for row in dataset:
for lang in languages:
yield {
"id": row["id"],
"source_language": "English",
"target_language": lang,
"source_text": row["english"],
"target_text": row[lang],
}
# Lazy iteration
for pair in to_long(ds):
print(pair)
break
Materialise as a Hugging Face dataset:
from datasets import Dataset
long_ds = Dataset.from_generator(lambda: to_long(ds, ["Hindi", "Kashmiri"]))
15. Filter, sample, and split
# Filter by provenance
subset = ds.filter(lambda r: r["source"] == "your_source_name")
# Reproducible sample
small = ds.shuffle(seed=42).select(range(10_000))
# Train / validation split
splits = ds.train_test_split(test_size=0.01, seed=42)
16. Export to CSV
ds.select_columns(["id", "english", "Telugu"]).to_csv("english_telugu.csv")
ποΈ Repository files
.
βββ data/
β βββ train-*.parquet # Viewer + load_dataset source
βββ jsonl/
β βββ english_to_22_indic_all_languages.jsonl # full wide JSONL
βββ full/
β βββ english_to_22_indic_all_languages.xlsx # full wide Excel workbook
βββ parquet/
β βββ english_to_22_indic_all_languages.parquet # single-file Parquet
βββ generation/
β βββ merge_manifest.json # packaging job manifest
βββ metadata.json # row counts, schema, sizes, hashes
βββ README.md
| Path | Use it for |
|---|---|
data/train-*.parquet |
Normal use: Dataset Viewer and load_dataset |
jsonl/β¦jsonl |
Line-by-line processing and preservation |
full/β¦xlsx |
Manual inspection in Excel |
parquet/β¦parquet |
Single-file Parquet for pandas, Arrow, DuckDB |
metadata.json |
Verifying row counts, schema, file sizes and hashes |
generation/merge_manifest.json |
Auditing how the packaging job ran |
β Suggested use cases
- Training and fine-tuning multilingual or Indic machine translation models
- Bootstrapping data for low-resource languages (Bodo, Dogri, Kashmiri, Santali, Manipuri and others)
- Distillation and synthetic-data pipelines
- Cross-lingual retrieval and alignment research
- Building evaluation sets after human review of a sampled subset
β οΈ Limitations
- Machine-generated. This is a translation corpus produced by a model, not a human-verified benchmark.
- Quality varies by language, domain, and sentence style. Lower-resource languages are likely to have more errors.
- No gold references. Do not treat these translations as ground truth for evaluation without human review.
- Inherited bias. Errors, biases, and artifacts of the source corpus and the translation model carry over.
- Provenance. The
sourcecolumn is preserved from the input English corpus; check it before redistributing subsets.
π Access
This dataset is currently private. Request access from the maintainer before use.
π Citation
@dataset{bharatsetu_22indic_1m,
title = {BharatSetu 22-Indic 1M: A Wide English-to-22-Indic Parallel Translation Dataset},
author = {Haq Nawaz Malik},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/Omarrran/BharatSetu-22Indic-1M}
}
π€ Maintainer
Omar (Haq Nawaz Malik) Β· Hugging Face Β· GitHub Β· Portfolio
Bridging English and India's languages, one row at a time. π
- Downloads last month
- 41
