You need to agree to share your contact information to access this dataset

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this dataset content.

πŸŒ‰ BharatSetu 22-Indic 1M

One English sentence β†’ 22 Indic languages. Nearly a million rows. One wide table.

Rows Languages Source Format Access Type

Quick start Β· Languages Β· Schema Β· Loading guide Β· Files Β· Limitations


πŸ“Š At a glance

Rows 984,764
Source language English (en)
Target languages 22 Indic languages
Layout Wide: one row per English sentence, one column per target language
Translation model family sarvamai/sarvam-translate
Viewer / load_dataset files data/train-*.parquet
Preserved artifacts JSONL, XLSX, single-file Parquet, merge manifest, metadata

🧭 How it's built

flowchart LR
    A[English source corpus<br/>CSV with provenance] --> B[sarvam-translate<br/>22 target languages]
    B --> C[Merge and packaging job]
    C --> D[data/ Parquet shards<br/>Viewer + load_dataset]
    C --> E[jsonl/ full export]
    C --> F[full/ XLSX workbook]
    C --> G[parquet/ single file]
    C --> H[metadata.json + manifest]

image

πŸš€ Quick start

from datasets import load_dataset

ds = load_dataset("Omarrran/BharatSetu-22Indic-1M", split="train")
print(ds)
print(ds[0]["english"], "->", ds[0]["Hindi"])

🌐 Languages

# Language Code # Language Code
1 Assamese as 12 Manipuri mni
2 Bengali bn 13 Marathi mr
3 Bodo brx 14 Nepali ne
4 Dogri doi 15 Odia or
5 Gujarati gu 16 Punjabi pa
6 Hindi hi 17 Sanskrit sa
7 Kannada kn 18 Santali sat
8 Kashmiri ks 19 Sindhi sd
9 Konkani kok 20 Tamil ta
10 Maithili mai 21 Telugu te
11 Malayalam ml 22 Urdu ur

Column names in the dataset use the full language name (e.g. Kashmiri), not the code.

🧱 Schema

Column Type Description
id int Row id aligned to the original source corpus
english string Source English text
source string Provenance field preserved from the input CSV
Assamese … Urdu string Machine translation into each of the 22 target languages
Full column list (25 columns)

id, english, source, Assamese, Bengali, Bodo, Dogri, Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Odia, Punjabi, Sanskrit, Santali, Sindhi, Tamil, Telugu, Urdu

Example row (shape only):

{
  "id": 0,
  "english": "…",
  "source": "…",
  "Assamese": "…",
  "Bengali": "…",
  "…": "…",
  "Urdu": "…"
}

πŸ“₯ Loading guide

1. πŸ€— datasets (recommended)

from datasets import load_dataset

ds = load_dataset("Omarrran/BharatSetu-22Indic-1M", split="train")
print(ds.column_names)

2. Streaming (no full download)

from datasets import load_dataset

stream = load_dataset("Omarrran/BharatSetu-22Indic-1M", split="train", streaming=True)
for i, row in enumerate(stream):
    print(row["english"], "|", row["Tamil"])
    if i == 4:
        break

3. One language, or a language pair

# English -> Kashmiri
ks = ds.select_columns(["id", "source", "english", "Kashmiri"])
print(ks[0])

# Hindi <-> Urdu pair
hi_ur = ds.select_columns(["id", "english", "Hindi", "Urdu"])

4. Load only the columns you need (faster, smaller)

from datasets import load_dataset

ds = load_dataset(
    "Omarrran/BharatSetu-22Indic-1M",
    split="train",
    columns=["id", "english", "Kashmiri"],
)

5. Pandas straight from the Hub

import pandas as pd

df = pd.read_parquet(
    "hf://datasets/Omarrran/BharatSetu-22Indic-1M/parquet/english_to_22_indic_all_languages.parquet",
    columns=["id", "english", "Hindi"],
)
df.head()

6. Polars

import polars as pl

df = pl.read_parquet(
    "hf://datasets/Omarrran/BharatSetu-22Indic-1M/data/train-*.parquet",
    columns=["id", "english", "Bengali"],
)
print(df.head())

7. DuckDB (SQL over the shards)

import duckdb

duckdb.sql("""
    SELECT id, english, Sanskrit
    FROM 'hf://datasets/Omarrran/BharatSetu-22Indic-1M/data/train-*.parquet'
    LIMIT 5
""").show()

8. PyArrow

import pyarrow.parquet as pq
from huggingface_hub import hf_hub_download

path = hf_hub_download(
    repo_id="Omarrran/BharatSetu-22Indic-1M",
    repo_type="dataset",
    filename="parquet/english_to_22_indic_all_languages.parquet",
)
table = pq.read_table(path, columns=["id", "english", "Marathi"])
print(table.num_rows)

9. Download a specific file

from huggingface_hub import hf_hub_download

jsonl_path = hf_hub_download(
    repo_id="Omarrran/BharatSetu-22Indic-1M",
    repo_type="dataset",
    filename="jsonl/english_to_22_indic_all_languages.jsonl",
)

10. Read the JSONL line by line

import json

with open(jsonl_path, encoding="utf-8") as f:
    for line in f:
        row = json.loads(line)
        print(row["english"], "|", row["Gujarati"])
        break

11. Snapshot the whole repo

from huggingface_hub import snapshot_download

local_dir = snapshot_download(
    repo_id="Omarrran/BharatSetu-22Indic-1M",
    repo_type="dataset",
    local_dir="./BharatSetu-22Indic-1M",
)

Only the Viewer shards:

snapshot_download(
    repo_id="Omarrran/BharatSetu-22Indic-1M",
    repo_type="dataset",
    allow_patterns=["data/*.parquet", "metadata.json"],
    local_dir="./bharatsetu-data",
)

12. CLI

huggingface-cli download Omarrran/BharatSetu-22Indic-1M \
  --repo-type dataset \
  --include "data/*.parquet" \
  --local-dir ./bharatsetu-data

13. Git LFS

git lfs install
git clone https://huggingface.co/datasets/Omarrran/BharatSetu-22Indic-1M

14. Wide β†’ long (translation-pair) format

Handy for training NMT models on (source, target, language) triples.

TARGET_LANGUAGES = [
    "Assamese", "Bengali", "Bodo", "Dogri", "Gujarati", "Hindi", "Kannada",
    "Kashmiri", "Konkani", "Maithili", "Malayalam", "Manipuri", "Marathi",
    "Nepali", "Odia", "Punjabi", "Sanskrit", "Santali", "Sindhi", "Tamil",
    "Telugu", "Urdu",
]

def to_long(dataset, languages=TARGET_LANGUAGES):
    for row in dataset:
        for lang in languages:
            yield {
                "id": row["id"],
                "source_language": "English",
                "target_language": lang,
                "source_text": row["english"],
                "target_text": row[lang],
            }

# Lazy iteration
for pair in to_long(ds):
    print(pair)
    break

Materialise as a Hugging Face dataset:

from datasets import Dataset

long_ds = Dataset.from_generator(lambda: to_long(ds, ["Hindi", "Kashmiri"]))

15. Filter, sample, and split

# Filter by provenance
subset = ds.filter(lambda r: r["source"] == "your_source_name")

# Reproducible sample
small = ds.shuffle(seed=42).select(range(10_000))

# Train / validation split
splits = ds.train_test_split(test_size=0.01, seed=42)

16. Export to CSV

ds.select_columns(["id", "english", "Telugu"]).to_csv("english_telugu.csv")

πŸ—‚οΈ Repository files

.
β”œβ”€β”€ data/
β”‚   └── train-*.parquet                          # Viewer + load_dataset source
β”œβ”€β”€ jsonl/
β”‚   └── english_to_22_indic_all_languages.jsonl  # full wide JSONL
β”œβ”€β”€ full/
β”‚   └── english_to_22_indic_all_languages.xlsx   # full wide Excel workbook
β”œβ”€β”€ parquet/
β”‚   └── english_to_22_indic_all_languages.parquet  # single-file Parquet
β”œβ”€β”€ generation/
β”‚   └── merge_manifest.json                      # packaging job manifest
β”œβ”€β”€ metadata.json                                # row counts, schema, sizes, hashes
└── README.md
Path Use it for
data/train-*.parquet Normal use: Dataset Viewer and load_dataset
jsonl/…jsonl Line-by-line processing and preservation
full/…xlsx Manual inspection in Excel
parquet/…parquet Single-file Parquet for pandas, Arrow, DuckDB
metadata.json Verifying row counts, schema, file sizes and hashes
generation/merge_manifest.json Auditing how the packaging job ran

βœ… Suggested use cases

  • Training and fine-tuning multilingual or Indic machine translation models
  • Bootstrapping data for low-resource languages (Bodo, Dogri, Kashmiri, Santali, Manipuri and others)
  • Distillation and synthetic-data pipelines
  • Cross-lingual retrieval and alignment research
  • Building evaluation sets after human review of a sampled subset

⚠️ Limitations

  • Machine-generated. This is a translation corpus produced by a model, not a human-verified benchmark.
  • Quality varies by language, domain, and sentence style. Lower-resource languages are likely to have more errors.
  • No gold references. Do not treat these translations as ground truth for evaluation without human review.
  • Inherited bias. Errors, biases, and artifacts of the source corpus and the translation model carry over.
  • Provenance. The source column is preserved from the input English corpus; check it before redistributing subsets.

πŸ” Access

This dataset is currently private. Request access from the maintainer before use.

πŸ“š Citation

@dataset{bharatsetu_22indic_1m,
  title     = {BharatSetu 22-Indic 1M: A Wide English-to-22-Indic Parallel Translation Dataset},
  author    = {Haq Nawaz Malik},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/Omarrran/BharatSetu-22Indic-1M}
}

🀝 Maintainer

Omar (Haq Nawaz Malik) Β· Hugging Face Β· GitHub Β· Portfolio

Bridging English and India's languages, one row at a time. πŸŒ‰

Downloads last month
41