Nawah-Parakeet-60M

A 62.7M-parameter Parakeet-style Token-and-Duration Transducer for Arabic speech recognition β€” FastConformer encoder, TDT objective β€” trained from scratch with no pretrained encoder and no external language model.

On the six test sets of the Open Universal Arabic ASR Leaderboard it reaches 31.34 average WER / 13.62 CER, which places it 8th of 38 evaluated systems β€” ahead of Voxtral-Small-24B, gemma-4-E4B, Qwen3-ASR-1.7B, nvidia-conformer-ctc-large-arabic and every Whisper model, at a fraction of their size. Every model above it is between 1B and 30B parameters, or a commercial API.

Leaderboard

Averages are recomputed over these six sets for every model so the ordering is consistent. Baseline figures are the leaderboard's published numbers.

# Model Avg WER Avg CER Params
1 audarai/Audar-ASR-V1-Turbo 23.17 9.22 2.35B
2 CohereLabs/cohere-transcribe-arabic-07-2026 25.87 11.80 β€”
3 omnilingual-asr/omniASR_LLM_7B 28.32 12.52 7B
4 omnilingual-asr/omniASR_LLM_1B 29.95 13.40 1B
5 omnilingual-asr/omniASR_LLM_3B 29.96 13.77 3B
6 CohereLabs/cohere-transcribe-03-2026 30.67 16.37 β€”
7 Qwen/Qwen3-Omni-30B-A3B-Instruct 30.71 13.66 30B
8 Nawah-Parakeet-60M (this model) 31.34 13.62 62.7M
9 audarai/Audar-ASR-V1-Flash 32.04 13.65 β€”
10 nvidia-conformer-ctc-large-arabic (lm) 32.91 13.85 0.12B
11 omnilingual-asr/omniASR_LLM_300M 32.96 14.84 300M
12 google/gemma-4-E4B-it 32.98 13.71 4B
13 Qwen/Qwen3-ASR-1.7B 33.69 12.33 1.7B
14 mistralai/Voxtral-Small-24B-2507 34.46 15.29 24B
16 openai/whisper-large-v3 36.86 17.21 1.55B
21 openai/whisper-large-v3-turbo 40.05 18.87 0.81B
33 openai/whisper-small 55.13 21.68 244M
38 speechbrain/asr-wav2vec2-commonvoice-14-ar 65.74 30.93 β€”

Full table in eval/ranking.json.

Per-set results

Scored with the leaderboard's own eval.py β€” its normalize_arabic_text and NeMo's word_error_rate β€” unmodified.

Test set Clips Hours WER CER board best board median
MASC clean 8,612 10.49 11.03 3.60 8.66 24.86
MGB-2 5,365 9.58 18.47 8.59 11.08 20.23
Common Voice 18 10,471 12.66 22.39 7.83 5.82 17.83
MASC noisy 9,394 8.92 30.45 13.27 19.01 35.64
SADA 6,186 10.75 43.05 20.38 28.92 60.11
Casablanca 6,818 7.85 62.67 28.03 47.02 69.37
Average 31.34 13.62 23.17 38.12

Test manifests were verified against the leaderboard's own datasets/*.json: reference text matches on 100% of Common Voice, MASC clean, MASC noisy and Casablanca clips, 99.2% of MGB-2 and 98.3% of SADA.

Scoring reproducibility

Two leaderboard models were re-run locally through this harness to confirm the scoring is faithful.

Audar-ASR-V1-Turbo (rank 1, 2.35B), using Audar's own reference inference code on their pinned transformers==4.57.6, seeded 600-clip subsets:

Avg SADA CV18 MASC clean MASC noisy MGB-2 Casablanca
leaderboard Space 23.17 28.92 8.09 16.73 27.19 11.08 47.02
Audar's own leaderboard.csv 24.78 29.41 8.60 19.60 28.35 11.13 51.58
measured here 24.93 32.26 8.69 18.00 28.57 12.21 49.86

The average reproduces Audar's own published figure to 0.15 WER, which confirms the audio preparation and the scorer. Note that Audar's two published sources differ by 1.6 average WER.

whisper-small (rank 33) does not reproduce per set: βˆ’15.25 on MGB-2, +14.76 on Common Voice, +12.51 on SADA, though its average lands within 0.61. The cause is measurable β€” under current transformers, whisper-small enters repetition loops on 6.5% of SADA clips and 3.5% of Casablanca clips, and excluding only those recovers 19.80 and 13.15 WER respectively. Published figures predate that version.

This model's TDT decoder caps symbols per encoder frame and shows 0 runaway hypotheses on Common Voice, MGB-2 and both MASC sets (0.02% SADA, 0.13% Casablanca), so it is not exposed to that failure.

Takeaway: absolute positions on the published table carry roughly 1–2 WER of uncertainty, more for models with unstable decoders.

Usage

pip install torch torchaudio tokenizers soundfile
from huggingface_hub import snapshot_download
import sys, soundfile as sf
from tokenizers import Tokenizer

d = snapshot_download("oddadmix/Nawah-Parakeet-60M")
sys.path.insert(0, d)
import parakeet_model as pm

tok   = Tokenizer.from_file(f"{d}/tokenizer_bpe1024.json")
model = pm.load(f"{d}/nawah_parakeet_60m.pt", device="cuda")

wav, sr = sf.read("clip.wav", dtype="float32")      # 16 kHz mono
print(pm.transcribe(model, wav, tok, device="cuda"))

config.json describes the architecture, feature front end and tokenizer in machine-readable form. It is not a transformers config: AutoModel.from_pretrained will not load this model β€” use parakeet_model.load() as above. Every value in it was read off the checkpoint rather than transcribed by hand.

The bundled tokenizer is the only one that will work. Any other 1024-unit BPE produces valid ids that mean different things, so a mismatch yields fluent-looking nonsense rather than an error.

Architecture

encoder FastConformer, 17 layers, d=384, 6 heads, FFN 1536, Γ—8 subsampling β†’ 12.5 Hz
decoder 2-layer LSTM predictor, 320-dim (2.09M params)
joiner additive, 384-dim β†’ 1024 vocab + 5 duration bins (0.40M)
total 62.71M
objective Token-and-Duration Transducer loss + auxiliary CTC (weight 0.2)
features 128-mel kaldi fbank, povey window, 25 ms / 10 ms, 16 kHz
tokenizer 1024-unit BPE, Metaspace pre-tokenizer and decoder, blank = id 0
training 2,404 h of Arabic speech, 2 epochs, ~32 h on one RTX 5090

No 30-second limit. The encoder uses clamped relative position bias rather than a fixed positional table, so audio of any length is transcribed in a single pass β€” unlike Whisper-family models, which need 30 s chunking. Measured on 2 CPU threads: 30 s in 0.3 s, 120 s in 2.1 s, 300 s in 10.8 s. Attention is O(TΒ²), so memory is the practical ceiling: ~3.2 GB at 5 minutes and ~11 GB at 15 minutes.

Limitations

  • Undiacritized output only. No tashkeel, no punctuation, no casing for Latin tokens.
  • Maghrebi remains the weakest region in absolute terms (Casablanca Morocco 69.15 WER, Mauritania 83.86, Algeria 77.04), even though it is where the model improved most.
  • The model does not emit Ψ©. Its training targets used Ω‡ throughout, so words that a reference writes with Ψ© are scored as errors on sets that use it. The leaderboard's normalizer folds the alef family but not Ψ©. Folding Ψ© on both sides changes MGB-2 by only 0.06 WER, so the residual effect is small.
  • Short clips can gain an invented trailing word β€” a property of the transducer's stop condition.
  • Single-pass long audio is out of distribution. Training clips were capped at 20 s; accuracy on multi-minute recordings has not been measured.
  • Audio must be 16 kHz mono; the kaldi fbank front end assumes it, and another rate misplaces every mel bin.

Files

nawah_parakeet_60m.pt            checkpoint (241 MB)
config.json                      architecture, features and tokenizer, machine-readable
tokenizer_bpe1024.json           the tokenizer this model requires
parakeet_model.py                standalone loader + TDT greedy decode, torch only
fastconformer.py                 encoder definition
parakeet_data.py                 fbank front end and encoder-length helper
eval/leaderboard_results.json    per-set scores for this model
eval/ranking.json                all evaluated models, per set
eval/lb_norm.py                  the leaderboard's scorer, ported
eval/run_bench.py                batch decoding to the leaderboard manifest format

Citation

Leaderboard and baseline figures: Wang, Alhmoud & Alqurishi, Open Universal Arabic ASR Leaderboard, arXiv:2412.13788.

Downloads last month
47
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using oddadmix/Nawah-Parakeet-60M 1

Paper for oddadmix/Nawah-Parakeet-60M