Nawah-Parakeet-60M
A 62.7M-parameter Parakeet-style Token-and-Duration Transducer for Arabic speech recognition β FastConformer encoder, TDT objective β trained from scratch with no pretrained encoder and no external language model.
On the six test sets of the Open Universal Arabic ASR Leaderboard it reaches 31.34 average WER / 13.62 CER, which places it 8th of 38 evaluated systems β ahead of Voxtral-Small-24B, gemma-4-E4B, Qwen3-ASR-1.7B, nvidia-conformer-ctc-large-arabic and every Whisper model, at a fraction of their size. Every model above it is between 1B and 30B parameters, or a commercial API.
Leaderboard
Averages are recomputed over these six sets for every model so the ordering is consistent. Baseline figures are the leaderboard's published numbers.
| # | Model | Avg WER | Avg CER | Params |
|---|---|---|---|---|
| 1 | audarai/Audar-ASR-V1-Turbo | 23.17 | 9.22 | 2.35B |
| 2 | CohereLabs/cohere-transcribe-arabic-07-2026 | 25.87 | 11.80 | β |
| 3 | omnilingual-asr/omniASR_LLM_7B | 28.32 | 12.52 | 7B |
| 4 | omnilingual-asr/omniASR_LLM_1B | 29.95 | 13.40 | 1B |
| 5 | omnilingual-asr/omniASR_LLM_3B | 29.96 | 13.77 | 3B |
| 6 | CohereLabs/cohere-transcribe-03-2026 | 30.67 | 16.37 | β |
| 7 | Qwen/Qwen3-Omni-30B-A3B-Instruct | 30.71 | 13.66 | 30B |
| 8 | Nawah-Parakeet-60M (this model) | 31.34 | 13.62 | 62.7M |
| 9 | audarai/Audar-ASR-V1-Flash | 32.04 | 13.65 | β |
| 10 | nvidia-conformer-ctc-large-arabic (lm) | 32.91 | 13.85 | 0.12B |
| 11 | omnilingual-asr/omniASR_LLM_300M | 32.96 | 14.84 | 300M |
| 12 | google/gemma-4-E4B-it | 32.98 | 13.71 | 4B |
| 13 | Qwen/Qwen3-ASR-1.7B | 33.69 | 12.33 | 1.7B |
| 14 | mistralai/Voxtral-Small-24B-2507 | 34.46 | 15.29 | 24B |
| 16 | openai/whisper-large-v3 | 36.86 | 17.21 | 1.55B |
| 21 | openai/whisper-large-v3-turbo | 40.05 | 18.87 | 0.81B |
| 33 | openai/whisper-small | 55.13 | 21.68 | 244M |
| 38 | speechbrain/asr-wav2vec2-commonvoice-14-ar | 65.74 | 30.93 | β |
Full table in eval/ranking.json.
Per-set results
Scored with the leaderboard's own eval.py β its normalize_arabic_text and NeMo's
word_error_rate β unmodified.
| Test set | Clips | Hours | WER | CER | board best | board median |
|---|---|---|---|---|---|---|
| MASC clean | 8,612 | 10.49 | 11.03 | 3.60 | 8.66 | 24.86 |
| MGB-2 | 5,365 | 9.58 | 18.47 | 8.59 | 11.08 | 20.23 |
| Common Voice 18 | 10,471 | 12.66 | 22.39 | 7.83 | 5.82 | 17.83 |
| MASC noisy | 9,394 | 8.92 | 30.45 | 13.27 | 19.01 | 35.64 |
| SADA | 6,186 | 10.75 | 43.05 | 20.38 | 28.92 | 60.11 |
| Casablanca | 6,818 | 7.85 | 62.67 | 28.03 | 47.02 | 69.37 |
| Average | 31.34 | 13.62 | 23.17 | 38.12 |
Test manifests were verified against the leaderboard's own datasets/*.json: reference text matches
on 100% of Common Voice, MASC clean, MASC noisy and Casablanca clips, 99.2% of MGB-2 and 98.3% of
SADA.
Scoring reproducibility
Two leaderboard models were re-run locally through this harness to confirm the scoring is faithful.
Audar-ASR-V1-Turbo (rank 1, 2.35B), using Audar's own reference inference code on their pinned
transformers==4.57.6, seeded 600-clip subsets:
| Avg | SADA | CV18 | MASC clean | MASC noisy | MGB-2 | Casablanca | |
|---|---|---|---|---|---|---|---|
| leaderboard Space | 23.17 | 28.92 | 8.09 | 16.73 | 27.19 | 11.08 | 47.02 |
Audar's own leaderboard.csv |
24.78 | 29.41 | 8.60 | 19.60 | 28.35 | 11.13 | 51.58 |
| measured here | 24.93 | 32.26 | 8.69 | 18.00 | 28.57 | 12.21 | 49.86 |
The average reproduces Audar's own published figure to 0.15 WER, which confirms the audio preparation and the scorer. Note that Audar's two published sources differ by 1.6 average WER.
whisper-small (rank 33) does not reproduce per set: β15.25 on MGB-2, +14.76 on Common Voice,
+12.51 on SADA, though its average lands within 0.61. The cause is measurable β under current
transformers, whisper-small enters repetition loops on 6.5% of SADA clips and 3.5% of Casablanca
clips, and excluding only those recovers 19.80 and 13.15 WER respectively. Published figures predate
that version.
This model's TDT decoder caps symbols per encoder frame and shows 0 runaway hypotheses on Common Voice, MGB-2 and both MASC sets (0.02% SADA, 0.13% Casablanca), so it is not exposed to that failure.
Takeaway: absolute positions on the published table carry roughly 1β2 WER of uncertainty, more for models with unstable decoders.
Usage
pip install torch torchaudio tokenizers soundfile
from huggingface_hub import snapshot_download
import sys, soundfile as sf
from tokenizers import Tokenizer
d = snapshot_download("oddadmix/Nawah-Parakeet-60M")
sys.path.insert(0, d)
import parakeet_model as pm
tok = Tokenizer.from_file(f"{d}/tokenizer_bpe1024.json")
model = pm.load(f"{d}/nawah_parakeet_60m.pt", device="cuda")
wav, sr = sf.read("clip.wav", dtype="float32") # 16 kHz mono
print(pm.transcribe(model, wav, tok, device="cuda"))
config.json describes the architecture, feature front end and tokenizer in machine-readable form.
It is not a transformers config: AutoModel.from_pretrained will not load this model β use
parakeet_model.load() as above. Every value in it was read off the checkpoint rather than
transcribed by hand.
The bundled tokenizer is the only one that will work. Any other 1024-unit BPE produces valid ids that mean different things, so a mismatch yields fluent-looking nonsense rather than an error.
Architecture
| encoder | FastConformer, 17 layers, d=384, 6 heads, FFN 1536, Γ8 subsampling β 12.5 Hz |
| decoder | 2-layer LSTM predictor, 320-dim (2.09M params) |
| joiner | additive, 384-dim β 1024 vocab + 5 duration bins (0.40M) |
| total | 62.71M |
| objective | Token-and-Duration Transducer loss + auxiliary CTC (weight 0.2) |
| features | 128-mel kaldi fbank, povey window, 25 ms / 10 ms, 16 kHz |
| tokenizer | 1024-unit BPE, Metaspace pre-tokenizer and decoder, blank = id 0 |
| training | 2,404 h of Arabic speech, 2 epochs, ~32 h on one RTX 5090 |
No 30-second limit. The encoder uses clamped relative position bias rather than a fixed positional table, so audio of any length is transcribed in a single pass β unlike Whisper-family models, which need 30 s chunking. Measured on 2 CPU threads: 30 s in 0.3 s, 120 s in 2.1 s, 300 s in 10.8 s. Attention is O(TΒ²), so memory is the practical ceiling: ~3.2 GB at 5 minutes and ~11 GB at 15 minutes.
Limitations
- Undiacritized output only. No tashkeel, no punctuation, no casing for Latin tokens.
- Maghrebi remains the weakest region in absolute terms (Casablanca Morocco 69.15 WER, Mauritania 83.86, Algeria 77.04), even though it is where the model improved most.
- The model does not emit Ψ©. Its training targets used Ω throughout, so words that a reference writes with Ψ© are scored as errors on sets that use it. The leaderboard's normalizer folds the alef family but not Ψ©. Folding Ψ© on both sides changes MGB-2 by only 0.06 WER, so the residual effect is small.
- Short clips can gain an invented trailing word β a property of the transducer's stop condition.
- Single-pass long audio is out of distribution. Training clips were capped at 20 s; accuracy on multi-minute recordings has not been measured.
- Audio must be 16 kHz mono; the kaldi fbank front end assumes it, and another rate misplaces every mel bin.
Files
nawah_parakeet_60m.pt checkpoint (241 MB)
config.json architecture, features and tokenizer, machine-readable
tokenizer_bpe1024.json the tokenizer this model requires
parakeet_model.py standalone loader + TDT greedy decode, torch only
fastconformer.py encoder definition
parakeet_data.py fbank front end and encoder-length helper
eval/leaderboard_results.json per-set scores for this model
eval/ranking.json all evaluated models, per set
eval/lb_norm.py the leaderboard's scorer, ported
eval/run_bench.py batch decoding to the leaderboard manifest format
Citation
Leaderboard and baseline figures: Wang, Alhmoud & Alqurishi, Open Universal Arabic ASR Leaderboard, arXiv:2412.13788.
- Downloads last month
- 47