Parakeet Ultra, ONNX for the browser

Want faster improvements? This project is entirely self-funded on my minimum-wage salary, and every training run competes for a single consumer GPU. If you or your organisation can donate an RTX 5090, or the money to buy one second-hand, it would directly speed up the next versions of the dataset and the fine-tuned models. Reach out via olicorne.org or open a discussion on this page.

ONNX export of moondream/parakeet-ultra, moondream's full-precision post-trained parakeet-tdt-0.6b-v3, same architecture and tokenizer, better than the original on every benchmark moondream reports.

Built for Parakeet Web (transcription entirely client-side) and onnx-asr, with the same layout and pipeline as Olicorne/parakeet-tdt-0.6b-v3-optimized-onnx. Built with Claude Code.

Files

folder contents size use
fp32/ encoder (2 shards) + decoder 2.5 GB reference; WASM (opt-in) and WebGPU
fp16/ encoder 1.2 GB WebGPU default (needs shader-f16)
int8/ encoder + decoder 650 MB + 18 MB CPU/WASM default (MatMulNBits 8-bit, block 64)
w4a8/ encoder 383 MB smallest download (MatMulNBits 4-bit, block 32)

vocab.txt, config.json and the nemo128.onnx preprocessor sit at the root, next to parakeet-tdt-0.6b-v3-ultra.nemo, the NeMo checkpoint the ONNX was exported from. *.zst files are precompressed copies for web servers. Only the int8 decoder ships besides fp32: on this model it is as accurate as fp32.

The int8 and w4a8 encoders are rounded the same way as the base v3 repo's, so they carry the same kind of quantization cost.

Batching. The encoder masks padded frames with a key-only attention bias, so onnx-asr's model.recognize([a, b, c]) works on clips of different lengths with no manual padding, and batch 1 is bit-identical to the previous unmasked export (the published numbers stay valid). On 1,000 UltiMed dictionary clips (UltiMed w4a8 encoder, fp32 decoder, CUDA, onnx-asr with a lockstep batched greedy TDT decode as in the ultra repo's scripts/drug-rules/transcribe_onnx.py), batches of 8, 16 and 32 took 0.06, 0.05 and 0.04 s per clip against 0.16 s one clip at a time, changed 3 to 4 transcripts out of 1,000, and kept WER at 3.26 % (3.26 % unbatched). Use the fp32 decoder when batching: the int8 decoder quantizes activations with one scale per batch, so a clip's transcript can depend on its batch neighbours. Stock onnx-asr decodes each clip on its own, so it gets the encoder speedup but not the decoder one.

How it was built

  1. scripts/hf-to-nemo.py (parakeet_web) maps the HF transformers checkpoint back onto the base v3 .nemo: a key rename only, every learned weight comes from moondream's safetensors, the VAD head is dropped.
  2. NeMo export_onnx.py --web-optimized, then scripts/build-from-nemo-export.sh (in this repository's scripts/): graph fold (bit-exact gate), 12-bit mantissa rounding, fp32 shards, decoder int8 + LSE/top-K surgery, fp16, pointwise Conv -> MatMul, MatMulNBits.

Checked: the ONNX transcribes the JFK clip word for word like moondream's PyTorch model run in NeMo.

scripts/bench/: benchmark tooling

The benchmark queue and publisher, written with Claude Code. This is the canonical copy: it moved here from the sunset optimized-onnx repository, which keeps an old copy only until its last runs finish. It fills the benchmark tables of this README and of Olicorne/parakeet-tdt-0.6b-v3-UltiMed-onnx.

  • wer-fleurs-validation.sh: per-language WER (and CER at beam > 1) of one or more models over FLEURS-layout sets, driving wer-quants.py and grid_search_benchmark.mjs from parakeet_web.
  • bench-overnight.sh: unattended queue of (model, precision, dataset, decoding) entries, one timestamped dir per launch under local/run_logs/; --queue FILE runs a one-off list.
  • bench-watchdog.sh: restarts a dead queue or publisher and kills a stalled entry.
  • collect-bench.py: aggregates the queue's results.jsonl files into tables.
  • bench-sanity.py: the gate that blocks a suspicious result from being published.
  • bench-publish.py, publish-voxpopuli.py: rewrite the <!-- BENCH-START: name --> blocks of the READMEs. bench-autopublish.sh loops publish, gate, commit and push. PUBLISH_TARGETS picks the repos: ultimed by default, ultra (this README) only when asked.
  • regex-rescore.mjs: re-scores a phrase-boost run with the drug-name fix rules applied.
  • bench_common.py: helpers and constants shared by the python scripts above (the EU11 language list, the README block splice); not run directly.

Machine paths (FLEURS, medical fixtures, parakeet_web) go in local/.env, gitignored: copy scripts/bench/env.example there. Run dirs and queue files are passed by path, so absolute paths to results kept elsewhere work.

scripts/drug-rules/: drug-name fix rules

Builds and checks the fix rules (drug_fix_rules.jsonl for drug names, term_fix_rules.jsonl for medical terms) shipped in the regex-fixes/ folder of Olicorne/parakeet-tdt-0.6b-v3-UltiMed-onnx, written with Claude Code. The rule learning itself lives in the UltiMed-ASR-FR-v1-scripts repository (08_drug_asr_rules/, expected next to this one); these scripts drive it over several models and measure the result.

  • transcribe.sh <model repo dir> <quant> <decoder quant> <out.jsonl> <manifest> ...: runs transcribe_onnx.py on CUDA with the pip CUDA and cuDNN wheels it needs, for any ONNX repo and any NeMo manifests. transcribe-drugs.sh is a wrapper around it.
  • shuffle-manifests.py <out.jsonl> <manifest> ... [-- <manifest> ...]: concatenates manifests into one shuffled manifest (seed 0, absolute audio paths), each -- group shuffled on its own and kept in order (e.g. val and test before train). Since transcribe_onnx.py keeps manifest order, repeated LIMIT=n runs then transcribe consecutive random chunks of n clips.
  • term-pilot-round.py <pilot dir> <k> [--chunks dictionary=13000,PARHAF=8000,PARROT=1549]: one round of the medical-term rule pilot. It builds rules (01 and 02 with --lexicon, the dictionary by default) from chunks 1 to k of every model and subset, applies them to the unseen chunk k+1 and appends to <pilot dir>/rounds.jsonl the coverage (share of chunk k+1's term errors fixed), the rule count, the singleton share of the errors seen so far and the correct labels of all splits the rules change. When one more chunk barely raises the coverage, more transcription stops paying. A dry run on 5,000 drug clips per model (UltiMed int8, base w4a8) gave 32 % and 37 % coverage and 0 labels changed, in about 4 min of CPU.
  • transcribe-drugs.sh <quant> <decoder quant> <out dir>: transcribes the UltiMed drugs subset (all splits) with one ONNX variant of this repository on CUDA, through transcribe_onnx.py (resumable, reuses parakeet_web's wer-quants.py loader). Env for transcribe_onnx.py: BATCH (clips per call, default 1; mixed lengths are fine since the encoder masks padding, and batches are decoded in lockstep), BATCH_SEC (cap on a batch's padded audio, clips x longest clip, default 400 s: 16 clips of ~80 s ran a 24 GB GPU out of memory), LOAD_WORKERS (audio loading threads, default 4: resampling the 24 kHz clips costs about 0.14 s each), LIMIT, USE_CUDA, DEC_QUANT (default fp32, which gives the same transcript whatever the batch).
  • build-rules.sh <out dir> NAME=<dir> ...: mines each model's drug errors, builds one rule set over all of them, then writes report.txt: test drug WER per model, overcorrections on every correct label of all subsets and splits, and the boost grids incl. FLEURS fr/en, each against the committed rules and no rules. SPLITS picks the clips and labels learned from: all splits by default (the released rules), "train val" to keep test held out. TEXTS adds outside correct text to the overcorrection check (one sentence per line, e.g. the VoxPopuli fr/en validation references). GUARDS (space-separated .txt files of correct French text, e.g. the Leipzig fra_wikipedia_2021_1M sentences) blocks any rule that would change them, like the labels do. term-pilot-round.py takes the same kind of files as --guard (blocking) and --check (counted only). NAMES (or --names for term-pilot-round.py) takes the name list insee-names.py <dir> writes (every INSEE surname and first name, about 248k): a rule whose variant is also a name is kept but does not fire right after a title, so le docteur Mirat keeps the surname while a bare mirat still becomes Humira.
  • bench/real-voice-bench.sh <audio dir> <out dir> then bench/real-voice-report.mjs: the real-voice stress test of the UltiMed README. The first transcribes a folder of recordings as is (references in its transcripts.md) with ultra fp32 and UltiMed fp32 / int8 / w4a8, beam 1 and 5, with and without the French medical phrase boost; the second applies the drug then medical-term rules, scores raw and normalised WER / CER and writes the page (bench/real-voice-intro.md is its introduction).
  • evalrules.py, compare-rulesets.sh: the two checks, also usable alone on any rule files. evalrules.py --refs <manifest> scores the hyps against any NeMo manifest instead of the test drug clips (e.g. FLEURS fr or Common Voice fr), a clip-level better/worse check on real speech that prints the clips a rule set makes worse. compare-rulesets.sh rescores finished boost cells with scripts/bench/regex-rescore.mjs; BENCH_LOGS / BENCH_GLOB point it at the bench runs.
  • rule-order-study.py <rules.jsonl> <hyps.jsonl ...> [--perms 5] [--out examples.jsonl]: checks a rule file against its own hyps. It lists the dead rules (never fire in file order: shadowed by an earlier rule, or matching no hyp), re-applies the rules in random orders to count the texts whose output depends on the order, and counts the replacements several variants write. On the drug rules (2026-10-04, 70,398 hyps): 51 dead rules (25 shadowed), 102 of 40,116 touched texts depend on the order, and the file order (longest variant first) gives the better output in the cases read, e.g. nicopassement de fraîcheur gives nicopass menthe fraîcheur in file order but nicopass menthe de fraîcheur when the shorter nicopassement fires first. On the term rules (78,850 rules, 1,126,610 hyps): 503 dead rules (281 shadowed), 1,288 of 386,035 touched texts depend on the order, again with the longer variant first giving the fuller fix (sarco myistiocytaire before myistiocytaire). Dead rules are only dead on these hyps: a shadowed nicopassement still fixes a text without the de of the longer variant, so they are kept.

Benchmarks

All numbers measured on CPU (onnxruntime, wer-quants.py from parakeet_web, with the int8 decoder for every encoder). Generated with Claude Code from the benchmark queue's output.

FLEURS, greedy, full validation split, 25 languages:

encoder macro WER macro CER micro WER micro CER sets scored
fp32 (reference) 11.55 % 3.45 % 11.16 % 3.38 % 25/25
int8 (WASM default) 11.57 % 3.45 % 11.18 % 3.38 % 25/25
w4a8 12.19 % 3.63 % 11.80 % 3.56 % 25/25

FLEURS, beam 5 (MAES), 60 clips per language:

encoder macro WER macro CER micro WER micro CER sets scored
fp32 (reference) 10.66 % 2.97 % 10.38 % 2.93 % 25/25
int8 (WASM default) 10.65 % 2.98 % 10.37 % 2.94 % 25/25
w4a8 11.41 % 3.17 % 11.11 % 3.13 % 25/25

Long audio: eight speeches (six of 390 s, two of 60 s), each transcribed in ONE pass. The reference is this repo's fp32 decoding the same audio in independent 60 s sections, so these figures measure long-pass drift, not absolute accuracy. Means across the eight files:

encoder overall WER overall CER last chunk WER worst chunk WER
fp32 4.7 % 1.2 % 7.2 % 10.3 %
int8 4.9 % 1.2 % 7.4 % 10.6 %
w4a8 7.3 % 2.7 % 10.2 % 12.9 %

Past 400 s: one 607 s French VoxPopuli clip in a single pass, same fp32-sectioned reference; it checks that a pass well past 400 s completes and stays close to sectioned decoding:

encoder WER CER proc/dur
int8 5.80 % 2.16 % 0.234
w4a8 6.85 % 2.50 % 0.215
fp32 5.72 % 2.14 % 0.234

The 607 s figures are drift against this backend's own fp32 reference (this run used the CPU backend, before the 2026-10-02 re-export that added the padding mask), so treat them as a stability check, not as a WER: reruns on CPU and CUDA on other dates put int8 between 5.8 % and 6.3 % (2026-10-04 recheck: CPU 6.02 %, CUDA 5.89 %).

Licence and attribution

CC-BY-4.0, like the original. Model by moondream, derived from NVIDIA's parakeet-tdt-0.6b-v3. This repo only converts and quantizes it.

Downloads last month
68
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Olicorne/parakeet-tdt-0.6b-v3-ultra-onnx

Quantized
(15)
this model