Instructions to use Prompt48/Qwen3-8B-MiMo-Music-GRPO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Prompt48/Qwen3-8B-MiMo-Music-GRPO with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Prompt48/Qwen3-8B-MiMo-Music-GRPO") model = AutoModelForCausalLM.from_pretrained("Prompt48/Qwen3-8B-MiMo-Music-GRPO", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qwen3-8B · MiMo Music GRPO
Code, logs and all results: https://github.com/PromptEngineer48/mimo-music-grpo
Qwen3-8B trained with GRPO (reinforcement learning) to write music in ABC notation, using the
music environment Xiaomi open-sourced with MiMo-V2.6: their prompts
(XiaomiMiMo/MiMo-V2.6-RL-oss, music subset)
and their rule-based grader
(XiaomiMiMo/verl · recipes/design/music/scorer).
A small, single-GPU replica of the method, made for a Prompt Engineer 48 YouTube video. Not a Xiaomi model.
Results (Xiaomi's strict grader, 40 held-out prompts × 4 samples, same seed)
| Qwen3-8B (base) | This model | |
|---|---|---|
| Playable, valid music | 1/160 (0.6%) | 19/160 (11.9%) |
| Prompts with ≥1 valid piece | 1/40 | 13/40 |
| Mean grader score | 0.002 | 0.041 |
| Avg answer length | 2171 tokens | 1256 tokens |
20× more valid music after 30 GRPO steps (1.7 h on 1× H100). Still only ~12% valid — a starting point, not a finished composer.
Training
- Base: Qwen3-8B, thinking off. LoRA r=64 (all linear), merged into this checkpoint. Adapter alone:
lora_adapter/. - GRPO via TRL 1.14 + vLLM: 8 prompts × 8 samples per step, 30 steps, lr 2e-5, no KL, clip 0.2, max 4096 tokens.
- Reward: Xiaomi's music score (0–1, "human-likeness" over 18 features) + Xiaomi's group-relative length penalty (MiMo-V2.6 report, Eq. 4).
- One shaping step: base models almost always put blank lines between sections, and in ABC a blank line ends the tune, so the grader rejects it. During training only, a piece that fails only because of blank lines gets 30% credit. Evaluation above uses the strict grader, unmodified.
- 456 English training prompts; 40 held out.
Use
Use long, detailed composition briefs like the training data (short prompts give short, repetitive output). Prompt for a piece and ask for ABC in a ```abc block. Disable thinking:
from transformers import AutoModelForCausalLM, AutoTokenizer
m = "Prompt48/Qwen3-8B-MiMo-Music-GRPO"
tok = AutoTokenizer.from_pretrained(m); model = AutoModelForCausalLM.from_pretrained(m, torch_dtype="auto", device_map="auto")
msgs = [{"role": "system", "content": "You are a composer. Reply with one complete piece in ABC notation inside a ```abc code block. Start with X:1 and include T:, M:, L:, Q:, K: headers."},
{"role": "user", "content": "Compose a Viennese waltz in Eb major, 181 BPM allegro, 3/4, about 96 bars and about three and a half minutes, scored for strings and harp in 4 voices. Texture: a slightly anticipated waltz bass under an elegant sweeping tune. Open with 8 bars of strings alone, then layer the other parts in one at a time. The main section runs on a 4-bar diatonic loop of I–IV–V–I, unchanged throughout. Thin the middle section to two voices for two descending phrases, then return with everything for the fullest passage. Dynamics start at p, push to mf in the middle and settle back to p for a sustained tonic close. Aim for the feeling of a rooftop above neon. Use no ornament symbols; write every note out. Output the complete piece as ABC notation."}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=False, return_tensors="pt").to(model.device)
print(tok.decode(model.generate(ids, max_new_tokens=2048, do_sample=True, temperature=1.0)[0][ids.shape[1]:]))
Render: abc2midi piece.abc -o piece.mid, then any MIDI player / fluidsynth.
GGUF (Ollama / LM Studio) — re-checked on the same test set: merged bf16 14.4% valid, Q8_0 18.8%, Q4_K_M 5.0% (use Q8_0): Prompt48/Qwen3-8B-MiMo-Music-GRPO-GGUF
Limitations
- ~88% of outputs still fail the strict grader (mostly blank lines / bar-length errors).
- The grader measures statistical human-likeness, not beauty. Listen before trusting a number.
Credits
Xiaomi MiMo team (environment, grader, method), Qwen team (base model), TRL / vLLM.
- Downloads last month
- 47