Qwen3-8B · MiMo Music GRPO

Code, logs and all results: https://github.com/PromptEngineer48/mimo-music-grpo

Qwen3-8B trained with GRPO (reinforcement learning) to write music in ABC notation, using the music environment Xiaomi open-sourced with MiMo-V2.6: their prompts (XiaomiMiMo/MiMo-V2.6-RL-oss, music subset) and their rule-based grader (XiaomiMiMo/verl · recipes/design/music/scorer).

A small, single-GPU replica of the method, made for a Prompt Engineer 48 YouTube video. Not a Xiaomi model.

Results (Xiaomi's strict grader, 40 held-out prompts × 4 samples, same seed)

Qwen3-8B (base) This model
Playable, valid music 1/160 (0.6%) 19/160 (11.9%)
Prompts with ≥1 valid piece 1/40 13/40
Mean grader score 0.002 0.041
Avg answer length 2171 tokens 1256 tokens

20× more valid music after 30 GRPO steps (1.7 h on 1× H100). Still only ~12% valid — a starting point, not a finished composer.

Training

  • Base: Qwen3-8B, thinking off. LoRA r=64 (all linear), merged into this checkpoint. Adapter alone: lora_adapter/.
  • GRPO via TRL 1.14 + vLLM: 8 prompts × 8 samples per step, 30 steps, lr 2e-5, no KL, clip 0.2, max 4096 tokens.
  • Reward: Xiaomi's music score (0–1, "human-likeness" over 18 features) + Xiaomi's group-relative length penalty (MiMo-V2.6 report, Eq. 4).
  • One shaping step: base models almost always put blank lines between sections, and in ABC a blank line ends the tune, so the grader rejects it. During training only, a piece that fails only because of blank lines gets 30% credit. Evaluation above uses the strict grader, unmodified.
  • 456 English training prompts; 40 held out.

Use

Use long, detailed composition briefs like the training data (short prompts give short, repetitive output). Prompt for a piece and ask for ABC in a ```abc block. Disable thinking:

from transformers import AutoModelForCausalLM, AutoTokenizer
m = "Prompt48/Qwen3-8B-MiMo-Music-GRPO"
tok = AutoTokenizer.from_pretrained(m); model = AutoModelForCausalLM.from_pretrained(m, torch_dtype="auto", device_map="auto")
msgs = [{"role": "system", "content": "You are a composer. Reply with one complete piece in ABC notation inside a ```abc code block. Start with X:1 and include T:, M:, L:, Q:, K: headers."},
        {"role": "user", "content": "Compose a Viennese waltz in Eb major, 181 BPM allegro, 3/4, about 96 bars and about three and a half minutes, scored for strings and harp in 4 voices. Texture: a slightly anticipated waltz bass under an elegant sweeping tune. Open with 8 bars of strings alone, then layer the other parts in one at a time. The main section runs on a 4-bar diatonic loop of I–IV–V–I, unchanged throughout. Thin the middle section to two voices for two descending phrases, then return with everything for the fullest passage. Dynamics start at p, push to mf in the middle and settle back to p for a sustained tonic close. Aim for the feeling of a rooftop above neon. Use no ornament symbols; write every note out. Output the complete piece as ABC notation."}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=False, return_tensors="pt").to(model.device)
print(tok.decode(model.generate(ids, max_new_tokens=2048, do_sample=True, temperature=1.0)[0][ids.shape[1]:]))

Render: abc2midi piece.abc -o piece.mid, then any MIDI player / fluidsynth.

GGUF (Ollama / LM Studio) — re-checked on the same test set: merged bf16 14.4% valid, Q8_0 18.8%, Q4_K_M 5.0% (use Q8_0): Prompt48/Qwen3-8B-MiMo-Music-GRPO-GGUF

Limitations

  • ~88% of outputs still fail the strict grader (mostly blank lines / bar-length errors).
  • The grader measures statistical human-likeness, not beauty. Listen before trusting a number.

Credits

Xiaomi MiMo team (environment, grader, method), Qwen team (base model), TRL / vLLM.

Downloads last month
47
Safetensors
Model size
8B params
Tensor type
BF16
·
Video Preview
loading

Model tree for Prompt48/Qwen3-8B-MiMo-Music-GRPO

Finetuned
Qwen/Qwen3-8B
Finetuned
(2160)
this model
Quantizations
2 models

Dataset used to train Prompt48/Qwen3-8B-MiMo-Music-GRPO