
YuE2 vs Suno v6: What the Benchmark Actually Says
Quick answer: YuE2-3B scores 6.9632 on WildSongBench, the highest mean observed — but that is a best-of-8 figure. In the standard two-candidate setting it scores 6.7316, fourth, behind Mureka 9 (6.9377) and Suno v5 (6.8721). It leads Suno v6 (6.5562) on the aggregate either way, and it is the strongest open-weights music model by a wide margin.
M·A·P published WildSongBench results for YuE2 on 12 September 2026: 192 prompts, 17 settings, automatic evaluation. The headline number has been widely quoted. Most of the detail underneath it has not, and the detail is where the useful conclusions are.
The full WildSongBench table
192 prompts, automatic evaluation, 12 September 2026. Open-weights status follows the dagger marking in M·A·P's own table. PER is phoneme error rate, where lower is better.
How to read the four metrics
The benchmark reports four numbers measuring genuinely different things. Reading only the first gives a misleading picture of the result.
- SongBench Avg — the headline aggregate judgement of song quality. Higher is better. This is the only column most articles quote.
- AudioBox PQ — production quality: how cleanly produced the audio is, independent of whether it matches the brief. Higher is better.
- MuLan — text-music alignment: how closely the output matches the prompt. Higher is better. For commercial work this often decides usability, because a beautiful track that ignores the brief is worthless.
- PER — phoneme error rate: how often the vocal mispronounces words. Lower is better, and it is the only column where the direction flips.
A model can lead one column and trail badly in another, which is exactly what happens here.
Neither YuE2 number is a single generation
Best-of-8 is openly labelled. The 6.7316 figure is often described as “single pass”, and it is not. M·A·P's benchmark documentation states that standard YuE2 selects the lower-PER candidate from two generations, and adds that this “is not equivalent to one unselected pipeline call”.
SettingGenerationsSelected bySongBench AvgYuE2 standard2Lower PER6.7316YuE2 best-of-88Musicality, then prompt control, then PER6.9632
Each candidate's PER is itself the lowest of four ASR passes. If you run the model once and keep whatever comes out, you should expect results below both published figures.
Suno v6 is the weakest Suno version tested
Read down the Suno rows and the version numbers stop behaving as you would expect:
Suno versionSongBench AvgSuno v56.8721Suno v5.56.7150Suno v4.56.6995Suno v66.5562Suno v6 Wild6.4195
Three older Suno releases outscore v6 on this aggregate. Anyone announcing that YuE2 has “beaten Suno” by pointing at the v6 row is leaning on Suno's lowest-scoring release.
The honest counterweight, which M·A·P supply themselves: Suno v6 scores higher than both YuE2 settings on SongEval, and both v6 variants record lower phoneme error rates and higher AllMusicCaps scores. The aggregate favours YuE2. Several components do not.
Where YuE2 actually wins outright
The open-weights comparison is the one that survives every caveat above.
A 0.41 gap to the next open model is far larger than anything separating YuE2 from the proprietary leaders, and unlike the Suno comparison it holds in the standard setting rather than needing best-of-8. If you need weights you can download, this is not a close contest.
MiniMax Music 3 belongs in this table. M·A·P classify it as public because its weights are released and the evaluation ran them locally.
LeVo 2 and why single-number rankings mislead
LeVo 2 records the highest AudioBox PQ in the entire table at 8.3966 — better production quality than YuE2, Suno or Mureka. It still lands ninth overall, because its MuLan score is 0.3542 and its phoneme error rate is 26.12%. It makes excellent-sounding audio that follows the prompt poorly and sings words indistinctly.
If you want an instrumental bed and do not care about lyrics or tight prompt adherence, LeVo 2 may serve you better than its ranking suggests. That is not a conclusion the aggregate column can express.
How far YuE2 came from YuE 1
A phoneme error rate falling from 36.38% to 8.44% is the difference between a vocal you have to excuse and one you can actually use. It is the largest single-generation improvement the table records.
The result that is missing from the main table
YuE2's zero-shot cover evaluation is reported separately and is arguably its most striking number. Across 948 works, full-score YuE2 reaches 0.647 CLEWS mAP against 0.006 without a score — a hundredfold difference — using the general song-generation checkpoint with no cover-specific fine-tuning.
That gap is the clearest evidence for the symbolic-planning approach in the whole release, and it appears nowhere in the rankings everyone quoted.
What the benchmark cannot tell you
M·A·P are unusually direct about the limits of their own results, and it is worth repeating rather than burying:
- The small gaps between the highest means do not establish statistical significance.
- The comparison is not matched-compute. Different systems ran under different candidate protocols; these are documented system comparisons, not controlled experiments.
- Scoring is automatic, not human preference. Nothing here measures whether listeners prefer the output.
- Both YuE2 settings used the YuE2-Vae-legacy benchmark decoder, not necessarily what you get by default today.
A 0.2 difference in an automatic aggregate, produced under unmatched protocols and disclaimed by its own authors, is a weaker signal than most of the coverage implied.
Should you switch?
Three situations, three different answers.
- You need commercial rights. YuE2 is not an option. The weights are CC BY-NC 4.0 and M·A·P intend the restriction to cover generated output, not just the model files.
- You need local, offline or unlimited generation. YuE2 is the strongest open option by a clear margin, and best-of-8 sampling costs you time rather than credits.
- You are choosing on quality alone and can use a cloud service. The top five settings sit within 0.25 of each other on an automatic metric its authors decline to call significant. Test on your own material instead.
Running YuE2 on Promptus
YuE2 is available in Promptus as a CosyFlow — a packaged ComfyUI workflow you open in the ComfyUI Canvas without installing anything, wiring nodes or hunting for model weights.
You can run it locally on your own GPU with no generation fees, or use Promptus cloud compute if you do not have a 24 GB card. That choice matters more for this model than for most, because the published scores depend on candidate selection: the standard figure picks the better of two generations and best-of-8 picks from eight. Sampling that way is cheap on hardware you own and expensive on metered credits.
One thing Promptus does not change is the licence. YuE2's weights are CC BY-NC 4.0 and M·A·P intend that restriction to cover generated output, so music you make with it is non-commercial wherever you generated it. If you need commercial rights, the licence article sets out what is and is not permitted, and ACE-Step — also available in Promptus, and in the table above — is Apache-2.0 and does allow commercial use.
Sources: M·A·P YuE repository, its docs/benchmarks.md and THIRD_PARTY_NOTICES.md, and the YuE2-3B model card. Figures checked 15 September 2026 against the WildSongBench run published 12 September 2026.
%20(2).avif)
%20transparent.avif)

