suno v6 vs yuE2
Jack
News

YuE2 vs Suno v6: What the Benchmark Actually Says

Promptus
September 16, 2026
Wiki 321
promptus ai video generator

YuE2 vs Suno v6: What the Benchmark Actually Says

Quick answer: YuE2-3B scores 6.9632 on WildSongBench, the highest mean observed — but that is a best-of-8 figure. In the standard two-candidate setting it scores 6.7316, fourth, behind Mureka 9 (6.9377) and Suno v5 (6.8721). It leads Suno v6 (6.5562) on the aggregate either way, and it is the strongest open-weights music model by a wide margin.

M·A·P published WildSongBench results for YuE2 on 12 September 2026: 192 prompts, 17 settings, automatic evaluation. The headline number has been widely quoted. Most of the detail underneath it has not, and the detail is where the useful conclusions are.

The full WildSongBench table

System / setting SongBench Avg ↑ AudioBox PQ ↑ MuLan ↑ PER ↓ Open weights
YuE2 (best-of-8) 6.9632 8.2714 0.5051 19.79% Yes
Mureka 6.9377 8.0226 0.4394 11.69% No
Suno v5 6.8721 8.1698 0.5428 8.10% No
YuE2 (standard, 2 candidates) 6.7316 8.2598 0.5068 8.44% Yes
Suno v5.5 6.7150 8.1955 0.5089 5.96% No
Suno v4.5 6.6995 8.2541 0.5022 5.80% No
Suno v6 6.5562 8.1296 0.4916 7.58% No
Suno v6 Wild 6.4195 8.1785 0.4999 7.45% No
LeVo 2 6.3247 8.3966 0.3542 26.12% Yes
MiniMax Music 2.6 6.3222 8.1711 0.4251 24.55% No
MiniMax Music 3 6.2830 8.2825 0.3928 6.27% Yes
HeartMuLa 6.2483 8.2933 0.3823 10.71% Yes
Muse 6.0349 8.0517 0.3937 33.42% Yes
ACE-Step 1.5 6.0118 8.0518 0.4372 7.46% Yes
DiffRhythm 2 5.2428 7.9782 0.3782 18.41% Yes
YuE 1 4.9165 7.8683 0.2623 36.38% Yes
SongBloom 4.2350 8.1539 0.2697 19.19% Yes

192 prompts, automatic evaluation, 12 September 2026. Open-weights status follows the dagger marking in M·A·P's own table. PER is phoneme error rate, where lower is better.

How to read the four metrics

The benchmark reports four numbers measuring genuinely different things. Reading only the first gives a misleading picture of the result.

  • SongBench Avg — the headline aggregate judgement of song quality. Higher is better. This is the only column most articles quote.
  • AudioBox PQ — production quality: how cleanly produced the audio is, independent of whether it matches the brief. Higher is better.
  • MuLan — text-music alignment: how closely the output matches the prompt. Higher is better. For commercial work this often decides usability, because a beautiful track that ignores the brief is worthless.
  • PER — phoneme error rate: how often the vocal mispronounces words. Lower is better, and it is the only column where the direction flips.

A model can lead one column and trail badly in another, which is exactly what happens here.

Neither YuE2 number is a single generation

Best-of-8 is openly labelled. The 6.7316 figure is often described as “single pass”, and it is not. M·A·P's benchmark documentation states that standard YuE2 selects the lower-PER candidate from two generations, and adds that this “is not equivalent to one unselected pipeline call”.

SettingGenerationsSelected bySongBench AvgYuE2 standard2Lower PER6.7316YuE2 best-of-88Musicality, then prompt control, then PER6.9632

Each candidate's PER is itself the lowest of four ASR passes. If you run the model once and keep whatever comes out, you should expect results below both published figures.

Suno v6 is the weakest Suno version tested

Read down the Suno rows and the version numbers stop behaving as you would expect:

Suno versionSongBench AvgSuno v56.8721Suno v5.56.7150Suno v4.56.6995Suno v66.5562Suno v6 Wild6.4195

Three older Suno releases outscore v6 on this aggregate. Anyone announcing that YuE2 has “beaten Suno” by pointing at the v6 row is leaning on Suno's lowest-scoring release.

The honest counterweight, which M·A·P supply themselves: Suno v6 scores higher than both YuE2 settings on SongEval, and both v6 variants record lower phoneme error rates and higher AllMusicCaps scores. The aggregate favours YuE2. Several components do not.

Where YuE2 actually wins outright

The open-weights comparison is the one that survives every caveat above.

Open-weights model SongBench Avg Gap to YuE2
YuE2 (standard) 6.7316
LeVo 2 6.3247 0.4069
MiniMax Music 3 6.2830 0.4486
HeartMuLa 6.2483 0.4833
ACE-Step 1.5 6.0118 0.7198
DiffRhythm 2 5.2428 1.4888
YuE 1 4.9165 1.8151
SongBloom 4.2350 2.4966

A 0.41 gap to the next open model is far larger than anything separating YuE2 from the proprietary leaders, and unlike the Suno comparison it holds in the standard setting rather than needing best-of-8. If you need weights you can download, this is not a close contest.

MiniMax Music 3 belongs in this table. M·A·P classify it as public because its weights are released and the evaluation ran them locally.

LeVo 2 and why single-number rankings mislead

LeVo 2 records the highest AudioBox PQ in the entire table at 8.3966 — better production quality than YuE2, Suno or Mureka. It still lands ninth overall, because its MuLan score is 0.3542 and its phoneme error rate is 26.12%. It makes excellent-sounding audio that follows the prompt poorly and sings words indistinctly.

If you want an instrumental bed and do not care about lyrics or tight prompt adherence, LeVo 2 may serve you better than its ranking suggests. That is not a conclusion the aggregate column can express.

How far YuE2 came from YuE 1

Model SongBench Avg MuLan PER
YuE 1 4.9165 0.2623 36.38%
YuE2 (standard) 6.7316 0.5068 8.44%
Change +37% +93% −27.94 points

A phoneme error rate falling from 36.38% to 8.44% is the difference between a vocal you have to excuse and one you can actually use. It is the largest single-generation improvement the table records.

The result that is missing from the main table

YuE2's zero-shot cover evaluation is reported separately and is arguably its most striking number. Across 948 works, full-score YuE2 reaches 0.647 CLEWS mAP against 0.006 without a score — a hundredfold difference — using the general song-generation checkpoint with no cover-specific fine-tuning.

That gap is the clearest evidence for the symbolic-planning approach in the whole release, and it appears nowhere in the rankings everyone quoted.

What the benchmark cannot tell you

M·A·P are unusually direct about the limits of their own results, and it is worth repeating rather than burying:

  • The small gaps between the highest means do not establish statistical significance.
  • The comparison is not matched-compute. Different systems ran under different candidate protocols; these are documented system comparisons, not controlled experiments.
  • Scoring is automatic, not human preference. Nothing here measures whether listeners prefer the output.
  • Both YuE2 settings used the YuE2-Vae-legacy benchmark decoder, not necessarily what you get by default today.

A 0.2 difference in an automatic aggregate, produced under unmatched protocols and disclaimed by its own authors, is a weaker signal than most of the coverage implied.

Should you switch?

Three situations, three different answers.

  • You need commercial rights. YuE2 is not an option. The weights are CC BY-NC 4.0 and M·A·P intend the restriction to cover generated output, not just the model files.
  • You need local, offline or unlimited generation. YuE2 is the strongest open option by a clear margin, and best-of-8 sampling costs you time rather than credits.
  • You are choosing on quality alone and can use a cloud service. The top five settings sit within 0.25 of each other on an automatic metric its authors decline to call significant. Test on your own material instead.

Running YuE2 on Promptus

YuE2 is available in Promptus as a CosyFlow — a packaged ComfyUI workflow you open in the ComfyUI Canvas without installing anything, wiring nodes or hunting for model weights.

You can run it locally on your own GPU with no generation fees, or use Promptus cloud compute if you do not have a 24 GB card. That choice matters more for this model than for most, because the published scores depend on candidate selection: the standard figure picks the better of two generations and best-of-8 picks from eight. Sampling that way is cheap on hardware you own and expensive on metered credits.

One thing Promptus does not change is the licence. YuE2's weights are CC BY-NC 4.0 and M·A·P intend that restriction to cover generated output, so music you make with it is non-commercial wherever you generated it. If you need commercial rights, the licence article sets out what is and is not permitted, and ACE-Step — also available in Promptus, and in the table above — is Apache-2.0 and does allow commercial use.

Frequently Asked Questions

On the headline SongBench Avg metric, yes: 6.7316 standard and 6.9632 best-of-8 against 6.5562. But M·A·P also report that Suno v6 scores higher than both YuE2 settings on SongEval, and that both v6 variants have lower phoneme error rates and higher AllMusicCaps scores. It leads on the aggregate, not on every measure.

Only in the best-of-8 setting, where its 6.9632 is the highest observed mean. In the standard two-candidate setting it scores 6.7316, behind Mureka 9 (6.9377) and Suno v5 (6.8721). M·A·P state the small gaps between the top means do not establish statistical significance.

Yes, and by the clearest margin in the table. 6.7316 against 6.3247 for LeVo 2, the next open-weights model — a gap far wider than anything separating YuE2 from the proprietary leaders.

Eight generations are produced and one is selected, ranked by SongBench Musicality, then prompt control, then phoneme error rate. It costs eight times the compute of a single generation for one usable track.

Phoneme error rate: how often the sung vocal mispronounces words. It is the one column in the benchmark where the direction flips, which is an easy way to misread the table.

It is documented rather than matched-compute. M·A·P say so directly: public baselines, Suno v6 and v6 Wild use two candidates with four ASR passes, while earlier proprietary systems retain their delivered-candidate protocols. These are system comparisons, not controlled experiments.

If you need downloadable weights, offline generation or unlimited local sampling, YuE2 is the strongest option available. If you need commercial rights, it is not an option at all — the weights are CC BY-NC 4.0 and the restriction extends to the output.

Yes. It is available as a CosyFlow — a packaged ComfyUI workflow — runnable on your own GPU with no generation fees, or on Promptus cloud compute if you do not have a 24 GB card. The CC BY-NC licence on the output applies either way.

Sources: M·A·P YuE repository, its docs/benchmarks.md and THIRD_PARTY_NOTICES.md, and the YuE2-3B model card. Figures checked 15 September 2026 against the WildSongBench run published 12 September 2026.

Written by:
Jack
A professional photographer captivated by Promptus, Jack integrates AI into his workflow to elevate his craft. He views AI as an invaluable tool and plans to continue leveraging its capabilities in his career.
Try Promptus Cosy UI today for free.
ai image generator

AI Generation Platform

Promptus AI is the easiest way to generate realistic photos, videos, 3D and ComfyUI workflows with artificial intelligence.

Our AI photo generator produces lifelike portraits, product images, and creative concepts in seconds, making it the perfect tool for creators and brands.

promptus ai video generator