YuE2-3B Review: English Songs and Editable Scores
By Evan Brooks
Verdict
YuE2-3B is a practical starting point for English song drafts and score-controlled composition experiments. All 12 works in this review completed without a runtime failure or length-limit flag. The full Heartland Rock baseline produced 226.64 seconds of stereo audio in 81.42 seconds of inference. The examples cover six styles, planning modes, a style change, stronger guidance, a supplied-score cover and a chord edit.
Overall, the listening assessment is positive: the songs work well as a whole, the melodies are pleasing, and noise is controlled to an acceptable level. Together with the measured generation speed and editable-score workflow, this makes YuE2-3B a useful option for English song creation and arrangement exploration. The YuE2-3B model card explains the official capabilities and non-commercial weight license, which should inform the choice of project. The practical appeal is the combination of good overall listening results and the option to revisit the score when developing another arrangement.
What I tested
This review evaluates 12 works: six baseline styles and six control variations, all using seed 831001. 12/12 completed without a planner or semantic limit flag, totaling 38.71 minutes of audio. Each setting has one candidate, so the results describe specific outputs rather than a multi-seed reliability estimate.
I used complete English inputs from the project's public demonstration collection, covering Heartland Rock, Cyber Metal, Industrial Hip-Hop, Country Pop, Cool Jazz, and Soul Blues. These examples offer a useful starting point for comparing different musical directions before writing your own prompt. Each work uses seed 831001. Country Pop uses direct generation, while the other five baseline examples begin with full score planning.
Each setting has one generated candidate. The examples therefore show what these particular requests produced, rather than an average across repeated attempts. For a first comparison, start with the style closest to your project, then listen to the related control variation. Keeping the lyric and seed fixed in those comparisons makes it easier to understand what a mode or prompt change means in practice.
Beyond the six style baselines, I compared full planning with melody and off modes, kept the rock lyrics and changed the style prompt to jazz, increased guidance for Cyber Metal, rendered the supplied heavy-metal Jingle Bells score, and edited the first chorus chords of the rock score. These comparisons examine how the available controls affect a song while reusing a baseline wherever possible.
What the score changes about the workflow
An ordinary text-to-music request leaves many composition decisions hidden. You describe a mood and instruments, supply words, and inspect the final recording. YuE2 can make some of those decisions visible first. In full mode it generates an ABC score with melody and chord annotations before generating music tokens and acoustic latents. A separate decoder produces the waveform. In melody mode, planning concentrates on melody; off mode bypasses symbolic planning.
ABC is a compact text notation. Its headers describe such things as key, meter, and tempo, while note symbols and durations describe musical events. Quoted chord names annotate harmony. You do not need to read every symbol to benefit from the separation: you can preserve a score, compare versions, and identify exactly which annotations changed before spending time on another render.
For the user, the score is a reusable starting point for another arrangement. Keep a version you like, change a small part, and render it again. The final recording still needs a listen because the same note text does not guarantee identical singing. This makes score editing useful for exploring variations, while a task that requires an exact repeated performance needs closer audio comparison.
The first English baseline
The initial Heartland Rock input contains a short style description and a complete lyric arranged into verses, choruses, and a bridge. I passed the complete input through the official Python pipeline, with full planning, default guidance, and no quantization. The request-preservation check compared the style and lyric strings received by the resulting plan against the submitted strings. Both strings matched the submitted input.
The Heartland Rock baseline lasts 226.64 seconds. Keeping its words and seed while switching to melody produces 223.44 seconds; off produces 157.32 seconds. These are mode comparisons, not repeated draws from the same setting. They show why I measure actual duration instead of deriving it from the lyric length. A faster request can also produce a shorter recording, so latency needs to be read beside output duration and RTF.
The baseline is a complete, decodable 48 kHz stereo recording saved as 24-bit FLAC, with no reported generation truncation. The lossless file provides an original to retain for later editing, while the MP3 preview makes it easy to listen in the table. The screenshot brings together the input settings and measured output, so you can connect the example you hear with the configuration that produced it.

Generated works and measured performance
| Case | Mode / CFG | Audio | Inference | Peak VRAM | Mean GPU use | RTF | Limit hit | Preview |
|---|---|---|---|---|---|---|---|---|
| A-heartland-rock-831001 | full / 1.0 | 226.64 s | 81.42 s | 9.98 GiB | 89.6% | 0.359 | No | |
| A-cyber-metal-831001 | full / 1.0 | 229.56 s | 77.31 s | 10.23 GiB | 95.8% | 0.337 | No | |
| A-industrial-hip-hop-831001 | full / 1.0 | 224.28 s | 73.65 s | 10.28 GiB | 95.6% | 0.328 | No | |
| A-country-pop-831001 | off / 1.01 | 153.76 s | 43.03 s | 10.67 GiB | 94.9% | 0.280 | No | |
| A-cool-jazz-831001 | full / 1.0 | 199.96 s | 63.89 s | 10.20 GiB | 96.9% | 0.319 | No | |
| A-soul-blues-831001 | full / 1.0 | 199.84 s | 62.02 s | 10.40 GiB | 98.5% | 0.310 | No | |
| B-melody-831001 | melody / 1.0 | 223.44 s | 74.36 s | 10.55 GiB | 98.2% | 0.333 | No | |
| B-off-831001 | off / 1.01 | 157.32 s | 43.95 s | 10.95 GiB | 97.1% | 0.279 | No | |
| C-jazz-831001 | full / 1.0 | 177.16 s | 56.78 s | 10.60 GiB | 94.0% | 0.321 | No | |
| D-cfg12-831001 | full / 1.2 | 234.28 s | 87.68 s | 12.01 GiB | 98.7% | 0.374 | No | |
| E-jingle-bells-heavy-metal-831001 | full / 1.0 | 69.48 s | 18.10 s | 10.18 GiB | 87.8% | 0.260 | No | |
| F-harmony-edit-831001 | full / 1.0 | 227.16 s | 62.11 s | 10.71 GiB | 97.5% | 0.273 | No |
All previews in the table come from these generated candidates. They are not the project's promotional audio. Duration is measured from each saved file, and inference time is a synchronized pipeline call. The table also keeps peak whole-device GPU memory, planning mode, guidance, and the runtime's truncation flags beside the player so that a sample can be assessed in context.
Real-time factor, abbreviated RTF, divides inference time by audio duration. A value of 0.35 means roughly 0.35 seconds of inference for each second of output, or about 2.86 seconds of audio per second of computation. Lower RTF is faster. Read it alongside the absolute waiting time: the first helps compare outputs of different lengths, while the second tells you how long you will wait for that candidate.
The first request also loads the model and decoder, so it can take longer than later requests. When planning a session, allow for that initial wait and then use the warm-request figures to understand subsequent generations. File saving is measured separately. These are generation timings; a complete workflow also includes preparing the lyric, listening to the result, and deciding whether another version is worth generating.
Across the retained requests, output duration ranged from 69.48–234.28 seconds and synchronized inference from 18.10–87.68 seconds. Median warm-request RTF was 0.319; these warm requests have different inputs and modes, so this is a descriptive summary, not a matched speed benchmark. Observed whole-device GPU peaks ranged from 9.98–12.01 GiB.

English lyrics and checking important lines
The Heartland Rock baseline transcription follows the main verse, chorus, bridge and final-chorus sequence. Its automated word discrepancy is 7.7% against a normalized reference of 169 words. The recognizer received no lyric hint. The transcript provides a useful reference for following the song and finding a particular line. Its discrepancy figure describes an automatic text comparison, rather than a score for the singing.
Those percentages describe the speech recognizer, reference normalization and recording together; they are not direct singing error rates. The transcripts contain suspicious introductory phrases and occasional substitutions that may originate in recognition. Singing stretches vowels, instruments can mask consonants, and performances include non-lexical vocalizations. A recognizer may miss a correctly sung word or invent speech during an instrumental passage. For an important line, listen to the corresponding passage alongside the original lyric.
The figures below measure automatic word discrepancy. They differ from the official benchmark's phoneme error rate (PER). WER counts word edits; PER counts phoneme errors. They are different measures even before differences in recognizer settings and candidate selection. Section labels and stage directions are excluded from the reference using documented rules, while the original input is preserved unchanged. Section labels are not sung lyrics, so they are excluded from the text comparison.
Transcription covers 12/12 outputs. The automated discrepancy range is 5.2%–100.0%, with a median of 13.6%. The highest value is D-cfg12-831001; it warrants inspection rather than an automatic claim that the model sang every mismatched word incorrectly. The full transcripts and normalized references are included in the evidence.
Automatic language detection classified 11 files as English and one as Russian: D-cfg12-831001. Its 100% text discrepancy cannot be interpreted as 100% incorrectly sung English words, because the recognition language itself is suspect. With English specified, automatic discrepancy across all candidates ranges from 5.2% to 88.2%, with a median of 13.6%. Both sets of figures are automatic text comparisons; use them to locate passages for listening.
| Case | Auto-language discrepancy | English-specified discrepancy |
|---|---|---|
| A-heartland-rock-831001 | 7.7% | 7.7% |
| A-cyber-metal-831001 | 5.2% | 5.2% |
| A-industrial-hip-hop-831001 | 44.0% | 44.0% |
| A-country-pop-831001 | 9.1% | 9.1% |
| A-cool-jazz-831001 | 10.9% | 10.9% |
| A-soul-blues-831001 | 30.7% | 30.7% |
| B-melody-831001 | 7.1% | 7.1% |
| B-off-831001 | 88.2% | 88.2% |
| C-jazz-831001 | 8.3% | 8.3% |
| D-cfg12-831001 | 100.0% | 47.0% |
| E-jingle-bells-heavy-metal-831001 | 16.3% | 16.3% |
| F-harmony-edit-831001 | 36.7% | 36.7% |
B-off-831001 illustrates why a transcript needs context. Its English-specified discrepancy reaches 88.2%, with a long sequence of repeated “oh” tokens in the recognition output. That number is not the percentage of missing lyrics. If a word or line matters to your project, use the transcript to locate the passage, then compare the recording directly with the intended words. Overall song appeal and exact lyric delivery answer different practical questions.
What the six style inputs reveal
Heartland Rock supplies the cleanest recurring comparison because its lyric is reused for mode, style and score checks. Cyber Metal supplies a longer and more descriptive prompt, including industrial textures and movement from spoken intensity toward a sung chorus. Industrial Hip-Hop increases the density of verbal and arrangement instructions. Together, they provide different starting points for testing how much detail to include in a musical prompt.
Country Pop is especially useful as a lyric diagnostic because the source material includes difficult verbal phrases. A simple genre prompt also puts less descriptive pressure on the input than the extended electronic-rock prompt. Cool Jazz and Soul Blues broaden the requested instrumentation and vocal phrasing. Their purpose is to expose differences that a single upbeat rock example might conceal.
The listening feedback is favorable overall: song quality and melody quality are good, and noise remains acceptable. That makes the collection worth exploring as material for English song drafts and arrangement ideas. To find a useful starting point, compare a complete verse and chorus from the examples closest to your intended direction. The available feedback is an overall assessment of the songs, rather than a separate score for each genre or instrument.
The base matrix contains 6 outputs. Their measured lengths span 153.76–229.56 seconds. The previews let you compare these starting points directly before choosing a style for your own lyric.
Full, melody, and off planning modes
The mode comparison keeps the Heartland Rock lyric and seed paired. Full asks for melody and chord planning, melody uses melody-only planning, and off removes the symbolic plan. This is a comparison of the native modes as delivered. In particular, the default semantic guidance differs slightly: full and melody use 1.0, while off uses 1.01. The comparison retains these native defaults.
Removing a planning stage may reduce work, but a fair timing interpretation must also account for how much audio each mode produces. A faster call that generates a shorter song is not the same improvement as a faster call producing an equally long song. RTF helps normalize that difference, while absolute latency remains relevant when you are waiting for a candidate.
Full mode is attractive when an inspectable score is part of the creative process: it provides a concrete artifact for retaining and revising harmonic decisions. Off mode is useful when the priority is obtaining a candidate with less planning overhead. The listening assessment of the reviewed material is favorable overall, but does not establish a quality order among these modes. I would therefore choose between them according to the need for score control and the measured wait, then select the recording that best fits the intended musical use.
full: 1 outputs, mean duration 226.64 seconds, mean inference 81.42 seconds, and mean RTF 0.359. melody: 1 outputs, mean duration 223.44 seconds, mean inference 74.36 seconds, and mean RTF 0.333. off: 1 outputs, mean duration 157.32 seconds, mean inference 43.95 seconds, and mean RTF 0.279.

Changing style without changing the words
The style experiment takes the Heartland Rock lyric and replaces only its style description with a fixed English jazz instruction. It requests an expressive lead, piano, tenor saxophone, upright bass, brushed drums, no guitar, and spacious harmony. The lyric, seed, full-planning mode, and other generation settings remain paired with the original candidate.
For someone who already has lyrics, this creates a straightforward way to explore another arrangement without rewriting the song's words. Listen to the rock baseline and the jazz variation back to back, focusing on the accompaniment, vocal space, and lyric delivery. The practical question is whether the new arrangement suits the intended song, especially if the original words need to remain prominent.
The prompt provides specific details to listen for: the requested rhythm section, space around the vocal, and the choice of instruments. Treat these as a checklist for selecting the version that fits your project. A broad label such as jazz is less useful than knowing which musical features you actually want. Both candidates are available in the table for that comparison.
Jazz-prompt variation: 1 outputs, mean duration 177.16 seconds, mean inference 56.78 seconds, and mean RTF 0.321.
Stronger guidance is a tradeoff to test
The guidance comparison changes Cyber Metal's full-mode CFG from 1.0 to 1.2 while preserving the original prompt, lyric, and seed 831001. CFG controls conditioning strength; it is not a general quality slider. The official implementation retains the same symbolic context in the conditioned and unconditioned branches for this mode, so stronger guidance changes the semantic generation work as well as its conditioning.
The relevant tradeoff is the extra wait for each candidate. If you are trying several arrangements, that cost accumulates across the session. Compare the two recordings alongside their timings, with particular attention to the part of the prompt you want to strengthen. A setting is useful when it improves the result you need enough to justify generating with it again.
CFG 1.2 is an alternative to compare with the default, rather than an automatic upgrade. The test uses the same seed for both settings, and the results do not establish a general quality advantage for stronger guidance. Listen for the balance of lyric clarity, vocal sound, and accompaniment when choosing between them. An improved transcript alone does not settle that choice.
CFG 1.2: 1 outputs, mean duration 234.28 seconds, mean inference 87.68 seconds, and mean RTF 0.374. Against the paired CFG 1.0 mean of 77.31 seconds, that is 10.37 seconds more per candidate and a 11.1% higher mean RTF. The recorded semantic branch count changes from one to two. This matched-seed comparison incurs the additional work; a quality benefit still requires assessment.
Covers from supplied scores
The supplied-score example is a heavy-metal arrangement of Jingle Bells, using complete English lyrics and ABC notation with seed 831001. It shows the workflow for starting from an existing score: supply the notation and style, then generate a recording. For users with a melody already written down, this is a different starting point from asking the model to compose everything from a lyric.
An existing audio file is a different input from an ABC score. Turning an arbitrary recording into this workflow requires a separate transcription step to obtain the notation. That audio-to-score process was not part of this test. If you only have a recording, account for that additional preparation; if you already have compatible notation, the supplied-score example is the relevant one to inspect.
The supplied score and style instruction both contribute to this recording. Its presence in the table does not establish that any arbitrary melody will be preserved equally well. A listener should compare the result with the supplied notation and lyric, while keeping arrangement changes separate from melody errors. Familiarity with Jingle Bells helps locate musical phrases, but recognition alone is not a note-level fidelity metric.
Supplied-score covers: 1 outputs, mean duration 69.48 seconds, mean inference 18.10 seconds, and mean RTF 0.260.
Editing harmony while preserving the score's notes
The edit experiment starts from the actual Heartland Rock baseline score. I changed eight chord annotations in its first chorus: G becomes G6 and D becomes D7sus4. Those substitutions accommodate the sustained melody pitches visible in that passage. I did not regenerate the score or change the lyric, style instruction or seed. The edited input and its original are retained so that the exact intervention can be inspected.
The edit leaves the note text, rhythm, tempo, section layout, and voice definitions unchanged; only the chord annotations differ. That gives the user a manageable version change to inspect before another render. Keeping the original score beside the edited one also makes it easy to return to the earlier input or try a different harmonic idea without rebuilding the entire song plan.
This comparison has an important timing limitation: the baseline includes score generation, while the edit request supplies a score directly. A shorter total call therefore does not establish that audio synthesis became faster. The baseline audio lasts 226.64 seconds and the edited result 227.16 seconds, but similar lengths do not prove identical melody performance. This pair involves both supplying a saved score and changing its chords, so not every audio difference can be attributed to the chord edit alone.
Edited-score renders: 1 outputs, mean duration 227.16 seconds, mean inference 62.11 seconds, and mean RTF 0.273. The original-note text constraint is verified; audible chord adherence is not rated.

Saving a version for further work
When you find a version you want to develop, keep the request and score alongside the audio. The seed is useful, but it is only one part of the settings: the lyric, style, planning mode, and model and decoder versions also matter. These twelve examples compare different inputs or settings with a common seed; they do not establish identical audio replay from a saved request.
Keeping request and runtime identities is valuable even when every result sounds acceptable. A decoder revision, sampling default, or tokenization change can alter a later reproduction. Saving only an MP3 and a seed discards the information needed to work out why. The official artifact bundle includes the score, exact token arrays, latent representation, effective settings, and weight identities; my measurements accompany those files.
Candidate selection changes the question a benchmark answers. The official standard YuE2 result uses candidate selection, while best-of-eight has a larger attempt budget. This review instead reports one candidate for each listed setting. That has a different attempt cost from generating several versions and selecting the best one. Keeping the attempt budget explicit is necessary before drawing a quality or efficiency comparison.
Audio format and preparation for editing
The original outputs are lossless FLAC files; MP3s are delivery previews. I check finite samples, duration, channel count, sample format, signal level, channel correlation, and low-level intervals. I also measure integrated loudness and true peak from the original files. These checks help identify file and level issues before taking a selected song into further editing.
The first Heartland Rock candidate has an integrated loudness measurement of -13.18 LUFS and a true-peak estimate of +0.37 dBTP. Its channels are not identical, with a correlation of approximately 0.832. The true-peak estimate suggests checking output headroom before further encoding or mastering. Before exporting, listen to the louder passages and leave headroom for subsequent encoding.
Waveforms are useful when locating a quiet interval or comparing the length and level envelope of two renders. Listen to those passages in context before editing them: a pause can be a musical choice. For selecting a song, the recording remains more useful than the plot. For preparing the selected file for another editor or an export, the level measurements provide an additional check.
Among 12 analyzed originals, true-peak estimates range from -0.40 to 0.56 dBTP; 11 exceed 0 dBTP. These peaks flag output headroom to check before export; listen to the relevant passages to assess audible distortion.
Resource requirements and practical waiting time
When exploring YuE2 AI, distinguish the native model's resource use from website performance. The timings in this review measure the Python generation pipeline and do not establish the website's response time.
The official starting point is a BF16-capable NVIDIA GPU with 24 GB of VRAM, Linux, and sufficient host memory; the documentation asks for 24 GB of available RAM. Allow room according to the official starting recommendation when choosing a configuration. Memory allocation policy, context length, other processes, and longer scores can change the room a request needs.
I recorded whole-device GPU memory as well as the Python framework's allocated and reserved peaks. Those are different views. Reserved memory can exceed active tensor storage, and whole-device readings can include allocations outside the framework. The sampling interval was 250 milliseconds, so a very brief spike could fall between samples. Exact machine and software identifiers are preserved in the evidence notes for reproducibility.
All twelve examples were generated one at a time, so their timings are most relevant to a person working through individual song candidates. A service handling several users at once also has scheduling and queueing to manage, which was not measured here. For your own session, distinguish the initial loading wait from later generation times and leave time to listen before requesting the next version.
YuE2-3B versus Suno and other local models
The official YuE2 benchmark is useful context, but this review is not a new head-to-head listening contest. I did not generate a matched set with Suno or ACE-Step alongside these files. The official published comparison uses its own prompts, evaluators, decoder, and candidate-selection rules. The results apply to those particular comparison conditions.
YuE2's clearest workflow distinction in this test is the inspectable score and the ability to supply an edited version. That is relevant when you want to preserve compositional intent while changing a limited part of the input. A hosted service may instead be attractive because it removes installation and resource management. These are different product tradeoffs, and neither tells me which recording a listener would prefer.
The ACE-Step 1.5 music generation and editing tools document another local workflow with different controls to consider. Its official repository is a reasonable place to examine alternative controls and supported platforms. I would compare the specific edit you need, installation burden, rights, and cost per acceptable candidate before choosing between tools. The most useful choice depends on the work you want to do after the first generation, particularly whether you need access to the score.
Setup and reproduction
| Component | Reproduction choice |
|---|---|
| Platform | Linux, Python 3.10+ |
| Official starting resources | 24 GB GPU memory and 24 GB available RAM |
| Inference package | yue2-infer 0.1.5 |
| Core libraries | PyTorch 2.10.0; Transformers 4.57.6; huggingface-hub 0.36.2 |
| Precision | BF16 main model; FP32 default VAE |
| Generation | Native defaults; full unless a case says otherwise; no quantization |
To reproduce this configuration, download the pinned model and decoder revisions from Hugging Face, install the inference package in an isolated uv environment, and load the weights from local paths. A slow package-download source was the main setup obstacle; switching transport and checking hashes resolved it without changing the generation settings. The included English request provides a starting point for your first song.
uv venv --python 3.10 .venv
uv pip install --python .venv/bin/python huggingface-hub==0.36.2
.venv/bin/hf download m-a-p/YuE2-3B --revision 29b3558dd46954a0cd9021dc76d5c91864a0f1c7 --local-dir model
.venv/bin/hf download m-a-p/YuE2-Vae --revision 9a94e1d0ea9f8087e98f77fa88df4a4068104d2a --local-dir vae
uv pip install --python .venv/bin/python ./model/yue2_infer-0.1.5-py3-none-any.whl
import json
from pathlib import Path
from yue2 import YuE2Pipeline
request = json.loads(Path("requests/A-heartland-rock-831001.json").read_text(encoding="utf-8"))
pipe = YuE2Pipeline.from_pretrained(
"model", vae="vae", local_files_only=True,
device="cuda", backend="torch", quantization="none",
memory_budget_gib=24,
)
song = pipe(**request)
song.save_artifacts("outputs/heartland-rock")
pipe.close()
Save the Python example as generate.py and run .venv/bin/python generate.py from the package directory. The referenced request is included with the review package. For timing work, synchronize device execution before and after the pipeline call, identify the first call, and retain the effective config. Quantization and runtime choices can change generation time.
Limitations and capability profile
| Capability | Evidence status |
|---|---|
| Native generation and file format | Measured in the retained outputs |
| Complete input preservation | Exact submitted style and lyric strings checked |
| Recoverable English lyric content | Automatic text comparison for locating lyric passages |
| Planning, CFG and external ABC controls | Executed requests and effective settings retained |
| Constrained score edits | First-chorus chord changes verified in text |
| Overall song quality | Positive qualitative human listening assessment |
| Melody quality | Favorable human listening feedback |
| Noise control | Acceptable in the listening assessment |
| Exact style, instrument and chord adherence | Not established by the general listening assessment |
| Arbitrary recording to cover | Not tested; source transcription is separate |
| Concurrency, quantized variants, legacy decoder | Not tested |
| Commercial suitability | Not established by generation success |
The review combines two kinds of evidence. Native-pipeline records establish input preservation, artifact integrity, timing, resource use and the exact scope of the score edit. Human listening provides a favorable assessment of overall song quality, melody quality and acceptable noise. Automated transcription adds a separate diagnostic for recoverable English words. These findings complement one another, but a general listening judgment does not establish exact note reproduction, chord-by-chord adherence or a standardized comparative music-quality score.
One candidate per setting cannot establish a success rate across random seeds, genres or lyric lengths. The official demonstration inputs are also a selected population and may be more favorable than arbitrary prompts. Their published recordings may come from different checkpoints or settings. Different versions and settings can produce different recordings from the same example input. The results support the listed cases and no broader statistical guarantee.
The weights are labeled CC BY-NC 4.0 on the model card. Treat commercial suitability as a separate licensing question, including any source lyric or recording rights. Nothing in a successful local generation resolves those questions. I would also keep a clear distinction between supplied-score covers and automatic audio-to-score workflows, between one-shot generation and a multi-candidate selection system, and between a verified text edit and an audible performance change.
FAQ
How do I install YuE2-3B, and what resources should I allow?
Start with the official model and default VAE repositories, install the pinned inference wheel in an isolated uv environment, and run the included complete English request through the Python pipeline. The official starting recommendation is 24 GB of GPU memory and 24 GB of available system RAM. The lower observed peaks in these cases are measurements, not a proven minimum hardware requirement.
Can YuE2-3B generate complete English songs?
The full Heartland Rock baseline generated 226.64 seconds of stereo audio from a complete English lyric, and its automatic transcript follows the main lyric sequence. The listening assessment also found the reviewed songs and melodies good overall, with acceptable noise. Together these results support both a working generation pipeline and a positive practical listening result, while reliability across other seeds is a separate question.
Does YuE2 need a separate decoder?
Yes. The main model and default YuE2-Vae were both fixed to explicit revisions for these tests. The default VAE used here differs from the legacy decoder used in another comparison setting. Keep the decoder identity beside your generated artifacts when comparing results.
Can I set an exact song duration?
This selection does not establish exact-duration control. The same Heartland Rock words produced different durations when the planning mode changed. Token limits can stop generation, but reaching a limit is not proof of a natural ending or of compliance with a target length. All twelve files were measured at their actual duration.
Does a lower automated word discrepancy mean better music?
No. It may indicate easier word recovery under that recognizer, but it can also reflect recognition errors. It says little about vocal tone, instrumentation, musical development, or emotional impact. Use the transcript to locate a possible issue, then inspect the recording.
Is YuE2-3B better than Suno?
These measurements do not answer that question. A fair comparison needs matched inputs, comparable candidate budgets, clear selection rules, and independent evaluation. The official leaderboard provides context under its own protocol, not a substitute for that work.
About the author
![]()
Evan Brooks writes practical reviews of AI models and open-source tools, covering output quality, setup, speed, and everyday use. Each review brings together test results and sample outputs to help readers decide whether a tool fits their needs.