What You'll Build
A song of any length the model will write — vocals and accompaniment, 48 kHz stereo — generated on one 16 GB RTX 4080 Super from a style prompt and lyrics, with the editable ABC score the model planned it from, so you can revise the melody or harmony and re-render the same song.
Hardware data: RTX 4080 Super (16GB VRAM, Ada Lovelace) · a full-context song was measured at 12.64 GiB in-process on a 16 GB card, with no OOM · See benchmark data
ℹ️ Where the measurement came from, and why song length is the question it answers. The vendor asks for a 24 GB card and publishes nothing for any 16 GB board, so on 2026-09-10 the site operator ran YuE2-3B four times on an RTX 5060 Ti 16GB. Two of those runs used the shipped semantic budget; the other two fill the model's context to the token, which matters because the context is the one setting the library refuses to change — leaving song length as the only lever a reader actually has. None of the four ran out of memory, the longest included. Those figures are that card's and are labelled so throughout. What transfers to an RTX 4080 Super is the memory verdict, because the tensors that decide it are sized from the model's
config.jsonand the library's own capacity-based clamp; what does not transfer is time, and this card is the faster of the two on memory bandwidth. If you run this pair, send us your numbers.
⚠️ The weights are non-commercial. YuE2's code is Apache 2.0, but the checkpoints this recipe downloads are CC BY-NC 4.0. The repository's
LICENSEscopes that to "the YuE2 checkpoint weights in model.safetensors, or the corresponding" model files. The licence text names the weights and does not address the audio you generate with them — if you intend to release or monetise output, read the full licence rather than this paragraph.
ℹ️ One runtime, and it is a Python wheel. YuE2-3B runs through the vendor's own
yue2package with CUDA. There is no llama.cpp, Ollama or LM Studio path; a GGUF conversion exists for a different engine and its status changed on the day this page was written — see Troubleshooting.
Requirements
| Component | Minimum | This recipe |
|---|---|---|
| GPU | 16GB VRAM, BF16-capable NVIDIA (the vendor's quick start asks for 24GB; the floor below is measured on a 16 GB card) | RTX 4080 Super 16GB — no run on this card exists. The fit comes from four measured runs on an RTX 5060 Ti 16GB, 9.03–12.64 GiB in-process with zero OOM, on a board of the same capacity and a later architecture |
| RAM | 24GB available host RAM (the vendor's figure) | — the reference rig had 23 GiB available to its WSL2 VM and none of its four runs was killed |
| Storage | 7.79 GB of weights and tokenizer files | 7.79 GB — byte counts read on 2026-09-10 from the Hugging Face tree API for the two repositories the loader downloads (YuE2-3B, YuE2-Vae) |
| Software | Linux, Python 3.10+, CUDA build of PyTorch 2.10 | Reference environment: WSL2 Ubuntu 24.04, Python 3.12.3, torch 2.10.0+cu128, CUDA 12.8, yue2_infer 0.1.5 |
The vendor states its hardware line as "Linux · Python 3.10+ · 24GB NVIDIA GPU with BF16 support." on the model card, and as "NVIDIA GPU with BF16 support and 24 GB VRAM." in the GitHub repository — a capacity in both cases, with no architecture attached. The 24GB is a recommendation rather than a floor the software enforces, and it is why somebody opened a request on a low-VRAM inference project the day the model shipped, asking "Any possibility to optimise within 16gb vram?" (Wan2GP issue #2283, 2026-09-10). Nothing has to be optimised. The next section derives the ceiling and shows that the longest song the code will produce cannot break it.
Why 16 GB holds, in bytes — and why a longer song cannot change it
The library caps the process, not the card. YuE2Pipeline.__init__ reads the device's total memory and sets budget = min((memory_budget_gib - 2) * 2**30, total - 2 * 2**30), then enforces it with set_per_process_memory_fraction (src/yue2/pipeline.py L161-165). memory_budget_gib defaults to 24 on every card, so on a 16 GB board the total - 2 term binds. total is what CUDA reports for your board — a little under 16 GiB, and not guaranteed identical between boards of the same nominal size; the measured card reported 15.93 GiB, so its ceiling was 13.93 GiB. Take yours from yue2 doctor in step 2 and subtract 2. Two consequences are universal: an out-of-memory error can arrive with memory physically free on the board, and the figure to compare against the ceiling is the process's own peak — not the device peak, which also includes whatever a display or monitoring tool holds, and which the clamp never sees.
The KV cache is sized by the budget you ask for, not by how many tokens the model emits. The graph runner preallocates, per layer and once each for keys and values, a tensor shaped (branches, capacity, num_key_value_heads, head_dim) where capacity = max(len(prefix)) + max_tokens (src/yue2/cuda_graph.py L70, L90-L92). From config.json — 28 layers, 8 key/value heads, head dimension 128, dtype bfloat16 — one token on one branch costs 28 × 8 × 128 × 2 bytes × 2 tensors = 114,688 B, i.e. 112 KiB. At the full context that is 2.625 GiB on one branch and 5.25 GiB on two. This is the sentence that makes song length the interesting variable: the cache is allocated for the budget, so it is already at full size before the model has emitted a note.
Two is the hard maximum, and so is the context. The sampler builds either [prefix] or [prefix, negative] — one branch at guidance exactly 1, two otherwise, and never more (src/yue2/sampling.py L92); cot="off" takes the two-branch path because its default guidance is 1.01 (src/yue2/protocol.py L106). A request whose prefix plus budget exceeds 24,576 tokens is rejected outright rather than quietly trimmed — "Prefix + requested generation budget exceeds 24576; no implicit truncation" (src/yue2/sampling.py L62-63) — and the context is not a setting you can trade away for memory: the generation config raises "Require context=24576 and midpoint with positive integer steps" for any other value (src/yue2/protocol.py L54-55).
So the worst case is bounded, and the bound is the case that was measured:
| Term | Bytes | GiB |
|---|---|---|
| AR/NAR checkpoint, bfloat16 | 7,261,441,640 | 6.763 |
| KV cache, full 24,576-token context on two branches | 5,637,144,576 | 5.25 |
| Sum | 12,898,586,216 | 12.013 |
Against a 16 GiB card that leaves 3.987 GiB for activations, the CUDA context and fragmentation — and the measurement lands where that predicts: the full-context two-branch run on the RTX 5060 Ti 16GB peaked at 12.64 GiB in-process, 0.627 GiB above the arithmetic floor, with 1.29 GiB still under that card's ceiling.
A longer song cannot exceed it, and this is the claim the round was run to test. Song length past the context limit is realised as more NAR and VAE chunks, not bigger ones: each chunk's width is computed from the context and the prefix, size = min((context - prefix_tokens - 3) // 2, CONTEXT) (src/yue2/protocol.py L141-145), and the function returns a list of such ranges over the frame count, processed one after another. So a ten-minute song is more chunks of the same width, not wider tensors. Combine that with the two bounds above — the cache is already at full capacity, and there is no third branch — and the peak in the table is not merely the largest thing anyone has measured; it is the largest thing the shipped code can allocate. More audio costs time.
Installation
1. Install the inference package into a dedicated virtual environment
The venv is not a style preference. The wheel pins all eight of its runtime dependencies with ==, torch==2.10.0 among them, and an == pin does not step aside for a newer version that is already installed — it replaces it. A ComfyUI user reported on 2026-09-10 that installing this wheel into a working portable build rewrote PyTorch 2.11+cu130 down to 2.10 and left the launcher failing to start. There is also no yue2, yue2-infer or yue2_infer package on PyPI, and the vendor's own setup notes name the hazard: "unverified package with a similar name from PyPI." is what not to install (skills/yue2-music/references/models-and-setup.md).
python3 -m venv .yue2 && source .yue2/bin/activate
python -m pip install -U pip huggingface-hub==0.36.2
hf download m-a-p/YuE2-3B yue2_infer-0.1.5-py3-none-any.whl --local-dir .
python -m pip install ./yue2_infer-0.1.5-py3-none-any.whl
Those middle three lines are the model card's quick start unchanged, and 0.1.5 is the version the card names — the same wheel the runs below were taken on (sha256 8801e2c0…, 66,117 bytes). The repository tree has moved on to 0.1.6; unpacking both and comparing them file by file leaves 12 of 14 modules byte-identical, the differences confined to the version string and to cli.py. Every engine module cited on this page — pipeline.py, cuda_graph.py, sampling.py, protocol.py, storage.py, quantization.py — is identical in both. Install PyTorch's CUDA build first if your environment would otherwise resolve a CPU-only wheel; beyond that there is no wheel selection to get right here.
2. Confirm the card, the dependencies, and your own ceiling
yue2 doctor
Read dependencies_ready and the cuda array in the JSON it prints. The array names your board, gives its total memory in GiB and reports its compute capability; subtract 2 from the total and you have the ceiling the clamp will hold this process to for the whole run, including the longest song you can ask for. doctor reports environment readiness and nothing more, which its own output says outright — "Environment readiness is not quality or real-24GB acceptance." (src/yue2/cli.py L78).
3. Let the loader fetch the weights
The first pipeline call downloads what it needs from m-a-p/YuE2-3B and m-a-p/YuE2-Vae. The loader uses an explicit allow-list (src/yue2/storage.py L32-40) rather than a full clone, so the demo audio and images in those repositories are skipped — that is why the Storage row is 7.79 GB rather than the 7.83 GB the two repositories hold in total. Weight files are hash-checked against weights_manifest.json as they load, which took 30.40–34.81 s across the four reference runs.
Running
Load the pipeline once, then generate:
import json
from pathlib import Path
from huggingface_hub import hf_hub_download
from yue2 import YuE2Pipeline
repo = "m-a-p/YuE2-3B"
demo = json.loads(Path(hf_hub_download(repo, "examples/tonight-awake.json")).read_text(encoding="utf-8"))
pipe = YuE2Pipeline.from_pretrained(repo, device="cuda")
song = pipe(style=demo["style"], lyrics=demo["lyrics"], cot="full", seed=demo["seed"])
song.save("song.flac")
song.save_artifacts("outputs/song") # ABC, tokens, latents, audio and settings
That is the exact call the two shipped-budget runs below were measured on, with the vendor's own example prompt — which ships inside the model repository and is fetched separately from the weights, because the loader's allow-list does not include it. To lengthen a song you raise the semantic budget rather than the context; the two context-filled runs drove the same pipeline one stage at a time (plan → generate_semantic → synthesize → decode) purely so that budget could be set to 24576 - len(plan.prefix) exactly, which is the largest value the sampler will accept. Every other parameter stayed at its shipped default.
song.flac is the finished stereo song. outputs/song holds the ABC score, the semantic tokens, the acoustic latents and the settings used — edit score.abc and pass it back as abc= to re-render the same song with a revised melody or harmony.
The same modes are available from the shell, which is the easier path when you want to queue several songs:
yue2 generate --cot full --style "Mandarin funk, nu-disco" --lyrics-file lyrics.txt --output outputs/song
cot="full" plans melody and chords and is the default; cot="melody" plans melody only and is what the vendor recommends for covers; cot="off" generates with no symbolic plan and, as derived above, is the mode that costs a second KV branch. Do not pass --budget on a 16 GB card: the board already clamps the default, so 16 changes nothing, and 12 or below lowers the ceiling to 10 GiB and halves the VAE decode tile (src/yue2/pipeline.py L148) — a real change to the decode path, and the wrong knob to reach for on a long song. Expect to run one song at a time; the vendor's own framing is "One song at a time.", and the batch subcommand queues requests rather than running them concurrently.
Results
Measured on a 16 GB card, and not on this one
All four rows were run by the site operator on 2026-09-10 on an RTX 5060 Ti 16GB (Blackwell, compute capability 12.0, driver 591.86, WSL2 Ubuntu 24.04, yue2_infer 0.1.5). Instrumentation was whole-device NVML sampled at 20 Hz — the same instrument the vendor used, and deliberately not torch.cuda.max_memory_allocated(), which excludes both the CUDA context and the allocator's reserved-but-unused blocks and therefore reads low. Each run was a separate process, because PyTorch's caching allocator returns nothing to the driver by itself: after del and gc.collect() the pipeline still held 7.92 GiB, and only torch.cuda.empty_cache() brought it to 1.14 GiB. A display was attached to that card and held 0.76–0.84 GiB throughout, so both a device peak and a peak-minus-idle figure are given; the second is the one the clamp applies to. One rig, one operator, one run per configuration.
| Run | Semantic budget | KV branches | Audio | Generate | Peak, device | Peak, in-process | Under that card's 13.93 GiB ceiling |
|---|---|---|---|---|---|---|---|
cot="full", shipped budget | 9,000 tokens | 1 | 201.04 s | 205.23 s | 10.05 GiB | 9.03 GiB | 4.90 GiB |
cot="off", shipped budget | 9,000 tokens | 2 | 165.40 s | 144.39 s | 10.27 GiB | 9.25 GiB | 4.68 GiB |
cot="full", context filled | 18,989 tokens | 1 | 299.96 s | 315.53 s | 11.51 GiB | 10.67 GiB | 3.26 GiB |
cot="off", context filled | 22,827 tokens | 2 | 345.68 s | 298.55 s | 13.48 GiB | 12.64 GiB | 1.29 GiB |
- VRAM usage, read by song length: in table order, the two shipped-budget runs produced 3.35 and 2.76 minutes of audio and peaked at 9.03 and 9.25 GiB in-process — note that the shorter of the two peaked higher, because it is the two-branch mode. The two context-filled runs produced 5.00 and 5.76 minutes and peaked at 10.67 and 12.64 GiB. The peak rises with the budget you request, not with the audio you get, which is why the last row is the ceiling for every longer song too. None of the four failed, and the closest approach to a limit was 1.29 GiB. This is the half that carries to an RTX 4080 Super: every input is the model's or the board's capacity, never its architecture.
- Speed: not stated for the RTX 4080 Super, on purpose. Seconds belong to the silicon. NVIDIA's comparison page lists this card as 16 GB GDDR6X on a 256-bit interface against the measured card's 16 GB GDDR7 on 128-bit (compare specs), and board-partner tech specs put the pins at 23 Gbps and 28 Gbps respectively (RTX 4080 Super, RTX 5060 Ti). That is 736 GB/s against 448 GB/s — 1.64×, the widest margin of any GDDR6X card in this capacity tier — so the measured times are an upper bound rather than an estimate, and the generation-old memory does not change the direction because the bus width doubles. An upper bound is still not a measurement, and the numbers a long song produces are exactly the ones nobody has for this card. Send us yours.
- The context-filled rows are the point of this page. The vendor reports that "maximum-context testing peaked at" 14.08 GiB on its own hardware — a figure that, read as a device peak against a 16 GB card's process ceiling, appears to rule out a maximum-length song. It does not: the two are measured against different things, and the run that filled the context to the token fit with 1.29 GiB to spare. The vendor never says what its maximum-context run was, so this page claims to have measured the worst case the shipped code can be asked for, not to have reproduced theirs.
- Quality notes: the vendor reports a 6.7316 SongBench average for YuE2 and 6.9632 for its best-of-8 setting across 192 WildSongBench prompts, using the legacy decoder — its own automatic metrics under its own candidate-selection protocol, not an independent evaluation. Pass
vae="m-a-p/YuE2-Vae-legacy"tofrom_pretrainedto reproduce that protocol; the default decoder ism-a-p/YuE2-Vae, which is what the rows above used.
The published speed figures, and why a long song is where they stop helping
The vendor's reference card is an RTX 4090 24GB, where it publishes 139.48 LM tokens/s, 214.85 s of audio in 71.04 s and an 11.18 GiB NVML peak in cot="full" mode, headlined as "A 3.6-minute song in 71 seconds on an RTX 4090." (model card), with a method note of "NVML records the full-run GPU peak." — the same instrument as above. Note the length: 3.6 minutes is a shipped-budget song, so the vendor's headline says nothing about the five-and-three-quarter-minute case that fills the context. The only third-party figure in existence has the same limit and less detail: the author of a C++ port put his engine at a real-time factor of 0.30 against "python 0.31" for the reference implementation on the vendor's example prompt (audio.cpp issue #499, 2026-09-10), naming no card in that comment, no mode and no method.
Nothing is filed at /check/yue2-3b/rtx-4080-super yet, so that page is where the first measurement of this pair will land — contribute yours.
Troubleshooting
The unquantized preset requires CUDA BF16 support
This is the pipeline's only hard hardware gate, and it tests BF16 rather than an architecture: the constructor calls torch.cuda.is_bf16_supported() and raises that message when it returns false (src/yue2/pipeline.py L158-160). Ada supports BF16, so an RTX 4080 Super passes. Nothing else in the shipped package tests compute capability on the default path.
An out-of-memory error with memory free on the card
That is the clamp, not the board. set_per_process_memory_fraction holds the process to total - 2 GiB while nvidia-smi still shows free memory: on the reference card the two-branch context-filled run peaked with 2.45 GiB physically free and only 1.29 GiB under its own 13.93 GiB ceiling. Two things make it worse and both look like helping. Passing --budget 12 or lower caps the process at 10 GiB and halves the VAE decode tile. Running a second generation inside the same Python process leaves the previous run's allocations resident, because the caching allocator does not hand them back to the driver — on the reference machine a loaded pipeline held 7.92 GiB, still 7.92 GiB after del and gc.collect(), and 1.14 GiB after torch.cuda.empty_cache(). If you are batching long songs, that second point is the one that will bite: call torch.cuda.empty_cache() between them, or use one process per song.
RuntimeError: USE_FLASH_ATTENTION was not enabled for build. on native Windows
The reference runs were made under WSL2 on Linux wheels; a native Windows install hits this instead. The graph runner's flash-attention eligibility test asks whether torch.ops.aten._flash_attention_forward exists and whether its schema carries seqused_k (src/yue2/cuda_graph.py L75-76) — a question about the operator's registration, not about whether the build compiled its kernel. On a wheel built without USE_FLASH_ATTENTION the operator is registered and then throws when called, which is precisely what a community ComfyUI wrapper documents for official Windows wheels: "The op exists and then crashes" (ComfyUI-YuE2 README). The vendor ships the escape hatch itself: backend="torch-eager" (or --backend torch-eager) is accepted at src/yue2/pipeline.py L128-129 and turns the CUDA-graph path off at L250, so the selection above never runs. Do not pip install flash-attn in response; the package declares no such dependency and its "FlashAttention" is PyTorch's own kernel. Two caveats keep this page on the default path: the eager route allocates its per-branch KV cache at len(ids) + max_tokens (src/yue2/sampling.py L76-79), so the bound derived above still holds and a long song cannot break it there either — but its peak has been measured by nobody, and neither has its speed. Requesting flash explicitly at least fails at construction: attention_backend="flash" raises "Pinned PyTorch variable-length CUDA FlashAttention is unavailable" (src/yue2/cuda_graph.py L82-83).
Memory budget must leave room for a 2GiB reserve
The same clamp on a card too small for it: budget comes out at zero or below once the device has 2 GiB or less. It is also why the floor here is 16 GB. A 12 GB card's ceiling works out to 10 GiB — both terms of the min evaluate there, so no --budget value raises it — against the vendor's own 11.18 GiB ordinary-case peak, with the context not reducible and the two candidate levers (--budget 12 --offload-ar) unmeasured by anyone. That is a card with no documented headroom, which is not the same statement as a proven failure, and it is why 12 GB stays out until someone runs it.
Raising cfg_scale costs a second KV branch
The model card's tuning table suggests cfg_scale=1.2 to strengthen text guidance, and any guidance other than 1 makes the sampler build a second, unconditional branch whose count is the leading dimension of the KV cache — 2.625 GiB more at full context, as derived above. On 16 GB it is affordable, and the two-branch case is measured rather than assumed: both cot="off" rows run two branches, and the context-filled one is the largest KV allocation the code permits, so raising cfg_scale cannot push that term past the table even on the longest song. What those rows cannot price is the second branch alone, because cot="off" also skips the ABC planning stage entirely. Change one thing at a time and watch the peak.
Experimental FP8 AR requires CUDA compute capability >=8.9
The package has an opt-in FP8 mode for the AR projections, selected with quantization="fp8", and it is the only place in the package that tests compute capability (src/yue2/quantization.py L71-72). NVIDIA's CUDA GPUs list does not enumerate the SUPER SKUs, and every Ada GeForce entry it does list reports 8.9 (CUDA GPUs), so on mechanism the gate is expected to pass here — but read the number out of yue2 doctor rather than trusting that inference. It makes no difference to the recommendation, which is to leave the mode off: its own status output calls its quality unvalidated, it is incompatible with the CUDA-graph decode path, and it keeps a second copy of the BF16 weights in host memory. The unquantized model fits, as measured.
The GGUF files on the Hub, and the C++ engine that reads them
A GGUF conversion of this model exists at audio-cpp/Yue2-3B-GGUF, and it is not a llama.cpp or Ollama artifact — its header declares general.architecture = "audiocpp", which is not among llama.cpp's registered architectures. It targets 0xShug0/audio.cpp, a separate ggml-based audio engine, and the port landed on that engine's dev branch while this page was being written: re-checked at head 3caeba87, committed 2026-09-10 at 17:09 UTC, 29 yue2 paths including src/models/yue2/ and docs/models/yue2.md. Its author reports a peak below 9 GB and a real-time factor of 0.23 to 0.28 through that engine's interface, on a 32 GB card, with the caveat that "The numbers are a bit noisy because of background workloads" (vendor issue #163). Before that sub-9 GB figure reads as an upgrade path: it is a different engine on a Q8 quantisation, self-reported by the port's author and unverified by us, and it is the deliberately smaller of his configurations — in the same thread he says a more aggressive cache would take the real-time factor below 0.2 at higher VRAM, which is the same trade this page's derivation describes in reverse. By his own announcement the model sits in a branch for community testing that "won’t be included in the prebuilt binaries before merging" (discussion 3). Build dev if you want it, and expect to be an early tester. This page documents the vendor's Python runtime, which is what every figure above was measured on.
Community reports so far
YuE2-3B was published on 2026-09-09. Re-checked on 2026-09-10 at 21:00 UTC: m-a-p/YuE2-3B carries three discussion threads, all opened the previous day — a question about getting reliably instrumental output, a ComfyUI node-pack announcement, and the audio.cpp port above — and m-a-p/YuE2-Vae carries none. The same query against the org's YuE v1 model returns ten threads, so the near-silence is the model's age rather than a broken check. A GitHub-wide issue search for YuE2 returns 22 items: the vendor's own release pull requests, three requests to port the model to other engines, and two bugs in a community ComfyUI wrapper. Reading the title and body of all 22, all three discussion threads in full, and the comment threads on the port requests: no SUPER-series card is named, and neither is any other 16 GB board — and neither the vendor nor any third party publishes what a full-context song costs in time on any card. The only such figure anywhere is our own, measured on an RTX 5060 Ti 16GB and reported on that card's page; it is a Blackwell 128-bit board, so it bounds this card's memory and not its clock. Every hardware-shaped issue in the vendor's tracker that a search for "YuE" surfaces — a ROCm report, Apple Silicon requests, GGUF requests — predates YuE2 and was filed against YuE v1, a different model with a different runtime; do not transfer their answers. Nothing in the current package mentions ROCm, HIP or Metal, and the vendor documents no AMD path.
If you hit something, please report it via the submission form.