# MoE LoRA Converter Converts an Unsloth-trained, MoE-expert-targeting LoRA adapter (one fused tensor per layer covering all experts, via PEFT's `target_parameters` mechanism) into vLLM's expected per-expert format (`experts.{e}.{gate_proj,up_proj,down_proj}`), so it can be served live through vLLM's FusedMoE LoRA support (PR #21229) instead of requiring a full merge-and-reload for every adapter switch. Developed and validated against Qwen3-30B-A3B (48 layers, 128 experts, r=16) — see the findings report below for the full story, including a genuine wrong turn in figuring out the tensor layout that's worth reading before trusting this on a different model shape. ## Contents - **`moe_lora_convert_vllm.py`** — the converter. `python3 moe_lora_convert_vllm.py ` - **`moe_lora_convert_validate.py`** — validates a converted adapter three ways (round-trip bit-exactness, per-expert delta match against Unsloth's real merge function, key coverage). Requires a CUDA GPU and `unsloth` installed. Edit `SRC`/`DST` at the top before running — the checked-in paths are from the run this was developed against, not a generic default. Both scripts assume the source adapter's expert-count/rank/layer-count via the constants at the top of each file (`E=128, r=16`, 48 layers) — adjust if converting a different model. --- # Findings: Serving Per-Author MoE LoRA Adapters via vLLM **Date:** 2026-09-08/09 **Status:** Resolved — Architecture A is fully reopened. ## Bottom line **vLLM can serve full-target (expert-covering) LoRA adapters on Qwen3-30B-A3B-MoE — correctly, live, with multiple adapters resident and routed properly, in the actual production shape this project uses, and at a cost that scales favorably to a 100-author roster.** This closes a question that had been treated as settled shut: "Architecture A" — a library of per-author LoRA adapters served live over one shared merged base — was believed unworkable because vLLM's MoE LoRA support couldn't load Unsloth's adapter format. It turned out to be a solvable data-layout conversion problem, not an architectural dead end. Joint training (the fallback this project had been using) remains a valid choice, but is no longer the *only* one. --- ## 1. The blocker, and why it wasn't what it first looked like vLLM added FusedMoE LoRA support (PR #21229) in October 2025. The first real attempt to use it here crashed immediately: ``` AssertionError: assert isinstance(lora_a, list) ``` Unsloth stores **one fused tensor per layer covering all 128 experts at once**. vLLM's MoE LoRA code expects a **Python list of 128 separate per-expert tensors**. This isn't a naming mismatch — it's a genuine data-layout conversion problem: someone has to slice the fused tensor into 128 per-expert pieces and re-key them to the naming vLLM's own error message already listed as supported (`experts.{e}.{gate_proj,up_proj,down_proj}`). ## 2. Finding the actual layout — and a wrong turn worth recording Writing that converter required knowing exactly how each expert's slice of the fused tensor is laid out. This took two passes, and the first one produced a wrong-but-plausible answer that's worth recording as a caution for future work like this. **First pass:** verified vanilla PEFT's `get_delta_weight()` method against a *live call to that same method* — internally consistent, and concluded a "mixed" layout (one tensor's expert-blocks contiguous, the other's strided). This felt like solid verification. It wasn't: it never checked against how real checkpoints are actually built. **The correction:** it turns out `save_pretrained_merged()` — the only merge path any real checkpoint in this project has ever gone through — bypasses vanilla PEFT entirely, via a separate hand-written implementation. Calling *that* real function directly, and cross-validating it against the independently-coded training-time forward pass, showed: | Comparison | Result | |---|---| | Real merge function vs. training-time forward path (independently coded) | agree to fp32 rounding noise | | Real merge function vs. vanilla PEFT's `get_delta_weight()` | disagree by a large, consistent margin | Two independently-written code paths agreeing with each other, and both disagreeing with the "verified" reference, is about as clear a signal as this kind of question ever produces. The actual layout: **both tensors' expert-blocks are contiguous**, not mixed. The lesson generalizes beyond this one adapter: when a real production code path exists and differs from a general library's own method, verifying a formula against the library's own method proves internal consistency, not correctness against what actually matters. ## 3. The converter `moe_lora_convert_vllm.py` implements the corrected layout: slices each layer's fused tensors into 128 per-expert pieces across all three projections, producing 18,432 (layer, expert, projection) pairs, re-keyed to vLLM's expected naming. Validated three ways before it was trusted with anything downstream: 1. **Round-trip** — reassemble the fused tensors from the converted pieces, compare to the originals. Bit-exact. 2. **Per-expert delta** — recompute each expert's delta from the converted tensors, compare to the real Unsloth merge function's output. Matches to fp32 noise. 3. **Key coverage** — exactly 18,432 pairs, no missing, no duplicates. The same converter was later run unchanged on three more adapters (two solo per-author CPT adapters, and a dual-author SFT LoRA) — one piece of code, four different real adapters, all correct. ## 4. The five gates | Gate | Question | Result | |---|---|---| | **1** | Does a converted adapter load at all? | **Pass** — no assertion error, engages vLLM's FusedMoE LoRA path | | **2** | Is it actually *applied* (live), not just loaded? | **Pass** — BPB matches its known reference to 2×10⁻⁴, 124× closer than to base's | | **3** | Do two resident adapters route correctly per request? | **Pass** — see below, this one took two attempts | | **4** | Does the real production shape (merged CPT base + SFT LoRA) work? | **Pass** — 4/4 generations correctly attributed by an already-validated classifier | | **5** | What does this cost, for sizing a 100-author roster? | **Done** — favorable on all three measurements | ### Gate 2 — liveness Scored the same held-out units through vLLM twice — once bare, once with the adapter attached — using the exact BPB methodology already established for the HF backend. | | vLLM | Known reference | agreement | |---|---|---|---| | Base (no adapter) | 0.727092 | 0.727071 | 2×10⁻⁵ | | Adapter attached | 0.701017 | 0.701228 | 2×10⁻⁴ | The base-case agreement validates the whole vLLM-side measurement approach; the adapter case is unambiguously closer to its own reference than to base's. ### Gate 3 — multi-adapter routing (a worthwhile detour) The first check (send the same prompt to two resident adapters, score the outputs with a style classifier) came back ambiguous — both continuations got attributed to the same author. Rather than re-running until it agreed with expectations, this was reported as what it was: most likely a single ~550-word sample continuing an already-strongly-styled prompt is a noisy, unfair test for a raw continuation adapter with no instruction to override the prompt's own register. A second, more decisive check reused Gate 2's own BPB approach instead — score each author's held-out set with their own adapter, then with the *other* author's adapter substituted: | Author | Own adapter | Reference | Wrong adapter substituted | |---|---|---|---| | Nuttall | 0.702174 | 0.702232 | 0.730985 | | Vox Day | 0.772137 | 0.772159 | 0.787210 | Own-adapter results match references almost exactly; misapplying the wrong adapter clearly degrades the fit past even the base model, every time. Decisive, and immune to the single-sample bias the style check had. ### Gate 4 — the shape that actually exists The original plan wanted per-author SFT LoRAs, which this project doesn't have — SFT here was always a second *merge* stage on top of the CPT merge, not a second LoRA. Redefined to the real shape: served base = the merged joint-CPT checkpoint, live adapter = the one dual-author scale-up SFT LoRA. Generated four held-out packages and scored them with the classifier this exact model was already validated against (0.949 Vox Day / 0.991 Nuttall separability on the HF backend) — not a freshly-fit instrument. | Packet | Author | Predicted | Margin | |---|---|---|---| | GRANDMASTER_Chapter-Sixteen_002 | Nuttall | Nuttall | −0.0673 | | GRANDMASTER_Chapter-Fourteen_001 | Nuttall | Nuttall | −0.1286 | | TOB_Lodi_004 | Vox Day | Vox Day | +0.1508 | | SOS_V2_UNTRAINED_Aulan_004 | Vox Day | Vox Day | +0.2120 | 4/4, all margins clearly signed. ### Gate 5 — what it costs **Throughput** (single-stream decode tok/s, one adapter active, one process per `max_loras` setting): | max_loras | tok/s | |---|---| | 0 | 26.13 | | 1 | 23.74 | | 2 | 23.95 | | 4 | 23.57 | Turning LoRA on costs ~9% once. Adding more *resident* adapters (not concurrently active) costs nothing further — the four numbers from 1 to 4 sit within a 1.6% band, i.e. noise. All at or above the llama.cpp baseline (22.7–23.8 tok/s), despite vLLM running clock-throttled here for GB10 stability against an unthrottled llama.cpp figure. **Residency memory** (vLLM's own logged KV-cache budget — measured, not estimated): | max_loras | KV cache available | Δ | |---|---|---| | 0 | 18.75 GiB | — | | 1 | 12.28 GiB | −6.47 GiB | | 2 | 8.93 GiB | −3.35 GiB | | 4 | 8.01 GiB | −0.46 GiB/slot | A large fixed cost to enable LoRA at all, then a **shrinking** marginal cost per additional resident adapter — the opposite of what would make 100 authors expensive. **Swap latency** (`max_cpu_loras`, adapters evicted from limited GPU slots but kept in CPU memory): the first pass at measuring this conflated two different things — some "swap" requests turned out to be a never-before-used adapter's first load, not a real swap, and those cost 20–27 seconds. The genuine swaps (an adapter used once, evicted, then re-requested) cost ~570ms — barely more than a warm hit (~523ms). **The real cost for capacity planning is the one-time disk load the first time any given author is ever served; everything after that is cheap.** ## 5. Infrastructure: a real, separate cost None of the above came for free on the hardware this was developed on (an NVIDIA DGX Spark / GB10, unified CPU+GPU memory). Getting Gate 2 to complete cost **four full-system hard reboots** — log-less, no OOM message, no kernel panic, just gone. This turned out to be a documented GB10/vLLM interaction, not a bug in this converter or its validation: `gpu_memory_utilization` above 0.8 is known to hang GB10 systems, and vLLM doesn't reserve host RAM headroom on unified-memory devices the way it would on a discrete-GPU box. The fix that held, applied to every run since: GPU clock locked to 2200MHz, `gpu_memory_utilization` capped at 0.65, `max_model_len` sized to the actual longest sequence in play (not a round-number default), and page cache dropped before a big load. Zero crashes across everything that followed — eleven more sequential vLLM engine loads, across five different workload shapes (BPB scoring, multi-adapter routing, long generation, short-context throughput, swap-latency cycling). One related, unresolved finding: vLLM's own CUDA-graph memory *estimator* is unreliable in these runs, overshooting actual usage by 400–1000% and, once, exhausting the entire memory budget before profiling could even begin. The workaround (`enforce_eager=True`, or a short `max_model_len`) is fine for correctness/liveness gates; a future throughput benchmark at longer context will need to budget around this directly rather than bypass it. If you're running this converter's output through vLLM on a discrete-GPU machine, none of the above should apply — it's specific to GB10's unified-memory architecture. ## 6. What's still open - **A throughput discrepancy, not reconciled.** The llama.cpp baseline used here (22.7–23.8 tok/s) came from an earlier report; a separate, more recent llama.cpp determinism-gate run measured ~44–46 tok/s for the same model. Worth resolving before quoting either number as authoritative for a roster-sizing decision. - **The memory/swap curves are only sampled at 0/1/2/4 adapters.** The shrinking marginal cost is promising but should be confirmed at 8–16 before sizing a real `--max-loras` value for 100 authors. - **Gate 4's classifier check used 4 generations.** Enough to confirm the shape works; not a substitute for a full multi-seed cross-eval if a rigorous quality comparison is ever needed through vLLM specifically. ## 7. Decision **Architecture A — a library of full-target per-author LoRA adapters served live over one shared merged base — is open.** All five gates pass, on both correctness and cost grounds. Joint training is no longer required to reach the 100-novel roster; it remains available as a design choice (e.g. for a genuinely shared cross-author capability), not as a workaround for a closed door.