Converts an Unsloth-trained, MoE-expert-targeting LoRA adapter into vLLM's expected per-expert format so it can be served live through vLLM's FusedMoE LoRA support, instead of requiring a full merge-and- reload per adapter switch. Includes: - moe_lora_convert_vllm.py -- the converter - moe_lora_convert_validate.py -- three-way validation (round-trip bit-exactness, per-expert delta match against Unsloth's real merge function, key coverage) - README.md -- usage plus the full findings report: the tensor-layout investigation (including a wrong turn worth recording), all five vLLM-serving validation gates, and the GB10 infrastructure lessons from developing this Developed and validated against Qwen3-30B-A3B (48 layers, 128 experts, r=16 LoRA). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015dwQp8JCqhM2MRFHYwMyxp
MoE LoRA Converter
Converts an Unsloth-trained, MoE-expert-targeting LoRA adapter (one fused tensor per layer
covering all experts, via PEFT's target_parameters mechanism) into vLLM's expected per-expert
format (experts.{e}.{gate_proj,up_proj,down_proj}), so it can be served live through vLLM's
FusedMoE LoRA support (PR #21229) instead of requiring a full merge-and-reload for every adapter
switch.
Developed and validated against Qwen3-30B-A3B (48 layers, 128 experts, r=16) — see the findings report below for the full story, including a genuine wrong turn in figuring out the tensor layout that's worth reading before trusting this on a different model shape.
Contents
moe_lora_convert_vllm.py— the converter.python3 moe_lora_convert_vllm.py <src_adapter_dir> <dst_adapter_dir>moe_lora_convert_validate.py— validates a converted adapter three ways (round-trip bit-exactness, per-expert delta match against Unsloth's real merge function, key coverage). Requires a CUDA GPU andunslothinstalled. EditSRC/DSTat the top before running — the checked-in paths are from the run this was developed against, not a generic default.
Both scripts assume the source adapter's expert-count/rank/layer-count via the constants at
the top of each file (E=128, r=16, 48 layers) — adjust if converting a different model.
Findings: Serving Per-Author MoE LoRA Adapters via vLLM
Date: 2026-09-08/09 Status: Resolved — Architecture A is fully reopened.
Bottom line
vLLM can serve full-target (expert-covering) LoRA adapters on Qwen3-30B-A3B-MoE — correctly, live, with multiple adapters resident and routed properly, in the actual production shape this project uses, and at a cost that scales favorably to a 100-author roster.
This closes a question that had been treated as settled shut: "Architecture A" — a library of per-author LoRA adapters served live over one shared merged base — was believed unworkable because vLLM's MoE LoRA support couldn't load Unsloth's adapter format. It turned out to be a solvable data-layout conversion problem, not an architectural dead end. Joint training (the fallback this project had been using) remains a valid choice, but is no longer the only one.
1. The blocker, and why it wasn't what it first looked like
vLLM added FusedMoE LoRA support (PR #21229) in October 2025. The first real attempt to use it here crashed immediately:
AssertionError: assert isinstance(lora_a, list)
Unsloth stores one fused tensor per layer covering all 128 experts at once. vLLM's MoE LoRA
code expects a Python list of 128 separate per-expert tensors. This isn't a naming mismatch
— it's a genuine data-layout conversion problem: someone has to slice the fused tensor into 128
per-expert pieces and re-key them to the naming vLLM's own error message already listed as
supported (experts.{e}.{gate_proj,up_proj,down_proj}).
2. Finding the actual layout — and a wrong turn worth recording
Writing that converter required knowing exactly how each expert's slice of the fused tensor is laid out. This took two passes, and the first one produced a wrong-but-plausible answer that's worth recording as a caution for future work like this.
First pass: verified vanilla PEFT's get_delta_weight() method against a live call to that
same method — internally consistent, and concluded a "mixed" layout (one tensor's expert-blocks
contiguous, the other's strided). This felt like solid verification. It wasn't: it never checked
against how real checkpoints are actually built.
The correction: it turns out save_pretrained_merged() — the only merge path any real
checkpoint in this project has ever gone through — bypasses vanilla PEFT entirely, via a
separate hand-written implementation. Calling that real function directly, and
cross-validating it against the independently-coded training-time forward pass, showed:
| Comparison | Result |
|---|---|
| Real merge function vs. training-time forward path (independently coded) | agree to fp32 rounding noise |
Real merge function vs. vanilla PEFT's get_delta_weight() |
disagree by a large, consistent margin |
Two independently-written code paths agreeing with each other, and both disagreeing with the "verified" reference, is about as clear a signal as this kind of question ever produces. The actual layout: both tensors' expert-blocks are contiguous, not mixed. The lesson generalizes beyond this one adapter: when a real production code path exists and differs from a general library's own method, verifying a formula against the library's own method proves internal consistency, not correctness against what actually matters.
3. The converter
moe_lora_convert_vllm.py implements the corrected layout: slices each layer's fused tensors
into 128 per-expert pieces across all three projections, producing 18,432 (layer, expert,
projection) pairs, re-keyed to vLLM's expected naming.
Validated three ways before it was trusted with anything downstream:
- Round-trip — reassemble the fused tensors from the converted pieces, compare to the originals. Bit-exact.
- Per-expert delta — recompute each expert's delta from the converted tensors, compare to the real Unsloth merge function's output. Matches to fp32 noise.
- Key coverage — exactly 18,432 pairs, no missing, no duplicates.
The same converter was later run unchanged on three more adapters (two solo per-author CPT adapters, and a dual-author SFT LoRA) — one piece of code, four different real adapters, all correct.
4. The five gates
| Gate | Question | Result |
|---|---|---|
| 1 | Does a converted adapter load at all? | Pass — no assertion error, engages vLLM's FusedMoE LoRA path |
| 2 | Is it actually applied (live), not just loaded? | Pass — BPB matches its known reference to 2×10⁻⁴, 124× closer than to base's |
| 3 | Do two resident adapters route correctly per request? | Pass — see below, this one took two attempts |
| 4 | Does the real production shape (merged CPT base + SFT LoRA) work? | Pass — 4/4 generations correctly attributed by an already-validated classifier |
| 5 | What does this cost, for sizing a 100-author roster? | Done — favorable on all three measurements |
Gate 2 — liveness
Scored the same held-out units through vLLM twice — once bare, once with the adapter attached — using the exact BPB methodology already established for the HF backend.
| vLLM | Known reference | agreement | |
|---|---|---|---|
| Base (no adapter) | 0.727092 | 0.727071 | 2×10⁻⁵ |
| Adapter attached | 0.701017 | 0.701228 | 2×10⁻⁴ |
The base-case agreement validates the whole vLLM-side measurement approach; the adapter case is unambiguously closer to its own reference than to base's.
Gate 3 — multi-adapter routing (a worthwhile detour)
The first check (send the same prompt to two resident adapters, score the outputs with a style classifier) came back ambiguous — both continuations got attributed to the same author. Rather than re-running until it agreed with expectations, this was reported as what it was: most likely a single ~550-word sample continuing an already-strongly-styled prompt is a noisy, unfair test for a raw continuation adapter with no instruction to override the prompt's own register.
A second, more decisive check reused Gate 2's own BPB approach instead — score each author's held-out set with their own adapter, then with the other author's adapter substituted:
| Author | Own adapter | Reference | Wrong adapter substituted |
|---|---|---|---|
| Nuttall | 0.702174 | 0.702232 | 0.730985 |
| Vox Day | 0.772137 | 0.772159 | 0.787210 |
Own-adapter results match references almost exactly; misapplying the wrong adapter clearly degrades the fit past even the base model, every time. Decisive, and immune to the single-sample bias the style check had.
Gate 4 — the shape that actually exists
The original plan wanted per-author SFT LoRAs, which this project doesn't have — SFT here was always a second merge stage on top of the CPT merge, not a second LoRA. Redefined to the real shape: served base = the merged joint-CPT checkpoint, live adapter = the one dual-author scale-up SFT LoRA. Generated four held-out packages and scored them with the classifier this exact model was already validated against (0.949 Vox Day / 0.991 Nuttall separability on the HF backend) — not a freshly-fit instrument.
| Packet | Author | Predicted | Margin |
|---|---|---|---|
| GRANDMASTER_Chapter-Sixteen_002 | Nuttall | Nuttall | −0.0673 |
| GRANDMASTER_Chapter-Fourteen_001 | Nuttall | Nuttall | −0.1286 |
| TOB_Lodi_004 | Vox Day | Vox Day | +0.1508 |
| SOS_V2_UNTRAINED_Aulan_004 | Vox Day | Vox Day | +0.2120 |
4/4, all margins clearly signed.
Gate 5 — what it costs
Throughput (single-stream decode tok/s, one adapter active, one process per max_loras
setting):
| max_loras | tok/s |
|---|---|
| 0 | 26.13 |
| 1 | 23.74 |
| 2 | 23.95 |
| 4 | 23.57 |
Turning LoRA on costs ~9% once. Adding more resident adapters (not concurrently active) costs nothing further — the four numbers from 1 to 4 sit within a 1.6% band, i.e. noise. All at or above the llama.cpp baseline (22.7–23.8 tok/s), despite vLLM running clock-throttled here for GB10 stability against an unthrottled llama.cpp figure.
Residency memory (vLLM's own logged KV-cache budget — measured, not estimated):
| max_loras | KV cache available | Δ |
|---|---|---|
| 0 | 18.75 GiB | — |
| 1 | 12.28 GiB | −6.47 GiB |
| 2 | 8.93 GiB | −3.35 GiB |
| 4 | 8.01 GiB | −0.46 GiB/slot |
A large fixed cost to enable LoRA at all, then a shrinking marginal cost per additional resident adapter — the opposite of what would make 100 authors expensive.
Swap latency (max_cpu_loras, adapters evicted from limited GPU slots but kept in CPU
memory): the first pass at measuring this conflated two different things — some "swap" requests
turned out to be a never-before-used adapter's first load, not a real swap, and those cost
20–27 seconds. The genuine swaps (an adapter used once, evicted, then re-requested) cost
~570ms — barely more than a warm hit (~523ms). The real cost for capacity planning is the
one-time disk load the first time any given author is ever served; everything after that is
cheap.
5. Infrastructure: a real, separate cost
None of the above came for free on the hardware this was developed on (an NVIDIA DGX Spark /
GB10, unified CPU+GPU memory). Getting Gate 2 to complete cost four full-system hard reboots
— log-less, no OOM message, no kernel panic, just gone. This turned out to be a documented
GB10/vLLM interaction, not a bug in this converter or its validation: gpu_memory_utilization
above 0.8 is known to hang GB10 systems, and vLLM doesn't reserve host RAM headroom on
unified-memory devices the way it would on a discrete-GPU box.
The fix that held, applied to every run since: GPU clock locked to 2200MHz, gpu_memory_utilization
capped at 0.65, max_model_len sized to the actual longest sequence in play (not a round-number
default), and page cache dropped before a big load. Zero crashes across everything that followed
— eleven more sequential vLLM engine loads, across five different workload shapes (BPB scoring,
multi-adapter routing, long generation, short-context throughput, swap-latency cycling).
One related, unresolved finding: vLLM's own CUDA-graph memory estimator is unreliable in these
runs, overshooting actual usage by 400–1000% and, once, exhausting the entire memory budget
before profiling could even begin. The workaround (enforce_eager=True, or a short
max_model_len) is fine for correctness/liveness gates; a future throughput benchmark at longer
context will need to budget around this directly rather than bypass it.
If you're running this converter's output through vLLM on a discrete-GPU machine, none of the above should apply — it's specific to GB10's unified-memory architecture.
6. What's still open
- A throughput discrepancy, not reconciled. The llama.cpp baseline used here (22.7–23.8 tok/s) came from an earlier report; a separate, more recent llama.cpp determinism-gate run measured ~44–46 tok/s for the same model. Worth resolving before quoting either number as authoritative for a roster-sizing decision.
- The memory/swap curves are only sampled at 0/1/2/4 adapters. The shrinking marginal cost
is promising but should be confirmed at 8–16 before sizing a real
--max-lorasvalue for 100 authors. - Gate 4's classifier check used 4 generations. Enough to confirm the shape works; not a substitute for a full multi-seed cross-eval if a rigorous quality comparison is ever needed through vLLM specifically.
7. Decision
Architecture A — a library of full-target per-author LoRA adapters served live over one shared merged base — is open. All five gates pass, on both correctness and cost grounds. Joint training is no longer required to reach the 100-novel roster; it remains available as a design choice (e.g. for a genuinely shared cross-author capability), not as a workaround for a closed door.