Add MoE LoRA converter (Unsloth fused format -> vLLM per-expert format)

Converts an Unsloth-trained, MoE-expert-targeting LoRA adapter into
vLLM's expected per-expert format so it can be served live through
vLLM's FusedMoE LoRA support, instead of requiring a full merge-and-
reload per adapter switch.

Includes:
- moe_lora_convert_vllm.py -- the converter
- moe_lora_convert_validate.py -- three-way validation (round-trip
  bit-exactness, per-expert delta match against Unsloth's real merge
  function, key coverage)
- README.md -- usage plus the full findings report: the tensor-layout
  investigation (including a wrong turn worth recording), all five
  vLLM-serving validation gates, and the GB10 infrastructure lessons
  from developing this

Developed and validated against Qwen3-30B-A3B (48 layers, 128 experts,
r=16 LoRA).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015dwQp8JCqhM2MRFHYwMyxp
This commit is contained in:
2026-09-09 09:19:28 -05:00
commit 7a66942242
3 changed files with 544 additions and 0 deletions

247
README.md Normal file
View File

@@ -0,0 +1,247 @@
# MoE LoRA Converter
Converts an Unsloth-trained, MoE-expert-targeting LoRA adapter (one fused tensor per layer
covering all experts, via PEFT's `target_parameters` mechanism) into vLLM's expected per-expert
format (`experts.{e}.{gate_proj,up_proj,down_proj}`), so it can be served live through vLLM's
FusedMoE LoRA support (PR #21229) instead of requiring a full merge-and-reload for every adapter
switch.
Developed and validated against Qwen3-30B-A3B (48 layers, 128 experts, r=16) — see the findings
report below for the full story, including a genuine wrong turn in figuring out the tensor
layout that's worth reading before trusting this on a different model shape.
## Contents
- **`moe_lora_convert_vllm.py`** — the converter. `python3 moe_lora_convert_vllm.py <src_adapter_dir> <dst_adapter_dir>`
- **`moe_lora_convert_validate.py`** — validates a converted adapter three ways (round-trip
bit-exactness, per-expert delta match against Unsloth's real merge function, key coverage).
Requires a CUDA GPU and `unsloth` installed. Edit `SRC`/`DST` at the top before running — the
checked-in paths are from the run this was developed against, not a generic default.
Both scripts assume the source adapter's expert-count/rank/layer-count via the constants at
the top of each file (`E=128, r=16`, 48 layers) — adjust if converting a different model.
---
# Findings: Serving Per-Author MoE LoRA Adapters via vLLM
**Date:** 2026-09-08/09
**Status:** Resolved — Architecture A is fully reopened.
## Bottom line
**vLLM can serve full-target (expert-covering) LoRA adapters on Qwen3-30B-A3B-MoE — correctly,
live, with multiple adapters resident and routed properly, in the actual production shape this
project uses, and at a cost that scales favorably to a 100-author roster.**
This closes a question that had been treated as settled shut: "Architecture A" — a library of
per-author LoRA adapters served live over one shared merged base — was believed unworkable
because vLLM's MoE LoRA support couldn't load Unsloth's adapter format. It turned out to be a
solvable data-layout conversion problem, not an architectural dead end. Joint training (the
fallback this project had been using) remains a valid choice, but is no longer the *only* one.
---
## 1. The blocker, and why it wasn't what it first looked like
vLLM added FusedMoE LoRA support (PR #21229) in October 2025. The first real attempt to use it
here crashed immediately:
```
AssertionError: assert isinstance(lora_a, list)
```
Unsloth stores **one fused tensor per layer covering all 128 experts at once**. vLLM's MoE LoRA
code expects a **Python list of 128 separate per-expert tensors**. This isn't a naming mismatch
— it's a genuine data-layout conversion problem: someone has to slice the fused tensor into 128
per-expert pieces and re-key them to the naming vLLM's own error message already listed as
supported (`experts.{e}.{gate_proj,up_proj,down_proj}`).
## 2. Finding the actual layout — and a wrong turn worth recording
Writing that converter required knowing exactly how each expert's slice of the fused tensor is
laid out. This took two passes, and the first one produced a wrong-but-plausible answer that's
worth recording as a caution for future work like this.
**First pass:** verified vanilla PEFT's `get_delta_weight()` method against a *live call to that
same method* — internally consistent, and concluded a "mixed" layout (one tensor's expert-blocks
contiguous, the other's strided). This felt like solid verification. It wasn't: it never checked
against how real checkpoints are actually built.
**The correction:** it turns out `save_pretrained_merged()` — the only merge path any real
checkpoint in this project has ever gone through — bypasses vanilla PEFT entirely, via a
separate hand-written implementation. Calling *that* real function directly, and
cross-validating it against the independently-coded training-time forward pass, showed:
| Comparison | Result |
|---|---|
| Real merge function vs. training-time forward path (independently coded) | agree to fp32 rounding noise |
| Real merge function vs. vanilla PEFT's `get_delta_weight()` | disagree by a large, consistent margin |
Two independently-written code paths agreeing with each other, and both disagreeing with the
"verified" reference, is about as clear a signal as this kind of question ever produces. The
actual layout: **both tensors' expert-blocks are contiguous**, not mixed. The lesson generalizes
beyond this one adapter: when a real production code path exists and differs from a general
library's own method, verifying a formula against the library's own method proves internal
consistency, not correctness against what actually matters.
## 3. The converter
`moe_lora_convert_vllm.py` implements the corrected layout: slices each layer's fused tensors
into 128 per-expert pieces across all three projections, producing 18,432 (layer, expert,
projection) pairs, re-keyed to vLLM's expected naming.
Validated three ways before it was trusted with anything downstream:
1. **Round-trip** — reassemble the fused tensors from the converted pieces, compare to the
originals. Bit-exact.
2. **Per-expert delta** — recompute each expert's delta from the converted tensors, compare to
the real Unsloth merge function's output. Matches to fp32 noise.
3. **Key coverage** — exactly 18,432 pairs, no missing, no duplicates.
The same converter was later run unchanged on three more adapters (two solo per-author CPT
adapters, and a dual-author SFT LoRA) — one piece of code, four different real adapters, all
correct.
## 4. The five gates
| Gate | Question | Result |
|---|---|---|
| **1** | Does a converted adapter load at all? | **Pass** — no assertion error, engages vLLM's FusedMoE LoRA path |
| **2** | Is it actually *applied* (live), not just loaded? | **Pass** — BPB matches its known reference to 2×10⁻⁴, 124× closer than to base's |
| **3** | Do two resident adapters route correctly per request? | **Pass** — see below, this one took two attempts |
| **4** | Does the real production shape (merged CPT base + SFT LoRA) work? | **Pass** — 4/4 generations correctly attributed by an already-validated classifier |
| **5** | What does this cost, for sizing a 100-author roster? | **Done** — favorable on all three measurements |
### Gate 2 — liveness
Scored the same held-out units through vLLM twice — once bare, once with the adapter attached —
using the exact BPB methodology already established for the HF backend.
| | vLLM | Known reference | agreement |
|---|---|---|---|
| Base (no adapter) | 0.727092 | 0.727071 | 2×10⁻⁵ |
| Adapter attached | 0.701017 | 0.701228 | 2×10⁻⁴ |
The base-case agreement validates the whole vLLM-side measurement approach; the adapter case is
unambiguously closer to its own reference than to base's.
### Gate 3 — multi-adapter routing (a worthwhile detour)
The first check (send the same prompt to two resident adapters, score the outputs with a style
classifier) came back ambiguous — both continuations got attributed to the same author. Rather
than re-running until it agreed with expectations, this was reported as what it was: most likely
a single ~550-word sample continuing an already-strongly-styled prompt is a noisy, unfair test
for a raw continuation adapter with no instruction to override the prompt's own register.
A second, more decisive check reused Gate 2's own BPB approach instead — score each author's
held-out set with their own adapter, then with the *other* author's adapter substituted:
| Author | Own adapter | Reference | Wrong adapter substituted |
|---|---|---|---|
| Nuttall | 0.702174 | 0.702232 | 0.730985 |
| Vox Day | 0.772137 | 0.772159 | 0.787210 |
Own-adapter results match references almost exactly; misapplying the wrong adapter clearly
degrades the fit past even the base model, every time. Decisive, and immune to the single-sample
bias the style check had.
### Gate 4 — the shape that actually exists
The original plan wanted per-author SFT LoRAs, which this project doesn't have — SFT here was
always a second *merge* stage on top of the CPT merge, not a second LoRA. Redefined to the real
shape: served base = the merged joint-CPT checkpoint, live adapter = the one dual-author scale-up
SFT LoRA. Generated four held-out packages and scored them with the classifier this exact model
was already validated against (0.949 Vox Day / 0.991 Nuttall separability on the HF backend) —
not a freshly-fit instrument.
| Packet | Author | Predicted | Margin |
|---|---|---|---|
| GRANDMASTER_Chapter-Sixteen_002 | Nuttall | Nuttall | 0.0673 |
| GRANDMASTER_Chapter-Fourteen_001 | Nuttall | Nuttall | 0.1286 |
| TOB_Lodi_004 | Vox Day | Vox Day | +0.1508 |
| SOS_V2_UNTRAINED_Aulan_004 | Vox Day | Vox Day | +0.2120 |
4/4, all margins clearly signed.
### Gate 5 — what it costs
**Throughput** (single-stream decode tok/s, one adapter active, one process per `max_loras`
setting):
| max_loras | tok/s |
|---|---|
| 0 | 26.13 |
| 1 | 23.74 |
| 2 | 23.95 |
| 4 | 23.57 |
Turning LoRA on costs ~9% once. Adding more *resident* adapters (not concurrently active) costs
nothing further — the four numbers from 1 to 4 sit within a 1.6% band, i.e. noise. All at or
above the llama.cpp baseline (22.723.8 tok/s), despite vLLM running clock-throttled here for
GB10 stability against an unthrottled llama.cpp figure.
**Residency memory** (vLLM's own logged KV-cache budget — measured, not estimated):
| max_loras | KV cache available | Δ |
|---|---|---|
| 0 | 18.75 GiB | — |
| 1 | 12.28 GiB | 6.47 GiB |
| 2 | 8.93 GiB | 3.35 GiB |
| 4 | 8.01 GiB | 0.46 GiB/slot |
A large fixed cost to enable LoRA at all, then a **shrinking** marginal cost per additional
resident adapter — the opposite of what would make 100 authors expensive.
**Swap latency** (`max_cpu_loras`, adapters evicted from limited GPU slots but kept in CPU
memory): the first pass at measuring this conflated two different things — some "swap" requests
turned out to be a never-before-used adapter's first load, not a real swap, and those cost
2027 seconds. The genuine swaps (an adapter used once, evicted, then re-requested) cost
~570ms — barely more than a warm hit (~523ms). **The real cost for capacity planning is the
one-time disk load the first time any given author is ever served; everything after that is
cheap.**
## 5. Infrastructure: a real, separate cost
None of the above came for free on the hardware this was developed on (an NVIDIA DGX Spark /
GB10, unified CPU+GPU memory). Getting Gate 2 to complete cost **four full-system hard reboots**
— log-less, no OOM message, no kernel panic, just gone. This turned out to be a documented
GB10/vLLM interaction, not a bug in this converter or its validation: `gpu_memory_utilization`
above 0.8 is known to hang GB10 systems, and vLLM doesn't reserve host RAM headroom on
unified-memory devices the way it would on a discrete-GPU box.
The fix that held, applied to every run since: GPU clock locked to 2200MHz, `gpu_memory_utilization`
capped at 0.65, `max_model_len` sized to the actual longest sequence in play (not a round-number
default), and page cache dropped before a big load. Zero crashes across everything that followed
— eleven more sequential vLLM engine loads, across five different workload shapes (BPB scoring,
multi-adapter routing, long generation, short-context throughput, swap-latency cycling).
One related, unresolved finding: vLLM's own CUDA-graph memory *estimator* is unreliable in these
runs, overshooting actual usage by 4001000% and, once, exhausting the entire memory budget
before profiling could even begin. The workaround (`enforce_eager=True`, or a short
`max_model_len`) is fine for correctness/liveness gates; a future throughput benchmark at longer
context will need to budget around this directly rather than bypass it.
If you're running this converter's output through vLLM on a discrete-GPU machine, none of the
above should apply — it's specific to GB10's unified-memory architecture.
## 6. What's still open
- **A throughput discrepancy, not reconciled.** The llama.cpp baseline used here (22.723.8
tok/s) came from an earlier report; a separate, more recent llama.cpp determinism-gate run
measured ~4446 tok/s for the same model. Worth resolving before quoting either number as
authoritative for a roster-sizing decision.
- **The memory/swap curves are only sampled at 0/1/2/4 adapters.** The shrinking marginal cost
is promising but should be confirmed at 816 before sizing a real `--max-loras` value for 100
authors.
- **Gate 4's classifier check used 4 generations.** Enough to confirm the shape works; not a
substitute for a full multi-seed cross-eval if a rigorous quality comparison is ever needed
through vLLM specifically.
## 7. Decision
**Architecture A — a library of full-target per-author LoRA adapters served live over one shared
merged base — is open.** All five gates pass, on both correctness and cost grounds. Joint
training is no longer required to reach the 100-novel roster; it remains available as a design
choice (e.g. for a genuinely shared cross-author capability), not as a workaround for a closed
door.

View File

@@ -0,0 +1,164 @@
"""Validates moe_lora_convert_vllm.py's output.
1. Round-trip: reassemble the fused tensors from the converted per-expert set
and compare against the originals. Must be bit-exact.
2. Per-expert delta: for a handful of experts, compute scaling * B_e @ A_e
from the CONVERTED tensors and compare against Unsloth's REAL production
merge function (unsloth_zoo/saving_utils.py -- the corrected ground truth,
NOT vanilla PEFT's get_delta_weight -- see the README's "wrong turn" note).
3. Key coverage: assert exactly 18,432 (layer, expert, projection) pairs,
every one present once, no orphans.
NOTE: SRC/DST below are the paths from the run this was developed against --
point them at your own source adapter and its moe_lora_convert_vllm.py output
before running. Requires a CUDA GPU (the real merge functions run on-device)
and `unsloth` installed (imported for its side-effecting patches before
`unsloth_zoo.saving_utils` is usable).
"""
import json
import re
import sys
import torch
import unsloth # noqa: F401
from safetensors import safe_open
from unsloth_zoo.saving_utils import LoraStats, _merge_moe_gate_expert, _merge_moe_up_expert, _merge_moe_down_proj_expert
SRC = "outputs/ladder_cpt/voxday_nuttall_joint_rung90000_90000/checkpoint-290"
DST = "outputs/ladder_cpt/voxday_nuttall_joint_rung90000_90000/checkpoint-290_vllm_converted"
E, r = 128, 16
with open(f"{DST}/adapter_config.json") as fh:
cfg = json.load(fh)
scaling = cfg["lora_alpha"] / cfg["r"]
print(f"scaling={scaling}")
ok = True
# ---------- 1. Round-trip ----------
print("\n=== 1. Round-trip (reassemble converted -> compare to original, bit-exact) ===")
with safe_open(f"{SRC}/adapter_model.safetensors", framework="pt") as fsrc, \
safe_open(f"{DST}/adapter_model.safetensors", framework="pt") as fdst:
src_keys = list(fsrc.keys())
expert_prefix_re = re.compile(r"^base_model\.model\.model\.layers\.(\d+)\.mlp\.experts\.(?:base_layer\.)?lora_A\.weight$")
layers = sorted({int(m.group(1)) for k in src_keys if (m := expert_prefix_re.match(k))})
layers_to_check = [0, 1, 23, 46, 47]
for L in layers_to_check:
prefix = f"base_model.model.model.layers.{L}.mlp.experts."
a_keys = [k for k in src_keys if k.startswith(prefix) and k.endswith("lora_A.weight")]
pairs = [(k, fsrc.get_tensor(k), fsrc.get_tensor(k.replace("lora_A.weight", "lora_B.weight"))) for k in a_keys]
pairs.sort(key=lambda p: p[1].shape[-1], reverse=True)
(_, gate_up_A, gate_up_B), (_, down_A, down_B) = pairs
hidden = gate_up_A.shape[1]
inter = gate_up_B.shape[0] // 2
# reassemble gate_up_A, gate_up_B, down_A, down_B from converted per-expert tensors
recon_gate_up_A = torch.zeros_like(gate_up_A)
recon_gate_up_B = torch.zeros_like(gate_up_B)
recon_down_A = torch.zeros_like(down_A)
recon_down_B = torch.zeros_like(down_B)
for e in range(E):
lo, hi = e * r, (e + 1) * r
base = f"base_model.model.model.layers.{L}.mlp.experts.{e}"
gA = fdst.get_tensor(f"{base}.gate_proj.lora_A.weight")
gB = fdst.get_tensor(f"{base}.gate_proj.lora_B.weight")
uA = fdst.get_tensor(f"{base}.up_proj.lora_A.weight")
uB = fdst.get_tensor(f"{base}.up_proj.lora_B.weight")
dA = fdst.get_tensor(f"{base}.down_proj.lora_A.weight")
dB = fdst.get_tensor(f"{base}.down_proj.lora_B.weight")
assert torch.equal(gA, uA), f"layer {L} expert {e}: gate_proj/up_proj lora_A should be identical (shared)"
recon_gate_up_A[lo:hi, :] = gA
recon_gate_up_B[:inter, lo:hi] = gB
recon_gate_up_B[inter:, lo:hi] = uB
recon_down_A[lo:hi, :] = dA
recon_down_B[:, lo:hi] = dB
g_a_ok = torch.equal(recon_gate_up_A, gate_up_A)
g_b_ok = torch.equal(recon_gate_up_B, gate_up_B)
d_a_ok = torch.equal(recon_down_A, down_A)
d_b_ok = torch.equal(recon_down_B, down_B)
all_ok = g_a_ok and g_b_ok and d_a_ok and d_b_ok
ok = ok and all_ok
print(f" layer {L}: gate_up_A={'OK' if g_a_ok else 'MISMATCH'} gate_up_B={'OK' if g_b_ok else 'MISMATCH'} "
f"down_A={'OK' if d_a_ok else 'MISMATCH'} down_B={'OK' if d_b_ok else 'MISMATCH'}")
# ---------- 2. Per-expert delta match against REAL Unsloth merge fn ----------
print("\n=== 2. Per-expert delta vs Unsloth's REAL production merge fn (saving_utils.py) ===")
with safe_open(f"{SRC}/adapter_model.safetensors", framework="pt") as fsrc, \
safe_open(f"{DST}/adapter_model.safetensors", framework="pt") as fdst:
for L in [0, 23, 47]:
prefix = f"base_model.model.model.layers.{L}.mlp.experts."
a_keys = [k for k in fsrc.keys() if k.startswith(prefix) and k.endswith("lora_A.weight")]
pairs = [(k, fsrc.get_tensor(k), fsrc.get_tensor(k.replace("lora_A.weight", "lora_B.weight"))) for k in a_keys]
pairs.sort(key=lambda p: p[1].shape[-1], reverse=True)
(_, gate_up_A, gate_up_B), (_, down_A, down_B) = pairs
hidden = gate_up_A.shape[1]
inter = gate_up_B.shape[0] // 2
stats_gate_up = LoraStats(module=None, lora_A=gate_up_A.cuda(), lora_B=gate_up_B.cuda(), alpha=scaling)
stats_down = LoraStats(module=None, lora_A=down_A.cuda(), lora_B=down_B.cuda(), alpha=scaling)
for e in [0, 1, 63, 127]:
base = f"base_model.model.model.layers.{L}.mlp.experts.{e}"
gA = fdst.get_tensor(f"{base}.gate_proj.lora_A.weight")
gB = fdst.get_tensor(f"{base}.gate_proj.lora_B.weight")
uB = fdst.get_tensor(f"{base}.up_proj.lora_B.weight")
dA = fdst.get_tensor(f"{base}.down_proj.lora_A.weight")
dB = fdst.get_tensor(f"{base}.down_proj.lora_B.weight")
from_converted_gate = scaling * (gB @ gA)
from_converted_up = scaling * (uB @ gA)
from_converted_down = scaling * (dB @ dA)
zero_gate = torch.zeros(inter, hidden).cuda()
zero_up = torch.zeros(inter, hidden).cuda()
zero_down = torch.zeros(hidden, inter).cuda()
real_gate = (_merge_moe_gate_expert(zero_gate, stats_gate_up, e, E, torch.float32) - zero_gate).cpu()
real_up = (_merge_moe_up_expert(zero_up, stats_gate_up, e, E, torch.float32) - zero_up).cpu()
real_down = (_merge_moe_down_proj_expert(zero_down, stats_down, e, E, torch.float32) - zero_down).cpu()
def rel_or_abs_err(conv, real, tag):
# Some experts are never routed during training and stay at LoRA's
# standard zero-init for lora_B -> delta is exactly zero for both
# sides. Relative error is 0/0 there; check absolute error instead.
real_norm = torch.linalg.norm(real)
if real_norm < 1e-12:
abs_err = torch.linalg.norm(conv).item()
return abs_err, abs_err < 1e-9, f"{tag}: untrained expert (both~0), abs_err={abs_err:.1e}"
err = (torch.linalg.norm(conv - real) / real_norm).item()
return err, err < 1e-3, f"{tag} rel_err={err:.2e}"
_, gate_pass, gate_msg = rel_or_abs_err(from_converted_gate, real_gate, "gate")
_, up_pass, up_msg = rel_or_abs_err(from_converted_up, real_up, "up")
_, down_pass, down_msg = rel_or_abs_err(from_converted_down, real_down, "down")
pass_ = gate_pass and up_pass and down_pass
ok = ok and pass_
print(f" layer {L} expert {e:3d}: {gate_msg} {up_msg} {down_msg} {'PASS' if pass_ else 'FAIL'}")
# ---------- 3. Key coverage ----------
print("\n=== 3. Key coverage ===")
with safe_open(f"{DST}/adapter_model.safetensors", framework="pt") as fdst:
dst_keys = list(fdst.keys())
expert_key_re = re.compile(r"^base_model\.model\.model\.layers\.(\d+)\.mlp\.experts\.(\d+)\.(gate_proj|up_proj|down_proj)\.lora_[AB]\.weight$")
found = set()
for k in dst_keys:
m = expert_key_re.match(k)
if m:
L, e, proj = int(m.group(1)), int(m.group(2)), m.group(3)
found.add((L, e, proj))
expected = {(L, e, proj) for L in range(48) for e in range(E) for proj in ("gate_proj", "up_proj", "down_proj")}
missing = expected - found
extra = found - expected
n_ok = len(found) == 18432 == len(expected) and not missing and not extra
ok = ok and n_ok
print(f" found {len(found)} (layer,expert,proj) triples (expect 18432): {'PASS' if n_ok else 'FAIL'}")
if missing:
print(f" MISSING (first 5): {list(missing)[:5]}")
if extra:
print(f" EXTRA (first 5): {list(extra)[:5]}")
print(f"\nOVERALL: {'ALL CHECKS PASSED' if ok else 'SOME CHECKS FAILED'}")
print("VALIDATE_DONE")
sys.exit(0 if ok else 1)

133
moe_lora_convert_vllm.py Normal file
View File

@@ -0,0 +1,133 @@
"""Convert an Unsloth fused-per-layer MoE LoRA adapter (target_parameters:
mlp.experts.gate_up_proj / mlp.experts.down_proj -- one tensor per layer
covering all 128 experts) into vLLM's expected per-expert format
(experts.{e}.{gate_proj,up_proj,down_proj}), per MOE_LORA_LAYOUT_PLAN.md §3.
Layout used (corrected 2026-09-09, see MOE_LORA_LAYOUT_REPORT.md §4):
lora_A [E*r, in]: expert e -> contiguous rows [e*r:(e+1)*r]
lora_B [out, E*r]: expert e -> contiguous columns [e*r:(e+1)*r] <- NOT strided
gate_up_proj's 2*inter-wide output is concatenated: gate=[0:inter], up=[inter:2*inter]
orientation [E, out, in], no transpose needed relative to the raw parameter
Usage:
python3 moe_lora_convert_vllm.py <src_adapter_dir> <dst_adapter_dir>
"""
import json
import re
import sys
from pathlib import Path
from safetensors import safe_open
from safetensors.torch import save_file
EXPERT_PREFIX_RE = re.compile(r"^(base_model\.model\.model\.layers\.(\d+)\.mlp\.experts\.)(?:base_layer\.)?lora_A\.weight$")
def find_layer_expert_keys(all_keys, layer_prefix):
"""Return {A_key: B_key} for the two expert-LoRA (gate_up, down) pairs in one
layer. Which key is which target_parameter is NOT resolved here -- PEFT's
nesting order between the two target_parameters isn't guaranteed, so the
caller disambiguates gate_up vs down by tensor shape instead."""
a_keys = [k for k in all_keys if k.startswith(layer_prefix) and k.endswith("lora_A.weight")]
if len(a_keys) != 2:
raise ValueError(f"expected exactly 2 expert lora_A keys under {layer_prefix}, found {len(a_keys)}: {a_keys}")
return {a_key: a_key.replace("lora_A.weight", "lora_B.weight") for a_key in a_keys}
def convert(src_dir: str, dst_dir: str, E: int = 128, r: int = 16):
src_dir = Path(src_dir)
dst_dir = Path(dst_dir)
dst_dir.mkdir(parents=True, exist_ok=True)
src_st = src_dir / "adapter_model.safetensors"
with safe_open(str(src_st), framework="pt") as f:
all_keys = list(f.keys())
expert_layers = sorted({int(m.group(2)) for k in all_keys if (m := EXPERT_PREFIX_RE.match(k))})
print(f"Found {len(expert_layers)} layers with expert LoRA: {expert_layers[:3]}...{expert_layers[-3:]}")
out_tensors = {}
# carry over everything that is NOT an expert-LoRA key unchanged (self_attn etc.)
expert_key_prefixes = {f"base_model.model.model.layers.{L}.mlp.experts." for L in expert_layers}
for k in all_keys:
if any(k.startswith(p) for p in expert_key_prefixes):
continue
out_tensors[k] = f.get_tensor(k)
print(f"Carried over {len(out_tensors)} non-expert tensors unchanged (self_attn etc.)")
n_pairs = 0
for L in expert_layers:
layer_prefix = f"base_model.model.model.layers.{L}.mlp.experts."
a_to_b = find_layer_expert_keys(all_keys, layer_prefix)
pairs = [(a_key, f.get_tensor(a_key), b_key, f.get_tensor(b_key)) for a_key, b_key in a_to_b.items()]
assert len(pairs) == 2, f"layer {L}: expected 2 expert-LoRA (A,B) pairs, got {len(pairs)}"
# Disambiguate gate_up vs down by A's in_features (last dim): whichever
# pair has the larger in_features is gate_up (in=hidden), the other is
# down (in=inter) -- true for standard transformer FFN shapes (hidden >= inter),
# and doesn't depend on PEFT's arbitrary nesting order for the two
# target_parameters.
pairs.sort(key=lambda p: p[1].shape[-1], reverse=True)
(_, gate_up_A, _, gate_up_B), (_, down_A, _, down_B) = pairs
E_r, hidden = gate_up_A.shape
out2, _ = gate_up_B.shape
inter = out2 // 2
assert E_r == E * r, f"layer {L}: expected E*r={E * r}, got {E_r}"
assert gate_up_B.shape[1] == E_r
assert down_A.shape == (E_r, inter), f"layer {L}: down_A shape {down_A.shape} != ({E_r},{inter})"
assert down_B.shape == (hidden, E_r), f"layer {L}: down_B shape {down_B.shape} != ({hidden},{E_r})"
for e in range(E):
lo, hi = e * r, (e + 1) * r
A_e = gate_up_A[lo:hi, :].contiguous() # [r, hidden] -- shared by gate & up
B_e = gate_up_B[:, lo:hi].contiguous() # [2*inter, r]
B_gate = B_e[:inter, :].contiguous() # [inter, r]
B_up = B_e[inter:, :].contiguous() # [inter, r]
A_dn = down_A[lo:hi, :].contiguous() # [r, inter]
B_dn = down_B[:, lo:hi].contiguous() # [hidden, r]
base = f"base_model.model.model.layers.{L}.mlp.experts.{e}"
out_tensors[f"{base}.gate_proj.lora_A.weight"] = A_e
out_tensors[f"{base}.gate_proj.lora_B.weight"] = B_gate
out_tensors[f"{base}.up_proj.lora_A.weight"] = A_e.clone()
out_tensors[f"{base}.up_proj.lora_B.weight"] = B_up
out_tensors[f"{base}.down_proj.lora_A.weight"] = A_dn
out_tensors[f"{base}.down_proj.lora_B.weight"] = B_dn
n_pairs += 3
if L == expert_layers[0]:
print(f" layer {L}: hidden={hidden} inter={inter} E*r={E_r} r={r}")
print(f"Converted {n_pairs} (layer, expert, projection) tensor pairs across {len(expert_layers)} layers")
print(f"Total output tensors: {len(out_tensors)}")
save_file(out_tensors, str(dst_dir / "adapter_model.safetensors"))
with open(src_dir / "adapter_config.json") as fh:
cfg = json.load(fh)
cfg.pop("target_parameters", None)
cfg["target_modules"] = ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]
cfg["_converted_from"] = str(src_dir)
cfg["_conversion_note"] = (
"Converted from Unsloth's fused per-layer MoE target_parameters format to "
"vLLM's per-expert experts.{e}.{gate_proj,up_proj,down_proj} format by "
"moe_lora_convert_vllm.py, per MOE_LORA_LAYOUT_PLAN.md section 3 "
"(contiguous-both A/B expert blocking, corrected 2026-09-09)."
)
with open(dst_dir / "adapter_config.json", "w") as fh:
json.dump(cfg, fh, indent=2)
for extra in ("tokenizer_config.json", "tokenizer.json"):
src_f = src_dir / extra
if src_f.exists():
(dst_dir / extra).write_bytes(src_f.read_bytes())
print(f"\nWrote converted adapter to {dst_dir}")
print("CONVERT_DONE")
if __name__ == "__main__":
if len(sys.argv) != 3:
print(f"Usage: {sys.argv[0]} <src_adapter_dir> <dst_adapter_dir>")
sys.exit(1)
convert(sys.argv[1], sys.argv[2])