Switch production checkpoint reference to epoch 3

Production was switched from ep2 to ep3 based on separability and
length-calibration evidence (JOINT_TRAINING_CLOSEOUT_REPORT.md). Updates
start-model.sh's default and the README's setup instructions. The Drive
link itself still points to the ep2 file pending re-upload -- flagged
inline in the README.
This commit is contained in:
2026-08-30 17:28:23 -05:00
parent 57ab1fa4d5
commit 2accd43cfe
2 changed files with 7 additions and 5 deletions

View File

@@ -24,11 +24,13 @@ The 208 training/holdout packages themselves are already embedded inside
### 1. Get the model file
Download `Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep2-Q8_0.gguf` (~31GB) from
Download `Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep3-Q8_0.gguf` (~31GB) from
Google Drive:
https://drive.google.com/open?id=1q1w_XMp6oRmfx8xqXDV-HvDx_krVxQtD
**Note (2026-08-30): the link above still points to the ep2 file.** Production was just switched to ep3 (separability + length-calibration evidence, see `JOINT_TRAINING_CLOSEOUT_REPORT.md`); the ep3 GGUF upload and link swap are pending.
Memory needed: ~31GB for the weights plus ~3GB of KV cache at the full
32768-token context, so ~35GB total. Confirmed working on a DGX Spark;
should comfortably fit a 48GB L40 as well. If a machine has less than that,
@@ -51,7 +53,7 @@ This produces `llama.cpp/build/bin/llama-server`, which is what
### 3. Start the model server
```
MODEL_PATH=/path/to/Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep2-Q8_0.gguf \
MODEL_PATH=/path/to/Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep3-Q8_0.gguf \
LLAMA_SERVER=/path/to/llama.cpp/build/bin/llama-server \
./start-model.sh
```