From 2accd43cfe6374d4216ec91cb243d37f5fd4848b Mon Sep 17 00:00:00 2001 From: Joey Grasty Date: Sun, 30 Aug 2026 17:28:23 -0500 Subject: [PATCH] Switch production checkpoint reference to epoch 3 Production was switched from ep2 to ep3 based on separability and length-calibration evidence (JOINT_TRAINING_CLOSEOUT_REPORT.md). Updates start-model.sh's default and the README's setup instructions. The Drive link itself still points to the ep2 file pending re-upload -- flagged inline in the README. --- README.md | 6 ++++-- start-model.sh | 6 +++--- 2 files changed, 7 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index ed1366e..da88c82 100644 --- a/README.md +++ b/README.md @@ -24,11 +24,13 @@ The 208 training/holdout packages themselves are already embedded inside ### 1. Get the model file -Download `Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep2-Q8_0.gguf` (~31GB) from +Download `Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep3-Q8_0.gguf` (~31GB) from Google Drive: https://drive.google.com/open?id=1q1w_XMp6oRmfx8xqXDV-HvDx_krVxQtD +**Note (2026-08-30): the link above still points to the ep2 file.** Production was just switched to ep3 (separability + length-calibration evidence, see `JOINT_TRAINING_CLOSEOUT_REPORT.md`); the ep3 GGUF upload and link swap are pending. + Memory needed: ~31GB for the weights plus ~3GB of KV cache at the full 32768-token context, so ~35GB total. Confirmed working on a DGX Spark; should comfortably fit a 48GB L40 as well. If a machine has less than that, @@ -51,7 +53,7 @@ This produces `llama.cpp/build/bin/llama-server`, which is what ### 3. Start the model server ``` -MODEL_PATH=/path/to/Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep2-Q8_0.gguf \ +MODEL_PATH=/path/to/Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep3-Q8_0.gguf \ LLAMA_SERVER=/path/to/llama.cpp/build/bin/llama-server \ ./start-model.sh ``` diff --git a/start-model.sh b/start-model.sh index 950ee22..ae584a5 100755 --- a/start-model.sh +++ b/start-model.sh @@ -3,17 +3,17 @@ # # Requires: # - llama.cpp built with CUDA support (see README.md) -# - the Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep2-Q8_0.gguf file (from the team +# - the Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep3-Q8_0.gguf file (from the team # Google Drive) # # Configure by setting environment variables before running, e.g.: -# MODEL_PATH=/data/models/Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep2-Q8_0.gguf \ +# MODEL_PATH=/data/models/Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep3-Q8_0.gguf \ # LLAMA_SERVER=/opt/llama.cpp/build/bin/llama-server \ # ./start-model.sh # or just edit the defaults below. set -e -MODEL_PATH="${MODEL_PATH:-$HOME/models/Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep2-Q8_0.gguf}" +MODEL_PATH="${MODEL_PATH:-$HOME/models/Qwen3-30B-A3B-VoxDay-Nuttall-SFT-ep3-Q8_0.gguf}" LLAMA_SERVER="${LLAMA_SERVER:-$HOME/llama.cpp/build/bin/llama-server}" HOST="${HOST:-0.0.0.0}" PORT="${PORT:-8200}"