Ran the full FTPO fine-tune to completion: 4h12m wall-clock, early-stopped at
step 380/750 by the configured chosen_win=0.85 threshold, train_loss 4.0->2.19.
Output is a merged 16-bit model plus LoRA adapter.
Also fixes an UnboundLocalError in core/orchestration.py when resuming a run
with generation disabled (--generation-step-enabled false), which is needed to
re-run just the finetune stage against an already-generated dataset.
Documents an unrelated platform incident: an uncapped multi-hour finetune
triggered a thermal shutdown mid-run with no checkpoint saved, requiring a
restart from step 0. Capping GPU clocks/power before the retry let it complete
without issue.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0141RdLKiMeXLMXsWyb4U3B5
- utils/vllm_manager.py: drop --disable-log-requests, removed in vLLM 0.26.0
- core/ftpo_trainer.py: pass token_type_ids to the 3 model forward calls in
compute_loss -- transformers 5.5.0's Gemma3 requires it during training
for causal-mask construction (Gemma3 is natively multimodal)
- configs/gemma-3-4b-it.yaml: lower vllm_gpu_memory_utilization 0.85->0.5,
since the DGX Spark's 121GB is unified CPU/GPU memory and the default
starved the OS, causing swap thrashing
- DGX_SPARK_SETUP.md: full writeup of the above plus the parts that don't
live in this repo (two-conda-env split to resolve a vllm/unsloth
transformers version conflict, flash-attn source build flags, torch/CUDA
version matching, ~/.triton/cache permissions)
The antislop-vllm submodule also needed a one-line fix (removing an invalid
reference_compile kwarg in utils/refusal_detector.py that was silently
disabling refusal filtering) -- documented in DGX_SPARK_SETUP.md rather than
committed as a submodule pointer change, since we don't have push access to
upstream's antislop-vllm repo.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>