|
|
|
|
@@ -131,8 +131,8 @@ serving 50 concurrent generation threads.
|
|
|
|
|
## 6. Code patches needed for current library versions
|
|
|
|
|
|
|
|
|
|
The repo (as of the commit we tested against) predates some of the exact library versions that
|
|
|
|
|
`pip install` resolves to today. Three small patches were needed — none are DGX-Spark-specific,
|
|
|
|
|
they'd bite on any platform once these versions are current on PyPI:
|
|
|
|
|
`pip install` resolves to today. Four small patches were needed — none are DGX-Spark-specific,
|
|
|
|
|
they'd bite on any platform in the same circumstances:
|
|
|
|
|
|
|
|
|
|
**`utils/vllm_manager.py`** — vLLM 0.26.0 removed the `--disable-log-requests` flag (replaced by an
|
|
|
|
|
opt-in `--enable-log-requests`, off by default already). Delete the line:
|
|
|
|
|
@@ -155,24 +155,93 @@ training, so all-zero/all-text is correct):
|
|
|
|
|
- the reference-model forward pass, `self.ref_model is None` branch (inside `null_ref_context()`)
|
|
|
|
|
- the reference-model forward pass, `self.ref_model is not None` branch
|
|
|
|
|
|
|
|
|
|
## 7. FTPO fine-tuning is compute-bound, not throughput-bound
|
|
|
|
|
**`core/orchestration.py`** — resuming a run with generation disabled
|
|
|
|
|
(`--generation-step-enabled false`, used to re-run just the finetune stage against an
|
|
|
|
|
already-generated dataset) crashes:
|
|
|
|
|
```
|
|
|
|
|
UnboundLocalError: cannot access local variable 'banned_ngrams_json_path' where it is not associated with a value
|
|
|
|
|
```
|
|
|
|
|
`banned_ngrams_json_path` / `banned_slop_phrases_json_path` are only ever assigned inside the
|
|
|
|
|
`if generation_enabled:` block, but a later resume-logging block references them unconditionally.
|
|
|
|
|
Guard that block with the same flag:
|
|
|
|
|
```python
|
|
|
|
|
if start_iter_idx > 0 and generation_enabled:
|
|
|
|
|
if banned_ngrams_json_path.exists(): ...
|
|
|
|
|
```
|
|
|
|
|
(There's a second, non-fatal instance of the same root cause — `iter0_output_file_for_dpo`
|
|
|
|
|
referenced before assignment when loading stats from a prior completed run — it's caught and
|
|
|
|
|
logged as a warning rather than raised, so it doesn't block anything, but it's the same bug shape.)
|
|
|
|
|
|
|
|
|
|
## 7. FTPO fine-tuning was slow because of fixed-length padding, not raw compute
|
|
|
|
|
|
|
|
|
|
With `finetune_batch_size: 1` / `gradient_accumulation_steps: 16`, we measured ~750 optimizer steps
|
|
|
|
|
at ~160-185s/step — a ~34 hour run for the full 12,000-example dataset. Bumping
|
|
|
|
|
`finetune_batch_size` to 4 (with `gradient_accumulation_steps` dropped to 4 to keep the same
|
|
|
|
|
effective batch size) made **no meaningful difference** — still ~160-175s/step.
|
|
|
|
|
|
|
|
|
|
Takeaway: this workload is compute-bound (two full forward passes per micro-batch — the model plus
|
|
|
|
|
a reference-model pass for the MSE tether loss term — over sequences up to
|
|
|
|
|
`finetune_max_seq_length: 4000` tokens), not limited by batch-size/scheduling overhead. Increasing
|
|
|
|
|
batch size doesn't reduce total FLOPs for a fixed effective batch size, so it doesn't help here.
|
|
|
|
|
If you need a faster run, the actual levers are:
|
|
|
|
|
- lower `finetune_max_train_examples` (fewer total steps, less data coverage)
|
|
|
|
|
- lower `finetune_max_seq_length` (less compute per step, truncates longer training examples)
|
|
|
|
|
- accept the long runtime and let it run in the background
|
|
|
|
|
Initial hypothesis was that this was inherent — genuinely compute-bound (two full forward passes
|
|
|
|
|
per micro-batch, over long sequences), not limited by batch-size/scheduling overhead. Measuring the
|
|
|
|
|
actual training data disproved that. `ftpo_trainer.py`'s collator pads every batch to a **fixed**
|
|
|
|
|
`finetune_max_seq_length` (4000 tokens) regardless of content:
|
|
|
|
|
|
|
|
|
|
We left `finetune_batch_size` at the default (`1`) since increasing it only costs more memory for
|
|
|
|
|
no speed benefit on this hardware.
|
|
|
|
|
```python
|
|
|
|
|
max_len = self.args.max_length # always 4000, never pad-to-longest-in-batch
|
|
|
|
|
prompt_ids = torch.full((batch_sz, max_len), pad_id, dtype=torch.long)
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
We tokenized all 12,000 training contexts with the real tokenizer to see how much of that 4000 was
|
|
|
|
|
actually needed:
|
|
|
|
|
|
|
|
|
|
| | tokens |
|
|
|
|
|
|---|---|
|
|
|
|
|
| mean | 529.9 |
|
|
|
|
|
| median | 509 |
|
|
|
|
|
| p90 / p99 | 953 / 1080 |
|
|
|
|
|
| **max across all 12,000 examples** | **1126** |
|
|
|
|
|
|
|
|
|
|
Not one example reaches even a third of the 4000-token padding target; the mean uses 13.2% of it.
|
|
|
|
|
Every forward pass — both the main model and the reference-model pass — was processing ~4000 tokens
|
|
|
|
|
of mostly padding, roughly 4-7x more than the actual content needs. This also explains why the
|
|
|
|
|
batch-size bump did nothing: total padded-token compute is invariant to how the effective batch of
|
|
|
|
|
16 gets split into micro-batches, so reshuffling batch/accum never touched the real cost. This is a
|
|
|
|
|
collator-design issue, not a hardware ceiling — it would waste the same proportion on any GPU.
|
|
|
|
|
|
|
|
|
|
**Fix:** lower `finetune_max_seq_length` to comfortably cover the real distribution, e.g. `1280`
|
|
|
|
|
(covers p99 with headroom, nothing in the dataset gets truncated) instead of `4000`. We left
|
|
|
|
|
`finetune_batch_size` at the default (`1`) since increasing it has no effect either way here.
|
|
|
|
|
|
|
|
|
|
**Measured, not just predicted:** re-ran training with `finetune_max_seq_length: 1280` for a 30-minute
|
|
|
|
|
validation window (46 steps, timing fully steady by the end — no drift):
|
|
|
|
|
|
|
|
|
|
| | `max_seq_length=4000` | `max_seq_length=1280` |
|
|
|
|
|
|---|---|---|
|
|
|
|
|
| steady-state step time | ~160-185s/step | **~40s/step** |
|
|
|
|
|
| memory used during training | ~82GB | ~25GB |
|
|
|
|
|
| speedup | — | **~4.3x** |
|
|
|
|
|
| extrapolated full run (750 steps) | ~34h | **~8.3h** |
|
|
|
|
|
|
|
|
|
|
Better than the ~3x predicted from the token-count ratio alone — the memory savings from shorter
|
|
|
|
|
sequences apparently helped beyond just the raw compute reduction. No errors across the validation
|
|
|
|
|
run.
|
|
|
|
|
|
|
|
|
|
If you need it faster still, the other lever is `finetune_max_train_examples` (fewer total steps,
|
|
|
|
|
less data coverage) — or just accept the ~8h runtime and let it run in the background.
|
|
|
|
|
|
|
|
|
|
## 8. Sustained full-GPU load can trigger a thermal shutdown
|
|
|
|
|
|
|
|
|
|
An unthrottled attempt at the full finetuning run caused the machine to shut down from heat partway
|
|
|
|
|
through (no OOM/crash signature in the logs — the process and GPU state simply vanished when the
|
|
|
|
|
box power-cycled). No checkpoint had been written yet, so the run restarted from step 0 — FTPO
|
|
|
|
|
doesn't checkpoint mid-run by default, so a mid-run shutdown here means losing all progress so far.
|
|
|
|
|
|
|
|
|
|
Fix: cap GPU clocks/power before starting a sustained multi-hour job (finetuning, not the bursty
|
|
|
|
|
generation step):
|
|
|
|
|
```bash
|
|
|
|
|
nvidia-smi -pl <lower_watts> # if supported on this platform
|
|
|
|
|
# or check nvidia-smi -q -d CLOCK,POWER for current caps and headroom
|
|
|
|
|
```
|
|
|
|
|
With clocks capped, the retry ran the full ~4h12m to completion without incident — GPU temperature
|
|
|
|
|
held in the 71-81°C range under sustained 90%+ utilization throughout.
|
|
|
|
|
|
|
|
|
|
## Validated results
|
|
|
|
|
|
|
|
|
|
@@ -180,9 +249,18 @@ Ran the full pipeline against `unsloth/gemma-3-4b-it` (2 iterations, 1200 prompt
|
|
|
|
|
|
|
|
|
|
- Iteration 0 (baseline, no bans): completed in 28m24s, `repetition_per_100k_chars` = 160
|
|
|
|
|
- Iteration 1 (with ban lists from iteration 0's analysis): completed in 1h31m43s (slower — active
|
|
|
|
|
backtracking around bans), `repetition_per_100k_chars` = 56 — a real, measured reduction in slop
|
|
|
|
|
- FTPO training: confirmed working end-to-end (750 steps, 12,000 preference pairs) after the
|
|
|
|
|
patches in §6; not run to completion due to the ~34h runtime (§7)
|
|
|
|
|
backtracking around bans), `repetition_per_100k_chars` = 56 — a real, measured reduction in slop.
|
|
|
|
|
Both numbers are from a run predating the `refusal_detector.py` fix in §6, so they reflect
|
|
|
|
|
n-gram/phrase-ban slop reduction only, not refusal-aware filtering.
|
|
|
|
|
- FTPO training: **ran to completion.** With the fixes from §6 and §7
|
|
|
|
|
(`finetune_max_seq_length: 1280`), trained on the 12,000-example preference-pair dataset (quota-
|
|
|
|
|
sampled from 58,369 raw pairs) at `finetune_batch_size: 1` / `gradient_accumulation_steps: 16`.
|
|
|
|
|
Stopped early by the configured `finetune_early_stopping_wins: 0.85` threshold — `chosen_win`
|
|
|
|
|
(fraction of chosen completions preferred over rejected) crossed 0.85 at step 380 of a possible
|
|
|
|
|
750 (1 epoch). `train_loss` fell from ~4.0 to 2.19 over the run. Total wall-clock time
|
|
|
|
|
(training + LoRA save + 16-bit merge): **4h12m32s**, steady-state ~40s/step throughout — matching
|
|
|
|
|
the §7 prediction. Output: LoRA adapter and a merged 16-bit model (8.1GB) under
|
|
|
|
|
`finetuned_model_ftpo_exp01/`.
|
|
|
|
|
|
|
|
|
|
## Quick-reference: full env setup
|
|
|
|
|
|
|
|
|
|
|