Commit Graph

2 Commits

Author SHA1 Message Date
Alpamys cba51e153b fix: close out the 3 deferred code-review items (EMA, hardware-fit, MoD)
These were documented as known limitations in the MEDIUM/LOW pass; now fixed.

1. reward_hack EMA smoothing window — smooth_signal("ema") folded only
   window[-1], so reward_hack_smoothing_window had no effect. Now a windowed
   EMA folds alpha over the whole retained window (oldest→newest) then the new
   sample, so a larger window incorporates more history; a 1-element window
   reduces to the old 2-tap form. Updated test_v07126 (0.3 → 0.275).

2. hardware_fit OOM gate wired into `soup train` — the analytical VRAM
   predictor was never called despite its docstring. Added
   _build_hardware_fit_input (SoupConfig → HardwareFitInput, best-effort;
   None when not statically predictable, e.g. batch_size="auto") and
   _hardware_fit_preflight, run after device detection. Refuses on predicted
   OOM (peak × 1.1 > available) unless the documented --allow-oom-attempt
   opt-out is passed; skips silently on CPU / unknown VRAM, and the flag is
   threaded through the --gpus re-exec.

3. MoD real token-dropping — mod_forward ran the full block on ALL tokens then
   masked (zero compute savings). Now the top-k tokens are gathered into a
   shorter sub-sequence, the block runs on ONLY those tokens (real saving),
   the gated result is scattered back, and unselected tokens pass through
   unchanged. Positional inputs (RoPE cos/sin, 4D-causal attention_mask,
   position_ids, cache_position) are gathered to the sub-sequence; any
   unsafe-to-gather case (positional forward args, KV cache, non-4D mask)
   falls back to the prior correct blend so attention can never be silently
   mis-computed. Validated on CPU (gather/scatter/passthrough/savings +
   fallback); the sub-sequence-attention numerics still warrant GPU validation
   at scale.

Adds tests/test_code_review_deferred.py (7 tests). ruff clean; full suite
14867 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:36:59 +05:00
Alpamys defd3151cf fix: remediate the MEDIUM/LOW code-review findings
- license_matrix: permissive ↔ weak-copyleft is now symmetric (MIT + LGPL no
  longer flagged incompatible regardless of order).
- formats.detect_format: check tool-calling before audio so an audio+tools row
  keeps its tool_calls instead of being classified audio.
- formats: reject null message content in the alpaca / sharegpt / vision text
  converters (routes the row to the drop path instead of literal None content).
  DPO only rejects an explicit null (chosen/rejected may be message lists).
- eval/custom.tool_call_args_subset: hallucinated args on a no-arg expected
  call now score 0.0 (was a dead `0.5 if ... else 0.5` ternary).
- monitoring/callback: SSE metric push uses `is not None` so a real 0.0 loss/lr
  is not reported as None.
- cans/schema.DeployTarget: reject Windows drive-absolute paths (C:\..., C:/...).
- commands/diagnose: reject a non-numeric evidence score with a clear
  BadParameter (was ValueError -> exit 1 with zero output).
- commands/generate: partial-save accumulated examples on a mid-run failure so
  paid API spend is not discarded.
- __init__.py: fix the byte-corrupted em dash in the package docstring.

Two MEDIUM/LOW items reverted to documented known limitations after they broke
existing behaviour locked by tests: (1) the reward_hack EMA smoother is a
recursive 2-tap by design — smoothing_window only affects `median`; (2)
`soup train`'s MoD compute-savings and the hardware_fit OOM preflight wiring
are architectural, GPU-validation work left as follow-ups.

Adds tests/test_code_review_medium_low.py (10 tests). ruff clean; full suite
14861 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 17:20:05 +05:00