Alpamys
ae6a18e7d7
test(v0.71.18): ANSI-strip the --minillm-on-policy --help assertion for CI FORCE_COLOR
...
TestTrainCliMinillmOnPolicy::test_flag_in_help removed newlines + spaces but
not ANSI codes, so under CI FORCE_COLOR the Rich-rendered long flag (ANSI codes
between the dashes) failed the substring check. Strip ANSI + remove all
whitespace before the check, matching the cloud/sandbox help tests + the
v0.71.17 precedent. Verified under FORCE_COLOR=1 (114 passed).
2026-06-09 00:05:36 +05:00
Alpamys
70fd5ee9f3
feat(distill,agent,cloud): on-policy MiniLLM + aligned ULD + agent sandbox eval + Modal cloud (v0.71.18)
...
Closes #257 , #258 , #110 , #16 .
- #257 MiniLLM true on-policy rollout: minillm_on_policy_rollout (Gu et al. §3.1
autoregressive teacher-mixed rollout, reverse-KL on full distributions,
grad-to-student-only) + on_policy_term + training.minillm_on_policy /
minillm_rollout_length, wired into DistillTrainer.compute_loss.
- #258 cross-tokenizer ULD wasserstein_aligned: align_token_sequences (difflib
char-span) + aggregate_aligned_logits + uld_aligned_loss for fully-disjoint
tokenizers, wired into DistillTrainer.
- #110 soup agent eval --sandbox: build_eval_stub (base64-embed-as-data) +
run_eval_in_sandbox (v0.25 RLVR isolation + SANDBOX_NETWORK_GUARD) +
classify_sandbox_outcome (ok/tool_error/timeout/arg_error).
- #16 soup train --cloud modal: render a Modal app from soup.yaml (config
base64-embedded), plan-only default, --cloud-submit token-gated; [modal] extra.
+114 tests (tests/test_v07118.py); 13656 -> 13770. Step-6 smoke on real input
(Windows + RTX 3050): Modal stub render, real subprocess sandbox scorecard,
on-policy distill (tiny-gpt2), cross-tokenizer aligned ULD (GPT-2 + Llama).
2026-06-08 23:49:14 +05:00