Full-scale retrain (stage1 31250 steps on 500k align, stage2 9375 steps on
100k sft) vs the prior small-scale run (5万/5k steps):
metric small-scale full-scale
stage2 EMA 0.81 0.75
model VUS-PR 0.14 0.22 (+57%)
model Aff-F1 0.21 0.25 (vs pure_llm 0.003 — TS modality helps)
model QA mean 2.45 2.41
QA vs pure_llm 2.8× 2.7× (stable lead, 6 categories)
Confirms the hypothesis that weak anomaly localization was mostly a
data-volume problem, not an algorithmic one: full-scale training lifted
VUS-PR 57% and the model now clearly beats the pure-LLM baseline on
affiliation-F1 (0.25 vs 0.003). VUS-PR still trails trivial baselines
(0.22 vs random 0.33) — precise localization remains the open frontier,
but the TS modality is demonstrably contributing (was unclear before).