cd6209119a
Full-scale retrain (stage1 31250 steps on 500k align, stage2 9375 steps on 100k sft) vs the prior small-scale run (5万/5k steps): metric small-scale full-scale stage2 EMA 0.81 0.75 model VUS-PR 0.14 0.22 (+57%) model Aff-F1 0.21 0.25 (vs pure_llm 0.003 — TS modality helps) model QA mean 2.45 2.41 QA vs pure_llm 2.8× 2.7× (stable lead, 6 categories) Confirms the hypothesis that weak anomaly localization was mostly a data-volume problem, not an algorithmic one: full-scale training lifted VUS-PR 57% and the model now clearly beats the pure-LLM baseline on affiliation-F1 (0.25 vs 0.003). VUS-PR still trails trivial baselines (0.22 vs random 0.33) — precise localization remains the open frontier, but the TS modality is demonstrably contributing (was unclear before).
1.4 KiB
1.4 KiB
Eval Report — 2026-07-01
Generated by scripts/run_eval.sh → tsmm.eval.report.
Discipline notes
- Main metric: VUS-PR (threshold-free).
- Point-F1 reported WITHOUT point-adjust. PA-F1 shown separately as a CONTROL only.
- Trivial baselines (random / constant / all-report) are included.
- LLM-judge backend:
stub.
Dataset: synth
| predictor | n | VUS-PR | Aff-F1 | AUC-PR | Point-F1 | PA-F1(control) | parse_rate |
|---|---|---|---|---|---|---|---|
| model | 361 | 0.2241 | 0.2474 | 0.3050 | 0.2186 | 0.5018 | 0.98 |
| pure_llm | 361 | 0.2391 | 0.0029 | 0.2940 | 0.0013 | 0.0071 | 0.00 |
| random | 361 | 0.3266 | 0.3970 | 0.2689 | 0.3266 | 0.9771 | 1.00 |
| constant | 361 | 0.2601 | 0.4302 | 0.2951 | 0.3914 | 1.0000 | 1.00 |
| all_report | 361 | 0.2601 | 0.4302 | 0.2951 | 0.3914 | 1.0000 | 1.00 |
PA = point-adjust; shown as control only, not headline.
QA judge (stub 0-5) — synth
| predictor | anomaly | compare | describe | event | forecast | root_cause | mean |
|---|---|---|---|---|---|---|---|
| model | 2.45 | 2.75 | 1.34 | 3.89 | 0.03 | 3.99 | 2.41 |
| pure_llm | 0.03 | 1.25 | 0.19 | 1.64 | 0.12 | 2.17 | 0.90 |
| random | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| constant | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| all_report | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
QA judge = offline stub token-overlap heuristic; swap --judge_backend openai for GPT-4o scoring.