Files
ts-as-modality/reports/eval-2026-07-01.md
T
张宗平 cd6209119a eval(m5): full-scale retrain report — VUS-PR 0.14→0.22, model beats pure_llm
Full-scale retrain (stage1 31250 steps on 500k align, stage2 9375 steps on
100k sft) vs the prior small-scale run (5万/5k steps):

  metric              small-scale    full-scale
  stage2 EMA          0.81           0.75
  model VUS-PR        0.14           0.22  (+57%)
  model Aff-F1        0.21           0.25  (vs pure_llm 0.003 — TS modality helps)
  model QA mean       2.45           2.41
  QA vs pure_llm      2.8×           2.7×  (stable lead, 6 categories)

Confirms the hypothesis that weak anomaly localization was mostly a
data-volume problem, not an algorithmic one: full-scale training lifted
VUS-PR 57% and the model now clearly beats the pure-LLM baseline on
affiliation-F1 (0.25 vs 0.003). VUS-PR still trails trivial baselines
(0.22 vs random 0.33) — precise localization remains the open frontier,
but the TS modality is demonstrably contributing (was unclear before).
2026-07-01 13:44:31 +00:00

1.4 KiB

Eval Report — 2026-07-01

Generated by scripts/run_eval.shtsmm.eval.report.

Discipline notes

  • Main metric: VUS-PR (threshold-free).
  • Point-F1 reported WITHOUT point-adjust. PA-F1 shown separately as a CONTROL only.
  • Trivial baselines (random / constant / all-report) are included.
  • LLM-judge backend: stub.

Dataset: synth

predictor n VUS-PR Aff-F1 AUC-PR Point-F1 PA-F1(control) parse_rate
model 361 0.2241 0.2474 0.3050 0.2186 0.5018 0.98
pure_llm 361 0.2391 0.0029 0.2940 0.0013 0.0071 0.00
random 361 0.3266 0.3970 0.2689 0.3266 0.9771 1.00
constant 361 0.2601 0.4302 0.2951 0.3914 1.0000 1.00
all_report 361 0.2601 0.4302 0.2951 0.3914 1.0000 1.00

PA = point-adjust; shown as control only, not headline.

QA judge (stub 0-5) — synth

predictor anomaly compare describe event forecast root_cause mean
model 2.45 2.75 1.34 3.89 0.03 3.99 2.41
pure_llm 0.03 1.25 0.19 1.64 0.12 2.17 0.90
random 0.00 0.00 0.00 0.00 0.00 0.00 0.00
constant 0.00 0.00 0.00 0.00 0.00 0.00 0.00
all_report 0.00 0.00 0.00 0.00 0.00 0.00 0.00

QA judge = offline stub token-overlap heuristic; swap --judge_backend openai for GPT-4o scoring.