# Eval Report — 2026-06-30 Generated by `scripts/run_eval.sh` → `tsmm.eval.report`. ## Discipline notes - **Main metric: VUS-PR** (threshold-free). - **Point-F1 reported WITHOUT point-adjust.** PA-F1 shown separately as a CONTROL only. - Trivial baselines (random / constant / all-report) are included. - LLM-judge backend: `stub`. ## Dataset: synth | predictor | n | VUS-PR | Aff-F1 | AUC-PR | Point-F1 | PA-F1(control) | parse_rate | |---|---|---|---|---|---|---|---| | model | 40 | 0.1414 | 0.1674 | 0.2343 | 0.1414 | 0.4296 | 1.00 | | pure_llm | 40 | 0.2195 | 0.0000 | 0.2357 | 0.0000 | 0.0000 | 0.00 | | random | 40 | 0.2859 | 0.3484 | 0.2241 | 0.2859 | 0.9746 | 1.00 | | constant | 40 | 0.2195 | 0.3771 | 0.2357 | 0.3395 | 1.0000 | 1.00 | | all_report | 40 | 0.2195 | 0.3771 | 0.2357 | 0.3395 | 1.0000 | 1.00 | _PA = point-adjust; shown as control only, not headline._ ### QA judge (stub 0-5) — synth | predictor | anomaly | compare | describe | event | forecast | root_cause | mean | |---|---|---|---|---|---|---|---| | model | 2.60 | 2.64 | 1.65 | 4.08 | 0.12 | 3.60 | 2.45 | | pure_llm | 0.04 | 1.21 | 0.17 | 1.57 | 0.13 | 2.18 | 0.88 | | random | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | | constant | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | | all_report | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | _QA judge = offline stub token-overlap heuristic; swap `--judge_backend openai` for GPT-4o scoring._