# Eval Report — 2026-07-02 Generated by `scripts/run_eval.sh` → `tsmm.eval.report`. ## Discipline notes - **Main metric: VUS-PR** (threshold-free). - **Point-F1 reported WITHOUT point-adjust.** PA-F1 shown separately as a CONTROL only. - Trivial baselines (random / constant / all-report) are included. - LLM-judge backend: `stub`. ## Dataset: synth | predictor | n | VUS-PR | Aff-F1 | AUC-PR | Point-F1 | PA-F1(control) | parse_rate | |---|---|---|---|---|---|---|---| | model | 361 | 0.2241 | 0.2474 | 0.3050 | 0.2186 | 0.5018 | 0.98 | | pure_llm | 361 | 0.2420 | 0.0023 | 0.2941 | 0.0013 | 0.0071 | 0.01 | | random | 361 | 0.3266 | 0.3970 | 0.2689 | 0.3266 | 0.9771 | 1.00 | | constant | 361 | 0.2601 | 0.4302 | 0.2951 | 0.3914 | 1.0000 | 1.00 | | all_report | 361 | 0.2601 | 0.4302 | 0.2951 | 0.3914 | 1.0000 | 1.00 | _PA = point-adjust; shown as control only, not headline._ ### QA judge (stub 0-5) — synth | predictor | anomaly | compare | describe | event | forecast | root_cause | mean | |---|---|---|---|---|---|---|---| | model | 2.45 | 2.75 | 1.34 | 3.89 | 0.03 | 3.99 | 2.41 | | pure_llm | 0.02 | 1.21 | 0.17 | 1.62 | 0.11 | 2.20 | 0.89 | | random | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | | constant | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | | all_report | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | _QA judge = offline stub token-overlap heuristic; swap `--judge_backend openai` for GPT-4o scoring._ ## Dataset: real | predictor | n | VUS-PR | Aff-F1 | AUC-PR | Point-F1 | PA-F1(control) | parse_rate | |---|---|---|---|---|---|---|---| | model | 39 | 0.2263 | 0.0000 | 0.4516 | 0.0000 | 0.0000 | 0.00 | | pure_llm | 39 | 0.2236 | 0.0001 | 0.4516 | 0.0000 | 0.0000 | 0.00 | | random | 39 | 0.2546 | 0.3070 | 0.2894 | 0.2546 | 0.9744 | 1.00 | | constant | 39 | 0.2263 | 0.3321 | 0.4516 | 0.3088 | 1.0000 | 1.00 | | all_report | 39 | 0.2263 | 0.3321 | 0.4516 | 0.3088 | 1.0000 | 1.00 | _PA = point-adjust; shown as control only, not headline._ ### QA judge (stub 0-5) — real | predictor | anomaly | describe | mean | |---|---|---|---| | model | 0.00 | 0.00 | 0.00 | | pure_llm | 0.28 | 0.00 | 0.14 | | random | 0.00 | 0.00 | 0.00 | | constant | 0.00 | 0.00 | 0.00 | | all_report | 0.00 | 0.00 | 0.00 | _QA judge = offline stub token-overlap heuristic; swap `--judge_backend openai` for GPT-4o scoring._