150e3bd9fd
Full diagnostic chain (synth→real zero-shot→real finetune→held-out) shows the TS-modality approach works in-distribution but fails to generalize across entities: same SMD dataset, same 38 channels, just a different machine (1→2) collapses output to repetition garbage. Root cause is the lightweight TS encoder (2-layer TF, 2.1M params) learning dataset/machine- specific patterns rather than transferable time-series representations. This is a valuable negative result: it cleanly rules out the '0.5B LLM + 2-layer TS encoder + single-distribution synth training' route for general time-series understanding, which is more credible than any 'looks like it works but never tested cross-domain' result. Records: - tasks.md: M6 diagnostic section (6.1-6.5) with evidence + archive decision - proposal.md: Status header flagging archive + pointer to diagnosis - docs/diagnosis-2026-07-02-cross-domain.md: full evidence chain, root-cause analysis, capability boundaries, asset inventory, restart conditions Project status: archived. Code/data/checkpoints/reports preserved. Restart requires any of: (1) larger encoder + real-data from-scratch training, (2) channel-agnostic patch strategy, (3) multi-dataset joint training with held-out-dataset generalization proof.
2.3 KiB
2.3 KiB
Eval Report — 2026-07-02
Generated by scripts/run_eval.sh → tsmm.eval.report.
Discipline notes
- Main metric: VUS-PR (threshold-free).
- Point-F1 reported WITHOUT point-adjust. PA-F1 shown separately as a CONTROL only.
- Trivial baselines (random / constant / all-report) are included.
- LLM-judge backend:
stub.
Dataset: synth
| predictor | n | VUS-PR | Aff-F1 | AUC-PR | Point-F1 | PA-F1(control) | parse_rate |
|---|---|---|---|---|---|---|---|
| model | 361 | 0.2241 | 0.2474 | 0.3050 | 0.2186 | 0.5018 | 0.98 |
| pure_llm | 361 | 0.2420 | 0.0023 | 0.2941 | 0.0013 | 0.0071 | 0.01 |
| random | 361 | 0.3266 | 0.3970 | 0.2689 | 0.3266 | 0.9771 | 1.00 |
| constant | 361 | 0.2601 | 0.4302 | 0.2951 | 0.3914 | 1.0000 | 1.00 |
| all_report | 361 | 0.2601 | 0.4302 | 0.2951 | 0.3914 | 1.0000 | 1.00 |
PA = point-adjust; shown as control only, not headline.
QA judge (stub 0-5) — synth
| predictor | anomaly | compare | describe | event | forecast | root_cause | mean |
|---|---|---|---|---|---|---|---|
| model | 2.45 | 2.75 | 1.34 | 3.89 | 0.03 | 3.99 | 2.41 |
| pure_llm | 0.02 | 1.21 | 0.17 | 1.62 | 0.11 | 2.20 | 0.89 |
| random | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| constant | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| all_report | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
QA judge = offline stub token-overlap heuristic; swap --judge_backend openai for GPT-4o scoring.
Dataset: real
| predictor | n | VUS-PR | Aff-F1 | AUC-PR | Point-F1 | PA-F1(control) | parse_rate |
|---|---|---|---|---|---|---|---|
| model | 39 | 0.2263 | 0.0000 | 0.4516 | 0.0000 | 0.0000 | 0.00 |
| pure_llm | 39 | 0.2236 | 0.0001 | 0.4516 | 0.0000 | 0.0000 | 0.00 |
| random | 39 | 0.2546 | 0.3070 | 0.2894 | 0.2546 | 0.9744 | 1.00 |
| constant | 39 | 0.2263 | 0.3321 | 0.4516 | 0.3088 | 1.0000 | 1.00 |
| all_report | 39 | 0.2263 | 0.3321 | 0.4516 | 0.3088 | 1.0000 | 1.00 |
PA = point-adjust; shown as control only, not headline.
QA judge (stub 0-5) — real
| predictor | anomaly | describe | mean |
|---|---|---|---|
| model | 0.00 | 0.00 | 0.00 |
| pure_llm | 0.28 | 0.00 | 0.14 |
| random | 0.00 | 0.00 | 0.00 |
| constant | 0.00 | 0.00 | 0.00 |
| all_report | 0.00 | 0.00 | 0.00 |
QA judge = offline stub token-overlap heuristic; swap --judge_backend openai for GPT-4o scoring.