aecccfab1d
- eval/ts_effectiveness.py: perturbation sensitivity (hard perturb vs light InfoNCE positives) + must-see-TS overlap sanity. - Measured on stage1_final.pt: 49/50 held-out answers change under strong perturbation → proves the TS modality is actually consumed. - Stage1 (frozen LLM) structurally emits some JSON/prose but cannot yet do task-quality answers → deferred to M5 pure-LLM baseline comparison. - Mark T3.3/T3.4 (M3 exit) done.