## Context Pi coding agent has no built-in mechanism for persisting a user-defined goal across a session. Skills can provide instructions, but cannot subscribe to events or inject state automatically. Extensions can, but have no declarative goal-tracking format. This change combines both layers. ## Goals / Non-Goals **Goals:** - Detect goal-setting phrases in user messages across languages - Persist goal state to a file that survives context compaction - Automatically inject goal context on every agent turn - Eliminate unnecessary user interaction (timed auto-defaults, auto-retry, auto-recovery) - Delegate coarse-grained goals to comet automatically - Verify completion against explicit acceptance criteria before declaring done **Non-Goals:** - Multi-goal concurrent tracking (only one active goal at a time) - Cross-session goal tracking beyond file persistence - Subagent orchestration (agent uses existing Pi tools) - Comet skill modifications ## Decisions 1. **Two-layer architecture (skill + extension)**: Skill provides instructions the agent reads; extension provides mechanical automation. This separates judgment (criteria quality, iteration work) from mechanics (trigger detection, state injection, timed confirms). 2. **File-based state (`.agents/goal/GOAL.md`)**: Markdown with YAML frontmatter. Survives context compaction. Human-readable and editable. Deleted on completion. 3. **Extension-driven state injection via `before_agent_start`**: Every agent turn re-injects goal context from the file. This is the primary mechanism for context-compaction resistance — even if conversation is compacted, the extension rebuilds goal context. 4. **Timed auto-default confirms (10s timeout)**: Stale state resolution, criteria confirmation, and other checkpoints use `ctx.ui.select({ timeout: 10000 })` with safe defaults. User can intervene but doesn't have to. 5. **Per-iteration exponential backoff (10 retries, ~37min)**: Tool failures and LLM API errors retry automatically. Counter resets per-iteration. After 10 retries, status becomes `blocked` and extension injects strategy-change prompt. 6. **Dual-signal stagnation detection**: Criteria passing count unchanged for 3 iterations AND no progress keywords in Progress Log. Both required. Extension auto-injects recovery prompt twice before notifying user. 7. **Granularity split with comet delegation**: Coarse-grained goals (file count > 5, multi-module, architectural) automatically invoke `comet-open`. Fallback to multi-stage mode if comet unavailable. 8. **Anti-trigger rules are mechanical**: Quote blocks, code blocks, reported speech, trivial-task keywords, and short descriptions without architectural keywords — all checked programmatically in the extension, not subjectively by the agent. ## Risks / Trade-offs | Risk | Mitigation | |------|------------| | Agent rationalizes skipping criteria | Red flags list in error-handling reference; criteria modification requires logged reason | | Sleep timeout in Pi bash tool | `timeout = delay + 30` parameter | | Extension bugs block the loop | Extension failures degrade gracefully — skill instructions still work without extension | | User loses ability to intervene | Explicit abandon and material criteria change still user-controllable | | Context tension detection imprecise | `ctx.getContextUsage() > 0.8` threshold with auto-save directive | | Mid-loop granularity upgrade needs skill-side signal | Skill can set `status: needs-upgrade` in GOAL.md; extension reads it on next turn |