19 KiB
Goal-Driven Loop Skill — Design
Date: 2026-06-10 Status: Approved Target: Pi coding agent (Agent Skills standard + Extension)
Purpose
A goal-driven loop skill + extension pair for Pi that keeps the current session focused on a user-defined goal with explicit acceptance criteria, iterating until all criteria pass or genuine stagnation is detected. The skill is the instruction layer; the extension is the automation layer that eliminates unnecessary user interaction.
Problem
Long-running coding sessions drift. The user states a goal, the agent works on it, gets sidetracked by intermediate questions, and loses sight of the original target. When multiple turns span context compaction, the original goal may be forgotten entirely. Pure skill-only designs force user interaction at every checkpoint (stale state, confirmation, stagnation, criteria modification), which suspends the loop and breaks the "set and forget" intent.
Solution
Two-layer architecture:
- Skill (
goal-driven-loop/SKILL.md) — instruction layer. Tells the agent what to do when a goal is active: criteria generation, iteration steps, verification, exit flow. - Extension (
goal-driven-loop/index.ts) — automation layer. Detects goal-setting phrases via theinputevent, injects goal state into context viabefore_agent_start, performs mechanical checks (context usage, stale state), and replaces infinite-wait confirms with timed auto-defaults.
The extension eliminates the most common causes of user-interaction hang-ups while preserving safety valves (explicit abandon, criteria modification).
Scope
In scope:
- A Pi skill named
goal-driven-loop(SKILL.md+references/error-handling.md) - A Pi extension
goal-driven-loop/index.tsthat automates the loop - Trigger detection (Chinese, English, other language variants) via extension
inputevent - Anti-trigger rules (quotes, trivial tasks, re-entry) via extension
- Goal state file
.agents/goal/GOAL.mdwith frontmatter + criteria + progress log - Auto-generation of acceptance criteria when not provided
- Per-iteration retry with exponential backoff (10 retries, ~37min coverage)
- Granularity detection that delegates coarse goals to comet
- Deslop pass and clean exit
- Extension-driven state injection (resists context compaction)
Out of scope:
- Tools for editing/running code (agent uses existing Pi tools)
- Comet skill implementation (already exists, only referenced)
- Cross-session goal tracking beyond file persistence
- Multi-goal concurrent tracking (only one active goal at a time)
Architecture: Skill + Extension
User message: "设定当前目标为..."
↓
[Extension: input event] detects trigger
↓
[Extension: before_agent_start] injects goal state + skill into context
↓
[Skill: SKILL.md] instructs agent on loop steps
↓
[Agent] runs iteration, updates GOAL.md
↓
[Extension: before_agent_start] on next turn re-injects state
↓
[Skill: completion] deslop + cleanup + output summary
The extension handles all mechanical concerns (detection, injection, context checks, auto-default confirmations). The skill handles all judgment-based concerns (criteria quality, iteration work, verification, exit flow). This separation lets the agent focus on substantive work while the extension guarantees the loop never gets stuck waiting.
Design
1. Skill: Trigger and Anti-Trigger
Trigger phrases (extension detects via input event):
- Chinese: 设定当前目标为 / 当前目标是 / 把当前目标设为 / 我的目标是
- English: set current goal to / current goal is / my goal is / the goal is
- Other: semantically equivalent phrasings + concrete deliverable
Uncertainty rule: If uncertain whether a phrase is a trigger, check for: agent-internal-commitment language + concrete deliverable description. If both are present, treat as a trigger.
The phrase must be followed by a concrete goal description (non-empty, not a short fragment).
Anti-trigger conditions (extension checks mechanically):
-
Mechanical quote detection:
- Trigger phrase inside markdown quote block (
>) - Trigger phrase inside code block (```)
- Trigger phrase in same sentence as reported speech markers (he said / she said / they said)
- Trigger phrase inside markdown quote block (
-
Trivial task red flags (any one triggers):
- Message contains keywords: hotfix, tweak, fix a bug, small change, quick fix, 修个bug, 小改
- Goal description ≤ 10 chars AND no architecture/module/system/refactor/integration keywords
- Agent estimates < 5 min AND goal description contains no architectural keywords
-
Re-entry: If
.agents/goal/GOAL.mdalready exists, see Section 7 (extension behavior).
Granularity split:
| Signal | Fine-grained (this skill) | Coarse-grained (delegate to comet) |
|---|---|---|
| Estimated file count | ≤ 5 | > 5 |
| Cross-module coordination | Single concern | Multi-module |
| Design decisions | None or minor | Architectural |
| Single-session completion | Likely possible | Likely impossible |
Coarse-grained goal → extension automatically invokes comet-open with the goal as input. This skill exits.
Fallback: If comet unavailable (project has no openspec config), downgrade to multi-stage mode — auto-split goal into phases, verify each phase.
2. Skill: Goal State File
Path: .agents/goal/GOAL.md
Format:
---
goal: "<goal text>"
created: "<ISO 8601 timestamp>"
updated: "<ISO 8601 timestamp>"
iteration: <int>
status: in-progress | blocked | complete
---
# Goal: <goal text>
## Acceptance Criteria
- [ ] <criterion 1>
- [ ] <criterion 2>
- [ ] <criterion 3>
## Progress Log
### Iteration 1
- <work done>
- <verification results>
## Blocked
- <blocker description, or delete section if empty>
Lifecycle:
- Created when trigger fires and goal is fine-grained
- Updated after every iteration (criteria checkboxes, progress log, iteration counter, updated timestamp)
- Set to
blockedwhen 10 retries exhausted or stagnation detected - Set to
completewhen all criteria pass - Deleted after deslop pass at exit
3. Skill: Error Handling and Interruption
Retry strategy (automatic, no human intervention):
| Parameter | Value |
|---|---|
| Max retries per iteration | 10 |
| Initial delay | 30s |
| Backoff factor | ×2 |
| Max delay | 300s (5min) |
Backoff sequence:
| Retry | Delay | Cumulative |
|---|---|---|
| 1 | 30s | 0.5min |
| 2 | 60s | 1.5min |
| 3 | 120s | 3.5min |
| 4 | 240s | 7.5min |
| 5 | 300s | 12.5min |
| 6 | 300s | 17.5min |
| 7 | 300s | 22.5min |
| 8 | 300s | 27.5min |
| 9 | 300s | 32.5min |
| 10 | 300s | 37.5min |
Execution: Skill instructs agent to run bash -c "sleep <delay> && echo done" with timeout = delay + 30. If sleep is interrupted, treat as tool failure, enter retry.
Retry counter scope: Per-iteration. Resets at the start of each new iteration.
After 10 retries exhausted: status = blocked, write Blocked section, extension automatically injects a brief blocker summary and auto-attempts alternative approach (see Section 7). If still blocked after auto-attempt, then notify user via UI.
Exception classification:
| Exception | Detection signal | Recoverable? | Response |
|---|---|---|---|
| Tool execution failure | bash returns non-zero exit | Yes | Exponential backoff retry |
| LLM API exception | subagent timeout/error | Yes | Exponential backoff retry |
| Context tension | ctx.getContextUsage() > 0.8 of context window |
No | Extension saves state, injects continuation message on next turn |
| User interruption | User sends new message mid-loop | No | Extension auto-pauses, responds to user message, resumes goal on next trigger |
| Stale state | GOAL.md exists with status = in-progress | No | Extension prompts: continue / replace / abandon (timed auto-default: continue) |
Stagnation detection (dual signal — both required):
- Acceptance criteria passing count unchanged for 3 consecutive iterations, AND
- Progress Log last 3 iterations contain no progress keywords (found / fixed / implemented / added / refactored / verified)
When triggered, extension auto-injects a strategy-change prompt rather than blocking. After 2 auto-strategy-changes still stalled, fall back to user notification.
Continuation prompt format (fixed):
🔄 目标状态已保存至 .agents/goal/GOAL.md
当前进度:X/Y criteria 通过,第 N 轮迭代。
Red flags — these do NOT count as completion:
- "LLM call failed, goal should be considered complete" → retry or blocked
- "Tests don't run but code looks fine" → must have test evidence
- "Context full, let me summarize" → save state, extension auto-injects continuation
- "Too many retries, skip this criterion" → exhausted retries → blocked, do not skip
- "This criterion is too hard to verify, skip it" → requires criteria modification path (user or auto-suggestion)
Explicit abandon: User says "abandon / don't want it / cancel" → extension deletes GOAL.md, outputs brief summary, terminates loop.
Criteria modification: Agent may suggest criteria modifications. Two paths:
- Auto-modify safe adjustments (e.g., "fix typo in criterion wording", "split compound criterion into two") — agent may apply, log the change
- Material change (e.g., "remove criterion", "lower acceptance bar") — extension prompts user with timed auto-default: keep current
4. Skill: Core Loop Flow
Step 1: Trigger and routing (handled by extension input event)
- Extension detects trigger phrase
- Extension performs anti-trigger check
- Extension performs granularity check
- Coarse → extension invokes
comet-open; fine → continue
Step 2: Initialize goal state (extension + agent)
- Extension checks
GOAL.mdexistence (stale state) - If stale: extension prompts user with timed auto-default (continue)
- If user provided criteria → use directly
- If not → auto-generate (verifiability red line below), no confirmation needed unless material uncertainty
Acceptance criteria verifiability red line:
- ❌ "Implement rate limiting" → not verifiable
- ✅ "Rate limiter middleware mounts on auth routes, excess requests return 429" → verifiable
- ❌ "Code quality is good" → not verifiable
- ✅ "Related files lint clean, typecheck passes" → verifiable
Step 3: Execute iteration
- Read
GOAL.md, identify unpassed criteria - Pick the criterion whose completion does not require other unpassed criteria to be met first
- Implement work for that criterion
- Run relevant verification (tests, lint, typecheck)
- Record verification results in Progress Log
Step 4: Verify and update
- Check each criterion against fresh evidence
- Update checkboxes,
iteration+1,updatedtimestamp
Step 5: Completion check
- All criteria ✅ → Step 6 (exit)
- Some unpassed → stagnation check, then loop back to Step 3
Step 6: Exit
- Deslop pass on modified files
- Deslop failure → does not block, note in summary
- Delete
GOAL.md - Output completion summary
Mid-loop granularity upgrade: If execution reveals the goal is more complex than estimated, extension automatically invokes comet-open with current GOAL.md state. User is informed via UI notification, not asked for confirmation.
5. Skill: Comet Integration
- Skill hands off to
comet-openfor coarse-grained goals - Skill does not manage goals inside comet cycles
- Mid-loop upgrade from fine to coarse is supported
- Extension triggers the hand-off automatically
6. File Structure
.pi/
├── skills/goal-driven-loop/
│ ├── SKILL.md # Instruction layer (≤500 words)
│ └── references/
│ └── error-handling.md # Detailed retry/recovery reference
└── extensions/
└── goal-driven-loop/
├── index.ts # Extension entry point
├── trigger-detector.ts # Trigger phrase + anti-trigger logic
├── state-injector.ts # State injection + context checks
└── ui-prompts.ts # Timed confirm/select helpers
Skill: instructions the agent reads when active. Extension: TypeScript module that runs in Pi's runtime, handles mechanical automation.
7. Extension Behavior (Automation Layer)
The extension is the primary mechanism for eliminating user interaction hang-ups. It runs autonomously in Pi's event loop and never blocks the user.
7.1 Trigger Detection (input event)
On every user input:
- Parse the input text
- Run anti-trigger checks (Section 1 conditions)
- If trigger detected and goal is fine-grained:
- Read or create
GOAL.mdskeleton - Inject the skill into the next turn via
before_agent_start - Auto-parse user message to extract goal text and any criteria
- If criteria absent, let the skill auto-generate
- Read or create
- If trigger detected and goal is coarse:
- Invoke
comet-openviapi.sendUserMessage()with goal as input - Notify user via
ctx.ui.notify()(non-blocking toast)
- Invoke
- If no trigger detected, pass through normally
7.2 State Injection (before_agent_start event)
On every agent start:
- Check if
.agents/goal/GOAL.mdexists - If yes and
status = in-progress:- Read the file
- Inject summary into the system prompt or as a persistent message
- Inject format:
[goal-driven-loop] ACTIVE GOAL
{goal text}
Progress: {X}/{Y} criteria passed, iteration {N}
Next step: read .agents/goal/GOAL.md and continue iteration
- If
status = blocked:- Inject blocker details
- Inject "change strategy" prompt
- If
status = complete:- Inject completion summary and clear the message
This is the key mechanism for context-compaction resistance: even if the conversation is compacted, the extension rebuilds goal context from the file.
7.3 Context Tension Detection
In before_agent_start:
- Call
ctx.getContextUsage() - If
usage.tokens / contextWindow > 0.8:- Read
GOAL.mdto ensure state is current - Inject a "save progress now" directive into the system prompt
- The directive tells the agent: if this turn's work would lose goal state, finalize
GOAL.mdfirst, then output the continuation prompt
- Read
7.4 Stale State Resolution
When GOAL.md exists with status = in-progress and a new trigger is detected:
- Use
ctx.ui.select()with timeout: 10000ms - Options: "继续当前目标" (default) / "替换为新目标" / "放弃当前目标"
- After 10s timeout, auto-select "继续"
- The user's choice triggers the corresponding skill branch
This replaces the "wait forever for user" pattern with a 10-second auto-default.
7.5 Criteria Auto-Generation Quality Gate
In before_agent_start, after skill auto-generates criteria:
- Parse the generated criteria
- Flag any criterion that fails the verifiability red line
- For flagged criteria, inject a "tighten this criterion" prompt
- The agent must rewrite or remove the flagged criterion before writing
GOAL.md
7.6 Stagnation Auto-Recovery
When the skill detects stagnation (dual signal from Section 3):
- Extension auto-injects a "change strategy" prompt:
Stagnation detected. Current approach not making progress. Try:
- Different tool or library
- Refactor before retry
- Split the failing criterion into sub-steps
- This is delivered via
pi.sendMessage({ deliverAs: "steer" }) - After 2 consecutive stagnation auto-recoveries with no progress, the extension then notifies the user via UI
7.7 Mid-Loop Granularity Upgrade
When the agent signals (via criteria update or system message) that the goal is more complex than estimated:
- Extension reads current
GOAL.md - Extension invokes
comet-openviapi.sendUserMessage() - Extension notifies user via
ctx.ui.notify()(non-blocking) - No user confirmation required
7.8 Abandon Detection
If user input contains "abandon" / "cancel" / "不要了" / "放弃":
- Extension deletes
.agents/goal/GOAL.md - Extension injects a brief completion summary
- No further user interaction required
8. User Interaction Inventory
After the extension layer is in place, the remaining user interactions are:
| Interaction | Trigger | User action needed? | Auto-default? |
|---|---|---|---|
| Stale state resolution | New trigger, GOAL.md exists | Optional | 10s timeout → continue |
| Criteria verifiability failure | Auto-gen produces non-verifiable | No (auto-rewrite) | N/A |
| Stagnation | 3-round stall | No (2x auto-recover) | After 2 recoveries, notify |
| Mid-loop upgrade | Complexity discovered | No (auto-delegate) | N/A |
| Explicit abandon | User says "abandon" | Required (must be user) | N/A |
| Criteria material change | Agent suggests removal/bar-lowering | Optional | Timed auto-default: keep |
The user's "set and forget" goal is preserved: they can leave the session, return, and the loop has continued autonomously.
Testing Strategy
Following writing-skills TDD approach:
RED Phase — Baseline Behaviors
Without the skill + extension:
- Agent ignores goal-setting phrases (treats as ordinary message)
- Agent gives up on transient tool failures, asks user
- Agent loses goal state on context compaction
- Agent blocks forever on user confirmations
GREEN Phase — Skill + Extension Implementation
Skill layer (Tasks 1-2):
- Write
SKILL.md(≤500 words) - Write
references/error-handling.md
Extension layer (Tasks 3-7):
trigger-detector.ts— phrase + anti-trigger parsingstate-injector.ts—before_agent_startstate injectionui-prompts.ts— timed confirm/select helpersindex.ts— wiring events together- Behavioral scenarios verifying all extension behavior
REFACTOR Phase — Close Loopholes
- Test scenarios for each anti-trigger pattern
- Test that context compaction preserves goal state
- Test that stale state auto-defaults work
- Test that stagnation auto-recovery kicks in
- Test that 2x recovery failure escalates to user
Risks and Mitigations
| Risk | Mitigation |
|---|---|
| Agent rationalizes skipping criteria | Red flags list, criteria modification requires logged reason |
| Sleep timeout in Pi | bash timeout = delay + 30 |
| Context exhaustion loses goal | Extension injects state on every before_agent_start; mechanical 0.8 threshold |
| Deslop pass fails | Does not block completion, noted in summary |
| Granularity misclassification | Mid-loop auto-upgrade to comet, no user gate |
| Re-entry confusion | Stale state uses timed auto-default, no continuation heuristic |
| Extension bugs block the loop | Extension failures degrade gracefully — skill instructions still work without extension |
| User loses ability to intervene | Explicit abandon and material criteria change still user-controllable |
Open Questions
None at design time. All major decisions resolved during brainstorming.