原始内容
pi-ultraplan
A provider-agnostic /ultraplan pattern:
fire a heavy "planning" agent off to the background, keep your terminal free,
poll for a structured plan, gate it on human approval with an interactive dialog,
optionally execute it in an isolated worktree.
Works as a Pi extension with full TUI integration, slash commands, and automatic plan export. Also runs standalone via CLI and HTTP API.
The reference provider is custom model fairy-tales-deepseek-openai running
deepseek-v4-pro — but the engine not include it. Swap the provider,
keep the engine.
Architecture
trigger (CLI / API / Pi extension) ← swappable surface
│
▼
Dispatcher (dispatcher.ts) ← 100% generic engine
├─ idempotency lock (set before any await)
├─ detached poll loop + timeout
├─ status-guarded delivery (late polls can't resurrect killed jobs)
├─ orphan reaping on error
└─ event bus: ready / progress / phase / gate / approved / failed / killed
│
▼
AgentProvider (types.ts) ← the ONLY thing you implement per backend
│
▼
DeepSeek Provider (providers/deepseek/)
├─ chat.ts ← OpenAI-compatible transport + cost metering
├─ planners.ts ← multi-lens fan-out (pragmatist, adversary, scale, maintainer)
├─ synthesis.ts ← structured JSON reconciliation with parse-retry
├─ repair.ts ← fold verification critique into plan
└─ index.ts ← AgentProvider orchestrator (launch/poll/approve/reject/execute)
│
▼
file-backed store (store.ts) ← every state claim verifiable on disk
The ultra pipeline (plan mode)
- Fan out — N planners in parallel, each a different lens (
pragmatist,adversary,scale,maintainer).PI_PLANNERScontrols N (default 3, max 4). One planner failing doesn't sink the run; all-fail →failed. - Synthesize — reconcile the drafts into one structured
StructuredPlanJSON with parse-retry and prose fallback. - Adversarial verify — 4-lens panel (correctness, security, feasibility,
completeness). Tie or majority refute triggers repair. Configurable via
PI_VERIFIERSandPI_VERIFY_ROUNDS. - Repair — fold critique back into the plan, re-verify. Loops up to
PI_VERIFY_ROUNDS(default 2) before gating regardless. - Gate — interactive dialog: Approve & Execute, Approve only,
Reject. Plans exported to
~/Documents/ultraplan-plans/on approve.
Pi Extension
Load automatically (symlinked at .pi/extensions/) or via -e:
pi -e extensions/ultraplan.ts
Slash Commands
| Command | Description |
|---|---|
/ultraplan <prompt> |
Dispatch a plan — fan-out → synthesize → verify → gate |
/ustatus [suffix] |
List recent sessions, filter by sessionId suffix |
/uview <suffix> |
View full plan details for a specific session |
Tools (LLM-callable)
| Tool | Description |
|---|---|
ultraplan_plan |
Start planning — returns sessionId immediately, runs in background |
ultraplan_status |
Read persisted state of any session |
ultraplan_approve |
Approve a plan at the gate |
ultraplan_reject |
Reject with a note (feeds cross-run memory) |
ultraplan_execute |
Execute approved plan → worktree PR/patch |
ultraplan_stop |
Halt a running job |
ultraplan_eval |
Run benchmark suite with LLM-judge scoring |
ultraplan_audit |
Read append-only audit log |
TUI Surfaces
- Footer — live session tracker, store-reconciled against disk state
- Widget — progress stream above editor (
planning 0/3 → synthesizing → verifying) - Status line — active job count
- Gate dialog —
select()with Approve & Execute / Approve only / Reject - Plan export —
.mdfiles written to~/Documents/ultraplan-plans/on approve
CLI
# Launch — returns immediately, prints plan at the gate
npx tsx src/cli.ts plan "Add Redis rate limiting to POST /api/login"
# Execute an approved plan
npx tsx src/cli.ts execute <sessionId>
# Inspect persisted state
npx tsx src/cli.ts status <sessionId>
# Human gate
npx tsx src/cli.ts approve <sessionId>
npx tsx src/cli.ts reject <sessionId> "use sliding window, not fixed"
# Cancel a running job
npx tsx src/cli.ts stop <sessionId>
# Run benchmark eval
npx tsx src/cli.ts eval benchmarks/sample.json
# Read audit log
npx tsx src/cli.ts audit
# HTTP API (SSE + web UI)
npx tsx src/cli.ts serve 8080
Swapping providers
Implement AgentProvider (see src/types.ts) against any backend — OpenAI
Assistants, Gemini, a queue+worker pod — and pass it to new Dispatcher(...).
Nothing else changes.
Env
| var | default | meaning |
|---|---|---|
PI_BASE_URL |
http://localhost:8000/v1 |
OpenAI-compatible base |
PI_API_KEY |
— | bearer token |
PI_MODEL |
deepseek-v4-pro |
default model |
PI_PLANNERS |
3 |
fan-out width (1–4) |
PI_VERIFIERS |
3 |
verification panel size (1–4) |
PI_VERIFY_ROUNDS |
2 |
max verify→repair loops (1–3) |
PI_MAX_TOOL_STEPS |
12 |
max tool calls per planner |
PI_COST_CEILING_USD |
— | hard cost cap per run |
PI_COST_JSON |
— | per-model pricing overrides |
PI_FALLBACK_MODELS |
— | comma-separated fallback chain |
PI_PLAN_DIR |
~/Documents/ultraplan-plans |
plan export directory |
PI_ALLOW_UNSOUND_EXEC |
— | set to 1 to execute plans that failed verification |
PI_SESSION_DIR |
~/.pi/ultraplan/sessions |
where session records persist |
Adding a Command
Commands live as individual modules in src/commands/. Each module exports a
handler with the signature (args: string[], flags: ParsedFlags) => Result | Promise<Result>.
Quick start
- Create the handler at
src/commands/mycommand.ts:
import { ok, fail } from '../result.js';
import type { Result } from '../result.js';
import type { ParsedFlags } from '../router.js';
export function myCommand(args: string[], flags: ParsedFlags): Result<any> {
if (args.length === 0) return fail('usage: mycommand <arg>');
return ok(`Got: ${args.join(' ')}`);
}
- Export from the barrel in
src/commands/index.ts:
export { myCommand } from './mycommand.js';
- Register in the router in
src/cli.ts:
router.register('mycommand', (a, f) => commands.myCommand(a, f),
'Describe what it does', 'mycommand <arg>');
- Write tests in
tests/— unit tests for the handler logic, and a golden snapshot integration test intests/smoke.test.ts. Add the import tosrc/commands/index.tsand register the command insrc/cli.ts.
If your command needs shared dependencies (e.g., the Dispatcher), inject them
through a deps object passed in the closure at registration time (see
planCommand or approveCommand for examples).
Result contract
Every command returns a Result<T> — a discriminated union:
| Variant | Shape | Meaning |
|---|---|---|
| Success | { ok: true, data?: T } |
Command succeeded; data is written to stdout |
| Failure | { ok: false, error: string } |
Command failed; error is written to stderr |
Use the factory functions:
ok(data?)— successfail(message)— failure- Throw
new CliError(message, exitCode)for early exits (caught by the bootstrap)
Streaming output
For commands that stream progress bars or incremental logs, write to
stderr (console.error). Only the final result data should be returned
through the Result. This keeps --json mode clean: progress goes to stderr, the
structured result goes to stdout.
The --json Flag
The --json flag can appear anywhere in argv:
npx tsx src/cli.ts --json help
npx tsx src/cli.ts plan --json "Add rate limiting"
npx tsx src/cli.ts status --json <sessionId>
When --json is present:
- Success output is a JSON object on stdout:
{ "ok": true, "data": ... } - Failure output is a JSON object on stderr:
{ "ok": false, "error": "..." } - Exit code is 0 for success, 1 for failure
Without --json, output is plain text: success data is printed to stdout,
error messages are printed to stderr.
The help Command
help is a first-class built-in command:
# List all commands
npx tsx src/cli.ts help
# Detailed usage for a specific command
npx tsx src/cli.ts help plan
# Machine-readable command list
npx tsx src/cli.ts --json help
The help output is automatically generated from the command registrations.
When a command is registered with a description and usage string, those
appear in the help text automatically.
Testing
The project uses vitest (v4) as its test runner with three categories of tests:
Unit tests (tests/*.test.ts)
Pure function tests for individual modules with no I/O:
| File | Tests | What it covers |
|---|---|---|
tests/pure.test.ts |
29 | extractJson, parseStructuredPlan, parseVerdict, parseScore, redact, redactObject, renderPlan |
tests/result.test.ts |
11 | Result discriminated union, ok()/fail() factories, CliError |
tests/router.test.ts |
15 | Command registration, dispatch, unknown command handling, generateHelp() |
tests/output.test.ts |
6 | writeOutput formatting (plain + JSON mode), formatUnexpectedError |
Integration / smoke tests (tests/smoke.test.ts)
Run the CLI as a child process (npx tsx src/cli.ts ...) and capture stdout,
stderr, and exit code. Tests exercise the full pipeline from argv parsing
through dispatch to output formatting.
Golden snapshots (tests/fixtures/golden/) establish a behavioral baseline
that cannot regress:
# Run all tests
npm test
# Update golden snapshots after intentionally changing output
UPDATE_GOLDEN=1 npx vitest run tests/smoke.test.ts
Test helper (tests/helpers.ts)
The runCli(args) helper spawns the CLI as a child process and returns
{ stdout, stderr, exitCode }. Use it in integration tests to assert
complete behavioral fidelity:
import { runCli } from './helpers.js';
const result = runCli(['help']);
expect(result.exitCode).toBe(0);
expect(result.stdout).toContain('Available commands');
CI
A GitHub Actions workflow (.github/workflows/ci.yml) runs on every push and
pull request to main:
- TypeScript type check (
npx tsc --noEmit) - Full test suite (
npm test) - Golden snapshot verification
- Tested against Node 18, 20, and 22