ultra-builder-pro

内容来源:README.md(说明文档) · 原始地址 · 查看安装指南

原始内容

Ultra Builder Pro

English · 简体中文

A production-grade engineering harness for Claude Code.

Six-command workflow · Sensor-driven hooks · Multi-agent review · Layered cross-session memory · Live project KB · Three-way AI verification.

Version Tests Hooks Agents Skills License

git clone https://github.com/rocky2431/ultra-builder-pro.git ~/.claude

Works wherever Claude Code does. macOS · Linux · Windows.

Why I Built This · What's Inside · How It Works · Why It Works · Architecture · Changelog


Why I Built This

I'm a Web3 + AI-native engineer. I don't write code — Claude Code does.

But getting code written and getting code to production are two different things. When you need:

  • Real code review, not "looks fine to me"
  • Real tests passing, not mocks pretending to be green
  • Specs that survive 50 rounds of drift, not silent semantic decay
  • Context that persists across sessions, not yesterday's decisions vanishing overnight
  • Cross-AI sanity checks, not one model marking its own homework

— LLM-by-conversation isn't enough. It forgets. It silently scope-reduces. It mocks core paths and calls them tested. It says "VIP free shipping" while implementing a 50% discount.

Ultra Builder Pro is the engineering harness that sits on top of Claude Code and makes it production-grade. Not one tool — an integrated substrate of six layers:

  1. Spec-driven 6-command workflow that takes you from idea to release
  2. Sensor-driven 15-hook chain that informs without obstructing (v7.0 doctrine)
  3. 7-agent parallel code review pipeline with semantic-drift detection
  4. Layered cross-session memory — claude-mem raw timeline + curated file-based knowledge
  5. Live bidirectional task ↔ code ↔ spec knowledge base (v7.1)
  6. Three-way AI verification (Claude + Gemini + Codex with consensus scoring)

It doesn't think for you. It keeps Claude honest while you do.

Other systems give you parts of this. BMAD has the workflow. claude-mem tried memory and exploded. GSD nailed the spec-driven discipline. Ultra Builder Pro is the integrated assembly — every piece reinforces the others.

rocky2431


Who This Is For

People shipping real products with Claude Code who want:

  • Production discipline by default — TDD enforcement, parallel review, atomic commits per task
  • A system that remembers — across files, tasks, sessions, AI providers
  • Drift caught early — when "free shipping" silently becomes "5% off", not after deploy
  • No theater — 9 commands instead of 30; no sprint planning, no story points, no Jira

If you want heavy enterprise process, use BMAD. For pure planning, use Speckit. For lighter context engineering, use GSD. Ultra Builder Pro is the maximalist option — pick it when you want every layer integrated.


What's Inside

Built on Claude Code. Adds six layers of production-engineering discipline.

1. Spec-Driven 6-Command Workflow

/ultra-init  →  /ultra-research  →  /ultra-plan  →  /ultra-dev  →  /ultra-test  →  /ultra-deliver

Each command does one thing well; the system handles complexity behind the scenes. Behind it: TDD red-green-refactor enforcement, automatic git flow with atomic per-task commits, walking-skeleton-first task ordering, integration checkpoints every 3-4 tasks.

/ultra-research is itself a 17-step-file architecture (BMAD-inspired) — each step has dense instructions, pre-written web search queries, structured output templates, and write-immediately discipline. Output: discovery.md, product.md, architecture.md, plus a token-efficient research-distillate.md for /ultra-plan.

→ See commands/ultra-*.md and docs/architecture.md.

2. Sensor-Driven Hook Harness (15 hooks)

v7.0 doctrine: blocks reserved for truly irreversible actions only. Hardcoded secrets, SQL injection, force-push to main, DB migration commits. Everything else — mocks, scope reduction, silent catches, TODO/FIXME, default-off feature flags — is advisory. The agent reads, decides, proceeds.

This inverts the pre-v7 over-correction loop, where blocked agents would silently edit tests and specs to escape blocks (worse outcome than no hooks). Sensor mode gives signal without distortion. Hooks now also inject goal context at decision time: when you start to edit a file, mid_workflow_recall.py injects the active task's acceptance criteria.

→ Hook table: docs/architecture.md#hooks-system.

3. Multi-Agent Code Review Pipeline (7 specialists, parallel)

Sequential reviews lose context as findings accumulate. /ultra-review fans out 7 specialists in parallel — each in a fresh 200k context — and a coordinator dedupes and prioritizes. Main session stays at 30-40% context usage even during deep review.

Specialists:

  • review-code — security, SOLID, forbidden patterns, scope drift
  • review-tests — mock violations, coverage gaps, boundary tests
  • review-errors — silent failures, swallowed errors, empty catches
  • review-design — type design, encapsulation, complexity
  • review-comments — stale, misleading, low-value comments
  • review-ac-drift (v7.1) — semantic alignment: reads spec text + diff together, catches "VIP free shipping → 5% off" structural lints can't see
  • review-coordinator — aggregate, deduplicate, generate SUMMARY

Verdict logic: P0 > 0 → REQUEST_CHANGES; P1 > 3 → REQUEST_CHANGES; P1 > 0 → COMMENT; else APPROVE. Branch-scoped session index with iteration chains for re-checks.

→ See skills/ultra-review/SKILL.md.

4. Layered Cross-Session Memory

Three non-overlapping layers replace the old self-built SQLite store (retired 2026-06):

  • L3 raw observations — the claude-mem plugin captures the timeline. At session start it injects a recent slice; query the rest on demand via its MCP tools (smart_search / observation_search / timeline).
  • L3 refined knowledgefile-based memory (projects/.../memory/: MEMORY.md + typed facts), curated by hand, injected at session start. This is the substrate for self-improvement.
  • L2 continuitysession_context.py injects pure git + active-goal state; historical_context_guard.py fences all injected history as reference only, not live instructions so a resume never re-runs a finished task.

The earlier self-built memory.db + /recall + skills/learned/ overlapped claude-mem and were removed.

→ See CHANGELOG v7.2.

5. Live Project Knowledge Base (v7.1)

Bidirectional task ↔ code ↔ spec index in .ultra/relations.json. Auto-derived wiki views in .ultra/wiki/{index,log}.md. Session facts folded into task contexts as a ## Session Trail section. Sessions without an active task still leave residue in .ultra/sessions/orphan-trail.md.

Three-layer architecture:

Layer 3 — Schema (immutable):     PHILOSOPHY.md, CLAUDE.md, harness rules
Layer 2 — Wiki (interpretation):  wiki/{index,log}.md, ## Session Trail, orphan-trail
Layer 1 — Facts (machine-kept):   relations.json, progress/*.json, git history

Wiki nodes never store facts; they only store interpretation. No silent staleness — when facts change, wiki regenerates. Edit a file owned by a task → stderr shows the task + first AC. Edit an unowned file → git context fallback (branch + last commit). Edit a non-Ultra project → silent.

→ See CHANGELOG v7.1.

6. Three-Way AI Verification

/ultra-verify spawns Claude + Gemini + Codex independently as parallel background tasks. Claude writes its answer to a file before reading the others — preventing contamination. Then Claude reads all three outputs and synthesizes with confidence scoring:

Outcome Confidence
3/3 agree Consensus
2/3 agree Majority
All differ No Consensus

Four modes: decision (architecture choices), diagnose (bug hypotheses), audit (code review), estimate (effort). Degrades gracefully — one AI fails → two-way capped at Majority; two fail → Claude-only with explicit warning.

Built on a shared ai-collab-base skill with synced collaboration protocol files. Eliminates ~90% structural duplication between gemini-collab and codex-collab.

→ See skills/ultra-verify/SKILL.md.


What's New — v7.1.0

Dynamic Project Knowledge Base — five additions on top of the v7.0 sensor-first foundation:

  • File→task reverse trace with git-context fallback (post_edit_guard.py)
  • 7th review specialist review-ac-drift for semantic alignment
  • Auto-derived wiki views with Recent Activity table
  • Session trail fold-back into task context
  • Orphan session handling for cross-task / no-plan / hotfix work

→ Full version history: CHANGELOG.md.


Getting Started

git clone https://github.com/rocky2431/ultra-builder-pro.git ~/.claude
cd ~/.claude && python3 -m pytest hooks/tests/
# Expected: 164 passed

In any project:

cd ~/your-project && claude

In Claude Code, run /ultra-init to initialize. Verify with /ultra-status.

Recommended: Skip Permissions Mode

claude --dangerously-skip-permissions

The harness is designed for frictionless automation. Stopping to approve git commit 50 times defeats the purpose. Granular alternative: whitelist specific commands in .claude/settings.json permissions.allow.


How It Works

The 6-command pipeline, end to end. Each command outputs files the next one consumes.

Step Command What happens Outputs
1 /ultra-init Auto-detect project type; copy templates from .ultra-template/; set up .ultra/ directory .ultra/{specs,tasks,docs}/, PHILOSOPHY.md, north-star.md
2 /ultra-research 17-step research pipeline; mandatory web search per step; structured output templates discovery.md, product.md, architecture.md, research-distillate.md
3 /ultra-plan Atomic task breakdown; mode (EXPAND/SELECTIVE/HOLD/REDUCE); walking skeleton first; integration checkpoints tasks.json, contexts/task-N.md per task
4 /ultra-dev TDD red-green-refactor; goal-always-present AC injection; per-task atomic commit; runs /ultra-review all at step 4.5 implementation, tests, commits
5 /ultra-review 7-agent parallel review; coordinator dedupes; SUMMARY.json + .md .ultra/reviews/<session>/SUMMARY.{json,md}
6 /ultra-deliver Pre-flight tests + review verdict APPROVE; CHANGELOG; version bump; tag; push release artifacts, git tag

Standalone gates (any time): /ultra-status, /ultra-verify, /ultra-test, /ultra-think. Memory: claude-mem MCP tools for cross-session recall; refined knowledge curated by hand in file-based memory.


Why It Works

Sensor-Not-Blocker Philosophy (v7.0)

Pre-v7 hooks blocked on every recoverable issue — agents responded by editing tests/specs to escape, worse than no hooks. v7 inverts: blocks for irreversibility only, advisories for everything else. The agent has the signal and the autonomy. PHILOSOPHY.md C3 (Sensors not Blockers) + C4 (Incremental Validation) codify this.

Bidirectional Traceability

Most tools maintain spec → task. Ultra Builder maintains all three:

  • task → spec section (trace_to)
  • spec section → tasks (referenced_by)
  • code path → tasks (files reverse map, v7.1)

Edit src/checkout/shipping.ts → system knows it's task-3 → traces to specs/product.md#vip-shipping → has 2 ACs. Visible in stderr the moment you edit.

Parallel Multi-Agent Review (Zero Context Pollution)

Sequential reviews lose context. Ultra Review fans out 7 specialists in parallel, each in fresh 200k context, writes findings to JSON files; coordinator dedupes. Main conversation never sees raw findings — only the deduplicated SUMMARY. Context usage stays at 30-40% even after a 7-way review.

Layered Memory, Clear Ownership

Three layers with non-overlapping jobs: claude-mem owns the raw timeline (recent slice injected at start, rest queried on demand), file-based memory owns curated knowledge (the self-improvement substrate), and session_context.py + historical_context_guard.py own continuity (pure git/goal state, fenced as reference-only so a resume never re-runs finished work). The old self-built SQLite + vector store was removed once it became a redundant duplicate of claude-mem.

Three-Way AI as Independent Reviewers

A single LLM marks its own homework. Three LLMs from different families (Anthropic, Google, OpenAI) provide actual independence. Confidence emerges from consensus, not assertion.


Commands & Skills

8 commands under commands/:

Family Commands
Workflow ultra-init ultra-research ultra-plan ultra-dev ultra-test ultra-deliver
Quality ultra-status ultra-think

17 skills under skills/:

Category Skills
Workflow ultra-research (17 step-files), ultra-review, ultra-verify
AI Collab ai-collab-base, gemini-collab, codex-collab
Agent-only checklists code-review-expert, integration-rules, security-rules, testing-rules
Utility agent-browser, find-skills, use-railway, market-research
Design / Output web-design-guidelines, guizang-ppt-skill, html-ppt
Vercel best practices vercel-react-best-practices, vercel-react-native-skills, vercel-composition-patterns

9 agents under agents/:

Type Agents
Interactive code-reviewer, debugger
Review pipeline (parallel) review-code, review-tests, review-errors, review-design, review-comments, review-ac-drift, review-coordinator

All agents have memory: project for per-project pattern accumulation.

→ Full reference: docs/architecture.md.


Configuration

Project state in .ultra/ (per-project, mostly gitignored). Global config in ~/.claude/settings.json.

Recommended .gitignore

.ultra/memory/
.ultra/reviews/
.ultra/compact-snapshot.md
.ultra/debug/
.ultra/workflow-state.json
.ultra/sessions/orphan-trail.md

Keep these in version control:

  • .ultra/specs/
  • .ultra/tasks/tasks.json, .ultra/tasks/contexts/
  • .ultra/relations.json
  • .ultra/wiki/{index,log}.md (auto-generated; useful for code review)

Sensitive File Protection

Add to ~/.claude/settings.json:

{
  "permissions": {
    "deny": [
      "Read(./.env)", "Read(./.env.*)",
      "Read(./secrets/**)", "Read(./**/*credential*)"
    ]
  }
}

Troubleshooting

Symptom Fix
Tests fail after install Run python3 hooks/system_doctor.py for deep audit
Hooks not firing Check ~/.claude/settings.json has hooks section; restart Claude Code
Stale wiki Edit any file in .ultra/specs/ or .ultra/tasks/ to retrigger; or run python3 ~/.claude/hooks/wiki_generator.py /your/repo
relations.json dangling trace_to Run /ultra-status; broken traces are highlighted
ultra-verify Gemini/Codex unavailable Install: npm i -g @google/gemini-cli @openai/codex; degrades to Claude-only with warning

License

MIT. See LICENSE.


Claude Code is powerful. Ultra Builder Pro makes it production-grade.