
forge: AI delivery pipeline for Claude Code
An open-source (MIT) plugin for Claude Code and Antigravity that turns a backlog into merged, reviewed, gated pull requests: 20 pipeline skills, a 12-role agent roster, GitHub Projects automation, mechanical ship gates, a learning loop and a graph-RAG code index. Built by its own pipeline.
About the project
forge installs a complete software-delivery pipeline into an AI coding host. You describe work as tickets on a GitHub Projects board; forge triages, plans, writes, reviews and gates each change, then opens a pull request for you to merge. Every lifecycle step is recorded on the ticket that drives it, so nothing happens as silent side-work.
It runs inside Claude Code and, through a generated native plugin package, on Antigravity. Alongside the plugin sits a cockpit: a local web app for the self-hosted GitHub Actions runners forge uses, with fleet control, logs, Claude usage and cost, machine load and a real terminal.
forge was built by its own pipeline: about 575 tickets and pull requests from July to September 2026, released from v0.x to v1.4.1, and open-sourced under MIT at v1.0 with zero non-permissive dependencies (checked by its own licence gate).
The landing-page shots are of the real site. The cockpit runs here with demo runner, usage and log data instead of its backend; its terminal replays a real run of forge's test suite.
Features
- Pipeline skills in 5 lanes: front of pipeline (ideate, brainstorm, spike, design, shape), build (plan, execute, deliver, ship, release), care (hotfix, respond, maintain), knowledge (distill, review, investigate) and scale (autopilot, triage, board).
- Agent roster: 12 role subagents (planner, scoper, implementer, test-architect, reviewer, security, designer, librarian and more), each with a role card and a narrow tool set.
- Board automation on GitHub Projects with a ticket-trail rule: every lifecycle moment is written back to the driving issue.
- Mechanical ship gates run as scripts: acceptance criteria proven by tests, plan drift, doc sync, test intent, dependency and licence policy, conventions and situation locks, plus model-driven review and security passes.
- Learning loop and graph RAG: lessons distilled from a work journal, and a SQLite code graph that answers reuse and blast-radius questions before new code is written.
- Autopilot: works the board ticket by ticket until it is clear.
- Cockpit: a local web app for the runner fleet (start, stop, re-provision, logs), Claude usage and cost, machine load, and a terminal.
Screenshots
Architecture
- A Claude Code plugin (skills, commands, agents, hooks, monitors) with a host-agnostic Node.js engine, so the same board automation, gates, graph RAG and learning loop run on Antigravity through an emitted native plugin package and MCP servers.
- Gates are plain Node scripts that fail closed. CI runs the test suite, a strict plugin validation, the licence gate, a workflow lint and a secret scan.
- Graph RAG indexes the codebase with ts-morph into SQLite via Node's built-in
node:sqlite(no native build step). - A denylist hook blocks a fixed set of known-catastrophic commands as a backstop.
- The cockpit is a FastAPI server bound to
127.0.0.1with a per-session capability token, wrapping framework-agnostic Python cores; its browser UI uses xterm.js over a WebSocket to a real PTY (ConPTY on Windows).
Challenges & what I learned
- Making the pipeline auditable, not just automated: every step has to leave a trail on the ticket, and a gate's "no" has to be a script result, not an opinion.
- Running one engine on two AI hosts while keeping unattended auto-merge on the host where the full safety stack is proven.
- Becoming licence-clean for the MIT release: the first cockpit was a Qt (LGPL) desktop app with a hand-rolled terminal, re-architected into a local web app with xterm.js.
- Keeping skill instructions within size limits as features grew, by moving history and edge cases out of the main skill files.
- An AI pipeline is only as trustworthy as its mechanical checks; reviews by models are valuable but advisory.
- Building the tool with itself (dogfooding every ticket) surfaced real gaps faster than any test plan.
- Writing decisions down as ADRs kept a fast-moving, agent-heavy project understandable.
