Two Months with GitHub's Spec-Kit and Claude Code: What Works, What Doesn't

✍️ OpenClawRadar📅 Published: May 15, 2026🔗 Source
Two Months with GitHub's Spec-Kit and Claude Code: What Works, What Doesn't
Ad

After two months of using GitHub's spec-kit for Spec-Driven Development (SDD) with Claude Code as the primary agent, a developer on r/LocalLLaMA reports on what works and what doesn't. The toolkit, available at github.com/github/spec-kit, enforces a five-phase workflow: Constitution, Specify, Plan, Tasks, Implement. The core idea: the spec, not the prompt, is the source of truth.

What's Actually Good

  • Agent-agnostic: Same spec works with Claude Code, Cursor, Codex, Gemini CLI, Copilot. The author generated code with Claude Code, then handed the spec to Cursor for test refactoring seamlessly.
  • Hard checkpoints between phases: The Plan phase shows the full proposed architecture before any code is written, catching bad decisions at a 5-minute fix cost instead of 5 hours.
  • Constitution file as quality gate: You define inviolable rules up front — test coverage minimums, dependency allowlists, perf budgets, typing strictness. The agent fails its own validation if it tries to violate them.
  • Improved determinism: Re-running the implement phase produces more consistent output than raw prompting, since the agent isn't filling in 30 implicit decisions.
Ad

What Annoys

  • Drift is real: Manual code edits without updating the spec cause fast desync. spec-kit has tooling but it's early.
  • Overhead for small changes: Bug fixes <50 LOC or trivial features feel ceremonial. The author's rule: only full SDD for new modules or features touching 200+ LOC.
  • Legacy migration painful: Retrofitting SDD onto a 30k-LOC codebase takes months.
  • Quality depends on agent: Claude Code (Sonnet/Opus 4.6+) handles it well; smaller models generate plans that compile but lack architectural reasoning.

Practical Setup

  • Install: uv tool install --from git+https://github.com/github/spec-kit.git specify-cli. Only the official repo is safe — PyPI has typosquatters.
  • Primary agent: Claude Code, with cross-validation on Cursor and Gemini CLI.
  • Local persistence: SQLite (easy to spec/validate, no cloud dependency).
  • Reusable constitution template: strict typing, pytest coverage >80%, explicit dependency allowlist, no cloud services unless required.

Open Questions

  • Can local models (Qwen, DeepSeek-Coder, GLM, Llama) handle Plan and Implement competently? The author found small models follow format but architectural reasoning fails.
  • Does multi-agent SDD work? Spec by one model, implement by another, audit by a third — theoretically better, but not measurably better than single-agent in practice.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

LLMs Leak Reasoning into Structured Output Despite Explicit Instructions
Tools

LLMs Leak Reasoning into Structured Output Despite Explicit Instructions

A developer building a tool that makes parallel API calls to Claude and parses structured output found that validation models intermittently output reasoning text before corrected content, despite explicit instructions to return only corrected text. The fix involved prompt tightening plus a defensive strip function that runs before parsing.

OpenClawRadar
ACO System: Multi-Agent AI Pipeline from GitHub Issue to Merged PR
Tools

ACO System: Multi-Agent AI Pipeline from GitHub Issue to Merged PR

ACO System is an open-source multi-agent framework where six specialized AI agents autonomously run the entire dev pipeline from GitHub Issue to merged PR, with a deterministic Architect gate that rejects bad stories before they reach developers.

OpenClawRadar
Hubcap Bridge: Persistent Two-Way Messaging Between CLI and Browser JavaScript via CDP
Tools

Hubcap Bridge: Persistent Two-Way Messaging Between CLI and Browser JavaScript via CDP

Hubcap Bridge is a new feature in the Hubcap CLI tool that creates a persistent two-way message channel between local processes and JavaScript running in browser pages via the Chrome DevTools Protocol. It enables Claude Code skills to interact with web apps through their internal JavaScript APIs without requiring public API access.

OpenClawRadar
Holaboss AI Runtime Moves to TypeScript, Implements Persistent MCP Ports
Tools

Holaboss AI Runtime Moves to TypeScript, Implements Persistent MCP Ports

The Holaboss AI local agent runtime has been refactored to use TypeScript exclusively, eliminating Python dependencies and reducing bundle size. It now persists MCP server ports in SQLite with UNIQUE(port) constraints to prevent collisions across restarts.

OpenClawRadar