TLTan LeBackend & automation · CalgaryLet’s talk
← All posts

Two Features, Two Days: What Spec-Driven Development with Spec Kit Taught Me

Shipping two real features to production with Claude Code and GitHub's Spec Kit, and where production proved the spec wrong.

For most of this year my way of working with an AI coding agent was simple: I wrote a planning document, asked the agent for step 1, reviewed it, then asked for step 2. It worked, but everything important lived in my head and in a chat history that disappeared when the session ended. This month I tried the opposite approach. I adopted Spec Kit, GitHub’s open-source toolkit for spec-driven development, and used it with Claude Code to ship two real features into production.

People have started calling this kind of workflow an ADLC, an agentic development lifecycle: the classic spec → design → build → verify loop, redrawn so that an agent does most of the writing and a human owns the decisions. This post covers what that looked like over two days. It covers what worked, what surprised me, and where production proved the spec wrong.

The setting

The project is a domain-name intelligence SaaS: an ASP.NET Core 8 API and a React 19 front end, in two repositories, on top of Azure SQL databases that are shared with an independent bot. I am effectively the only engineer. The front end deploys through CI/CD. The API could too, but with a single environment I keep its pipeline switched off and deploy it by hand, because I want to control exactly when it goes out. And a bad SQL statement against a shared table cannot be undone. That last constraint shaped everything below.

The features were two pages of a new “Drops” module:

  • Pending Drops: a searchable list of domains scheduled to be deleted by the registry, with filters on name, drop date and TLD.
  • Available Drops: a live grid of the names confirmed available on drop day, refreshing every minute, with the same filters, sorting, export and data gating as our main domain grid.

How Spec Kit works, briefly

Spec Kit installs a set of slash commands into your agent, plus templates and a few scripts. Each command writes a specific document into specs/NNN-feature-name/:

  1. /constitution: the project’s non-negotiable rules, written once and versioned.
  2. /specify: the what and why: user stories with priorities, functional requirements, measurable success criteria, edge cases, assumptions, out of scope.
  3. /clarify (optional): the agent asks up to five targeted questions and writes the answers back into the spec.
  4. /plan: the how: technical context, research decisions, data model, API contracts, a quickstart for verification, and a gate check against the constitution.
  5. /tasks: a dependency-ordered task list grouped by user story, so each story can ship on its own.
  6. /analyze (optional): a read-only consistency check across spec, plan and tasks.
  7. /implement: the agent works through the task list and ticks items off.

That is the whole mechanism. What matters is how you use it.

Start with the constitution. It is the best part.

Before writing a single spec I spent an evening on the constitution, and it turned out to be the most valuable hour of the whole experiment. It is not a style guide. It is a list of rules the agent is required to check its own plan against, each with a short rationale. A few of mine:

  • The agent never runs schema or data changes against a database. Every DB change is delivered as an idempotent .sql file with a header, for me to run by hand in SSMS, dev first, then prod. Every plan must state which scripts run before the deploy, which run after, and which only verify.
  • Time is stored and compared in UTC. Display conversion goes through a single utility on the front end. We had already been hurt by this when the registry changed its drop hour.
  • Feature keys are stable contracts. The string in an endpoint’s feature gate is the same key in the database, the token and the UI. You never rename a released key, and a new key ships with its seed script.
  • “It builds” does not count as tested. UI changes need a golden-path run in a real browser, with the console and network tab checked.
  • The neighbouring bot is out of scope. It may only be influenced through shared tables, via a script that states its effect on the bot.

Every plan.md then contains a Constitution Check table: one row per principle, how the plan satisfies it, PASS or an exception I have to approve. In practice this is where the agent explains itself. For example, it noted that a date-only column is rendered without a timezone conversion on purpose, “recorded explicitly so the PR reviewer does not flag it”. I cannot get that kind of pre-emptive explanation out of a chat prompt.

The constitution is also a living document. On day two I added a rule that the agent must never write to our project-management tool. It went through the amend command, was versioned 1.0.1 → 1.1.0 with a sync report, and the tasks and checklists that referenced the old behaviour were updated the same day. Treating process rules as versioned code was new to me, and I will keep doing it.

The loop, with real numbers

Feature one (Pending Drops) went from an empty folder to a merged pull request across one evening and the following morning. Feature two (Available Drops) took a single morning. Its spec, plan, tasks and implementation commits all landed within about an hour, because feature one had already done the foundational work.

  • Pending Drops: 3 user stories, 27 tasks in 6 phases (28 after analysis), 16 new unit tests, full suite 91/91 green.
  • Available Drops: 4 user stories, 22 functional requirements, 9 success criteria, 33 tasks, full suite 152/152 green.
  • Documentation: about 2,600 lines of Markdown across the two spec folders (spec, plan, research, data model, contracts, quickstart, tasks, checklists).

That last number is not a vanity metric. It is the real cost of the approach, and I come back to it below.

What worked

/analyze caught a real bug before any code existed

I almost skipped the optional analysis step. On feature one it reported 0 critical, 2 high, 3 medium and 4 low findings, with 100% requirement-to-task coverage. One of the high findings was a genuine defect in the design. The research doc had chosen SQL Server bracket escaping ([%]) for the “domain contains” filter, but the in-memory database provider used by the unit tests does not understand bracket escaping. The tests would have tested something different from production. The fix, an explicit escape character, went into the research decision and the affected tasks before implementation started. The other high finding was simpler but just as practical: no task created the feature branches in either repository.

A requirement justified a refactor that “minimal change” would have blocked

My constitution says to make minimal changes and not refactor unrelated code. Feature two needed the same filters, sorts, export and gating as our main grid, and the obvious shortcut was to copy about a thousand lines. Instead, the spec stated it as a requirement (one shared implementation, so a filter added to one grid is available on both). The plan then made the refactor explicit and scoped: lift the generic grid query and presenter code into shared helpers, and leave per-organisation logic where it was. The tasks put that extraction in its own foundational phase, backed by “dual-entity” tests proving both grids produce the same results. The refactor happened because a requirement asked for it and a test guarded it, not because the agent felt like tidying up.

Artifacts that outlive the chat

The documents I use most turned out not to be the spec. They were the quickstart and the deploy checklist. The quickstart contains the verification SQL and, after the dev run, the expected counts. When the API returned exactly the same row counts as the SQL (536,603 pending names on the dev snapshot), I knew the filters were right without having to trust anyone’s word, including the agent’s. The deploy checklist lists every deploy step, in the constitution’s fixed order: scripts → API → UI → verify, dev then prod. A week from now, neither of us will remember any of this, and the folder will.

Human checkpoints were built into the task list

Some tasks were simply blocked until I did something: run the seed script on dev, or walk the golden path in a browser. The agent implemented everything else, marked those tasks open, and stopped. That is how I want an agent to behave around shared databases. It should not hunt for credentials or reset a test password. It should tell me it is waiting.

Where reality won

A spec is only as good as its assumptions

Both specs came back with “no clarification needed”, and sensible defaults were recorded under Assumptions. I accepted that and skipped /clarify both times. That was a mistake. Two of those bets went wrong within 24 hours of launch:

  • An assumption carried over from my earlier planning doc said a drop day would move “a few thousand” names into the main table. The real number was about 93,000. The nightly rollover timed out four times and rolled back completely each time. The fix was a batched rollover (4,000 rows per transaction, under the lock-escalation threshold). The backlog then cleared in 24 batches.
  • The Available Drops spec assumed the page should show whatever rows were in the day table, including yesterday’s. That was fine while the rollover worked. When it failed, the page showed two days of names at once. The page now reads only the latest drop date.

My lesson: the Assumptions section is not boilerplate. It is a list of bets, and each one deserves a question: what happens if this is off by 10x? I now read it first, and I run /clarify whenever an assumption involves volume, timing or another system.

The spec did not model production

The most expensive problems came from outside anything the spec described. Within hours of the deploy, a post-deploy review using real telemetry found three things. A third-party availability API took anywhere from 6 to 120 seconds per request. A second provider whitelists callers by source IP, while our App Service sends traffic from dozens of outbound IPs, so every request was rejected. And a third-party API key was showing up in plain text in request logs. None of these was a spec problem. They were integration problems, and they were caught by reading logs and telemetry after the deploy, not by the planning loop.

What I did about them is telling, too. The key leak got a focused fix branch the same hour. The IP-whitelisted job was removed from the API the same day and rebuilt as a small bot on a fixed-IP server, with a hand-off document written into the spec folder. None of this went through /specify.

Spec Kit is for features, not for firefighting

Looking at the git history afterwards, the pattern is clear. The two features have spec: commits followed by feat: commits. Everything after that is fix: and chore: branches with no spec folder at all: the log redaction, the batched rollover, a statistics problem that briefly made a 70-million-row table slow to query. Running a seven-step ceremony for a hotfix would have been absurd. What those fixes did reuse was the constitution (the script rules, the UTC rule, the Hangfire rules), and the fix commits cite the spec folders for context. So the constitution applies to everything, while the full spec loop only makes sense for planned features.

The reading is the real work

2,600 lines of Markdown for two features is a lot to review, and reviewing it is my job, not the agent’s. The spec and the constitution check deserve careful reading. The research and data-model docs deserve a skim. The task list deserves a sanity check on ordering and on which tasks are marked parallel. If you do not read the documents, you are just generating paperwork, and the agent will happily implement a plan built on an assumption you never looked at.

My ADLC, as it stands today

After two features, this is the lifecycle I actually follow:

  1. Read the reference docs first. The constitution requires the spec to cite them.
  2. Specify, then clarify. Always clarify when an assumption involves volume, timing or an external system.
  3. Plan, with the constitution check as a hard gate. Every exception is approved by me and recorded.
  4. Generate the tasks, then analyze. The analysis is cheap and has already paid for itself.
  5. Implement on a numbered branch in each repo. Keep the build and tests green, and block tasks on human steps.
  6. Verify by hand against the quickstart: expected counts, a golden path in the browser, a regression pass.
  7. Deploy from the checklist, with scripts before code, dev before prod.
  8. Review after the deploy using real logs and telemetry. This is where the spec gets corrected, and it was the step I had not planned for.
  9. Update the docs and the agent guide, so the next session starts from the truth.

What I would tell someone starting today

  • Write the constitution before the first spec, and give every rule a rationale. It guides the agent on every change, including the ones that never get a spec.
  • Encode your irreversible boundaries as rules, not prompts. “The agent delivers SQL files; a human runs them” did more for my peace of mind than any amount of careful prompting.
  • Do not skip /analyze. One real bug caught before implementation pays for every run.
  • Treat Assumptions as risks. Ask “what if this is 10x bigger?” for each one.
  • Make the quickstart concrete. Real queries and real expected numbers let you check the agent’s work without taking its word for it.
  • Plan a post-deploy review step. The spec describes what you intended. Production shows you what you missed.
  • Keep the ceremony for features. Use a normal branch and PR for fixes, with the constitution still in force.

Spec-driven development did not make the agent smarter. What it gave me was a place to store decisions somewhere other than my memory and a chat window, and a set of documents I can check the agent’s work against. For a one-person team shipping into shared production databases, that was worth all the extra reading.


More case studies in Systems in daily production, or reach me at [email protected].

First published on tanldt.blogspot.com on Sep 26, 2026.