[ Not vibe coding ]

Your product lead and one engineer ship enterprise‑grade code at 10× the output.

Your product lead directs CraftHQ's AI agents on Claude Code. A separate agent writes the tests, and your engineers approve every merge.

[ Scroll to watch one idea go live in four days ]The company and its people are made up, to tell the story.

The company and its people are illustrative.

  1. 01 / 07Mon 8:40

    Dana makes a call

    A big customer is waiting on one feature. Dana, the CEO, makes it this week’s top job.

  2. 02 / 07Mon 9:12

    Emily asks for it

    Emily, the product lead, asks for the feature in a normal chat message. CraftHQ asks her two quick questions.

  3. 03 / 07Mon 9:15

    CraftHQ makes a plan

    CraftHQ’s Planner AI turns her message into a clear plan, and adds lessons from past mistakes.

  4. 04 / 07Mon–Tue

    AI builds and tests

    Builder AIs write the code. Tester AIs write the checks, without ever seeing the code.

  5. 05 / 07Tue 17:40

    It all comes together

    All the work comes back as one simple report: what changed, how it was checked, and the risks.

  6. 06 / 07Wed–Thu

    Sam checks and approves

    Sam, an engineer, reads the report and asks for one more check. Then he says yes, and it goes live.

  7. 07 / 07Fri

    Next time is faster

    Everything this request taught is saved, so the next one starts smarter.

[ That was one request ]

Emily and Sam ship like a team ten times their size.

10×+Output per person vs ISBSG median
734PRs in two months
3Claude Code subscriptions
0.19%Revert rate

Output, reverts and test share: one two-person BetaCraft team, measured over 30 days. PRs: across BetaCraft client repos, mid-Jul to mid-Sep 2026. The story is illustrative.

MADE-UP COMPANY WORK CHECKED PERSON DECIDES MEMORY
Built with CraftHQ for
NexdigmZiplines EducationiSoftpull
How it works

Three steps. Your engineers keep the final say.

A product lead asks, three CraftHQ agents do the work, an engineer decides.

  1. 01

    Point it at a repo

    Pull-request access to one repo. It learns your architecture, conventions and test setup once, then cross-examines every ticket before building.

    109 contradictions caught pre-code
  2. 02

    It builds and tests from the spec

    Builders write the change. Separate testers write checks from the requirement, never from the code.

    0.19% rolled back · 30 days
  3. 03

    Your engineers review and merge

    A normal PR in your CI, risk-tagged (UI-only · logic · core), with the requirement, tests and results attached.

    Every merge human-approved
Proof

Measured from git history, the same way you'd measure your own team.

Output is counted in function points (FP), the ISO-standard unit of delivered functionality. It's independent of language and hard to pad with AI volume.

Output per person
10×+
0150 FP
function points per person-month: 140+ vs. an ISBSG median of ≈ 14
two-person BetaCraft team · 30 days · IFPUG-counted
Revert rate
0.19%
of shipped changes rolled back: fewer than 2 in 1,000
same team · repo's last 30 days
Time to production
12 wks
~780 FP platform · ~9 months for six people at the ISBSG median
one client platform · 18 Jun → 14 Sep 2026
  • 734 PRs merged in two months
  • 3 Claude Code subscriptions
  • Every PR approved by a human
Get your own baseline report, free. Read-only, one repo, NDA first.Get my baseline report
Not another coding assistant

Keep Claude Code and Cursor. CraftHQ adds the proof around them.

Out of the box, an assistant doesn't know your business rules or what broke last quarter, and it grades its own work.

What CraftHQ adds compared with a coding assistant used out of the box
CapabilityCoding assistant, out of the boxWith CraftHQ
Your codebaseKnows what's in the context windowFitted per repo: architecture, conventions, test setup
TestsOften written in the same session as the codeWritten from the spec by a separate agent that never sees the implementation
MemoryRules files you maintain by handA ledger of every defect class, enforced on every later change
ReviewDiffs, reviewed like any other codeRisk-tiered PRs with requirement → tests → results attached
ProofNot measured by defaultFunction points, reverts and test share from your own git history
It compounds

CraftHQ specializes to your codebase.

Every lesson becomes part of your harness: a ledger of rules that CraftHQ keeps alive, with a check for each new rule.

Starts as
the same harness every team gets
Grows with
every defect caught, rule clarified, domain quirk
Lives in
plain, readable files in your repo
Enforced
a check per rule, run on every build

And the ledger is yours. Passing checks are review evidence, not a guarantee that every business rule is met.

Ledger · defect classes on fileyour repo
  • R-041webhook retried → customer charged twicecaught in review
  • R-041.1rule: every payment handler is idempotentchecked every build
  • R-042tenant ID missing from export querychecked every build
  • R-043usage rollup at timezone month-endchecked every build
  • R-044your next onepending
Security

Built for the security review your team will run.

Your repo, a dedicated machine, your model subscription. Only your engineers merge.

Repo-scoped token

Contents + pull requests only, verified by you, with branch protection on.

Security brief first

Token permissions, hosting region, retention and who can reach the machine, before any access.

Your own cloud, if you prefer

We'll assess your cloud account before the pilot is agreed.

  • NDA before read access
  • ISO/IEC 27001 certified
  • Third-party VAPT
  • SonarCloud gates on every build
  • Every change traceable
  • DPIA for personal data
For engineering teams

The questions your engineers will ask.

Do you need write access to our repo?

We need a repo-scoped token with contents and pull-request permissions: it can push feature branches and open pull requests. You turn on branch protection for main and your release branches, so the token can't merge or push to main. That rule is enforced by GitHub, not by trusting us.

Nothing reaches main unless one of your engineers approves it.

Why not just give our developers Claude Code or Cursor?

You should, and many teams already have. A coding assistant helps developers write code faster. On its own it doesn't carry your business rules or check its work independently.

CraftHQ is the layer around the model that fixes that:

  • A per-repo understanding of your architecture and conventions.
  • Tests written independently from the spec, never from the code.
  • A ledger of every defect class and business rule, enforced on every later change.

That combination is what we measure, and the pilot tests it against the workflow you run today.

Won't this replace our engineers?

It changes what they spend time on. The harness handles building and first-pass testing. What it can't supply is judgment: architecture, what a requirement really means, and whether a change is safe to ship.

The measured team was one product owner and one engineer, with the engineer as reviewer, architect and quality gate. That's the most senior part of the job. The pilot runs on your existing team and measures what changes for them.

We'll be buried in PRs. Who reviews all of this?
  • Every PR carries an automatic risk level: UI-only, logic or core.
  • Your team reviews every PR at first. Once the track record is in, you can choose to fast-track UI-only PRs to a lighter review. That's your policy, set in your branch rules, and a human still approves every merge.
  • Each PR arrives with the requirement it satisfies and the spec-derived tests that check it, so review means checking evidence, not reading code cold.

Nothing merges without a human.

What is a function point, and why measure with it?

A function point is the ISO-standard unit of delivered functionality, counted with the IFPUG method. It measures what the software does (inputs, outputs, queries, stored data) regardless of language, framework or how many lines it took.

We use it because lines of code, commits and story points can all be inflated, and AI makes that easy. Function points are hard to pad, and anyone trained in IFPUG can recount them.

We report the everyday metrics too: PRs merged, revert rate and test share.

What does the pilot need from us?
  • Pull-request access to one repo, with branch protection on main.
  • A CI/CD pipeline that runs your tests on every pull request. The target is 95%+ test coverage; if you're below it, we close the gap before the pilot month starts.
  • Claude Code subscriptions on your organisation's account. For scale, the 734 PRs shown above shipped on 3.
  • A real backlog of tickets, so there are no toy tasks.
  • One or two engineers reviewing PRs, with a reviewer-time budget and a PR submission limit agreed up front. We measure actual review hours and slow the submission rate if the budget is exceeded.
  • A success measure agreed before day one, including reviewer hours.
Still have questions? Book a 30-minute technical session with the engineer who'd run your pilot.Book a session
The one-month pilot

One month. Your repo. Your numbers.

The pilot measures your team operating CraftHQ, and logs every time we step in.

  1. 01BeforeBaselineYour team's last 30 days
  2. 02Week 1Setup & tuning5–10 PRs, reported separately
  3. 03Weeks 2–4MeasuredSame repo, comparable tickets, same team
For the CEO

What you get

  • A measured before-and-after on your own delivery data, not our benchmark.
  • A go/no-go decision in 30 days, backed by numbers your CTO signed off on.
  • Everything we build stays yours, including tuning work and unmerged PRs.
For the CTO

What you control

  • Pull-request access to one repo, branch protection on main.
  • Your CI and review: nothing merges without your engineers.
  • The success measure, agreed before day one, reviewer hours included.

Before you start: a CI/CD pipeline and 95%+ test coverage. Below that, we write the missing tests from your specs first.

Scope your pilot

A 30-minute call agrees the fee, scope and success measure.

Start here

Don't take our numbers. Measure yours.

One repo, one month: delivery speed, review effort and rework, measured from your own git history against your own baseline.