HomeServices › The Delivery Baseline
The Delivery Baseline

Get the number before
you buy the agent.

Every vendor in this category will quote you a multiple. Three-and-a-half times the leverage. Ten times faster releases. Nobody will tell you what those numbers are for your codebase. That is what this is: two weeks, three measurements, one honest answer — including "do not do this."

Three numbers, taken from your repository.

Not a survey of your team, and not a maturity model with five levels and a spider chart. Three measurements taken from your commit history, your tracker, and a real pipeline run on your own code.

The measurements

  • Cycle time. Ticket opened to pull request merged. Median and worst case, because the worst case is what your team remembers.
  • Rework rate. How much merged work comes back. The number most teams have never counted.
  • Cost per merged pull request. What an agent actually spends to land one change on your codebase, at your volume — not a vendor's demo repo.

What you leave with

  • Your three numbers. Yours to keep, hire me or not.
  • The method, written down. Re-run it next quarter without me.
  • An honest read. Including "this is not worth doing."
  • A shape and a price — only if the numbers justify one.

Three reasons not to hire me.

This category is full of people who will tell you that you need what they sell. Here is where you genuinely do not need this, written down so you can check it in ten minutes and keep your money.

You have not tried what you already pay for

  • GitHub's coding agent can be assigned directly to an issue and will open a pull request. If it is in your plan and nobody has tried it, try it this week.
  • Google's Jules is free for a meaningful number of tasks a day. It costs an afternoon to find out whether it is sufficient.
  • If either one lands work on your codebase, you are done. Keep the money.

Your problem is composition, not delivery

  • If the real pain is eighteen teams maintaining eighteen versions of the same component library, that is an architecture problem.
  • Platforms built around a shared component graph solve that shape properly, and I do not.
  • Different problem, different vendor. I will say so on the call.

Your bottleneck is review, not writing

  • If changes sit in review for six days and take two hours to write, generating them faster changes nothing.
  • More agent output makes a review queue worse, not better.
  • The baseline will show you this, which is the cheapest way to find it out.

When it is worth the two weeks

  • You bought the tools and the work is not landing. Licences are paid, adoption is not happening.
  • The codebase is old and nobody wants to touch it. The repos that will never be re-architected are exactly where this pays.
  • Someone above you wants a number. A measured one, before more budget goes out.

Measured on a pipeline that already ships.

The measurement is not theoretical. It comes out of running the thing — real tickets off real backlogs, through ingest, root cause, tests and review, merged by engineers who did not write them. Every step traced and replayable, which is why the cost-per-pull-request number exists at all.

50+
Pull requests merged
4
Ticket sources
0
Self-merges

The baseline is the product. The rest is optional.

Most engagements in this category start with a proposal and work backwards to a justification. This one starts with a measurement and stops there unless the numbers argue otherwise.

If they do argue otherwise, there is a pipeline behind this — one that takes a ticket, runs root-cause analysis against the actual repository, writes the failing test before the fix, and opens a pull request a human still has to approve. It runs sandboxed, isolated per client, and never against your live environment. We would scope that after the numbers, not before. If you want the reasoning behind how it holds together, I wrote up agents that ship pull requests.

Common questions.

What does it cost?
$6,000 fixed, for two weeks. Not hourly, no change orders, and the number does not move if the work turns out to be harder than expected. If the measurement says you should not spend anything further, you have still bought the cheapest possible answer to that question.
What do I actually get at the end?
Three numbers measured on your repository — ticket-to-merge cycle time, rework rate, and cost per merged pull request — plus the method used to get them, written down so your team can re-run it without me. The numbers are yours either way.
We already pay for Copilot. Why would we pay you as well?
Possibly you would not. If nobody on your team has tried assigning an issue to the coding agent already included in your GitHub plan, do that first — it is free to you and it might be enough. This is for teams who have bought the tools and are not seeing the work land.
Will you just tell us to buy your pipeline?
The deliverable is the measurement, and it is priced so that it stands alone. Roughly a third of the honest answers in this category are "your bottleneck is review capacity, not code generation" — which no pipeline fixes. You will get that answer if it is the true one.
Do you need access to our production systems?
No. Read access to the repository and the ticket history is enough. Nothing runs against your live environment, and no code leaves your infrastructure for the measurement.
How is this different from a consultant writing a report?
A report is an opinion with a header. These are three measurements taken from your commit history, your tracker, and a real run on your own code. You can check every one of them yourself, which is the point.

Think we might be a fit?

I work with a small number of early-stage startups at a time. Tell me where your product is — if I'm not the right call, I'll say so.

Book a 30-min call