Anthropic says AI writes most of its code now. The next sentence is the one that matters.
Anthropic published some numbers about itself this year that are worth a firm owner’s attention. More than 80% of the code they merge is now written by Claude, up from low single digits before early 2025. Their typical engineer merges about eight times the code they did in 2024. Impressive, and the kind of stat that makes a partner wonder if they’re being left behind.
That’s not the sentence that matters. This is: human review has become the bottleneck on their own AI development. The model writes faster than people can check the work.
Read that as a confession about where the real constraint sits. It isn’t the AI’s output. That part got cheap. It’s the human who has to look at the output and say yes. The company with the best models in the world, pointed at its own code, ran straight into the same wall you will: the work is easy to produce and hard to trust, and trust is a person’s job.
Two ways firms will misread this
“So the AI can just do the work now.” The 80% number reads like permission to hand the work over. It isn’t. Eighty percent authored is not eighty percent unsupervised. Every one of those lines still passed a human review on the way in. What Anthropic automated was the typing. What they didn’t automate, and say they can’t, is the checking. If you take the headline as license to skip the gate, you’ve copied the one part of their operation they’d tell you not to.
“That’s a big-tech problem, not a us problem.” The instinct is to file this under things that happen at AI labs and not at a twenty-person firm. But the shape is identical, and you hit the wall sooner. The moment AI gives your firm real leverage, your constraint moves from “can we produce the work” to “can we stand behind it.” For Anthropic that’s code review. For you it’s a partner’s name on a filing, an opinion, a return. Their bottleneck is a slowdown. Yours has a client and a regulator attached to it.
Both misreadings come from the same place: treating review as friction to remove instead of the thing the work is actually for.
For your firm, the bottleneck is the product
Here’s the turn. When an AI lab says review is the bottleneck, they mean it as a cost to drive down. They’d automate the reviewer too if they could. That’s the right goal for them, because for them the review is overhead on shipping code.
It is the exact opposite for you. The review is not overhead on the work. It is the work. A client isn’t paying for the draft, they’re paying for the judgment that a named professional stood behind it. The sign-off is the service. The accountability is the product. So the thing the model’s own maker is trying to shrink is the thing your firm exists to sell.
That changes what a good AI setup looks like inside a firm. You don’t build to minimize the human in the loop. You build so the human’s judgment lands exactly where it’s worth the most, and every step below it is prepared, permissioned, and written down. The gate isn’t the part slowing you down. It’s the part that turns a fast machine into work a partner can put their name on. That’s also where your record comes from: a review that happened on purpose leaves a trail you can show a client without flinching.
I run my own businesses this way, with a fleet of AI agents doing the production work across them. Over the past year the output climbed steadily while my own hours stayed flat. The one thing that never scaled, on purpose, was the time I spend checking what the fleet did before it ships. That ceiling isn’t a failure of the setup. It’s the setup working. I keep the running record of it in public.
The one rule
If you hold the firm to one thing as AI comes in, hold it to this:
AI prepares the work. A person decides anything that can’t be undone.
The assistant can draft, find, summarize, and propose all day, and it should, because that’s where the leverage is. The moment something becomes a decision with consequences, a sent filing, a moved file, a client commitment, a person signs off. Anthropic is trying to widen that gate as fast as it safely can. Your firm’s job is to keep it exactly where it belongs and prove it held.
What good looks like
- AI carrying the production work, so your people spend their hours on judgment, not assembly.
- Human review on every action that can’t be reversed, by design, not as an afterthought.
- A record of what the AI did and who approved it, in writing, that you could hand to a client or a regulator.
- Someone who owns the system after launch, the way you’d own any hire trusted with sensitive work.
- The whole thing scoped to one workflow first, proven, then widened.
Start with one workflow
You don’t need to match Anthropic’s numbers. You need the part they’d tell you not to skip. Pick the workflow that costs your best people the most time, map how the work actually moves, decide where a human has to stay in the loop, and prove AI can carry the rest with review and a record you trust. Then widen. One workflow, one AI Team, one measurable outcome.
That’s the work we do in a workflow diagnosis. No tools installed, no data moved. We map how the work moves through your firm, where the lines have to hold, and where AI would actually earn its place. You walk out with the map whether or not you ever hire us.
This is what a workflow diagnosis maps for your firm. No tools installed, no data moved.
Book a workflow diagnosis →