Thought LeadershipThought Leadership

Every AI Dollar Is Doing the Job of Twelve. We Think It Can Do a Hundred.

JB Joe BarnesonHead of Engineering

We're 16 engineers producing the output of a 60-person team. That's not a feeling — it's a measurement. Every dollar we spend on AI is currently 12x more efficient than a dollar spent on a senior engineer or an offshore contractor, and we think that number gets to 100x.

I want to explain how we know that, because "how are you measuring ROI" is the question I get asked most — and it's the question a lot of engineering leaders can't answer right now, which is why some of them are quietly pulling spend back. We went the other direction. Here's why we could.

The number comes from Canary

Canary is an agent we built that evaluates every merge to production — application code, rating calculators, the hundreds of prompts that govern our agentic systems. All of it lives in source control, so all of it gets measured.

On every merge, she scores three things: effort, business value, and risk. Effort is where the ROI math lives. She calculates what each piece of work actually cost — the human hours, the AI usage, the token spend — and converts it into a dollar figure split between human and AI contribution. Business value ties the output back to our annual objectives. Divide one by the other and you have a real efficiency number for every dollar, human or AI, on every piece of work we ship.

That data is public inside our platform. Anyone at MGT can look at it. No committee compiles it, no one massages it for a quarterly deck — it's just continuously produced. When I sit down with leadership, we can see roughly 60% of engineering effort flowing to our official bets and exactly what that investment returned. Same lens on the unglamorous side: production support, log monitoring, hardening. We can ask "are we getting enough movement here?" and answer with numbers instead of instinct.

Most organizations do the opposite. They estimate effort before work starts — planning sessions, debates over story points — and then never measure what the work actually cost or what it actually delivered. Effort becomes a proxy for value, and the estimate was a guess to begin with. We skipped estimation entirely. We build, then we measure what really happened.

What the measurement changes

It changes people, not just spreadsheets. Canary posts an assessment to Slack on every merge. A while back, a pattern started showing up: when an engineer consistently trends high-effort and high-risk in aggregate, it almost always means they're building in large batches instead of shipping small, frequent, reviewable changes. And large-batch work, in aggregate, produces measurably less business impact.

That pattern now triggers a conversation — a manager and an engineer looking at real data together, usually a coaching moment, occasionally something more. No gate slammed shut, no process police. The work still ships. But the signal surfaces before the habit becomes a problem, and it surfaces from measurement, not from a manager's hunch.

The measurement is also what makes the scale manageable. On a given day I'm running 10 to 20 active AI sessions. The agents coordinate through what we call the water cooler — ephemeral markdown bulletins where they broadcast what they're working on, spot overlap, and resolve conflicts themselves. It's not built for humans, though we can read it. I don't distribute a roadmap or referee collisions. The system handles the coordination; Canary handles the accounting.

The part that keeps me honest

Our AI spend now gets treated like payroll. It isn't as large as payroll yet — but it's bigger than almost every vendor we have, nearly on par with our data spend, and growing fast. There's real anxiety in that, and I won't pretend otherwise.

Here's what I've learned: you cannot evaluate that spend on a micro basis. A single quote, a single PR — the numbers will scare you or fool you. It only makes sense in aggregate, where you can see where it's working brilliantly and where it isn't, and make resourcing decisions from the whole picture. The anxiety doesn't go away because you spend less. It goes away because you can see what the spending produces. That's the entire reason Canary exists.

Where the curve is heading

We're about twelve months into this. The last six to nine months have accelerated hard, and I expect the next six to nine to be steeper. 12x today. 100x is where I think this goes — not because the models get incrementally better, but because we get better at managing them. Our engineers are making the transition from doing the work, to coworking with agents, to managing agents that do most of the work while they guide, train, and improve them. Each stage compounds.

Which raises the obvious question: how do you trust output you didn't write line-by-line, at this speed? The honest answer is that trust at this velocity has to be engineered — evaluation frameworks that regression-test our agents against golden datasets, quality gates that score every change before and after it ships. Because "we move fast" means nothing to a regulator, a broker, or a policyholder unless you can prove fast isn't reckless. That story deserves its own post, and it's coming.

The efficiency isn't confined to engineering, either. The same measurement discipline is showing up in underwriting, where one underwriter with the right agents behind them operates at a multiple that would have sounded like fiction two years ago. That story is coming too.

The point of all this was never a bigger team. It's a team of 16 that stays small while the business scales past it — no management pyramid, no onboarding churn, the same people getting better together for years. Every AI dollar doing the job of twelve makes that possible. Getting it to a hundred makes it inevitable.

Recent News from MGT Insurance

Stay current with the latest appetite updates from MGT and news from thought leaders in the industry.

×

Contact Broker Success

×

Form Submitted Successfully