Somebody on your team swapped AI models last month because of a leaderboard result, and nobody who understands what that leaderboard measures was in the room when they did it.

That’s not a research problem. Nobody read the wrong benchmark. The decision had no owner, so it got made by whoever was paying attention that week.

Here’s the pattern I keep seeing at Seed to Series A companies. A model tops a leaderboard. Someone on the team, often a good engineer with no bandwidth to dig deeper, sees the result and swaps the model into production. Two weeks later something breaks in a way that’s hard to trace back to the swap, because nobody logged the decision as a decision.

The standard answer is: public leaderboards lie, build your own eval. True, and useless to a founder with no engineering leadership. You can’t build what you don’t have the team or the time to maintain.

The gap between leaderboard and reality bit people in public this July

A widely shared benchmark this summer showed a leaderboard-topping model badly underperforming established models the moment it hit real engineering tasks. It wasn’t a subtle gap: a model with a great public score did visibly worse work than models ranked below it. It made the rounds because it confirmed what engineering leads already suspected: the leaderboard and the job are not the same test.

This isn’t a one-off either. Research out of Cohere, Stanford, MIT and the Allen Institute analysed roughly two million head-to-head battles across 243 models and 42 providers, and found major labs routinely test dozens of private variants before release and publish only the best-scoring one. Their estimate: that data access asymmetry alone can inflate a leaderboard score by up to 112% on the benchmark’s own turf. The leaderboard measures the best private version of a model, on questions picked to flatter it, not the version you’ll ever run.

So “build your own eval” is the right instinct. It’s also expensive in a way that doesn’t show up until you try it. A single LLM-judge-based benchmark run can land north of £7,000, and then it decays: the moment your product moves into a new workflow, your eval set stops covering what you actually ship, and a stale eval creates false confidence, which is worse than admitting you don’t have one. Most companies at this stage genuinely cannot build and maintain one. Telling them to is like telling someone to build their own credit rating agency because the public score is noisy.

This used to be a one-time decision. It isn’t anymore.

Model selection used to look like an architecture choice: pick a provider, build around it, revisit maybe once a year. That’s gone. Menlo Ventures surveyed 150 technical decision-makers earlier this year and found only 11% had switched LLM vendors in the past twelve months. But 66% had upgraded models within their existing provider, and fast. Claude 4 captured 45% of Anthropic’s user base within a month of launch, and Sonnet 3.5 usage dropped from 83% to 16% in the same window. As Menlo put it, switching vendors is relatively easy, but increasingly rare.

Read that carefully. The churn isn’t vendor-hopping. It’s low-key, monthly, model-level re-picking inside a relationship you already have. Nobody’s rearchitecting. They’re swapping the engine on a Tuesday because a new one shipped. It happens in a pull request, not because anyone’s careless but because nobody’s job is to own it.

Who’s actually in the room

More than half of UK startup founders rate technical talent, especially engineering leadership, as the hardest thing to hire for, according to Tech Nation’s 2025 report. That tracks. Most companies at this stage have strong individual engineers and no one whose job is to sit above the model-selection decision and ask what “reliable” actually means for this product, this workload, this customer.

That’s the actual gap: not a missing benchmark, a missing owner.

What a fractional CTO does here isn’t run an eval suite personally. It’s treat model selection like any other vendor decision with stakes: define what reliable means for your product before you’re choosing between options, test the real candidates against your actual workload instead of a generic leaderboard, and be the person accountable when it breaks in production instead of a Slack thread trying to reconstruct who decided what.

That process doesn’t go stale when the next model ships next month, because it was never about which model. It was about who’s answerable when the model choice costs you a customer.

If nobody at your company owns that, book a call and we’ll find out where it’s sitting exposed.

Share: