
Our AI benchmark couldn't detect a 7-point regression
Our eval harness compared leaderboard means across 3 trials, which made anything smaller than a 10-point change invisible. Here's how paired statistics and a hierarchical bootstrap turned it into something we can actually gate model switches on.










