Frontier AI models are improving rapidly at coding, reasoning and autonomous agent tasks while potentially becoming worse at some forms of general writing. The author, Wizard Labs founder Maz Ahmadi, says his team observed measurable regressions on client-specific writing benchmarks as models were upgraded. The bigger issue is that public benchmarks usually measure broad capabilities, while businesses depend on very specific tasks and workflows.
This creates an important enterprise AI paradox. According to figures cited in the article, 44% of organizations now report scaling AI across the enterprise, yet only 37% report an enterprise-level EBIT impact, even though 80% say AI has improved individual productivity. In other words, AI adoption is spreading faster than measurable business value. A model can become objectively stronger on public tests while becoming less useful for a particular company's writing, customer service, research or decision-making requirements.
The proposed solution is to stop choosing models primarily by leaderboard rankings. Companies should create their own evaluation suites before development begins, test competing models against real business tasks, and continuously repeat those tests as models are upgraded. This also means reconsidering the assumption that the biggest proprietary model is automatically the best choice; increasingly capable open-weight models can sometimes provide a better combination of performance, control, security and cost depending on where they run and how they are integrated.
The broader takeaway is that the AI race is shifting from raw model intelligence to engineering fit. Businesses should not ask vendors, “Which model is the best?” but rather, “Which model performs best on our actual work?” That shift could become increasingly important as models evolve rapidly and their optimization priorities change. The companies that gain the most from AI may ultimately be those that continuously test, compare and replace models based on measurable business outcomes rather than hype or benchmark rankings.