1. It is happening
What has been coming for a while is now happening. The frontier models (ChatGPT, Claude, Gemini) are loosing ground and LLMs are becoming a commodity. For the first phase of generative AI, intelligence was scarce. Only a handful of companies had the capital, compute and engineering capability to build models at the frontier. For enterprises, the logical response was simple: choose one of those providers, connect through an API and start building.
That assumption is now beginning to break down. Not because frontier models have stopped improving, but because open and open-weight alternatives are improving almost as quickly, inference costs are falling rapidly, and increasingly capable models can be deployed independently. The strategic consequence is significant: it becomes increasingly difficult to justify building an entire enterprise AI architecture around one proprietary frontier model.
2. The performance gap is closing fast
Research by Frank Nagle of MIT and Daniel Yue of Georgia Tech provides clear evidence of this shift. Between May and September 2025, they analysed model usage through OpenRouter, representing approximately 1% of global LLM inference spending.1 2 Closed models still dominated, accounting for roughly 80% of token consumption and almost 96% of revenue. But they were also dramatically more expensive. The average inference price was approximately $1.86 per million tokens for closed models versus $0.23 for open models. In other words, open models were around 87% cheaper.
More importantly, the performance gap was relatively small. The researchers found that open models typically entered the market at around 90% of frontier closed-model performance and often approached parity within months.3 The economics of intelligence are therefore beginning to change faster than enterprise architectures are changing with them.
3. Near-frontier intelligence at a fraction of the cost
Current benchmark data makes the shift even more tangible. Artificial Analysis currently scores Claude Fable 5 at 62 on its Intelligence Index and Kimi K3 at 60. That means Kimi delivers approximately 97% of the measured intelligence of Fable 5. Yet the benchmark cost per task is approximately $3.15 for Fable 5 compared with $0.86 for Kimi K3. That is roughly 97% of the measured intelligence for 27% of the cost. DeepSeek V4 Flash pushes this further. With an Intelligence Index score of approximately 52, it delivers around 84% of Fable 5’s measured intelligence at a benchmark cost of roughly $0.03 per task.4
The implication is not that every enterprise should select the cheapest model.
The implication is that companies can increasingly choose the amount of intelligence appropriate for each task instead of paying frontier-model economics for everything. That is a fundamental change.
4. The model becomes interchangeable
The broader ecosystem reinforces this direction. Llama has seen hundreds of millions of downloads through Hugging Face. Hugging Face itself now hosts millions of public models and derivatives. Ollama has brought local model deployment into mainstream development workflows. Meanwhile, families such as Llama, Qwen, Kimi, DeepSeek, Muse and an expanding generation of smaller specialised models are rapidly increasing the number of credible alternatives.5 6
This does not mean cloud frontier models disappear. There will always be tasks where the absolute highest level of general reasoning capability is valuable. But that is very different from making one frontier model the architectural foundation of enterprise AI.
If model performance changes every few months, prices differ by factors of ten or one hundred, and new models can increasingly be deployed locally, then the model itself should become a replaceable component. Companies should be able to upgrade it, replace it, route between models or combine several models without rebuilding their intelligence infrastructure.
5. The scarce asset is company context
This is where the value equation changes. General-purpose intelligence is becoming increasingly abundant. Company intelligence is not. A frontier model understands language and general knowledge, but it does not inherently understand what your company means by a critical supplier, a production constraint, an inventory exception or a customer commitment.
That context is proprietary. This is why we separate the model from the context layer. It does not know which ERP fields are authoritative, how your factories operate, what contractual rules apply, how products relate through a bill of materials or which planning constraints cannot be violated.
The durable intelligence infrastructure consists of the company’s terminology, documents, structured operational data, business rules, relationships, industry knowledge and knowledge graph. Models operate against that environment, but they do not own it. The company context remains stable while the underlying model can change. That creates technological independence.
6. From frontier models to Local Intelligence
We therefore believe enterprise AI will increasingly move toward smaller, specialised and locally deployable models. Not one model trying to know everything. Different models performing specific tasks within an environment that supplies exactly the context they require.
A planning model needs to understand planning, inventory, capacity and the company’s network. A maintenance model needs asset structures, failure modes, manuals and work-order history. A supply-chain agent needs suppliers, customers, locations, BOMs, constraints and operational relationships. It does not need to be the world’s best general-purpose model. This is what we refer to as a vertical local language model.
Vertical because its intelligence is focused. Local because company context and inference can remain within the enterprise perimeter. And language model because natural language becomes the interface to a much richer system of structured data, documents, relationships and business logic.
The strategic question is therefore no longer which frontier model should we build on?
It is:
How do we build an intelligence infrastructure that can use whichever model is best tomorrow?
That is the shift now underway. General-purpose AI intelligence is becoming a commodity. Company intelligence is not.
References
- Brian Eastwood, AI open models have benefits. So why aren’t they more widely used?, MIT Sloan, 20 January 2026. ↩︎
- Intelligence Index:63 (Fable max) versus 60 for KIMI K3 Max and 54 Deep Seek V4 (on 11/08/2026). See https://artificialanalysis.ai/#intelligence/. The Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR ↩︎
- The Latent Role of Open Models in the AI Economy, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5767103. Frank Nagle, Massachusetts Institute of Technology (MIT); The Linux Foundation, Daniel Yue, Scheller College of Business. November 18, 2025 ↩︎
- Cost per Intelligence Index Task. Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better ↩︎
- Meta AI, Llama adoption and Hugging Face download statistics. ↩︎
- https://huggingface.co/blog/huggingface/state-of-os-hf-spring-2026 ↩︎
