The highest-scoring open-weight model this week ran on downloadable weights, and the cheapest verified coding run cost 82 cents, which means the capability question for most enterprises has become a hosting and cost question.
The most consequential fact in this week's input is not a closed-lab launch. It is that Artificial Analysis ranked Kimi K3 (max) and GLM-5.2 (max) as the highest-intelligence open-source models, and both are downloadable. Kimi K3 shipped at 2.8 trillion parameters with a 1M-token context window and recorded 1,456,459 downloads. GLM-5.2 carries the same 1M-token context and posted 2,488,397 downloads. Alibaba's Qwen3.6-35B-A3B-FP8 recorded 8,984,811 downloads. These are not previews or waitlists. The weights are on Hugging Face now, and organizations can run them where their data already sits.
Set that against price. VulcanBench published a coding result showing Claude Sonnet 5 (low) completing 25 of 26 sandboxed tasks for $0.82 total. OpenAI's GPT-5.6 rate card puts Luna at $0.20/$1.20 per million input/output tokens after an 80% cut on July 30, with Terra at $2/$12 after a 20% cut, and cached input billed at 10% of standard rates. Artificial Analysis ranked Claude Opus 5 (max) and Opus 5 (xhigh) highest for intelligence and GPT-5.6 Luna lowest cost per task. The frontier of measured capability and the frontier of measured price are now separated by a choice, not a gap in what is available.
What corroborates the shift is that buyers with the most to lose are building rather than renting. Microsoft deployed an in-house AI system to take over most of its security work, reportedly at lower cost than relying on frontier third-party models, and unveiled its Azure Cobalt 2 chip claiming a 3x improvement over its predecessor, though independent benchmarks and production timelines are not yet published. Salesforce reported Agentforce deployments more than doubled year over year. The signal underneath both is that the model layer is commoditizing faster than the layer that makes a model useful.
That last layer is where the counter-evidence lives. Pinecone's Nexus, an agent using a knowledge engine rather than a bare model, posted the top score on Sierra's τ-Knowledge benchmark, beating agents built on frontier models from OpenAI, Anthropic and Google. The reading here is that raw model intelligence is necessary but no longer sufficient to win a task, and the durable advantage is moving toward retrieval, grounding and orchestration. A cheap open-weight model wired into good context can beat an expensive closed one that is flying blind.
The policy and spending rails tell you where the friction and the money are. US federal AI dollars went overwhelmingly to research and to the physical infrastructure that lets any of this run, not to model licenses. In the UK, Moody's warned that banks' reliance on a small number of AI and cloud providers is a systemic dependency risk, which is the concentration argument that open weights partly answer. The week's argument is that the affordable choice and the top-ranked choice can now be the same choice, and the harder work has moved to what you wrap around the model.
$0.82 Claude Sonnet 5 (low) completing 25 of 26 VulcanBench coding tasks
$0.20/$1.20 GPT-5.6 Luna per million input/output tokens after 80% cut
8,984,811 downloads of Alibaba's open-weight Qwen3.6-35B-A3B-FP8
1,456,459 downloads of Moonshot AI's open-weight Kimi K3, 2.8T parameters
2,488,397 downloads of Zhipu AI's open-weight GLM-5.2, 1M-token context
$671.2M US federal R&D physical/engineering/life sciences AI spend, 47 awards
$278.9M US federal computer systems design services spend, Interior lead
$2/$12 GPT-5.6 Terra per million input/output tokens after 20% cut
Signal Pro adds the evidence behind the argument, how this week sat against previous weeks, the policy read for every market and what it changed for planning.