Two independent leaderboards now disagree about which model is best, and in the same week the price of matching last generation halved, which moves the buying question from capability to cost per passed task
Artificial Analysis ranks Claude Fable 5.1 as the highest-intelligence model, with GPT-6 Astra behind it. LLM Stats ranks GPT-6 Astra first overall at an index score of 59.6, ahead of Claude Fable 5.1 at 55.7. Both are independent aggregators, neither sells the models, and they have swapped the top two places between them. That disagreement is the most useful frontier fact of the week, because it says the measured gap at the top is now inside the margin between two credible measurement methods.
What sits underneath it is a price move. xAI released Grok 4.6 at $2.00 per million input tokens and $6.00 per million output tokens for prompts under 200K, rising to $4.00/M input and $12.00/M output above that threshold, with a 500K-token context window. OpenAI opened a limited preview of the GPT-5.6 series, adding Sol, Terra and Luna, and states that Terra matches GPT-5.5 performance at 2x cheaper. Two different labs, the same movement: yesterday's capability at roughly half of yesterday's price, with the tiering shifted onto context length rather than onto the model name.
Put those together and the procurement question changes shape. If the top of the leaderboard is contested and last generation is now half price, the decision worth making this quarter is not which model is best but which is the cheapest that clears a fixed internal task set. The VulcanBench cost figure, unchanged again this week on its 26-task suite, is the format to copy: a bounded set of tasks, a pass count and a dollar total. A ranking answers a question no buyer actually has.
The adoption rail cuts against the simplest version of that story, and it does so awkwardly, because every named organization on it this week is a seller. Anthropic released 11 open-source plugins for its Claude Cowork agent across legal, finance, sales and marketing, and markets read that as AI eating vertical software, with a reported selloff of nearly $1 trillion in software market capitalization inside a week. Salesforce answered with numbers: Agentforce and Data 360 at nearly $3.9 billion annualized revenue, up over 240% year on year, and raised full-year guidance. Atlassian, Snowflake and MongoDB reported acceleration in the same direction. These are vendor disclosures about what is being sold, not evidence of what has been deployed and held.
Governments spent the week working on the off switch rather than the price. California's governor issued an executive order to accelerate independent oversight of AI companies and advance an 'AI kill switch' mechanism; Rep. Subramanyam introduced H.R.9925, the FRONTIER Act, for federal oversight of frontier development and deployment; Australia opened consultation on national AI standards that would require companies to report 'rogue' AI incidents; OpenAI urged UK lawmakers to introduce stricter regulation. Four jurisdictions, one shared premise: that the binding control is the ability to stop a system after it is running. None of that touches the cost curve, which is the variable actually changing what organizations can afford to build.
59.6 LLM Stats index score, GPT-6 Astra ranked first overall
55.7 LLM Stats index score, Claude Fable 5.1 ranked second
$2.00 per million input tokens Grok 4.6 pricing below 200K-token prompts
$6.00 per million output tokens Grok 4.6 output pricing below 200K-token prompts
$3.9 billion Salesforce Agentforce and Data 360 annualized revenue, vendor-reported
240% Salesforce year-on-year growth rate on that annualized figure
20% more Jira work items completed by Atlassian Rovo customers
400% Atlassian MCP interface call growth, quarter on quarter
69 USPTO published patent activity, neural networks and deep learning
Signal Pro adds the evidence behind the argument, how this week sat against previous weeks, the policy read for every market and what it changed for planning.