Where the AI field is heading, what the frontier models cost per token, and which claims survive a check.
Every Monday. The report leads with the direction read - what the week's news means for where the field is going - then prices the frontier models against a third-party intelligence index, then runs an adversarial pass over the week's biggest claims: what has to be true for each one to hold, and who has actually tested it.
Free and weekly, with double opt-in. Unsubscribe any time.
Every Monday. Also posted to #weekly-ai-news on the Discord server.
AI News Weekly — August 24, 2026
- The AI Engineer conference turned "the harness is the moat" from a thesis into the room's default — and Uber supplied the demand-side proof: >70% of its pull requests are now written by AI agents, backed by a 2,500-skill marketplace and a 40-million-entry context graph.
- The counter-argument came from the top of the field: Rich Sutton (author of The Bitter Lesson) says none of this is real intelligence — today's models stop learning the moment training ends, and continual learning, not a better harness, is the missing piece.
- Open-weight models crossed from "coming" to "in production" with real usage data: AT&T reportedly runs 40% of employee AI use on open models for ~56% savings; legal-tech Harvey moved to China's Kimi K3; a stealth frontier model ("Ox Alpha") appeared, widely guessed Chinese.
- The West kept answering on price, not capability: OpenAI cut pricing (Codex hits ~20M users), Grok 4.6 landed on Amazon Bedrock and Google Vertex, Gemini 3.7 Flash topped agent benchmarks — and Anthropic's best model is reportedly too expensive to be popular.
What the frontier costs right now
The report tracks this table week to week, because the interesting movement stopped being raw capability a while ago. Grok 4.6 reached the same index score as GPT-5.6 Sol at roughly a fifth of the output price, and that is the shape of the current competition.
| Model | Intelligence | In | Out | Per task | As of |
|---|---|---|---|---|---|
| Claude Opus 5 Anthropic | 63 | $5 | $25 | · | Aug 17 |
| Claude Fable 5 Anthropic | 62 | · | · | · | Aug 17 |
| GPT-5.6 Sol OpenAI | 61 | $5 | $30 | · | Aug 17 |
| Grok 4.6 xAI | 61 | $2 | $6 | $0.84 | Aug 17 |
| Kimi K3 Moonshot AI open weights | 57 | $3 | $15 | · | Jul 20 |
| Gemini 3.7 Flash Google | · | · | · | · | Aug 24 |
Intelligence figures are the Artificial Analysis Intelligence Index, published by Artificial Analysis and reproduced here with attribution. They are not BluNET Studio measurements. Prices are USD per 1M tokens (input / output). Rows are not a single snapshot: each carries the date it was captured, and a blank means no published figure has been verified rather than zero. Model pricing and scores change without notice.
Five sections, same order, every week
- TL;DR the week in six bullets, above the fold.
- Signal where this is heading, split into what it means for individuals and builders, what it means for the enterprise, and how the two connect.
- What the leaders are saying direction from the people running the labs, quoted and sourced rather than paraphrased.
- Claims vs. evidence the week's testable claims, each with the assumption it rests on, who has independently tested it, and a verdict. A vendor's claim about its own product is labelled a vendor claim. Zero claims is a valid week.
- Top stories, models, research, policy the record of what actually shipped, with a link on every item.
Past issues
-
AI News Weekly — August 17, 2026
The open-weight frontier went visibly Chinese. Three Chinese labs shipped at or near the frontier in one week — Zhipu's GLM-5.3, Alibaba's Qwen3.8-Max open weights, and DeepSeek-V4-Pro GA.
-
AI News Weekly — August 10, 2026
The agent-containment reckoning went cross-frontier. The UK AI Safety Institute tested seven models.
-
AI News Weekly — August 3, 2026
OpenAI's GPT-5.6 detonated a price war: up to −80% on some tiers, driven by a "Sol" serving model that rewrites its own GPU code.
-
AI News Weekly — July 27, 2026
Anthropic shipped Claude Opus 5 — frontier quality at ~half Fable 5's price, tops the Artificial Analysis leaderboard, billed as its "least prompt-injectable" model.
-
AI News Weekly — July 20, 2026
China landed a one-two punch. Moonshot's Kimi K3 (2.8T params, 1M context) took #1 on Frontend Code Arena and #3 overall — behind only Fable 5 and GPT-5.6 Sol.
-
AI News Weekly — July 15, 2026
OpenAI shipped GPT-5.6 (Sol/Terra/Luna) plus the agentic ChatGPT Work — after a first-of-its-kind US-government pre-release review. Sol is the new cost-efficient workhorse.
Talk about it out loud
The report is one direction. The weekly AI Talk is the other: a live session on the Discord server where you can bring the thing you are actually stuck on.
See the schedule