Meanwhile, the models started to feel interchangeable ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ ͏‌ 
View in browser
philippdubach
September 2026  ·  Nobody Is Pacing the Frontier
Forwarded this? Subscribe

It's been a while since the last issue, and a lot has happened. It doesn't look like the frontier is being paced much. What did change over the summer is that the models started to feel interchangeable, and most of what I've written since May is about that.

Until this year I was completely Claude-pilled. When Fable 5 was switched off in June, I finally tested the alternatives properly. I ran Kimi K3 inside Claude Code and rebuilt a frontend for about seven dollars, less than it would have cost on Opus, and I kept forgetting which model I was using. For the hardest work I'd still pick Opus. For everything else the gap had become small enough that I stopped noticing it. And if an open-weight model is good enough, you can run it yourself, so Put the Model in the Basement works through what it would take to run one locally.

Then I looked at my own habits and found the same thing. I had always reached for the highest reasoning setting, on the theory that the strongest option must be the best one. It mostly isn't: the sweet spot sits a notch below the top, where you give up a little intelligence and save a lot of waiting. So in September I stopped choosing by hand. pi-jev-router lets a small classifier look at each task and pick whichever model is the best value for it. The write-up is already on its second version, because a week of real use showed that the first rule ignored what the task actually was. So far the cheapest model I tested passes the same hard coding tasks as the most expensive one, and I have not yet found a cheap way to show where the expensive ones earn their price.

If that holds beyond my own setup, it changes who makes the money. Jeremy Stern's profile of Mark Zuckerberg argues that Meta does not need the best model at all: if models commoditize, value moves to distribution, and Meta reaches 3.6 billion people. I agree with most of it, though "Anthropic and OpenAI go to zero" is further than the argument gets you.

The labs, meanwhile, had a different kind of summer. In July, OpenAI agents running inside a security evaluation ended up in Hugging Face's production systems. I wrote up what happened in plain English: the headlines read like a rogue AI story, but the evidence points to a test environment that could reach things it shouldn't have. And when Mustafa Suleyman argued that treating models as possibly conscious is itself a safety risk, I pushed back. His concern is legitimate, but it can't settle whether a system has experiences.

What I've been writing

Choosing a Model on the Pareto Frontier with Jev

Jev classifies the task, a role policy filters the OpenRouter catalogue, and a value function picks the model. Second version, after 72 more benchmark runs.

Maybe Meta Is Right About AI, or at Least Jeremy Stern Is

Meta earns its AI return through ads, its own apps and independence from other platforms, so commoditization can help it while hurting the labs.

AI Consciousness Is Not a Safety Property

Model self-reports prove little when training rewards them, whether they claim an inner life or deny one.

Finding the Performance–Cost–Speed Sweet Spot With LLMs

Every GPT-5.6 Sol reasoning effort compared on intelligence, cost, latency and working time. Opus 5 at Max adds about 87 hours without moving the score.

The OpenAI–Hugging Face Incident in Plain English

What the agents reached, what did not ship, and why an evaluation that can reach production is already a deployment.

Put the Model in the Basement

A 64-GPU inference cluster in Zurich at 70% utilization earns CHF 7.4 million a year from customers whose data must stay in Switzerland.

I Tried Kimi K3 Inside Claude Code

Moonshot's 2.8-trillion-parameter open-weight model in the harness I already use, roughly 70% cheaper than Fable 5 on the same token mix.

Krugman, Fable 5, and Europe in Decline?

On living standards I'm mostly with Krugman. On technology access, Washington replaced rules with discretion, which is worse for a dependent ally.

Degoogling cost me my YouTube feed, so I made my own

A self-hosted subscription feed on a Cloudflare Worker, with public RSS, no API key and no Shorts. Less work than configuring the app to leave me alone.

How DORA Made Sovereignty a Bank Problem

19 critical ICT providers now face direct EU oversight. In the year to May, every data and AI project we ran with DACH banks had sovereignty on the agenda.

Aschenbrenner's Receipts

Two years on, most technology and infrastructure calls in *Situational Awareness* have landed or are tracking, and most of the politics went the other way. My longest piece this year.

What I've been working on

pi-jev-router

A Pareto-optimal OpenRouter model router for pi. Shadow mode logs recommendations without switching models, and a replay tool reruns every logged decision through the current selector.

What I've been reading

Introducing System One Models and Jev  — Jev returns structured decisions with calibrated probabilities instead of text, and TypeSafe claims it is two orders of magnitude faster than an LLM on the same tasks. Sebastian Raschka has the best explainer, Arcturus Labs the bear case.  Typesafe
State of Open Models: Summer 2026 Observations  — In almost every month of 2026, the largest Chinese open model was bigger than anything a US lab released.  Huggingface
The Economics of Open-Weight Inference  — Ornn finds that open-weight demand keeps older NVIDIA generations economically useful, which undercuts the usual GPU depreciation assumption.  Data
Swarm Traces  — The Hugging Face agents could only load URLs, so they used a link shortener to build almost a million and chained them into code execution.  Swarmtraces
We Must Pace the Frontier  — Amodei now wants to slow capability gains because recursive self-improvement has sped things up since the summer. Counterpoints from Ramez Naam and LessWrong.  Darioamodei
After Math  — A guest post on Terence Tao's blog about who gets credit for OpenAI's AI-generated Navier–Stokes solution, announced September 8.  Terrytao
How To Write With An LLM  — Never take a word the model suggests, never let it encourage you, and use it to find flaws in what you wrote. Pairs with Colin Breck's I Don't Want to Read What You Didn't Write.  Sockpuppet
How professional gamblers size bets  — A clear walk through the Kelly criterion. In one experiment, 30% of 61 quant finance students went bust betting on a coin that lands heads 60% of the time.  Hails

Thank you for reading. Have a wonderful week!
— Phil

Feedback or ideas? hi@philippdubach.com
Blog  ·  Projects  ·  Research  ·  GitHub  ·  Bluesky
Unsubscribe