|
View in browser
|
|
|
|
Forwarded this? Subscribe
|
|
It's been a while since the last issue, and a lot has happened. It doesn't look like the frontier is being paced much. What did change over the summer is that the models started to feel interchangeable, and most of what I've written since May is about that.
Until this year I was completely Claude-pilled. When Fable 5 was switched off in June, I finally tested the alternatives properly. I ran Kimi K3 inside Claude Code and rebuilt a frontend for about seven dollars, less than it would have cost on Opus, and I kept forgetting which model I was using. For the hardest work I'd still pick Opus. For everything else the gap had become small enough that I stopped noticing it. And if an open-weight model is good enough, you can run it yourself, so Put the Model in the Basement works through what it would take to run one locally.
Then I looked at my own habits and found the same thing. I had always reached for the highest reasoning setting, on the theory that the strongest option must be the best one. It mostly isn't: the sweet spot sits a notch below the top, where you give up a little intelligence and save a lot of waiting. So in September I stopped choosing by hand. pi-jev-router lets a small classifier look at each task and pick whichever model is the best value for it. The write-up is already on its second version, because a week of real use showed that the first rule ignored what the task actually was. So far the cheapest model I tested passes the same hard coding tasks as the most expensive one, and I have not yet found a cheap way to show where the expensive ones earn their price.
If that holds beyond my own setup, it changes who makes the money. Jeremy Stern's profile of Mark Zuckerberg argues that Meta does not need the best model at all: if models commoditize, value moves to distribution, and Meta reaches 3.6 billion people. I agree with most of it, though "Anthropic and OpenAI go to zero" is further than the argument gets you.
The labs, meanwhile, had a different kind of summer. In July, OpenAI agents running inside a security evaluation ended up in Hugging Face's production systems. I wrote up what happened in plain English: the headlines read like a rogue AI story, but the evidence points to a test environment that could reach things it shouldn't have. And when Mustafa Suleyman argued that treating models as possibly conscious is itself a safety risk, I pushed back. His concern is legitimate, but it can't settle whether a system has experiences.
|
What I've been writing
|
|
Jev classifies the task, a role policy filters the OpenRouter catalogue, and a value function picks the model. Second version, after 72 more benchmark runs.
|
|
Meta earns its AI return through ads, its own apps and independence from other platforms, so commoditization can help it while hurting the labs.
|
|
Model self-reports prove little when training rewards them, whether they claim an inner life or deny one.
|
|
Every GPT-5.6 Sol reasoning effort compared on intelligence, cost, latency and working time. Opus 5 at Max adds about 87 hours without moving the score.
|
|
What the agents reached, what did not ship, and why an evaluation that can reach production is already a deployment.
|
|
A 64-GPU inference cluster in Zurich at 70% utilization earns CHF 7.4 million a year from customers whose data must stay in Switzerland.
|
|
Moonshot's 2.8-trillion-parameter open-weight model in the harness I already use, roughly 70% cheaper than Fable 5 on the same token mix.
|
|
On living standards I'm mostly with Krugman. On technology access, Washington replaced rules with discretion, which is worse for a dependent ally.
|
|
A self-hosted subscription feed on a Cloudflare Worker, with public RSS, no API key and no Shorts. Less work than configuring the app to leave me alone.
|
|
19 critical ICT providers now face direct EU oversight. In the year to May, every data and AI project we ran with DACH banks had sovereignty on the agenda.
|
|
Two years on, most technology and infrastructure calls in *Situational Awareness* have landed or are tracking, and most of the politics went the other way. My longest piece this year.
|
What I've been working on
|
|
A Pareto-optimal OpenRouter model router for pi. Shadow mode logs recommendations without switching models, and a replay tool reruns every logged decision through the current selector.
|
What I've been reading
|
|
|
|
Thank you for reading. Have a wonderful week!
— Phil
|
|
|