Skip to main content
12 min read

The volume doesn't match the output

This week a researcher quit Anthropic warning AI could kill us all. The same week, a company promised a whole game from one sentence. The noise is deafening. The software mostly can't back it up.

ByJames Dodd

Note filed under:

This week a researcher at Anthropic resigned and said the quiet part out loud. Jacob Coxon wrote that "the people building AI earnestly believe that it could kill us all by the end of the decade," and, to head off the obvious dismissal, "this is not a marketing stunt... I hear the same people express fear privately." His alignment-lead colleague Evan Hubinger replied in public with a number: "we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." Over at OpenAI, another researcher, Zoë Hitzig, quit the same week with an essay titled, more or less, we are making the mistakes Facebook made.

Jacob Coxon's post on the day he resigned. It passed ninety million views inside a day.

Ten per cent. The end of everyone. From the people with the best view of the machine. I want to take that seriously, and then put it down.

If any other company held something with a credible one-in-ten chance of ending all human life, what would a government do? It wouldn't hold a hearing. It would send people to the building and take the thing away. We regulate anthrax labs and nuclear fuel and the water supply on exactly this logic. The gap between what these researchers say they believe and how the world is behaving is the part that should unsettle you, whichever way you lean. Either the risk is real and we are strolling past it, or the risk is overstated and some of the most serious people in the field are frightening the public for reasons of their own.

I lean towards the second, and I want to be honest that it's a lean, not a proof. Here's what pushes me there.

Start with the basics, because that's the part the end-of-the-world talk skips over. These same models, the ones we're told might wipe us out, will quietly refuse to do ordinary work. Last week I asked one to update a single row in a database, a small, essential, entirely ordinary change, and it refused. Not couldn't. Wouldn't. It decided it shouldn't. Somewhere a guardrail written by someone I'll never meet sat between me and a job I understood and it didn't. I made the change myself in a few seconds.

Two ways to read that. Either the caution is real and clumsy, the labs so nervous about what an agent might break that the brakes now catch ordinary work too. Or the refusal is doing a different job. A model that solemnly declines to touch your database looks careful, and careful looks powerful: see how capable this thing is, we had to hold it back. In an industry that sells its products by telling you how dangerous they are, a well-placed refusal is on message. I can't prove which it is, and my money's on the second. But hold the two thoughts together either way: a machine credited with ending civilisation, that also won't edit a database row when told to.

We hit a smaller version on this very site. We asked a model to add an image to a blog post. It added the wrong one, at the wrong aspect ratio. A trivial task, confidently botched. Neither of these is a catastrophe. Both are the same reminder: the thing does not know what it does not know, and it is nowhere near as capable as the billboards insist.

Go past the basics and the picture doesn't improve, it just gets more specialised. I build games as a hobby. Nothing serious: a Zelda-ish 3D adventure, a point-and-click tycoon thing, a 2D platformer, a first-person shooter that will never see daylight. Over the last year I've put the best models available to work on all of them, the ones the headlines are about.

They cannot reliably reason about 3D space. I've watched a model get the Z axis backwards, place things behind the camera, build a room inside out. Once it spent the better part of an hour explaining, with great confidence, why something was broken. The fix took me ten seconds. It was a basic thing.

That isn't a knock on the technology as such. On ordinary software, Python and TypeScript, the common web frameworks, these models are genuinely strong, and we use them every day. The trouble is that everything built that way is starting to look the same, because the models learned from what already exists and now write more of the same back. Game code is different. There's less of it in the open, less to learn from. Maybe that changes. Right now the machine is confident and wrong, and the only reason I know it's wrong is that I can see the game running and it can't.

So here is my worry, and it isn't the horror film.

The danger isn't a machine that's too clever. It's a machine that's wrong, and a room full of people who've been told it isn't.

The extinction talk and the wrong image are the same failure seen from opposite ends: everyone believes the thing can do more than it can. Somebody builds a system on that belief. Somebody wires it to something that matters. Don't press the button. Oh. You pressed the button.

A red STOP button on the grab pole of a bus, lit by warm side light.
The button we press without thinking.Photo by Pedro Oliveira on Unsplash

Which raises the real question, quieter than extinction and more useful. Why do we keep handing it the button in the first place? Why are companies so keen to give the thing more reach than it has earned?

The honest answer is usually money: growth to show investors, a launch to out-shout the launch before it. That's not a conspiracy, it's an incentive, and it's worth naming plainly because it's the engine under all of this. Which leaves the rest of us with a choice we don't often say out loud. We can keep buying from, and building on, whoever ships fastest and claims most. Or we can spend our money and our attention on the people being straight with us, and give the honest end of the market room to grow. Most of us feel trapped between the two. We are less trapped than we think.

The oversell in miniature

If you want the whole pattern in one product, look at what launched the same week as the resignations. Higgsfield put out Games 2.0, running on OpenAI's newest model. A VP there put the promise on LinkedIn: "We turned GPT-6 into your personal game producer. Describe the game you want and it handles the rest: mechanics, world, characters, props, in any genre, 2D or 3D." Multiplayer, he added, "is a toggle." One sentence in, a game out.

Go to the launch material and here's what you won't find: a game. There's the interface and there's the claim. No footage of the thing actually being fun to play. The examples that are around are the sort of basic, template-shaped games a hobbyist would build in a weekend, which is fine, except that isn't what "your personal game producer" is selling.

I've done the honest version of this experiment myself, on my own games, with the same class of model. The result: the models make no real difference to how fast the work goes, and none at all to whether the result is any good. They're trained on open-source examples, so what you get back is a remix of other people's uncredited work. And the one thing they cannot do is the whole job. Nobody is building a Grand Theft Auto from a sentence. What's actually gone viral this year isn't a playable game at all: it's turning a selfie into a fifteen-second Vice City trailer, or a GTA-style poster. Cover art. The surface. Never the game.

You can watch the seam. The fake "gameplay leaks" that flooded YouTube this year, the ones Kotaku spent a piece teaching people to spot, all give themselves away the same way: they run five to fifteen seconds because nothing longer holds together, the on-screen text wobbles, and the little map in the corner doesn't match the road being driven. AI can fake the trailer. It can't run the world the trailer is pretending to show.

The fairest test of the claim is someone who actually tried it in the open. Matt Shumer pointed a fleet of AI agents at a single brief, "match a modern Call of Duty," and published the lot: the full project, the prompt, and his own scorecard. Fifty-five thousand lines of code came out. His verdict on the result is one line: "The goal was to match a modern Call of Duty. It does not." His own critics scored it around five out of ten and marked most of it "amateur." And the most useful thing he found has nothing to do with graphics: when he set six agents loose in parallel, each owning a piece, they "kept breaking each other's assumptions" and the defects went up, not down. He was honest about all of it, which is more than the billboards manage. But honest or not, the output is not a game you'd want to play.

Claude of Duty running: a minute of the fifty-five-thousand-line output. Watch it for a moment and the seams show. It moves like a shooter. It isn't one you'd choose to play.Uploaded to YouTube by Dolgovec14
A game is fun or it is nothing. And AI can't make fun.

That's the crux of it. The machine can look at thousands of games and produce what it calculates fun should look like, the way it calculates what a sentence should say next. Sometimes the guess is close. It is never the thing.

Fun is the hard part of a game. It was always the hard part. Speeding up the coding, which the machine doesn't reliably do anyway, was never the barrier. No one finishes a game and says: what I loved was how quickly this was coded, and how clean the code was underneath. Nobody has ever said that. They say it was fun, or they ask for a refund.

Look at what Shumer's experiment actually was, though, and there's a lesson in it that points somewhere better. He took the biggest, most general model available and asked it to be a whole game studio. It's the same instinct we see everywhere: the summary running on the frontier model, or the AI chatbot dropped onto a website where a simple form that asks a few questions and follows the answers would be faster, more accurate and cheaper. The reflex is always to reach for the biggest, cleverest model and trust that the size will carry the job. But most work doesn't need a giant general mind. It needs a small tool that does one narrow thing well, and that doesn't even need to be an LLM.

Small and specialised turns out to be where a lot of the technology is quietly heading, and that matters for more than cost, because the loudest fears about AI all assume the opposite. The land grabs. The data centres. The stories about AI drinking a town's water to stay cool. Every one of those assumes a single future: everything routed through enormous models on enormous hardware, forever, and the planet billed for it.

Yet for all the fear mongering, a different future is already coming into focus, and there's a small model that shows both halves of the story at once. TwIL-LM3 landed in August to a wave of coverage. The headline webAI put on it was hard to miss: a family of formal-logic models "that outreason a 120B model and run on an iPhone." Three billion parameters, squeezed down to under two gigabytes, on a phone, outscoring something forty times its size on the tests they chose. Their launch graphic sells it as "the capabilities behind reliable tool calling, code generation, structured outputs, and agentic workflows." In the announcement, webAI's chief executive David Stout says he uses it "every day for writing, tool calling, and general reasoning."

webAI's launch graphic for TwIL-LM3, headed "Outreasoning the Data Center. From an iPhone", with bar charts comparing it to larger models and a line naming the capabilities behind reliable tool calling and agentic workflows.
webAI's own launch graphic. Read the line under the headline: "the capabilities behind reliable tool calling... and agentic workflows." That is the claim I went to test.webAI

So I ran a quick test. Zero tool calls landed. It mangled the JSON it was supposed to hand back (the structured text a program reads to know what to do next). It couldn't write to a file. The big models, the ones in the data centres, do all of this reliably now. This small model designed to run on a phone doesn't, nor does it on a MacBook Pro.

That's the distinction the marketing quietly steps over. Tool calling is a real capability, and the frontier models have it. A three-billion-parameter model on a handset mostly doesn't, not yet. The launch graphic sells "the capabilities behind reliable tool calling"; the company's chief executive says he uses it for tool calling every day. Then you read the model's own card on Hugging Face, where in the driest language you'll find anywhere the same company writes that its tests "do not cover code generation or tool use... so this release makes no claim about those." It is, they add, "not a general assistant." The shop window and the label on the box disagree, and the fans repeating the headline never reach the footnote. My afternoon just found the space between the two.

Here's the part that matters, though, and it's why I'm not sneering. Strip the marketing off and what's left is genuinely good. The reasoning it does in your pocket, on a low single-digit percentage of the battery, would have needed a rack of servers a few years ago, back before any model could call a tool at all. No shed of servers in a drought-hit county, no slice of a country's water to keep it cool. That's the promising bit, and the honest way to read it: this is the worst a phone-sized model will ever be, and it's already useful.

There's a charitable reading of the oversell, too, and I think it's the right one. A small team with a genuinely good small model is trying to be heard over an industry that runs on volume: the doom, the superintelligence just around the corner, the game from a single sentence. Volume is for attention, and attention is what you convert, later, into a share price. That's not a sneer, it's the shape of the business, and it does mean the loudest voices are the least useful guide to what any of this is for. If you're webAI, shouting a bit is just how you get a look at all. I'd rather back the people betting on small, honest and yours than the ones betting on enormous, frightening and theirs.

AI was meant to make the work lighter and save us time. Somewhere we started asking it to be the whole future instead.

So it's worth asking, plainly, what we wanted from any of this. We've watched it try to write a game and fail at the only part that matters. We've watched it promise to end the world by Friday. We've watched it shrink to the size of a phone. Set those side by side and the honest answer is small: not a mind that reasons like a person or frightens anyone about the end of the species, but a tool that quietly does the dull thing we asked and nothing more. The half-remembered line you know is buried in a document somewhere, found, the way you hunt for your keys but without the hunting. Flagship power in your pocket that sips a battery instead of a reservoir, taking the boring parts so a person is left free for the ones only a person can do.

That's a smaller ambition than the one on the billboards. It's also the bigger prize: not how people work, but how much of their lives they get back. The way out isn't fear and it isn't faith. It's the boring middle: find the one job the tool is genuinely good at, check that it's good at it before you trust it with anything, and don't hand it the button just because someone on stage said it was ready. That's most of what we do for the businesses we work with. It's less exciting than the end of the world. It's also true.

Written by

James Dodd

Founder of moralai.co. A design led problem solver, with a photojournalism background, who has spent the last decade building software, brands and products for small businesses and the third sector.

If this got you thinking, let's talk.

A short call, no prep, no sales deck. Bring a question, a half-formed idea, or just your thoughts, and we'll help you find what actually works, AI or not.