Are humans cheaper?

On Tokenmaxxing And The Thermodynamics Of Work


I caught myself last week looking at my own token usage and asking whether it actually said anything about my contribution to my business.

It is a strange question to sit with as someone who runs an AI infrastructure company. I use different models for different tasks and the tools are part of how I think now. So when I started questioning whether the meter on the side of my desk was a measure of anything at all, I knew I had stumbled into something worth following.

And then the receipts started arriving.

Last week, Bryan Catanzaro, NVIDIA's VP of applied deep learning, told Axios that "for my team, the cost of compute is far beyond the costs of the employees." Around the same time, The Information reported that Uber's CTO Praveen Neppalli Naga had already exhausted the company's entire 2026 AI budget on token costs alone, four months into the year. His own words: "I'm back to the drawing board because the budget I thought I would need is blown away already." Anthropic has raised its pricing to manage demand. Goldman Sachs ran a survey and found that large companies are overrunning their AI budgets by orders of magnitude. Gartner is now forecasting that more than 40 percent of enterprise AI agent projects will be shut down by the end of 2027, citing escalating costs and unclear business value.

This is not the story anyone was sold. The pitch for enterprise AI was relentless cost reduction. Faster, cheaper, more scalable than human labour. Companies cut headcount on the assumption that the replacement would be a fraction of the price. The numbers coming in suggest something different. The replacement is, in some cases, more expensive than the workforce it displaced.

So a question that would have sounded absurd eighteen months ago is now sitting on quarterly earnings calls…

Are humans cheaper?

Then there is Meta, which has taken this question and bent it into something you almost cannot believe is real until you read it twice.

In April, an internal Meta dashboard called "Claudeonomics" surfaced publicly. It is a leaderboard that ranks the company's 85,000 employees by how many AI tokens they consume. The top users get titles like Token Legend or Session Immortal and even a Cache Wizard. In a thirty-day window, Meta employees collectively burned more than 60 trillion tokens. The single highest individual user consumed 281 billion tokens, which at current pricing translates to compute costs north of a million dollars for one person.

Meta's CTO Andrew Bosworth has publicly endorsed the practice. He has said his best engineer spends the equivalent of his salary on tokens but is '5x to 10x more productive,' and described the spend as easy money with no limit. Jensen Huang, around the same time, said he would be "deeply alarmed" if a $500,000 engineer was not burning at least $250,000 worth of tokens a year. Meanwhile, employees have reportedly been caught running bots in loops overnight to climb the rankings. An investor called the practice "incredibly stupid" and compared it to the discredited 1980s habit of measuring engineers by lines of code written.

I want to sit with this because I think it is the strangest part of the whole picture.

For a century, productivity meant output per unit of input. Less effort, more result. Think lean manufacturing or Six Sigma where every management theory of the modern era has rewarded the worker who solved the problem in fewer moves. Meta's leaderboard inverts the entire logic. The engineer who solves a task in two prompts is failing. The engineer who runs an agent in a loop for six hours is winning. We have built a corporate culture, briefly, where wasting electricity is the metric.

And the electricity is not theoretical.

According to cognitive biologist Ladislav Kováč, a human brain runs on roughly 20 watts. About the same as a dim lightbulb, burning in your head whether you are solving a problem or staring at a wall. And according to the US Congressional Research Service, a modern hyperscale data centre exceeds 100 megawatts, enough to power around 80,000 homes. That means about five million times the power, doing what until recently a human brain did for almost free.

So the cost calculation that drove much of the AI labour wave was incomplete. It compared the salary of a person to the subscription price of a tool, and concluded the tool was cheaper. It did not account for what the tool would cost when usage scaled the way agents make usage scale. It did not account for what the tool would cost when providers stopped subsidising prices. It did not account for the energy. And it did not account for the fact that a human being is, by any reasonable thermodynamic measure, an extraordinarily efficient piece of infrastructure.

I want to be careful here, because the conclusion is not that AI is bad or that the leaders who deployed it were wrong.

The productivity gains in the right contexts are tangible like at Uber, nearly 95 percent of engineers now use AI tools every month, close to 70 percent of committed code comes from AI, and around 11 percent of live backend updates are written by AI agents with no human in the loop. So, the tools are working. The budget overrun is due to the tools succeeding faster than anyone modelled.

The question I think business leaders should be asking is more specific than "is AI worth it." The question is what kind of AI, used how, for what.

The reason a token bill scales the way it does is that most general-purpose agents are paying to think the same thoughts again. Every time you ask a general agent to do a job it has done a thousand times before for a thousand other companies, it sits down, reads the question, considers its options, sketches a plan, tries something, checks the result, and tries again. The plan was knowable and the token bill is what it costs to rediscover something that was already known. A narrower, more deliberately built tool starts from the answer which means the work has been done in advance, by people who understood the problem. What looks like a smaller, cheaper agent is actually the same agent with the thinking pre-paid.

The reason tokenmaxxing exists at Meta is that the tools are doing thinking that, in many cases, has already been done. The engineer is paying for the agent to rediscover something a more deliberately designed system would already know.

So the question for a leader making AI decisions this quarter is no longer just "should we adopt." It is whether the AI you are adopting is the kind that thinks every time, or the kind that has been taught. One of those scales. The other one bills you.

If you have rolled out general-purpose agents across an organisation without knowing your unit economics, your variable cost structure, your decommissioning plan, or your fallback if the provider raises prices or sunsets the model, you have taken on a liability you have not priced. The Meta leaderboard is funny because it is happening to someone else. The same maths is happening on a smaller scale in companies that cannot absorb a Goldman-Sachs-survey level of overrun.

I started this piece by asking whether my own token consumption was a measure of my contribution. I do not think it is. I think it is a measure of how much work my tools are doing, which is a different question, and I think conflating the two is the error at the centre of the current moment.

We did not replace labour. We simply changed the meter.

And the cheapest worker in the building?

Still runs on toast.

Next
Next

You Asked For This