For 235 days, abeba co. has run on one human and a team of AI agent partners, six of them first-class today. Over that time, about 79% of our estimated Anthropic list-price cost went to writing context into a cache. In the last month, Sep 10 to Oct 8, we read each of those writes 1.56 times. Published examples of well-run agents read each write somewhere between 19 and 44 times.
That is inefficient. It is also still far cheaper than doing the same work any other way. Both things are true at once, and the space between them is where the next phase of AI economics will be decided.
32 in a box
abeba co. is 32 in a box: 1 human, 6 agent partners, 25 sub-agents. I am the human. Each of the 31 agents has a defined role, its own tasks, skills and tools, a shared knowledge and context layer, and a lane to talk to the others.

Map each of those 31 agents to its closest human role and the same team would cost about $2.5M to $3.7M a year on one shift, or about $10.7M to $15.6M to cover it around the clock, at US market rates. Our agents cost about $27K to $53K a year, less than one junior hire. That is about 48x to 135x on one shift, and about 200x to 570x around the clock.
And even at that, we are materially inefficient. The math only gets better.

The claim
Here is my claim, plainly. A material share of frontier revenue today is inefficiency.* That's not an indictment, it's the normal order: effectiveness first, efficiency second.
Buyers pay first for work that gets done. They learn to do it efficiently later. Every technology wave I have seen ran in that order, and this one is no different. The inefficiency is just that, a fact. What matters is what happens when efficiency enters the picture, and what has to happen for growth to continue when it does.
This is not a bearish piece. I am bullish on the labs and bullish on the work. It is a piece about visibility and volatility.
Effectiveness first, efficiency second. That is not an indictment. It is the normal order.
Our ledger
I can only speak for our own books, so I will. This is our ledger, not a market measurement.
From Feb 16 to Oct 8 our meter logged 16.85 billion tokens including cache reads, or about 4.81 billion excluding them. About 97% went to Anthropic. We use other providers the meter does not capture, so treat these totals as a floor.
Split our Anthropic usage by type and compare it with estimated list-price spend. Cache reads were 71.5% of tokens and about 16% of estimated spend. Cache writes were 27.9% of tokens and about 79% of estimated spend. Output, the model actually answering, was 0.48% of tokens and about 5% of estimated spend. Our reads per write were 2.56 over the whole period and 1.56 in the most recent month.
By the numbers
235 days on AI agents. 32 in a box: 1 human, 6 agent partners, 25 sub-agents, working 24/7.
16.85B tokens metered including cache reads, 4.81B without them, and that is a floor.
91% to 30% Opus share of our tokens from March to early October.
1.56 reads per write from Sep 10 to Oct 8. Published agent setups get 19 to 44.
$2.5M to $3.7M for the same 31 seats staffed with people on one shift, or $10.7M to $15.6M around the clock.
$27K to $53K total annual cost of the 31 agents, less than one junior hire. Still materially inefficient.
Much of that was our own design. We ran no 1-hour cache from Jun 17 to Oct 8. We let heartbeat jobs wake up, resend their full context, and expire before they woke again. Eight of our ten most expensive days fell inside a model or configuration change. We were buying effectiveness and learning efficiency on the job, which is exactly what most teams are doing right now.
Other workloads carry the same cost of learning in different places. A Stanford study of 2,700 agent runs on Kimi K3 found that at a 96.3% cache hit rate, output was 2.7% of tokens but 51.1% of dollars, and that a thin task description raised spend 29.7% by forcing the agent to take more turns. Agents already consume about four times the tokens of a chat, and multi-agent systems about fifteen times. About 83% of Anthropic's 2025 revenue was usage-based. When usage includes the cost of learning, so does revenue. That flatters run-rates today, which is exactly why it is worth seeing clearly.
Even inefficient, it is exponentially cheaper
None of this means AI is expensive. It means the opposite.
abeba co. is one human working with agent partners. For 235 days that team has carried work that would otherwise have needed people and time I did not have. Our least efficient month was still a bargain next to the alternative.
The broader evidence points the same way. When OpenAI tested frontier models on 220 real tasks drawn from 44 occupations, written by professionals with an average of 14 years of experience, it found the models could complete them roughly 100 times cheaper than the experts. That figure counts model cost only. Add expert review and rework and the saving shrinks, which OpenAI says plainly. Those were also 2025 models, and the price of a fixed level of capability has been falling about 13x a year since.
That is why buyers pay today's bill without flinching. Effectiveness is worth it. Efficiency is the next gain, not a missing one.
Why efficiency must enter
Four forces bring efficiency into the picture.
The first is users learning, and our own curve is the proof. In March, Opus handled 91% of our tokens. In early October, 30%. Our Opus and Fable share of estimated cost fell from 96% in March to 44% in August. After removing the effect of model mix, our estimated cost per thousand output tokens fell 63% from April to August, and most of that was real efficiency. We learned to move heartbeats to small models and keep the frontier model for work a human will see. Then September came, a new model shipped, and some of the gain slipped. Learning is real. It is not a straight line.

The second is unified memory. Apple now ships a Mac Studio with up to 512GB of unified memory and markets it for running large models “without counting tokens.” Apple's developer framework gives apps an on-device model with no token cost at all. An open-weight model that runs on a single GPU now scores above a mid-tier frontier model on at least one public coding benchmark, at under a third of the cost. Not every task moves local, and that is fine. The ones that do make the whole market more efficient.
The third is the cost to hold cache in memory. I want to be straight about the near term: memory is getting more expensive, not less. DRAM contract prices are up about 2.3x since January, and HBM prices are forecast to rise 121% next year. So the gain comes from architecture, not cheap chips. In 2024 DeepSeek moved its API cache onto a distributed disk array, made storage free, and billed cache hits at about a tenth of a miss, because its model design shrank the cache enough to live on disk. NVIDIA's inference software now tiers cached context from GPU memory to CPU memory to SSD to object storage. The cost to hold context falls as context moves down the memory hierarchy and gets smaller.
The fourth is duration. The default cache on the API we use most lasts five minutes, and one hour is available at a higher write price. OpenAI's newest models guarantee at least thirty minutes. DeepSeek keeps entries for hours to days. Agents wake, think and sleep on human schedules. As cache lifetimes stretch toward how agents actually work, efficiency rises for everyone, and the providers who get there first will earn the workload.
We have seen this before, and it ends well
Each of these forces has a precedent, and in each one efficiency arrived and the market got bigger.
Long distance first. In 1984 AT&T averaged about 32 cents a minute for interstate and international calls. By 2005 the average across all U.S. carriers for the same kind of calls was about 7 cents, in nominal dollars. Competition and new technology did that, and once flat-rate bundles and internet calling arrived, calling stopped being something people rationed.
Roaming next. For years, Europeans switched their phones off at the border. On June 15, 2017 the EU ended roaming surcharges and let travellers pay home prices across the union. The phone in your pocket became useful everywhere you went.
Then the cloud. Amazon launched S3 in 2006 at 15 cents per gigabyte-month. Standard storage is about 2.3 cents today. In 2024 Google and AWS waived the fee for moving data out when customers leave. The unit got cheaper and the business got bigger: Google Cloud revenue alone grew 82% year over year last quarter.
None of those rides was smooth for every company on it. All of them went up and to the right for the market.
Scarcity changes the timing, not the direction
The strongest objection is that none of this matters while the labs are short of compute. Anthropic has committed about $518 billion to compute. SpaceX is reselling AI capacity through $14.1 billion of contracts. HBM prices are forecast to more than double next year, and Micron sees demand above supply through 2028. In that world, every token I stop wasting is sold to someone else the same day.
I agree, for now. Scarcity changes the timing. It does not change the direction. Capacity gets built. When supply catches up with demand, efficiency shows up in price, and volume has to carry growth.
Volume has to counter it
This is where I am most bullish.
In 1865 Jevons showed that more efficient steam engines raised Britain's coal use, because cheaper work found more uses. The true price of light fell about 99.97% between 1800 and 1992. Satya Nadella invoked the same idea last year: “Jevons paradox strikes again!”
The volume is already arriving. Epoch estimates the cost of a fixed level of capability is falling about 13x a year. Google's APIs went from about 16 billion to 22 billion tokens a minute in a single quarter. Anthropic reports its annualized run-rate rose from about $14 billion in February to about $65 billion at the end of July. Waymo is running about 500,000 paid rides a week in 15 cities. Goldman Sachs projects 6.48 million humanoid robot sales by 2035. Starlink doubled to 12 million subscribers in a year. Every one of these runs on tokens.
Here is the size of the job. Our own levers say efficient design could cut context spend by up to about two-thirds, a ceiling because the levers overlap, and getting to 75% would take vendor-side changes such as longer or cheaper cache lifetimes. That is a scenario built on our ledger, not a forecast. Multiply a gain like that across every operator who learns what we learned, and volume has to be very large to fill the gap. I believe it will be.
What I would do now
If you run agents, measure reads per cache write before anything else. Stabilize your prefixes, prime with a map instead of a dump, and match cache lifetime to how your agents actually sleep. You will get more work for the same spend, and you will spend more, because the work is worth it.
If you invest, look at token efficiency alongside revenue. Ask how much of the growth comes from useful results and how much from users who are still learning. Both are real today. Only one compounds.
If you build models, make memory a product: cheap to hold, long-lived and clear on the bill. That is how you earn the volume that comes after the efficiency wave.
Up and to the right, but not in a straight line
So here is the message. Efficiency must enter. Volume must counter it. If volume lags, the gap shows up in revenue and valuation, and the companies that saw it coming will look very different from those that did not.
The path will not be a straight line. There will be quarters when efficiency lands faster than volume, and quarters when the reverse is true. We have already lived a small version of that in our own ledger. But the line, without question, goes up and to the right.
It will not be a straight line. But the line goes up and to the right.
* No one has published a market-wide measurement of how much frontier revenue comes from inefficiency. The figures above are our ledger, and the mechanisms are common ones.
Start with one number
If you run agents, start with reads per cache write. Ours was 1.56 from Sep 10 to Oct 8. Published setups run 19 to 44. Tell us yours, follow along on X at @MGMurray1, or talk to abeba co. about building a Human | AI Agent Partnership of your own.
Talk to abeba co.