All articlesHalf a billion tokens in three days: who's going to pay for this?

Half a billion tokens in three days: who's going to pay for this?

By Allan Clempe

ai-agentscostspricingopinion

Last week, one of our developers at Blastin burned through 574 million tokens in three days of work. Not on some runaway batch job. Just normal, day-to-day development: 47 Claude Code sessions, 9.3 hours of agent runtime, steering agents across three client projects while logging just under 32 hours of his own time.

At current API prices, that came to US$413.86. Run the same workload through a top-tier model like Opus and you're looking at roughly US$500 for a single developer, in a single burst of work.

We know this because Clocktopus now tracks agent costs alongside the human timeline it already builds from git activity. Every commit ties the agent sessions behind it to a ticket, a project and a person. Which means for the first time, we can put an actual dollar figure next to the question everyone in this industry has been quietly avoiding:

Is any of this sustainable?

The subscription maths doesn't add up

Most developers today aren't paying API prices. They're on flat-rate subscriptions of US$20, US$100, maybe US$200 a month. Our developer consumed US$400+ of inference in three days. Extrapolate that over a month of steady agent-heavy work and the gap between what heavy users cost and what they pay is enormous.

The AI labs are wearing that gap right now, funded by investors who will, at some point, want their money back. The rate limits are already tightening. The "unlimited" plans are already gone. Subscription price rises aren't a possibility, they're a certainty. The only question is how steep and how soon.

When that correction lands, it lands hardest on the businesses that have rebuilt their delivery model around agents. Businesses like ours.

The market is already splitting

You can see the pressure playing out in the tools themselves. At the premium end sits Claude Code, which is what our team runs day to day. The frontier models behind it are the best at agentic coding, and they're priced like it. That's where our US$400-in-three-days figure comes from.

At the other end, OpenCode is attacking on price. Their Go plan bundles a roster of open-weight models (DeepSeek, GLM, Qwen, Kimi) for a flat US$10 a month, with the first month at US$5. The bet is straightforward: Chinese open-weight models now cost a fraction of a cent per thousand tokens, they're closing the gap on coding benchmarks, and for a lot of day-to-day work "good enough at 20x cheaper" wins.

There's a catch, though, and it's not technical. Those cheap models are Chinese: DeepSeek, Zhipu's GLM, Alibaba's Qwen, Moonshot's Kimi. Plenty of enterprise and government clients simply won't allow them anywhere near their codebase, whatever the hosting arrangement. Anyone who has filled out an enterprise security questionnaire knows the model provider is about to become a line item on it. So the cheap option isn't equally available to everyone: a shop serving startups can run its whole pipeline on DeepSeek, while one serving banks or government is locked into frontier pricing whether it likes it or not. Compliance, it turns out, has a token price too.

If the cheap bet pays off, the sustainability question changes shape. Maybe the frontier models become what senior developers are today: expensive, reserved for the hard problems, while the cheap open-weight models do the grinding. Which raises its own commercial question: do you quote a client frontier-model rates, or DeepSeek rates?

The questions nobody has answered yet

For software development houses, this creates a set of commercial questions that the industry simply hasn't worked through:

How do we absorb the cost? If a developer's tooling bill goes from $50 a month to $2,000, that comes straight out of margin unless something else changes. For a consultancy running lean, that's not a rounding error.

How do we price it upfront? Fixed-price proposals assume you know your input costs. Token consumption varies wildly by task, by model, by how well the agent is steered. Do you estimate tokens the way you estimate hours? Pad the quote and hope? Nobody has a mature answer.

Do customers start paying for tokens? Time and materials billing has been the backbone of professional services for decades. Are we heading toward T&M&T (time, materials and tokens) where clients see agent consumption as a line item next to human hours? Some clients will accept that. Others will ask why they're paying for your tools at all.

You can't answer any of this without data

We don't have answers to those questions yet. What we do have is the data to start reasoning about them, and that's the part we could actually build.

Clocktopus now shows agent time and spend at the task level. In our own report from those three days, one ticket cost $62.81 in agent spend across 10 commits, while another ticket on the same project got done for $7.34. One client project ran at US$16.49 of agent spend per human hour; another at US$9.83. We can even see US$4.80 of "unattributed" spend: agent sessions that ran outside any repository mapped to a project. Unsteered work, burning money with nothing to show for it.

The reports filter by person, by agent, and by language model. That last one matters more than it sounds: when you can compare what different models cost against what actually shipped, you're no longer choosing models on vibes and benchmarks. You're choosing them on delivered features per dollar. Same for developers. Not to rank people on a leaderboard, but to understand which combinations of human, agent and model are actually productive, and which are just expensive.

That's the level of visibility development shops will need before they can absorb these costs sensibly, price them into proposals, or pass them through to clients with a straight face.

Where do you think this lands?

Our bet is that agent spend becomes a first-class line item in software delivery, estimated, tracked and billed with the same rigour as human hours. But we could be wrong. Maybe inference gets so cheap it never matters. Maybe clients refuse to pay for tokens and consultancies quietly eat the cost forever. Maybe fixed-price work dies entirely because nobody can predict consumption.

So, an open question for anyone running a development shop, or paying one: when the subscription prices correct, who pays? The lab, the dev house, or the client?

We'd genuinely like to hear where you think this is going.

Start measuring what your team ships

Output, cost and what your agents burn. Read from the commits, pull requests and agent runs you already have.

Start free

Free for single developers.