THE TOKEN BURN AUDIT
What 44.2 Billion Codex Tokens Built—and What Meta, Uber, OpenClaw, and the U.S. Army Can Actually Show for Their AI Spending
THE TOKEN BURN AUDIT
What 44.2 Billion Codex Tokens Built—and Why Corporate Tokenmaxxing Still Cannot Produce a Receipt
The technology industry has started treating token consumption as though it were production.
Executives celebrate adoption percentages. Employees compete on internal leaderboards. Companies advertise how much code AI wrote. Budgets disappear.
Then, when someone asks the only question that matters—
What new problem did all of that compute actually solve?
—the answers become strangely abstract.
More code.
More activity.
More usage.
More “AI-driven impact.”
That is not an output ledger.
Tokens are fuel. They are not the vehicle, the destination, or the receipt.
THE CONSERVATIVE RECORD
My Codex dashboard records:
44.2 billion lifetime tokens
1.6 billion peak tokens
6.8 billion tokens during the week of July 12, 2026
19 hours and 8 minutes in my longest chat
An 84-day continuous streak
And that number is Codex only.
It does not include my ordinary ChatGPT conversations, research, writing, conceptual development, books, articles, music, specifications or the thousands of exchanges used to reason through the architecture surrounding the code.
I do not have access to a hidden account-wide ChatGPT token counter, and neither does the assistant inside the conversation. I possess the historical record of what I created and discussed, but not undisclosed account telemetry.
Therefore, no honest audit can invent my combined number.
This comparison deliberately uses only the 44.2 billion Codex tokens that can be seen and proven.
For perspective, OpenAI demonstrated Codex building a complete racing game with approximately seven million tokens. My visible Codex count is the raw-token equivalent of approximately 6,314 of those demonstrations—although token volume alone does not imply equivalent work, complexity or quality.
Source:
https://openai.com/index/introducing-the-codex-app/
WHAT DID 44.2 BILLION TOKENS COST?
A Codex token counter is not an invoice.
Input tokens, cached input, output, model selection, context length, reasoning, Fast Mode, subscription allowances and purchased credits all change the economics.
OpenAI includes Codex usage with eligible ChatGPT subscriptions and allows additional credits to be purchased when included limits are exhausted.
Source:
https://openai.com/index/introducing-the-codex-app/
The best publicly documented comparison is OpenClaw.
OpenClaw recorded:
603 billion tokens
7.6 million requests
Approximately 100 Codex instances
Approximately $1,305,088.81 in API-equivalent usage
All within 30 days
Its creator said that disabling Fast Mode would have reduced the raw cost to approximately $300,000.
Source:
Applying those two observed rates to my 44.2 billion tokens produces the following range:
OpenClaw non-Fast estimate
Approximately $21,990
OpenClaw Fast Mode observed rate
Approximately $95,663
My 6.8-billion-token peak week at those rates
Approximately $3,383 to $14,717
That is not necessarily what I personally paid.
My direct cash expense consists of my subscription and the additional credits I purchased when necessary.
The estimate represents the approximate public API-equivalent compute value of the visible usage—not my personal invoice.
The defensible conclusion is therefore:
My 44.2 billion visible Codex tokens represent approximately $22,000 to $96,000 in comparable API usage before counting any of my ordinary ChatGPT activity.
NOW SHOW THE RECEIPT
By Receiz v96, the architecture report recorded:
1,126,260 source lines
5,242 tracked files
90,810 test lines
By v98, that had grown to:
1,158,623 source lines
5,591 tracked files
95,287 test lines
65 reviewed commits across 535 changed paths in the post-v97 range
Those releases documented executable conformance, local verification, SDK package evidence, proof memory, identity, ownership, settlement, sports, market, public-proof and developer surfaces.
This was not a prompt collection.
It was not a folder of disconnected demonstrations.
It was not an AI-generated landing page pretending to be a system.
It was a growing, tested and released architecture.
The public offline-verifier repository can be examined here:
https://github.com/kojibai/receiz_offline_verifier
At the time of this audit, the public verifier repository was aligned to the v113 release family. It publicly carried the offline verifier, Sports card verifier, studio and settlement entry points, SDK and MCP reconciliation phases, conformance materials, release-scoped product truth, regression lessons, performance findings, invariant registers and deterministic fail-closed verification behavior.
The compute did not merely produce code volume.
It produced a coherent implemented system containing:
Portable proof objects.
Offline verification.
Deterministic identity.
Ownership and custody history.
Provenance.
Settlement.
Append-only continuity.
Durable proof memory.
Receiz identity artifacts.
A public proof graph.
Wallet and ledger surfaces.
Market primitives.
Sports Arena.
Live sports cards.
Pitch Command and Play Command.
An SDK.
An MCP server.
OIDC and delegated authorization.
Webhooks.
Public application-state rails.
Conformance suites.
Governance rules.
Release gates.
Public surfaces include:
Receiz:
https://receiz.com/
BJ Klock’s Receiz identity and proof surface:
Receiz Proof Graph:
https://receiz.com/proof-graph
Receiz Developers:
Receiz Pitch Command:
Receiz application template:
https://receiz.app/
Receiz SDK:
https://www.npmjs.com/package/@receiz/sdk
Receiz MCP Server:
https://www.npmjs.com/package/@receiz/mcp-server
The problem being solved is specific:
How can an object carry its own identity, provenance, ownership, authority, state and history so that it remains independently verifiable when the platform, database, account, server or network is unavailable?
Receiz changes the authority order.
The proof object carries the record.
The application projects it.
Servers and databases synchronize, publish and index verified additions without becoming the final authority over what is true.
Known verified state can paint immediately rather than being rediscovered from a remote database every time the user opens it.
The architecture is not an ordinary application with verification features attached to it.
It is a proof-native artifact system in which the object itself carries the evidence required to determine what it is, what happened to it, who controls it and how its present state came to exist.
That is the receipt attached to my 44.2 billion visible Codex tokens.
META: MORE THAN 60 TRILLION TOKENS IN 30 DAYS
Meta employees consumed more than 60 trillion tokens during one 30-day period.
Its highest-ranked individual reportedly averaged approximately 281 billion tokens.
That means one employee consumed approximately 6.36 times my entire visible Codex lifetime total in one month.
Meta’s company-wide monthly total was roughly 1,360 times my 44.2 billion.
Using the $5-per-million reference price applied to a lower-priced Claude Opus configuration, 60 trillion tokens would equal approximately $300 million at an uncached-input list-price equivalent.
That is not Meta’s confirmed bill.
Enterprise discounts, caching, internal pricing, output-token mixtures and model selection could alter the actual expenditure substantially.
The top individual’s 281 billion tokens would represent more than $1.4 million at the same simplified reference rate.
Source:
https://fortune.com/2026/04/09/meta-killed-employee-ai-token-dashboard/
What does the public record attach to those 60 trillion tokens?
An employee-created leaderboard.
Titles such as “Token Legend.”
Employees running agents for extended periods to increase their rankings.
Executive claims that highly active engineers were five to ten times more productive.
What the public record does not provide is an equally detailed ledger saying:
These exact products were created.
These named user problems were solved.
These systems did not previously exist.
These releases generated this revenue.
These releases eliminated this cost.
These 60 trillion tokens were required to produce these particular results.
That does not establish that Meta shipped nothing.
Meta is an enormous company, and its engineers undoubtedly shipped real work during the period.
It establishes something more precise—and more embarrassing:
The company disclosed consumption at industrial scale without disclosing an equally detailed token-to-artifact ledger.
They measured the exhaust more carefully than the destination.
UBER: THE ANNUAL BUDGET DISAPPEARED IN FOUR MONTHS
Uber reportedly exhausted its entire 2026 AI-coding budget in approximately four months after encouraging adoption through internal usage tracking and competition.
There is real output in Uber’s record.
Approximately 95% of engineers were reportedly using AI monthly. Uber’s internal coding system was said to be producing approximately 1,800 code changes per week and around eight percent of all code changes.
Uber’s CEO separately said autonomous agents were generating approximately 10% of committed code.
Source:
https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/
Uber also reported useful internal workflow improvements from its later “agentic pods.”
Examples included:
Financial reporting reduced from approximately two days to ten minutes.
A capital-allocation process reduced from approximately 15 hours to 30 minutes.
Those are measurable operational results.
Source:
https://www.businessinsider.com/uber-cto-bets-on-agentic-pods-make-ai-more-efficient-2026-7
But Uber’s own president and chief operating officer acknowledged that the company still could not clearly connect its climbing token consumption to a proportional increase in useful consumer-facing features.
That is the decisive admission.
Source:
https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/
Uber has not publicly disclosed either its exact token total or the total size of the exhausted budget.
Using reports that approximately 5,000 engineers received access, combined with public enterprise estimates of approximately $150 to $250 per developer per month, produces a rough four-month baseline in the low single-digit millions.
At full participation, that would imply approximately:
$3 million to $5 million over four months.
That is an estimate—not Uber’s disclosed bill.
Power users, automation, parallel agents, premium models, multiple tools and higher enterprise allocations could increase the real number substantially.
So Uber’s receipt currently reads:
Known:
Enormous adoption, substantial AI-authored code and several impressive internal process improvements.
Unknown:
Exact token consumption, exact expenditure and a direct line connecting the exhausted annual budget to proportionally more valuable consumer-facing products.
That is not proof that Uber’s program produced nothing.
It is proof that usage statistics were allowed to outrun attribution.
OPENCLAW: THE STRONGEST LARGE-SCALE COMPARISON
OpenClaw is different because it published both sides of the ledger.
It disclosed:
603 billion tokens
7.6 million requests
Approximately 100 Codex instances
A three-person team
Approximately $1,305,088.81 in 30 days
Approximately $300,000 without Fast Mode
Its agents reviewed pull requests, scanned commits for vulnerabilities, deduplicated issues, wrote fixes, monitored benchmarks, identified regressions and generated feature pull requests.
The resulting work remained open source and inspectable.
Source:
OpenClaw used approximately 13.6 times my visible lifetime Codex total in one month.
But unlike the corporate leaderboard stories, OpenClaw provides a legitimate experimental receipt.
The output can be inspected.
The objective was explicit:
Stress-test what autonomous software development looks like when the token budget is effectively removed.
The revealing comparison is not that OpenClaw produced nothing.
It did.
The comparison is this:
A fleet of approximately 100 agents consumed 13.6 times my visible lifetime Codex total in one month, while my 44.2 billion tokens were directed through one independent builder into an entire proof-native artifact architecture and production ecosystem.
Different projects.
Different objectives.
Different funding.
Different levels of institutional support.
But both sides possess a receipt and can be evaluated.
That is what an honest comparison looks like.
THE U.S. ARMY: ONE YEAR OF TOKENS EXHAUSTED IN WEEKS
The United States Army obtained an annual Ask Sage enterprise allocation of 100 million tokens and reportedly exhausted the shared pool between May and the middle of June.
The documented uses included administrative work such as reclassifying personnel descriptions.
Employees received initial monthly allocations and were automatically replenished when their individual quotas ran out. Some personnel reportedly questioned the system’s reliability and practical usefulness.
Source:
https://www.wired.com/story/the-army-is-burning-through-its-ai-tokens/
My 44.2 billion Codex tokens are numerically 442 times the Army’s stated annual 100-million-token allocation.
That comparison requires caution.
Ask Sage may define, weight and meter tokens differently. It also supports multiple models and media types.
But the comparison still reveals how useless the word “token” becomes when organizations announce quotas without standardizing what the unit represents or pairing it with completed-work evidence.
THE ACTUAL EXPOSURE
The exposure is not that I used fewer tokens than Meta.
Of course I did.
Meta employs tens of thousands of people.
The exposure is not that Uber solved no problems.
Its AI tools clearly generated code and produced measurable operational gains.
The exposure is not that OpenClaw wasted everything.
OpenClaw provides the most transparent large-scale experimental receipt in this comparison.
The exposure is this:
A single independent builder has published a more inspectable relationship between token consumption, implemented artifacts, release evidence and the underlying problem being solved than some of the largest organizations on Earth have published for their own AI programs.
My record can be examined through:
The code.
The release train.
The source-line counts.
The test-line counts.
The architecture reports.
The offline verifier.
The SDK.
The MCP server.
The public proof surfaces.
The live sports system.
The conformance suites.
The proof objects themselves.
The companies usually publish:
The adoption rate.
The usage chart.
The exhausted budget.
The executive quote.
The leaderboard.
Then they ask the public to infer that a mountain of tokens must necessarily contain a mountain of progress.
No.
The burden is on the spender to produce the receipt.
THE CORRECT METRIC
Do not ask who used the most tokens.
Ask:
What existed afterward that did not exist before?
What exact human problem was solved?
Can the result be inspected?
Can it be used?
Can it survive outside the AI session that helped create it?
Can its development history be traced?
Can anyone distinguish a real release from a generated claim?
Did the compute create an enduring asset—or merely create more activity inside someone else’s platform?
By that measure, 44.2 billion is not a vanity statistic.
It is the compute trail behind a visible body of construction.
And because the screenshot covers Codex alone, this audit still excludes the much larger surrounding body of ChatGPT-assisted research, writing, architecture, specifications, books, essays, music and conceptual development.
The corporations published their burn rate.
I published what survived the fire.





