AI economics
Why Cost Per Token Is The Wrong Executive Metric
The useful unit of AI economics is an accepted business outcome; cheap output that creates rework is expensive.
Cost per token is useful to an engineer selecting a model. It is incomplete for an executive deciding whether an AI system creates value.
A token tells us that a model processed or produced a fragment of information. It does not tell us whether the work was necessary, grounded, verified, adopted, safe, or accepted. It does not tell us whether the result prevented an hour of labor or created a week of rework.
Cheap Output Can Be An Expensive Operating Model
Imagine two workflows. The first uses an inexpensive model to generate ten implementation proposals. Eight violate project constraints, six repeat abandoned ideas, and all require a senior engineer to reconstruct the missing context. The second spends more per model call, loads reviewed context, produces one bounded proposal, and passes its evidence and constraints into implementation.
The first workflow wins a token-price comparison. The second may win the business.
The useful economic unit is not model activity. It is an accepted outcome whose total cost and residual risk can be understood.
Calculate The Complete Cost
The cost of an accepted AI-assisted outcome includes:
- Model and embedding consumption
- Search, storage, tool, and infrastructure costs
- Human framing and review time
- Verification, testing, security, and compliance work
- Rework caused by drift, missing context, or weak evidence
- Operational recovery when a result fails after deployment
- Opportunity cost when the organization follows the wrong recommendation
This is why CPF uses resource governance. Tokens, tool calls, human attention, deployment time, production risk, and financial runway belong in one design problem.

Useful economics include infrastructure, Human attention, verification, rework, recovery, and the consequence of a wrong decision.
Choose a Business Unit Metric
A support organization might measure cost per correctly resolved case. A research team might measure cost per accepted decision brief. An engineering group might measure cost per verified change. A sales operation might measure cost per qualified opportunity that reaches a human owner with complete evidence.
The measure should connect technical consumption to a business unit that leadership already understands. Tokens remain an input metric; they stop pretending to be the outcome. This follows the broader unit-economics practice of relating technology cost to a measurable unit of business value.2
Route Capability By Consequence
Not every task needs the largest model, every tool, the entire repository, and maximum autonomy. Classification, routing, extraction, formatting, contradiction testing, synthesis, and approval are different capability lanes.
A small model may classify the task. Deterministic tools may retrieve current records. A stronger model may compare ambiguous alternatives. Automated verification may check citations and calculations. A human may own the final decision. The workflow spends expensive intelligence only where consequence justifies it.

Different steps deserve different resources; the system should reserve expensive reasoning and review for uncertainty that matters.
Price Verification Deliberately
Verification has a cost; failure has a cost too. The correct verification budget depends on consequence. A draft internal summary does not need the same evidence as a production migration, financial recommendation, customer communication, or public safety decision.3
The executive question is not “Can we remove review?” It is “Where does review change expected loss, and what evidence lets the reviewer decide efficiently?”
Measure Before And After Comparable Work
AI transformation claims become credible when the organization compares the same classes of work before and after adoption. Measure model cost, human time, cycle time, first-pass acceptance, rework, failure, and business value. Keep simulated savings separate from observed results. NIST likewise recommends documenting expected benefits and costs against appropriate benchmarks, including nonmonetary costs from AI errors.1
That discipline shapes the NanoRes transformation measurement model. Historical values are not invented simply because the formula is persuasive.
The Executive Metric
The metric I want on the dashboard is cost per accepted outcome, accompanied by quality, cycle time, and residual risk. That metric creates the right conversation between engineering, finance, product, operations, and leadership.
A system may spend more tokens and create more value. It may spend fewer tokens and create more rework. The objective is not to make the model cheap; it is to make the complete work system useful.
Research Record
References and Evidence
The cost model in this essay is a proposed executive measurement approach, not a published accounting standard. The references support lifecycle measurement, risk-adjusted costs, and business unit economics; they do not independently validate CPF or a universal return on investment. Sources were reviewed on August 14, 2026.
- AI Risk Management Framework CoreGovernment Guidance · National Institute of Standards and Technology
NIST calls for documented benefits, monetary and nonmonetary costs, appropriate benchmarks, uncertainty, and lifecycle measurement. This supports measuring the system in deployment context; it does not prescribe my proposed cost-per-accepted-outcome metric.
- FinOps Framework: Unit EconomicsIndustry Framework · FinOps Foundation
The current framework distinguishes resource-efficiency metrics, including cost per token, from business-unit metrics such as cost per customer, transaction, or resolved case. It supports connecting technology spending to business value; it does not prescribe the AI measurement formula proposed here.
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileGovernment Guidance · National Institute of Standards and Technology
The profile describes governance, monitoring, incident response, feedback, and evaluation practices that can create costs beyond model inference. It is voluntary guidance and provides no benchmark for the savings claimed by any specific AI implementation.