The Price Is Not the Price
Updated: 6 hours ago
A technology leader's framework for AI cost, risk, and right-sizing.
Why the price Is not the price
For a growing share of enterprise workloads, a frontier language model is the most expensive and highest-risk route to an inferior result. The macro data already shows it. An NBER study published in February 2026 found 89 percent of firms reporting no impact of AI on productivity over the past three years, and more than 90 percent reporting none on employment. [1] A METR randomized controlled trial found experienced developers 19 percent slower with AI assistance while believing they were faster. [2] AI spend moved to the center of enterprise cost agendas over the same period. [3] The NBER survey’s executives forecast a 1.4 percent productivity lift over the next three years, [1] and narrow, well-scoped deployments show real returns. The returns concentrate where the workload fits the tool. The problem is not model quality. It is that AI spend is approved against a price that is not real and a cost model that is not complete.
Three facts, developed below, belong under every AI budget line. The price is subsidized: the industry’s economics run on investor money and off-balance-sheet leverage, so today’s price list is promotional. Some vendors will not survive its correction, and the enterprises that built on them will inherit costs they did not create. The invoice is the smallest cost line: evaluation, human review, rework, and compliance dominate the true bill, and in the illustrative example below review labor runs 75 times the inference charge. And cheaper substitutes exist: classical models, rules engines, search, and qualified people beat a frontier model on cost and reliability for much of the portfolio.
The fix is a per-workload decision discipline: a five-question test you can run in an afternoon, and a two-week inventory that sorts the portfolio into workloads to re-platform and workloads to govern. Most AI strategies predate this evidence. They were built when the price list looked real and the cost model looked complete.
The price list is a promotional rate
Today’s AI price list is a promotional rate, and the exposure behind it splits in two.
The frontier labs earn positive gross margin on metered inference but lose money as companies: training runs, capacity commitments, and free tiers that dwarf paid usage sit on top of the profitable serving business. [4] Part of their reported demand is also vendor-financed: chip and cloud suppliers invest in the labs, and the labs spend the money back on capacity. Much of the buildout debt is parked in special purpose vehicles off the sponsors’ balance sheets. The structure echoes telecom vendor financing before the 2001 collapse, when the financed demand evaporated and the equipment makers wrote off billions in loans to failed customers. [5] The distortion is the same now as then: reported demand looks stronger than it is, and the debt load looks smaller than it is. If the funding turns, the likely path is restructuring under the chip and cloud suppliers, who are both the labs’ creditors and their hosts: service continues, the roadmap does not. Plan for frozen models, deprecations, and normalized prices, not a blackout.
The layers above the labs are priced with investor money more directly, and the correction there is already underway: major providers have moved to usage-based pricing, and heavy users are seeing costs multiply. [6] Flat-rate seat plans priced for chat now carry agentic workloads that consume many times what the plan assumed. The venture-funded tool layer (orchestration, enterprise search, guardrails, agent security) has no floor at all: it resells frontier capability at prices covering neither its API bill nor its payroll, and its failure mode is extinction. Prompts, evaluation suites, fine-tunes, and workflow glue are tuned to one provider’s behavior, and the tuning repeats at every layer of the stack. So when a pre-profit vendor fails instead of repricing, migration becomes an emergency on the vendor’s timeline. Frontier exposure realizes the same way at smaller scale: model deprecations force migrations on the lab’s schedule. The layers most likely to fail outright are the ones least likely to be on the risk register: the second- and third-tier tools, not the frontier lab.
The standing counterargument is token deflation: per-token prices have fallen by an order of magnitude and keep falling. Cheaper tokens have not produced cheaper bills, because agentic volume grows faster than prices drop. The corrections arriving in practice are plan conversions, caps, and metering rather than token-price spikes. The Anthropic and Cursor moves are the shape of it. [6]
Cloud-era lock-in fears never had this dimension. Cloud capex bought long-lived, general-purpose assets, so scale defrayed cost; AI capex buys hardware that depreciates in a few years and training runs whose competitive life is measured in months. The correction, when it comes, lands downstream: repricing, deprecation, and forced migration. A business case built at subsidized prices is not a business case.
Every scenario above ends the same way, including a rescue: a bailed-out lab changes owners, then normalizes prices, and the buyer’s exposure survives the bailout. The useful posture is mobility rather than prediction. Portable prompts, workload-specific substrates, and decoupled architecture turn repricing, deprecation, or vendor transitions from emergencies into routine procurement events.
The invoice is the smallest line
Five cost lines determine what an LLM workflow costs. The vendor invoice shows the first one; the third one dominates.
1. Inference at workflow multipliers. A chat session is one model call. Production is not chat. In my client work, agentic patterns commonly issue 10 to 20 model calls per user task, and retrieval-augmented generation inflates context windows three to five times. Uber’s CTO reported the company exhausted its annual AI coding budget in roughly four months. [7]
2. Evaluation and testing. A non-deterministic system, one that can return a different output for the same input, cannot be tested once and trusted. Model version updates silently change behavior, so every material change re-triggers evaluation. Budget permanent engineering capacity for it.
3. Review labor. Non-deterministic output in a consequential workflow requires qualified human review before it acts on the world. Review is senior labor at senior rates. The reviewers’ salaries are part of the cost of this guardrail. When reviewing an output takes as long as producing it, the model has added a cost layer, not removed one.
Review also carries a supply-side exposure the other lines do not. Reviewers are made by doing the work the model now automates: skill forms in the working relationship between novices and experts, and delegating the novice tier’s work to a machine severs it. [8] In my client work, and in 25 years of building engineering teams, the mid-level people an AI deployment cuts are the ones who train the next bench. The narrowing is already visible where automation concentrates: employment of workers aged 22 to 25 in the most AI-exposed occupations sits 19 percent below the path of their less-exposed peers, driven by reduced hiring rather than layoffs. [8] That scarcity raises the price of the qualified reviewers who remain. The review rate is not a constant, and a workload priced on today’s reviewer supply understates the line that already dominates.
4. Error and rework. These are the failures that get past review (escapes), priced per failure times the escape rate. In regulated or customer-facing workflows a single escape can exceed the annual inference bill. The same line carries the security failure modes a deterministic system does not import: data leakage, prompt injection, and IP contamination. Most deployment analyses never price them, and the exposure compounds as workflows move from chat to agents with tool access. The controls that contain them are the subject of a companion whitepaper, The Control Stack.
5. Compliance. The EU AI Act’s transparency obligations have applied since August 2026. High-risk obligations follow in December 2027 for standalone systems and August 2028 for AI embedded in regulated products. Penalties for high-risk violations run up to EUR 15 million or 3 percent of global annual revenue (EUR 35 million or 7 percent for prohibited practices). [9] Classification, documentation, and conformity work attach to model-based systems in ways they do not attach to a rules engine. US state-level requirements add a patchwork on top. Compliance allocation follows the cost of one bad output: high-risk classification spend on documentation and conformity is justified only where a single error is material to the business.
Review is the swing line. The fully loaded annual cost of an LLM workflow:
(inference per task × calls per task × annual tasks) + evaluation and monitoring + (review minutes × loaded review rate × annual tasks) + (escape rate × cost per failure × annual tasks) + compliance allocation
The pricing exposures in the promotional-rate section enter as risk adjustments, not annual lines: stress the inference and tooling terms at a multiple of current prices, price a forced migration for any layer supplied by a pre-profit vendor, and test the review-rate term against a thinning reviewer supply.
Consider an illustrative example with round numbers. A document triage workflow handles 200,000 items per year. Inference at agentic multipliers runs $0.04 per item: $8,000. Evaluation and monitoring take a quarter of an engineer: $60,000. Review takes two minutes per item at a $90 loaded hourly rate: $600,000. Escapes run 0.5 percent at $200 per failure: $200,000. Total: roughly $868,000, of which inference is under one percent (compliance allocation excluded from the example because it is workload-specific). Stress the assumptions and the shape holds: quintuple the inference price and halve the review time, and review is still half the total cost and more than seven times the inference charge.

Review is set by the cost of a bad output, and it is where most upside-down cost-benefit ratios come from. It also points at the fix: if a template-based extractor handles the 90 percent of items that follow a consistent format, the model handles only the residue. Review shrinks with it, and the workflow’s economics change class. The cheapest token is the one a deterministic system made unnecessary.
The substitutes nobody prices
“Model or nothing” is a false choice. For much of the portfolio a cheaper, more reliable substitute exists, and one of them is hiring.
Workload class | Reflex choice | Frequently better | Why |
Tabular prediction (churn, credit risk, demand, fraud scoring) | LLM | Gradient boosting or regression | Orders of magnitude cheaper, lower latency, deterministic, auditable, mature tooling, familiar to regulators |
Fixed-logic decisions (eligibility, routing, pricing rules) | LLM | Rules engine | Testable, explainable, zero hallucination, lower latency; changes are code review, not prompt archaeology |
Information lookup | LLM chat | Search and retrieval without generation | Returns the source; cannot fabricate one |
Structured extraction from consistent document formats | Frontier LLM | OCR plus templates, or a small fine-tuned model | Format stability rewards purpose-built tools at a fraction of the cost |
Low-volume, high-judgment decisions | LLM | A qualified person | Accountability, context beyond the prompt, learns from a single mistake |
High-volume unstructured language (summarization, drafting, triage, translation) | LLM | LLM, governed | Where models earn their keep, provided review cost per item stays low relative to value per item |
Between the frontier API and the alternatives sits a middle band: self-hosted open-weight models and small task-specific models. They trade capability for cost control, determinism of spend, and data locality. Their economics are the subject of a companion brief, The Middle Band (available on request). The table’s point stands either way: the substitution set exists, and most AI roadmaps never price it.
Headcount belongs on the same decision sheet as build-versus-buy. The example above shows why: when review labor is nearly 70 percent of workflow cost, you are already paying for people. The open question is whether they review a model’s output or do the work directly. Hiring has frictions of its own: time-to-fill, scarce specialties, ramp time, and the budget politics that make a req harder to approve than tooling opex. Put those on the sheet too. They narrow the gap without closing it.
People carry advantages no current model matches: accountability (a person can sign), single-example learning (one corrected mistake changes future behavior), context that never made it into any prompt, and immunity to vendor repricing. Models carry advantages people cannot match: marginal cost per additional unit, latency, and around-the-clock availability. High-volume work with cheap review favors the model. Low-volume work with expensive review favors the hire, and the break-even is closer to the middle than most 2024-era business cases assumed.
Treating hiring as an HR track and AI as a technology track guarantees nobody prices the comparison.
Five questions, answered with numbers
Ask five questions per workload. Answer them with numbers.
1. Volume. Is there enough throughput to amortize evaluation and governance overhead? The fixed cost of running a model responsibly does not scale down for small workloads.
2. Cost of one bad output. A wrong summary of an internal meeting costs approximately nothing. A wrong figure in a regulatory filing is material. This number sets the review burden, which sets the dominant cost line.
3. Determinism requirement. Does the workflow need the same input to produce the same output, for audit, reconciliation, or repeatability? If yes, a model is the wrong substrate: no amount of prompting makes it deterministic. The table’s rules-engine and extraction rows are the usual landing.
4. Review economics. This is calculated as minutes of qualified review per output multiplied by the loaded review rate, compared against the value per output. When that ratio approaches one, review alone consumes the output’s value.
5. Substitute check. Can a rules engine, a classical model, search, or a hire deliver the outcome at equal or lower fully loaded cost, build and time-to-production included? If yes, the burden of proof shifts to the model. The substitution table is this question’s checklist.
Question 3 disqualifies on its own: a hard determinism requirement ends the evaluation. For the rest, a workload that fails three or more questions gets re-platformed, and one that passes all five gets governed and scaled. A question fails when its number comes back against the model: too little volume to amortize the overhead, an error cost demanding review the output’s value cannot carry, a review ratio near one, a cheaper substitute at equal outcome.

Price the re-platforming itself before moving. The switching cost has four parts: building the substitute, running both systems in parallel until the substitute’s error rate is proven, reworking the integrations, and retraining the workflow around deterministic behavior. The model’s speed advantage is real, and it is how most misfits got here: a prompt reaches production in weeks, a rules engine or trained classifier in a quarter or more, so the model won every build-versus-wait decision by default. The decision number for a failing workload is payback: the gap between its loaded annual cost and the substitute’s, against the one-time cost to move. The wider the gap, the faster the payback. Marginal failures can be cheaper to govern than to move.
The strongest objection to the test is option value: running model workloads now, even at negative return, builds the capability to exploit better models later. The objection is real and the remedy is scope. Option value is bought with a bounded pilot at pilot cost, not a production workload at production cost. A workload that fails the test can fund a capability sandbox with a fraction of its savings. Re-running the test when models or prices move catches the moment a failed workload starts to pass.
If the workload passes the test
A workload that survives the five questions still needs the control stack for non-deterministic systems in consequential roles: explicit role definition (the model drafts, a human decides), mandatory review gates, immutable logging, continuous testing against hallucination and drift, and monitoring with alerting. The Control Stack covers that stack in depth. For financial services, LLM Workloads in the Model-Risk Gap maps it to the interagency model-risk expectations that replaced SR 11-7 (both available on request).
The portfolio outcome to aim for is fewer model workloads, each defensible in front of a CFO and an auditor. That posture also survives a market correction: if AI budgets tighten, the workloads chosen this way are the last ones cut.
The two-week inventory
A two-week inventory answers most of it. Enumerate every AI workload, deployed and proposed. Score each against the five questions. Compute fully loaded cost per workload using the formula above. Map each workload’s vendor stack and mark every layer whose price cannot plausibly cover its cost to serve. Those entries are repricing or evaporation exposures, not stable inputs. The output is two lists: workloads to re-platform onto cheaper substrates, each priced with its migration cost and payback, and workloads to govern and scale.
I owe two disclosures on incentives. I sell no software and take no vendor margin, so a recommendation to shrink your stack costs me nothing. A reseller cannot say the same. The check on that claim: everything needed to run the inventory is in this paper, and engaging me buys speed and pattern recognition from prior inventories, not access to the method. The inventory runs as a fixed-scope engagement, typically paired with an AI spend audit. The posture it buys: AI-forward where it pays, disciplined where it doesn’t, with the numbers to show which is which.
Companion papers and inventory inquiries: dave@dave-nix.com
Notes
Yotzov, I., Barrero, J. M., Bloom, N., et al., “Firm Data on AI,” NBER Working Paper No. 34836, February 2026 (revised March 2026). https://www.nber.org/papers/w34836. Survey of nearly 6,000 executives across the US, UK, Germany, and Australia, fielded November 2025 to January 2026; the paper reports 89 percent of firms with no impact of AI on labor productivity (sales per employee) over the past three years and more than 90 percent with no impact on employment.
Becker, J., Rush, N., Barnes, E., Rein, D., “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/. RCT, 16 experienced open-source developers, 246 real tasks, February-June 2025 tooling; METR characterizes the result as specific to early-2025 tools and workflows.
FinOps Foundation, 2026 State of FinOps (1,192 respondents, $83B+ annual cloud spend under management): 98 percent of FinOps practices now manage AI spend, up from 31 percent in 2024; FinOps for AI is the top forward-looking priority; organizations report self-funding AI investment through optimization savings.
Barclays Equity Research (Ross Sandler et al.), unit economics of AI labs and hyperscalers, late August 2026; client-distributed. Accessible coverage: 36Kr, “For Every $100 AI Model Company Earns, Cloud Providers Take Nearly $40,” https://eu.36kr.com/en/p/396310896
0910724. Paid inference margins rose from low double digits in 2025 to 50 to 65 percent or higher in 2026; subscription products are the lowest-margin line as labs subsidize token costs for retention; Barclays estimates API inference margins above 80 percent.
“Tech groups shift $120bn of AI data centre debt off balance sheets,” Financial Times, December 24, 2025. https://www.ft.com/content/0ae9d6cd-6b94-4e22-a559-f047734bef83. Morgan Stanley, “Bridging a $1.5tr Data Center Financing Gap”: of roughly $2.9 trillion in global data center capex through 2028, about $1.5 trillion requires external financing, with private credit projected to supply roughly $800 billion, led by asset-based finance. https://www.morganstanley.com/content/dam/msdotcom/en/assets/pdfs/Research_Bridging-Data-Center-Gap.pdf. Vendor-investment arrangements: Nvidia-OpenAI; Microsoft-OpenAI; Amazon and Google with Anthropic. Precedent: telecom equipment vendor financing (Lucent, Nortel) before the 2001 collapse; nine suppliers had extended an estimated $25.6 billion by end-2000 (McKinsey, per CFO.com, March 2003), and Lucent alone took $3.5 billion in bad-debt provisions across 2001-02 (Lazonick and March, “The Rise and Demise of Lucent Technologies”).
McLaughlin, K., and Holmes, A., “Anthropic Changes Pricing to Bill Firms Based on AI Use as Demand Jumps,” The Information, April 14, 2026; accessible corroboration: PYMNTS (April 15, 2026) and Gizmodo (April 2026), including a licensing advisor’s estimate that the change doubles or triples costs for heavy Claude Enterprise users. Cursor (Anysphere) converted its $20 Pro plan from fixed request allotments to usage-based credit pools on June 16, 2025; heavy frontier-model users saw effective costs rise by half or more, and CEO Michael Truell issued a public apology and refunds on July 4, 2025 (TechCrunch).
Uber CTO Praveen Neppalli Naga, April 2026 disclosure to The Information: the company exhausted its planned 2026 AI coding budget in the first four months of the year. Uber subsequently capped agentic coding tool spend at $1,500 per employee per tool per month, per Bloomberg, June 2, 2026 (“Uber Caps Usage of AI Tools Like Claude Code to Cut Costs”).
Brynjolfsson, E., Chandar, B., Chen, R., “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence,” Stanford Digital Economy Lab working paper, revised August 12, 2026. https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/. ADP payroll data covering millions of US workers through June 2026; the 19 percent gap for workers aged 22 to 25 in AI-exposed occupations operates primarily through reduced hiring rather than separations, and concentrates in occupations where AI usage substitutes for rather than complements human work; the authors present the results as descriptive indicators, not causal estimates. On the skill-formation mechanism: Beane, M., The Skill Code: How to Save Human Ability in an Age of Intelligent Machines, HarperCollins, 2024.
EU AI Act (Regulation (EU) 2024/1689), Article 99 (penalties). Article 50 transparency obligations applicable since 2 August 2026. The Digital Omnibus on AI (Regulation (EU) 2026/1744; provisional agreement 7 May 2026; Parliament 16 June; Council final approval 29 June 2026; in force 27 July 2026) deferred high-risk obligations to 2 December 2027 (Annex III standalone systems) and 2 August 2028 (Annex I embedded systems), with a grace period to 2 December 2026 for content-marking on systems marketed before 2 August 2026.
© 2026 David Nix. Licensed under CC BY-NC-ND 4.0: share and quote with attribution; no modifications, no commercial redistribution.