Do Not Let Cheap AI Become Expensive Infrastructure
AI inference prices may fall, rise, or fragment. Build workflows with budgets, routing, fallbacks, and evidence before dependency becomes expensive.
Product perspective
Private Model Infrastructure
AI is unusually easy to adopt before anyone understands its long-term operating cost. A team buys a few seats, connects an API, automates a queue, and discovers that work which once waited for a person can now run continuously. The first invoice looks small beside the apparent productivity gain. Six months later, prompts are longer, models are more capable, agents take several passes, and the organisation has quietly redesigned its operating habits around a service whose future price it does not control.
The temptation is to assume that today's price describes the underlying economics. It does not. An API tariff is a commercial decision made inside a market competing for developers, usage, distribution, and enterprise commitments. The provider also carries research, training, inference hardware, networking, energy, datacentres, reliability, safety, support, and the cost of capacity that must exist before demand arrives. Public pricing tells a customer what a request costs today. It does not reveal the provider's fully allocated margin or guarantee what the same class of capability will cost after the market matures.
This is not an argument against implementing AI. It is an argument for treating inference as a volatile operating dependency. Prices may rise as introductory offers end and investors seek durable returns. They may also fall dramatically as models become smaller, chips improve, competition increases, and open alternatives spread. A responsible workflow should survive both futures. If the economics improve, it should capture the saving. If they deteriorate, the business should still know how to work.
The Economic Unknown
The Price of a Token Is Not the Cost of the AI Industry.
Inference is the repeated computation that turns an input into a model response. Its immediate cost includes accelerator time, memory, energy, and the infrastructure required to serve requests with acceptable latency. But a commercial model provider has a wider bill: training new systems, acquiring chips, building datacentres, reserving capacity, employing researchers, operating safety systems, supporting customers, and funding failed experiments. Microsoft reported a $20.1 billion year-over-year increase in additions to property and equipment in fiscal 2025 while describing continued investment in cloud and AI infrastructure. Alphabet said it invested $91 billion in capital expenditure during 2025. Those figures cover more than inference, but they show the scale of the platform beneath the apparently weightless prompt box.
It is therefore plausible that some free tiers, consumer subscriptions, and introductory API offers are priced partly to accelerate adoption rather than maximise current profit. What we cannot responsibly claim is that every inference request is sold below its marginal cost. Providers do not publish enough model-level cost and margin information to prove that. Some requests may be profitable, some subscriptions may be limited by usage caps, and expensive research may be funded by profitable cloud or advertising businesses. ‘The market is subsidised’ is a useful risk hypothesis, not an audited fact.
Businesses should make decisions from the uncertainty itself. A workflow does not become safe because a commentator predicts a price increase, and it does not become safe because a provider predicts efficiency. The relevant question is whether the workflow still creates value when the input price, model mix, latency tier, context size, and request volume move away from the assumptions in the pilot.
Marginal Cost
The compute and infrastructure consumed by another request, which outsiders can only estimate.
Fully Allocated Cost
Inference plus research, training, capacity, support, safety, and the wider organisation required to supply it.
Market Price
What the provider chooses to charge now, shaped by competition, packaging, utilisation, and growth strategy.
The Dependency Risk
Cheap Adoption Can Arrive Before Durable Monetisation.
Technology markets often reduce friction first and monetise dependency later. Streaming is a familiar consumer example. Netflix described price increases across several markets in its 2019 shareholder letter, and later discussed price as one of the levers available after engagement and retention had been established. AI does not have to follow the same path, but the analogy matters: once a product becomes part of a household habit or a business process, a supplier has more room to test what that dependency is worth.
AI offers more direct examples. Google currently describes pricing for a Gemini Flash model as promotional, publishes an expiry date, and states the later input and output rates. Its free Gemini developer tier can use submitted content to improve products, while the paid tier says content is not used for that purpose. This is a transparent exchange: access, adoption, product learning, and future monetisation can be different parts of the same ladder. A business should read the ladder before building on its first step.
The data argument also needs precision. It is incorrect to say that every business prompt is automatically used to train a provider's models. OpenAI says API and business-product inputs and outputs are not used for training by default; Anthropic makes a comparable commitment for commercial services unless the customer opts in. Consumer and free products can have different terms, controls, or improvement programmes. Data policy should therefore be checked for the exact product and account type, then recorded as an architectural requirement rather than inferred from the brand name.
The larger source of leverage may be operational dependence rather than training data. If employees stop practising the manual process, internal knowledge moves into prompts and vendor-specific tools, integrations assume one model's behaviour, and customers expect instant AI-mediated service, switching becomes expensive even when an alternative API is cheap. The supplier does not need to trap the customer deliberately. Habit, integration depth, and organisational forgetting can create the lock-in on their own.
Beyond the Demo
Agentic Workflows Multiply Usage in Ways Seat Prices Conceal.
A single chat has a legible cost. An operational workflow is a chain. It may retrieve documents, send a large context, ask a model to plan, call tools, inspect their output, retry a failed action, ask another model to review the answer, and preserve a summary for the next run. Providers can charge for input, output, cached context, storage, search grounding, tool calls, containers, priority capacity, or provisioned throughput. The visible answer is only the final edge of a larger inference graph.
The monthly cost is not inherently exponential, but it can compound quickly: cases per month multiplied by model calls per case, tokens per call, retry rate, review passes, and the price of the selected tier. A team that automates more work because each unit appears cheap can experience the same paradox seen in other efficiency gains—lower unit cost encourages enough additional consumption that total spending rises. Autonomous loops intensify this risk because software can consume tokens faster than a person can notice that the marginal work has stopped being useful.
This changes the correct unit of measurement. Cost per million tokens is useful for procurement, but cost per accepted outcome is useful for operations. A cheaper model that creates more retries, manual correction, or customer harm may be expensive. A premium model that completes the task once may be economical. The workflow should record token usage and provider charges beside completion, rejection, latency, escalation, and business value so the team can see which intelligence actually pays for itself.
Dependence becomes dangerous when the organisation cannot answer a simple counterfactual: what happens if this workflow costs three times as much next quarter? If the only answer is to accept the bill because nobody remembers the prior process, then the AI integration has become financial infrastructure without financial controls.
Volume
How often the workflow runs, including agent retries, evaluations, background jobs, and abandoned attempts.
Intensity
The context, reasoning, output, tools, and model tier consumed during each completed or failed case.
Acceptance
The proportion of results that become useful work without costly correction, escalation, or repetition.
The Counterargument
The Strongest Evidence So Far Points Toward Cheaper Intelligence.
Caution should not become a story in which price increases are inevitable. Stanford's 2025 AI Index found that the inference price of a system performing around GPT-3.5 level on one benchmark fell more than 280-fold between November 2022 and October 2024. Smaller models improved, hardware costs declined, energy efficiency increased, and open-weight models narrowed parts of the performance gap. The economic force pushing prices down is at least as real as the investor pressure pushing providers toward sustainable returns.
Competition also gives customers options that did not exist at the beginning of the generative-AI cycle. Providers offer smaller model families, cached input discounts, asynchronous batch processing, and lower-cost capacity tiers. OpenAI's Batch API advertises a 50 percent discount for work that can complete within 24 hours. Google offers batch and flexible inference discounts and explains caching as another large saving for repeated context. Many business tasks do not need a frontier model responding at premium interactive latency.
The likely future is not one universal price curve. Commodity classification, extraction, transcription, and drafting may become much cheaper, while the most capable reasoning, long contexts, real-time voice, computer use, premium reliability, and scarce capacity retain higher prices. A workflow tied to ‘the best model’ can get more expensive even while equivalent intelligence becomes cheaper elsewhere. Architecture determines whether the business can move down the cost curve.
The NFT comparison is useful only at the boundary. If an AI use case creates value solely because access is temporarily cheap, its demand may disappear when the real price arrives, just as technically interesting products can collapse when speculation no longer supports them. But AI already performs many ordinary tasks with measurable value, so the technology is unlikely to vanish as one market did. The vulnerable part is the extravagant workflow: impressive in a demo, weak at full price, and too entangled to remove cleanly.
Cost-Resilient Implementation
Every AI Workflow Needs an Economic Exit Route.
A cost-resilient implementation starts with a budget per outcome and a ceiling per run. Deterministic code should handle stable rules before a model is called. Repeated instructions and documents should be cached where the provider permits it. Offline work should use batch or flexible capacity. Small models should attempt bounded, reversible tasks, with stronger models reserved for exceptions whose additional quality is worth the premium. Agent loops need maximum turns, token limits, timeouts, and a stopping condition connected to useful completion.
Provider abstraction matters, but a generic interface alone is not portability. Prompts, tool schemas, safety behaviour, context limits, structured output, and evaluation results vary between models. Maintain a small cross-provider evaluation set made from real workflow cases. Run it when prices or models change. Keep the business state outside the model, store source evidence in portable formats, and ensure a human can complete consequential work when inference is unavailable or over budget.
Private Model Infrastructure can be one part of the exit route when privacy, predictable utilisation, or supplier control justifies it. It is not automatically cheaper: hardware, hosting, quantisation, observability, updates, and specialist operation become the customer's responsibility. The useful comparison is a fully loaded managed-service cost against a fully loaded private deployment, measured at the required quality and throughput—not an API invoice against the price of a bare GPU.
Finally, assign ownership. Someone should review monthly spend, accepted-outcome cost, model changes, data terms, failure rates, and the manual fallback. Procurement should negotiate pricing protections for high-volume dependencies. Product owners should know which capabilities can degrade gracefully. Operators should be able to disable low-value automation without stopping the business. The human process should not be preserved out of nostalgia; it should remain recoverable until the AI workflow has demonstrated durable economics.
Measure
Track provider spend, tokens, retries, latency, acceptance, correction, and value at the workflow level.
Route
Use deterministic software, smaller models, caching, batch capacity, and premium reasoning according to need.
Port
Keep state and evidence outside the provider, then test real cases across credible managed and private alternatives.
Recover
Retain spending limits, kill switches, human escalation, and a workable procedure for continuing without inference.
Conclusion
Implement AI, but Refuse an Unpriced Dependency.
The present AI market contains competing truths. Serving frontier systems requires extraordinary capital. Free access and promotional pricing can accelerate adoption. Investors will eventually expect durable economics. At the same time, inference at a fixed capability has become dramatically cheaper, competition is intense, and new optimisation methods keep widening the range of viable models. Nobody can responsibly promise which force will dominate the price of the capability a particular business needs.
That uncertainty is the reason to build, not the reason to wait. Put AI into workflows where it creates measurable value, but keep the unit economics visible and the operating knowledge recoverable. A business should be able to change models, reduce quality deliberately, route work to cheaper capacity, move selected workloads into private infrastructure, or return a consequential decision to a human. The safest AI workflow is not the one with the lowest invoice today. It is the one that remains worth operating when today's price stops being true.
Supplier Independence
Self-Hosting Without SaaS Subscriptions
Compare subscription convenience with the real infrastructure and operating responsibilities of owning the stack.
Read the Self-Hosting EssayPrivate Deployment
Private Model Infrastructure
Explore controlled model deployment, data boundaries, evaluation, observability, and managed operation.
Explore Private Model InfrastructureOperational Ownership
Implementation Partners as Software Gets Cheaper
See why lower creation costs increase the importance of integration, maintenance, governance, and accountable operation.
Read the Implementation EssayResearch notes
Sources and Supporting Material
These references support factual claims in the article. Brownsmith's interpretation and forward-looking analysis remain editorial judgement rather than vendor promises.
- Stanford HAI: 2025 AI Index Report
- Microsoft: 2025 Annual Report
- Alphabet: 2025 Q4 earnings call
- Google AI: Gemini API pricing and promotional terms
- Google AI: Inference optimisation options
- OpenAI: Business data privacy and training defaults
- Anthropic: Commercial data and model training
- OpenAI: Batch API pricing model
- Netflix: Q1 2019 shareholder letter
