Venkat Mangudi

Writing on cybersecurity, AI, resilience, leadership, and risk.

Follow

The Price of Cheap Intelligence

Artificial Certainty, an Elytra Security newsletter. Notes on AI, overconfidence, and the new machinery of persuasion. Venkat Mangudi, Founder and CEO, Elytra Security.

Edition 13

The cost of capable machine intelligence has collapsed in under three years. Businesses are now being designed around the assumption that it keeps falling.

A few years ago, access to anything resembling today’s generative AI required specialist teams, expensive infrastructure and a fairly serious reason for doing it. Now a person can open a browser and ask a capable model to analyse a contract, write software, explain a scientific paper, translate a document, review a spreadsheet or help plan a business decision. Sometimes for nothing. For developers, the change is starker. What once demanded substantial computing resources can be bought by the token, often for an amount small enough to disappear inside the cost of the application built around it.

Stanford’s 2025 AI Index quantified part of that change. The cost of querying a model performing at roughly GPT-3.5 level on one widely used benchmark fell from about $20 per million tokens in late 2022 to seven cents by late 2024, a reduction of more than 280 times in under two years. The models got better over that period. Hardware improved, inference became more efficient, models became smaller without becoming useless, competition increased, open-weight models improved, and providers introduced caching, batching and tiered service.

The market now runs from very inexpensive systems built for high-volume work to considerably more expensive frontier systems built for difficult reasoning, with prices changing regularly as efficiency and competition shift. That is good news for anyone who wants to use AI. It also creates an unusual planning problem, because entire businesses are being designed around the assumption that intelligence will keep getting cheaper.

What happens if the cheapest intelligence and the intelligence a business ends up depending on turn out to be two different things?

The Prompt

Software taught us to read price in familiar units. A licence costs so much per user, cloud storage so much per gigabyte, a virtual machine so much per hour, a transaction so much to process. AI adds another unit to the vocabulary: the token. It sounds precise. A provider publishes a price for input tokens and another for output tokens, a developer estimates usage, multiplies by expected volume and arrives at a number. For many applications that number looks remarkably small.

The market is pushing it lower still. OpenAI, Anthropic, Google and others offer several models at different price and capability points. Smaller models handle work that once needed much larger ones. Caching makes repeated context considerably cheaper. Batch processing reduces prices where the work can wait. Open-weight models give organisations another route entirely, particularly where workloads are predictable enough to justify running them independently.

All of which produces a seductive assumption. If intelligence is getting cheaper this fast, its cost will eventually stop mattering. A company can put it everywhere. A software product can include it by default. An agent can make repeated model calls because it is inexpensive. Documents can be analysed repeatedly, every customer can get a personalised interaction, every employee an assistant, every workflow some intelligence inside it. At low enough prices, arguing over whether an inference costs one cent or half a cent starts to feel petty.

That is where the economics begin, because cheap inputs change behaviour. When storage became inexpensive, we stopped deleting things. With unlimited email storage, people hardly delete emails. When was the last time you attempted to clean up your email inbox? When bandwidth became abundant, websites stopped being designed for dial-up. When computing got cheaper, software consumed more of it. In my opinion, AI will follow the same pattern. When anything is charged like a utility, it starts being accounted for properly.

A cheaper token does not mean a smaller AI bill. It can mean many more tokens.

A system that once made one model call may make ten. An agent may reason, search, call a tool, inspect the result, ask another model, reconsider and try again. A coding system may read an entire repository rather than a single function. A research tool may examine hundreds of pages before returning three paragraphs. The unit gets cheaper and the appetite grows.

The Mirage

The attractive number in AI economics is the price of the model call. There is no single price for intelligence. Capability, speed, context length, reliability and the manner of consumption all carry different prices. Even within one provider the figure varies substantially depending on whether the customer wants an inexpensive general-purpose model, a frontier model for difficult work, cached inputs, batch processing or guaranteed low-latency service.

The price can also be temporary. During 2026 providers have reduced prices, made introductory prices permanent and offered promotional rates for particular models or tiers. Google currently publishes some Gemini API rates explicitly scheduled to change at the start of 2027. OpenAI has temporary promotional pricing on GPT-5.6 Sol. Anthropic made the introductory price of Sonnet 5 permanent after originally planning to raise it. None of that is unusual. It is what competitive technology markets do.

Today’s price is a commercial decision as much as a technical fact. It reflects hardware efficiency, model efficiency, utilisation, competition, product strategy, available capacity and the provider’s reading of the market. A token price does not tell an enterprise what the underlying infrastructure cost, what the provider’s margin is, what the same capability will cost in three years, or what a business will pay once AI is embedded deeply enough that switching becomes difficult.

One can easily be taken in to the illusion today because AI feels unusually cheap at the point of consumption. A user can hold a long conversation with a sophisticated model without ever seeing the data centre behind it. A developer can send an API request without knowing which accelerators processed it, how heavily utilised they were, how much electricity was consumed, how the provider financed the infrastructure, or whether the request was economically attractive to the company serving it. That abstraction is useful, and one of the main reasons adoption has moved so fast. It can also create a false sense of security that the price on the screen represents something permanent.

Technology prices can fall for decades while the amount customers spend keeps rising. Computing, storage and network capacity all did it. Cheaper units enable new uses, and the new uses consume vastly more units.

The Reality Check

There is a strong case that useful intelligence really will keep getting cheaper. Hardware price-performance has improved steadily, with Stanford’s AI Index reporting roughly 30 percent annual improvement in hardware price-performance and around 40 percent annual improvement in energy efficiency in the data it examined. Model efficiency has improved alongside it. Distillation, quantisation, better training methods, improved architectures and better inference software keep reducing the compute needed for useful work.

Open-weight competition adds pressure. An organisation with the right skills and workload may decide it does not need a frontier API for every task, because a smaller model can classify a document, extract structured information, answer a narrow question or run a repetitive internal workflow perfectly well. Expensive models then have to justify their premium.

What emerges may resemble other mature technology markets. Routine intelligence becomes cheap, while specialised, fast, highly reliable or jurisdiction-bound intelligence costs more. Very long context can cost more, reserved capacity can cost more, and frontier capability may keep commanding a premium where the work genuinely needs it.

The word AI hides all of those distinctions, which is why a company saying it needs AI has told us almost nothing about its future cost. What kind, for which task, how often, at what latency, with how much context, at what level of reliability? Does the work require the strongest available model, or one merely good enough to perform a narrow function repeatedly? Can it wait and use cheaper batch processing? Can common context be cached? Can a smaller model handle eighty percent of requests and send only hard cases to a frontier model? Could an open-weight model handle part of it locally?

These are architecture questions, not questions for model enthusiasts, and architecture will decide whether cheap intelligence stays cheap after adoption scales. A badly designed application repeatedly sends enormous context to an expensive model for work a small locally hosted model could have done. A well-designed one routes by difficulty, uses deterministic software where the answer is already known, retrieves only what is needed, caches repeated context and reserves expensive reasoning for the few cases where it changes the outcome. The future cost of AI may depend less on finding the cheapest model than on learning when intelligence is actually required.

The Leadership Question

There is a simple stress test for any business case built on low AI prices. Change the price. Double it. Triple it. Does the application still make sense? This is not a prediction that prices will triple, because in many areas they may keep falling fast. The exercise exposes something more useful: whether AI is creating genuine economic value or merely exploiting a temporarily attractive input price. If a workflow saves a thousand rupees of real cost every time it consumes ten rupees of AI, the economics will not collapse when the model cost becomes fifteen. If the business works only because the AI input costs fractions of a cent, the company has built something highly sensitive to its provider’s decisions.

The test runs in the opposite direction too. Replace the frontier model with something weaker and cheaper, and see what breaks. If very little changes, the organisation was paying for intelligence it did not need. If quality collapses, frontier capability is genuinely part of the value proposition and its future price deserves closer attention.

This is most acute for products whose own customers pay a fixed subscription. Imagine selling a service for a thousand rupees a month while the AI behind it is usage-based. At launch the average customer consumes forty rupees of inference. Then behaviour changes. People learn to use the product harder, agents make more calls, context windows grow, customers upload larger files, new features bring image, voice or video processing. The customer still pays a thousand rupees and now consumes two hundred and fifty rupees of AI. The product has become more useful and its economics have changed.

Cloud software companies have handled versions of this for years. AI complicates it, because usage is not always visible to the user and agentic systems generate their own demand for inference. One human request can trigger a chain of machine activity. That makes architecture part of pricing strategy, and businesses should know which model calls create customer value, which merely add convenience, and which exist because nobody designed a cheaper path.

Then there is bargaining power. Switching models is becoming easier in some circumstances, since APIs are similar, open models exist and abstraction layers can route between providers. But models do not behave identically. Prompts need changing, evaluations have to be repeated, safety behaviour differs, tool use differs, outputs differ, and customers may notice. A model that began as a replaceable component gradually becomes part of the behaviour of the product. At that point the cheapest intelligence in the market becomes an important part of the decision. The price of the intelligence the product has learned to depend on might become less important.

The Closing Signal

Cheap intelligence is a major development of this decade. The cost of reaching capable models has fallen at a speed that would have sounded implausible a few years ago, and the consequences are good. Small companies can access capabilities once reserved for technology giants. Students get explanations that once required a tutor. Developers use sophisticated models without owning any AI infrastructure. Researchers cover material faster. Organisations in countries without enormous domestic compute can still build sophisticated applications. Smaller and open models widen all of that further.

But cheap intelligence also changes what we build. When something becomes inexpensive enough we stop rationing it. We design around abundance, allow it into more workflows, create products whose economics assume it stays available, and let one request trigger ten more. Eventually the price of intelligence becomes part of the architecture of the economy around it.

That is why falling token prices deserve attention beyond the immediate saving. We are making long-term decisions against a price curve moving extraordinarily quickly. The curve may keep heading down. Efficiency may outpace demand, open models may force prices lower, new hardware may make today’s inference costs look absurd. Demand may also expand precisely because intelligence is getting cheaper, with workloads that do not exist today consuming quantities of inference that would currently seem extravagant, while premium capability stays expensive and capacity, latency, jurisdiction and reliability acquire prices of their own.

The market will sort much of this out. Businesses still have to survive while it does. The sensible response is neither to avoid AI because its future economics are uncertain nor to assume today’s prices have revealed the permanent cost of machine intelligence. Use the cheap intelligence, enjoy the competition, build with it, and understand which parts of the business depend on it staying cheap. The most important number may turn out to be neither the cost of a token nor the price of the model, but the amount of the organisation that was designed around the assumption that intelligence would always cost less tomorrow. If this intelligence stopped getting cheaper, would we still have built the same business around it?

The Closing Signal

The token got cheaper. The dependency did not.


Discover more from Venkat Mangudi

Subscribe now to keep reading and get access to the full archive.

Continue reading