Venkat Mangudi

Writing on cybersecurity, AI, resilience, leadership, and risk.

Follow

The Model Chase

Artificial Certainty, an Elytra Security newsletter. Notes on AI, overconfidence, and the new machinery of persuasion. Venkat Mangudi, Founder and CEO, Elytra Security.

Edition 04

Every release resets the market’s confidence. It should not automatically reset the decision.

There is a familiar rhythm to the AI market now. A new model appears. The benchmark charts arrive. Demo videos begin to travel. Someone posts a comparison thread. Someone else says the old model is finished. A product team announces a larger context window, lower latency, better coding, better reasoning, better multimodal ability, or the first serious step toward agents that can do real work.

Then the calls begin inside companies. Should we switch? Pause the current pilot? Is the tool we picked last quarter already outdated, and is the model we’re building around still the right one?

I understand the anxiety. In ordinary software, teams are used to version changes. In AI, every major model release feels as if the ground itself may have shifted. A decision that looked reasonable in March can feel incomplete by June. That happens because the market has learned how to turn model progress into decision pressure.

I have seen several teams ask themselves, mid-quarter, why they had not already switched to whatever had just made headlines. The honest answer usually had nothing to do with the model.

Model progress is real. Better models do matter. A stronger model can make a previously weak use case practical. A cheaper model can change the economics of a workflow. A model with better context handling can make some forms of document analysis more useful. A coding model that makes fewer basic mistakes can reduce review load for experienced developers.

The mistake is assuming that better automatically means better for us.

Between market progress and organizational readiness sits a gap. Beyond technical improvement, the model chase is a search for confidence in a market that keeps moving faster than most organizations can evaluate.

The Prompt

The useful question is whether the organization is evaluating model progress, or being carried by it.

A good AI decision should begin close to the work. What changed for the user? Is it faster once someone actually reviews the output, not before? Is it cheaper once training, checking, and exception handling are counted, and does the person who has to live with the result trust it more?

These questions are less exciting than a model launch. They are also where most of the truth sits.

The model chase creates a different atmosphere. Every decision starts to feel temporary. Some teams hesitate because they fear a better model is around the corner. Others rush because the latest release appears too important to miss. A vendor demo introduces one impressive new capability, and suddenly the previous evaluation feels stale. A leadership group that approved one direction starts wondering whether it has backed yesterday’s stack.

That is how news FOMO gets into the room. The story around a model can move faster than the model’s relevance to a particular use case. A release becomes a signal. A signal becomes pressure. Pressure becomes a budget discussion. Soon the organization is reacting to the market’s excitement more than to its own work.

Reacting to the market’s excitement, instead of your own work, does not build durable capability.

A company can watch model progress closely without becoming captive to it. The real discipline is to ask what the new model changes in the specific job being considered. If the answer is not clear, the launch may be important to the market, but not yet important to the business.

The Mirage

The mirage is that the newest model will solve the adoption problem by itself.

It is easy to believe this because model improvements are visible. The answers read better, the code looks cleaner, the system now handles images and calls tools without being asked twice. Each improvement invites people to assume the hard parts of adoption are being steadily removed.

Sometimes they are. Many earlier frustrations do get solved by better models. A task that was too slow may become practical. A tool that needed too much correction may become useful. A workflow that felt awkward may begin to make sense.

But many failures in AI adoption are not model failures. The process was unclear, the data was messy, and nobody had defined what good output looked like. The pilot got chosen because it looked impressive, not because the work mattered, and the tool ended up somewhere that needed trust when it had only earned supervision.

A stronger model can make a good use case better. It can also make a weak use case more expensive, more convincing, and harder to challenge.

The frontier model chase adds another layer. The newest and most capable models create excitement for good reason. They make headlines. They pull developer attention. They make older systems feel dull. They promise access to the edge of what is possible. For a serious research task, a complex reasoning workflow, advanced coding, or a high-value expert assistant, that edge may be worth pursuing.

But frontier models bring dependency risk. Access, pricing, usage limits, and safety behavior can all change without warning. A provider can alter the model, withdraw a feature, restrict a geography, or change the terms under which the system can be used.

The Fable 5 and Mythos 5 episode is one example. For many teams, the news created urgency because those models represented frontier capability. Then access was disrupted because of a government directive. Later, the restrictions were reversed. This is not a debate about that policy decision. The concern is narrower. If a workflow depends too heavily on one frontier model, the model is more than a tool. It has become part of the organization’s operating risk.

I have watched a team rebuild an entire support workflow over a weekend because one frontier provider changed its terms without notice. Nobody on that team had done anything wrong. They had simply bet the whole workflow on a single model that was never theirs to depend on.

At that point the model chase stops being a technical preference. It becomes a dependency decision.

The Reality Check

The better question is rarely, “Which model is best?” The better question is, “Which model fits this job?”

That sounds less dramatic, but it is closer to how real systems work. The best model for a complex research assistant may be too expensive for repetitive internal classification. The best model for code reasoning may not be the best model for customer support. The best model for open-ended analysis may not be the best model for structured extraction. A smaller, cheaper, more predictable model may be better for a controlled workflow than a frontier model that attracts attention but introduces cost, latency, access, or compliance concerns.

Fit depends on the work. For customer support, fit means tone, escalation, and knowing when to step aside for a person. For software development, it means maintainability and review effort, not just a clever suggestion. For document review, it means finding the one clause that matters, not producing a smooth summary of the whole file.

A frontier model may be the right answer. In some cases, anything less may be false economy. In other cases, chasing the frontier adds little to the outcome and a great deal to the dependency.

There is also a physical footprint to this, even in an edition about market pressure rather than infrastructure. Model races are not weightless. They rely on chips, data centers, energy, cooling, water, supply chains, batteries, and hardware that may be refreshed sooner than older infrastructure cycles. That does not mean advanced models should not be built or used. It means the work should deserve the machinery behind it.

If the output is a better diagnosis, better access to education, better research, safer software, less drudgery, better language inclusion, or clearer decisions, the case can be strong. If the output is more disposable text, more synthetic noise, more features nobody asked for, or more dashboards that nobody reads carefully, then the question becomes uncomfortable in a useful way.

What are we consuming all this capability for?

The Leadership Question

The leadership question is whether the organization is chasing capability, or building competence.

Capability can be bought quickly. A company can subscribe to a better model, enable a new feature, sign with a larger platform, add a copilot, integrate an API, or switch vendors. These are visible moves, and sometimes they are necessary.

Competence is slower. It means knowing which work should be assisted, automated, or left entirely to a person. It means training people to challenge outputs instead of admiring them, and understanding the real cost of review, rework, and dependency.

A company chasing capability keeps asking, “What is the best model?” A company building competence asks a more useful question.

“What is the right model for this work, under these constraints, with this level of risk?”

Asking it that way changes the decision. It brings model selection back into the operating reality of the company. What happens if access or pricing changes tomorrow, can the workflow move to another model, and does the company understand what quality means before the tool is allowed to produce volume?

Frontier models make this harder to ignore. The newest model may be the best choice, but it should earn that place. It should not win only because it has the most attention. News FOMO is a poor architect. It rewards movement, not fitness. It makes the organization feel late before it has even defined the problem.

The best leadership teams I have sat with do something unglamorous. They write down, in one sentence, what the tool is actually for, before they let anyone argue about which model should run it.

There is a better way to think about model choice. Use the strongest model where the work genuinely needs it, a smaller or local model where cost, control, and privacy matter more, and human judgment where the cost of being wrong is too high to hide behind automation.

Using AI effectively means choosing the best-fitting model for the use case, with the right human system around it.

The Closing Signal

The model chase will continue. There will be another release, another benchmark, another impressive demo, another frontier claim, another reason to reopen yesterday’s decision. Fast-moving fields work that way.

The task is not to ignore model progress. The task is to stop treating every new model as a command.

A better model may improve the answer. It does not decide the question. It does not know which customer promise matters. It does not know which internal process carries quiet risk. It does not know which human skill should be protected. It does not know whether the problem deserved AI at all.

That judgment still belongs to people.

I have made this mistake myself: chasing a model because the market made it feel urgent, not because the work needed it. The correction was never technical. It was slowing down long enough to ask the right question first.

The organizations that benefit from AI will not simply be the ones that chase the newest model fastest. They will be the ones that learn how to match models to work, work to value, value to risk, and risk to human judgment.

Before the next model headline changes the room, ask one plain question: are we choosing this model because it is newer, or because it is the best fit for the work we actually need to improve?

The Closing Signal

The model will keep changing. The judgment about fit is still ours.


Discover more from Venkat Mangudi

Subscribe now to keep reading and get access to the full archive.

Continue reading