Venkat Mangudi

Writing on cybersecurity, AI, resilience, leadership, and risk.

Follow

Bad Deployment

Artificial Certainty, an Elytra Security newsletter. Notes on AI, overconfidence, and the new machinery of persuasion. Venkat Mangudi, Founder and CEO, Elytra Security.

Edition 08

Some problems deserve AI. That does not mean every AI system built for them deserves trust.

Too much of the AI conversation gets trapped between two lazy positions. One side treats AI as inevitable progress. The other treats it as inflated theatre. Both miss the more useful middle, where real organizations and real people are trying to decide where AI belongs and how it should be used.

There are places where AI can genuinely help: healthcare access, language support, disability assistance, education, customer service, cybersecurity, public administration, small business operations, research, logistics, agriculture, legal access, everyday knowledge work. These needs are full of friction, delay, confusion, shortage, paperwork, cost, and uneven access.

A well-designed system can make a real difference in such places. It can help a patient describe symptoms before a consultation, help a student practise in a second language, help a public servant answer routine questions faster, help a small business prepare better documents, help a security team triage noise, help a farmer or borrower reach the next step without waiting for a person who is already overloaded.

Some public good uses are obvious because the need is large and the human capacity to meet it is not. Disaster and emergency alerting can reach people in a language and format they actually understand, in the minutes that matter. Rural telemedicine triage can tell a patient whether a symptom needs a clinic today or can wait, in places where the nearest doctor is hours away. Legal aid intake can help someone without a lawyer understand what kind of help they qualify for before they ever reach a queue. Financial inclusion tools can assess a first-time borrower who has no credit history but a real ability to repay. Assistive technology for disability, screen readers, voice interfaces, real-time captioning, can turn a system built for the average user into one that also works for everyone else.

That possibility is real, and it should not be dismissed.

The danger is that a good problem can make a weak deployment look more noble than it is. Once the purpose sounds important enough, people stop examining the design closely. The use case is real, so the solution is assumed responsible. The need is urgent, so the deployment is rushed. The dashboard improves, so the experience is declared better, while the person using the system may feel delayed, misunderstood, or quietly pushed aside.

A healthcare tool can be introduced to reduce waiting time and still increase anxiety if it cannot escalate at the right moment. A learning assistant can be introduced to support students and still weaken effort if it gives answers before helping them think. A customer service bot can be introduced to improve speed and still damage trust if it blocks the path to a human being.

Good intent does not guarantee good use.

The Prompt

The useful question is what happens when AI is applied to a real problem without enough care.

Adding AI where it is not needed is a different problem, and an easier one to spot: a simple form gets a chatbot, a clear search problem becomes a conversational maze. The harder case is where the need is genuine and the deployment still goes wrong, because a real need does not tell you anything about whether the system was built to fail safely.

Take a public helpline that screens callers before routing them to a counsellor. The need for triage is real: counsellors are scarce, and most calls are routine. The failure lies in what the system does with an ambiguous case, not in wanting triage. A model tuned to keep average handling time low will learn, quietly, to resolve borderline calls itself rather than escalate them, because escalation is what the metric penalizes.

The system was never told to ignore the hard cases. It was only told to be efficient, and efficiency found the cheapest way to satisfy that instruction.

The Mirage

The mirage is that a valid use case guarantees a responsible outcome. It does not, and the reason is specific: the system optimizes whatever was measured, and the thing that mattered to the person was usually not the thing that was measured.

A loan-screening tool built for financial inclusion is judged on approval speed and default rate. Both numbers can look excellent while the tool quietly declines every applicant whose income is irregular, which describes most informal workers, the exact group the tool was meant to reach. Nobody instructed it to exclude them. The pattern in the training data did that on its own, and the dashboard has no line item for the person who never got a fair look.

A telemedicine triage line built to reduce hospital load is judged on how many cases it resolves without a doctor. That number improves every time the model is a little more confident than it should be. The cost of that overconfidence does not show up in the triage system’s own metrics. It shows up later, in a clinic, in a case history nobody connects back to the tool.

The failure is not visible in the system that causes it. It shows up downstream, in someone else’s numbers, which is exactly why it survives review after review.

The Reality Check

If the failure hides downstream, the fix has to start by naming, in advance, who is downstream and what a failure costs them. Not the average user. The person the system is most likely to fail.

For the helpline, that is the caller who is ambiguous rather than clearly routine or clearly urgent. The design question is less whether the model can classify most calls correctly than what threshold sends an unclear case to a person instead of a script, and who set that threshold, and whether they were rewarded for setting it generously or stingily.

For the loan tool, that is the applicant with irregular income. The design question is whether the model was tested against that group specifically, before launch, with someone accountable for the result, rather than discovered by the applicants themselves after the fact.

For the triage line, that is the borderline symptom. The design question is whether the system’s confidence threshold for self-resolving a case was set by a clinician thinking about harm, or by an engineer thinking about throughput.

Naming the person the system can fail, before launch, turns an abstract risk into a specific design decision someone has to own.

The Leadership Question

The leadership question is narrower than it sounds: whose job is it to find out, after launch, whether the group most likely to be failed actually was?

In most organizations, no one owns that question, because it belongs to no existing metric. The product team owns adoption and satisfaction scores. Operations owns cost and handling time. Nobody owns the applicant who was quietly declined, the caller who was quietly self-resolved, or the symptom that was quietly reassured. If leadership does not assign that ownership explicitly, the honest answer is that no one is looking, no matter how good the launch numbers are.

That ownership has to come with authority, not just a title. The person responsible for the downstream harm needs the standing to pause the system, change the threshold, or route more cases to a human, even when doing so makes the dashboard look worse for a quarter.

A metric that cannot be overruled by a person is a number the system was trained to satisfy, not a safeguard.

The Closing Signal

None of this argues against using AI in healthcare, lending, legal aid, or public services. The need in those places is real, and the potential is real. The argument is narrower: a good cause is not a substitute for the specific, unglamorous work of finding out who the system is most likely to fail, and giving someone the authority to act on the answer.

Before the next AI system is approved for a genuine public need, ask one plain question: who has checked what happens to the person this system is most likely to get wrong, and who can stop it if the answer is bad?

The Closing Signal

A real problem deserves a real design. Confidence alone is not one.


Discover more from Venkat Mangudi

Subscribe now to keep reading and get access to the full archive.

Continue reading