Venkat Mangudi

Writing on cybersecurity, AI, resilience, leadership, and risk.

Follow

What Happens When the Model Is Off

Executive Signals. Strategic foresight for boardrooms and investors. Venkat Mangudi, Founder and CEO, Elytra Security.

Issue 7 · September 2026

The pitch that gave itself away

Recently an email arrived telling me that an AI agent had found Elytra Security, identified me through public information, verified my address and written the message itself. A company called Explee had built it using Claude 5.1, the email said. Having completed those apparently impressive tasks, the agent offered to land high-paying customers for me, fully automated. All I had to do was reply with a single word.

I blocked the sender. I could not verify any of the claims, and I had no interest in the offer. This was spam, and being told that a machine had produced it gave me no extra reason to trust it. I would not call it a cyberattack. What stayed with me was the sales logic underneath it: the absence of a human being was presented as the selling point.

I have the same reservation about the enthusiasm for AI versus AI in cybersecurity. Before I accept that proposition, I want to understand what the proposed defence does, why it needs a model to do it, and what happens when the model is unavailable or wrong.

IPart I

Three Reports, Read Carefully

Google, Anthropic and Adalat AI published within roughly 48 hours of each other. Different activities, different timelines, different purposes. The proximity of the publication dates is not evidence of a single event.

Three reports on the 8th and 9th

Three reports published on 8 and 9 September give us useful material for this discussion. Google’s threat intelligence report, Anthropic’s alignment assessment and Adalat AI’s account cover different activities on different timelines. They arrived within roughly 48 hours of one another, and a great many people have since treated them as three parts of a single incident. They are not.

ReportWhat it actually describes
Google, threat intelligenceAn attacker used an AI tool and agents to build and run a credential-harvesting campaign in under six hours, with scanning and troubleshooting handled autonomously. Offensive capability, in an attacker’s hands.
Anthropic, alignment assessmentFour incidents, three disclosed in July and one dating from January 2026. Models ran in sandboxes that allowed Internet access and had normal cyber safeguards switched off.
Adalat AI, defensive investigationA human-directed investigation using agents. An exposed research service led to files containing credentials, which led to source code and production infrastructure.

Read separately, each one tells us something specific and useful. Read as a single wave, they tell us mainly that three organisations chose the same week to publish.

Six hours, and the sentence after it

Google describes an incident in which an attacker used an AI tool and agents to build and execute a credential-harvesting campaign in under six hours. The agents managed scanning, troubleshooting and other operational tasks on their own. That is clear evidence of capability in an attacker’s hands, and I have no argument with it.

Google Threat Intelligence · One documented campaign · Sept 2026

< 6hFrom build to execution of a credential-harvesting campaign
AutonomousScanning and troubleshooting
Human-directedObjective and campaign design
Not observedFully autonomous pipelines in the wild

Google also says it has not yet observed threat actors deploying fully autonomous pipelines against targets in the wild. That may well be true of this incident, and it does not mean the agent-enabled operations documented elsewhere in the same report can be waved away. Reading both findings together gives us a more accurate picture than picking whichever sentence produces the better headline.

An assessment, and a laboratory that was not sealed

Anthropic’s report examines four incidents, three of them disclosed in July and one dating from January 2026. The sandboxed environments allowed Internet access, and the models were running without their normal cyber safeguards. That is a significant miss. Whatever the models did inside that environment, the environment itself was not secured to the standard the work required.

A laboratory that allows outbound access and has its safeguards turned off is not testing containment. It is assuming it.

Anthropic has since revised its earlier explanation that the models appeared to believe they were operating in a simulation. Its assessment now identifies biased reasoning, and continued pursuit of the assigned task despite evidence of potential real-world harm. I give Anthropic credit for correcting the record in public. The first explanation of an AI incident is often the most comforting one, and often the one that later changes.

Serious failure

Adalat’s account, co-authored with the researcher, describes a defensive investigation directed by people and carried out with agents. An exposed research service provided access to files containing credentials, which opened the way to source code and production infrastructure. Those are serious failures.

The access chain, as described in the report

Exposed research serviceReachable, unintended
Files containing credentialsStored in the clear
Source code and productionReached

The report also states that confidential documents and audio were not accessed, and that the infrastructure credentials did not supply the keys needed to decrypt stored court documents. It further acknowledges that no controlled comparison was run to quantify the agents’ speed advantage. Parallel investigation was demonstrated. The size of the improvement over another approach was not measured. That is a useful restraint, especially when speed is being used to argue that existing defensive approaches are inadequate.

A capability is not a verdict

My concern is the conclusion these reports may encourage. A capability demonstrated under particular conditions does not establish that every organisation is equally exposed, and it certainly does not establish that resistance is futile. For an executive, the useful discussion concerns whether similar conditions exist in their own organisation, and what can be done about them.

Resistance being futile is a conclusion somebody sells. It is not a finding in any of these three reports.

An experimental system should have a defined purpose and a deliberately limited set of connections and permissions. Credentials available to it should be restricted to what the work requires. It should never have access to production. Network restrictions and identity controls should be enforced independently of whatever a model has been instructed to do. OWASP’s guidance on excessive agency makes the same architectural distinction: downstream systems should enforce authorisation, rather than relying on a model to decide whether an action is allowed.

Organisations can do this work. It requires knowing their systems, deciding what access is acceptable, and verifying that the restrictions operate as intended. AI involvement gives us a reason to examine those arrangements carefully. It does not make them irrelevant.

IIPart II

The Engine and the Repertoire

AI belongs in the defender’s repertoire. I do not want it to be the engine on which essential protection depends. That engine should be deterministic, testable and engineered for reliability.

Useful in the repertoire, wrong as the engine

AI is useful in the defender’s repertoire. I do not want it to be the main engine on which essential protection depends. There is plenty of room for it further down the chain, in analysis, investigation and reporting, where its contribution can be examined and where an unavailable answer does not suspend the underlying protection.

Calling the core deterministic does not make it reliable by declaration. Deterministic software fails too, and it will execute a bad rule consistently and at speed. Availability, capacity and recovery still have to be engineered and tested. What I am asking for is behaviour that can be understood and verified, including during failure, instead of accepting the phrase AI-powered as an explanation of how protection is maintained.

A model that produces a plausible answer quickly has not necessarily shortened the time to a correct decision.

A local, open-weight model can be a good choice, provided it has enough relevant context and is demonstrably faster than a deterministic solution for the task at hand. I would be comfortable using one on that basis. Local deployment changes the dependency arrangement, and I would still want tests for unavailable inference, incomplete context and incorrect output.

Disable the model, then look again

During a product demonstration for AI versus AI, ask what happens if the model powering the tool is disabled. Then find out which detections continue, whether evidence collection is affected, and whether restrictions remain in force. The demonstration should also show how the system resumes without disrupting the flow of evidence.

FunctionInference unavailable
Network restrictions at the boundaryHOLDS
Identity and authorisation enforcementHOLDS
Signature and rule-based detectionsHOLDS
Log and evidence collectionHOLDS
Model-generated triage and narrativeSTOPS
Model-driven autonomous remediationSTOPS
The continuity ledger

This is the shape I want to see on the vendor’s own product. Anything in the STOPS column that an executive believed was in the HOLDS column is a gap somebody will discover during an incident.

Some products offer a switch to an alternate model when one is unavailable. That has to be tested and demonstrated, including the conditions under which the fallback also fails.

A defined state without always-on inference

I would require that essential protection have a defined operating state that does not depend on always-on model inference. Keeping established functions available during a dependency failure is a familiar resilience principle. AWS describes it as static stability. Applying that principle to AI-dependent defence is a reasonable engineering requirement, not an unusual one.

The measurement should cover the whole job. Time spent assembling context, checking the answer and correcting mistakes belongs in the comparison. For a purchasing decision, I would ask for representative tasks, comparable requirements, and results the customer can reproduce.

Questions for the vendor, before the commercial conversation

  • Which detections continue when the model is disabled?
  • Is evidence collection affected, and for how long?
  • Do network and identity restrictions remain in force?
  • How does the system resume without breaking the evidence flow?
  • Under what conditions does the fallback model also fail?

Manual does not mean slow

I also object to the assumption that entirely manual work must be slow. Manual describes how the work is performed. It tells us very little about the time required or the quality of the result. An experienced person with the right information and the authority to act may reach an appropriate decision promptly. A poorly designed automated process may delay it.

Test the response against the requirement. Do not assign performance characteristics to the label.

The organisation needs to test its own response against its own requirements. That test settles the argument, and it costs less than most of the products sold on the strength of the assumption.

IIIPart III

Finding It and Fixing It Are Different Jobs

Reliability includes the consequences of the defensive system’s own actions. Identifying an exposure and safely changing a production environment are separate responsibilities, and competence at the first is no qualification for the second.

The tool reports success

This is why I am wary of tools that promise to find vulnerabilities and fix them automatically. I would not grant a tool authority to change production merely because it had demonstrated competence at discovery.

Tightening an over-permissive firewall may address the specific finding while interrupting traffic a legitimate application depends on. The change can be technically correct for the security issue and entirely wrong for that environment. The tool sees something that should change, applies the prescribed change, and reports success. The business experiences an outage.

As the tool reports itAs the business experiences it
Finding closed. Over-permissive rule tightened. Exposure reduced. Ticket resolved within the service level.Service down. Traffic a legitimate application depended on is now blocked. Nobody connects the outage to the security change for several hours.

Counting the finding as closed would be an inadequate measure of that result. I would expect verification of both reduced exposure and continued operation of the affected service.

Approval that means something

Conventional automation deserves the same scrutiny. Automatic remediation is safe when the procedure has been approved and tested for the relevant conditions, with a defined scope, checks before execution, and health monitoring that can halt a rollout when something goes wrong. Those safeguards are part of the remediation capability.

For critical actions, examine how approval works

  • Does the person approving a change receive enough information to assess its consequences?
  • Is the tool’s permission limited to the approved action?
  • Can the permission be withdrawn independently of the model?

If the system behind the prompt retains broad authority, asking for confirmation achieves nothing. Ask for a demonstration of the rollback process as well. Review changes involving databases, schemas or other stateful components that may be difficult to reverse. Find out what restoration entails, what happens to work performed after the change, and how long the service could be disrupted.

Urgency is not unlimited permission

Patching is necessary maintenance, and delay carries its own risk. NIST’s patch-management guidance connects that work directly to an organisation’s ability to achieve its mission. The decision concerns how to reduce exposure under the circumstances, with security and business owners jointly answerable for the outcome.

Organisations should establish who can authorise emergency changes. Safe-deployment guidance does not treat urgency as unlimited permission to alter production systems.

A security programme has to account for attacks that succeed, dependencies that vanish, and interventions that go wrong.

This is where resilience deserves the same attention as prevention. NIST’s cyber-resilience guidance addresses the ability to withstand, recover from and adapt to adverse conditions. The objective is to keep delivering the organisation’s purpose under conditions less favourable than those assumed in normal operation.

Start with one essential service

For a management review, I would begin with an essential business service and examine what happens when one of its dependencies is lost. I would want to know what can continue safely, who decides to restrict or suspend operations, and what evidence supports the recovery plan. The exercise should include the security tools themselves.

A network boundary that limits access, an identity that cannot administer unrelated systems, and a recovery procedure that has actually been exercised each give you something tangible to examine. None of them is an absolute guarantee, and their effectiveness depends on implementation and verification. They do give executives a practical basis for deciding where attention and investment are needed. Network security guidance continues to emphasise isolation, filtering and controls at multiple boundaries for exactly this reason.

I would apply the same standard to an AI-based proposal. Show the improvement on a task that matters to the organisation. Explain the dependencies and the authority being granted. Demonstrate what happens when the component is unavailable, returns an unusable answer, or recommends an inappropriate change. These are reasonable questions for any system entrusted with protecting a business.

The closing view

The three reports warrant careful reading and, where relevant, action. They do not warrant abandoning judgement, and they do not mean an attacker’s use of AI should dictate the defender’s architecture. My preference is to invest in cybersecurity and resilience that can be demonstrated in the environment being protected.

AI has a place wherever it improves that work. The test I would put to any vendor, and to my own engineering, is the same one the spam email failed: tell me what the thing does, why it needs a model to do it, and what is left standing when the model is not there.

The closing signal

The organisation should remain protected and able to operate when it does not.

Sources and references

  1. Google. Threat intelligence report on AI-enabled credential harvesting. September 2026.
  2. Anthropic. Alignment assessment of four sandbox incidents. September 2026.
  3. Adalat AI. Account of a defensive investigation conducted with agents. September 2026.
  4. OWASP. Guidance on excessive agency.
  5. NIST. Patch management guidance.
  6. NIST. Cyber resilience guidance.
  7. AWS. Static stability.

Discover more from Venkat Mangudi

Subscribe now to keep reading and get access to the full archive.

Continue reading