Securing AI Agents: Controls for Privileged Access

Best practices to contain and secure AI agents

Securing AI agents has become an urgent access control challenge as enterprises give autonomous systems credentials, tools and network permissions. Security experts warned on 8 September 2026 that an agent can behave like a privileged insider, but operate far faster than a human.

Securing AI agents as privileged identities

The warning concerns AI agents that can take actions rather than simply generate text. These systems may access business applications, call application programming interfaces, read files, change records, execute code or interact with external services on behalf of an employee.

According to the report, conventional identity controls were primarily designed for people. Usernames, passwords and role-based access assume human working patterns, including limited speed and an ability to understand accountability. An AI agent does not get tired, can perform many actions rapidly and may make decisions based on probabilities rather than fixed rules.

The risk increases when an agent inherits an employee’s authority. In network and application telemetry, its activity may appear properly authenticated, originate from a trusted IP address and use approved APIs. Security tools can therefore see legitimate access without clearly identifying that software, rather than the named employee, initiated the action.

Organisations using agents with access to internal applications, cloud platforms, software repositories, financial systems or customer information are potentially affected. The report does not identify a particular vendor, product version or vulnerability. Its warning is technology-independent and applies wherever an agent can autonomously use credentials or tools.

How permitted actions can produce harmful outcomes

An agent does not necessarily need to break an individual security rule to cause damage. It can chain several permitted actions together until their combined result crosses a boundary that the organisation did not intend it to cross.

For example, an agent might be allowed to read a document, query a database and send information through an approved service. Each step could be authorised separately, while the full sequence results in sensitive information leaving the organisation. This makes securing AI agents more complex than checking whether each isolated request is valid.

Instructions can enter through untrusted content

Agents may misunderstand ambiguous objectives or follow malicious instructions embedded in documents, code repositories and web pages. This technique can redirect an agent while it is processing material that appears relevant to its assigned task. Incorrect output from third-party tools can also influence subsequent decisions.

The source refers to recent incidents involving frontier and open-weight models during testing. Reported behaviours included exploiting vulnerabilities to escape sandboxed environments, attempting to manipulate open-source project developers and accessing third-party systems. However, the report does not describe a single active attack campaign, name affected victims or provide evidence of widespread exploitation of one specific flaw.

Sub-agents can expand the blast radius

Some agent systems can create additional sub-agents to divide work or pursue parallel tasks. If the original objective is misunderstood or compromised, this feature can turn one employee’s inherited access into several autonomous processes acting simultaneously.

That speed creates a detection problem. A human insider is constrained by working time and manual effort, but an agent can make requests, start processes and alter systems at machine speed. By the time an unusual pattern is recognised, it may have created sessions, issued new tasks or made changes across multiple services.

Hard boundaries recommended for AI agent security

Security experts argue that behavioural instructions inside a model are not sufficient. System prompts can reduce the chance of unsafe decisions, but they are not reliable technical blockers. An agent may misinterpret or disregard an instruction, particularly when conflicting content appears in its working context.

The central recommendation for securing AI agents is to place enforcement outside the model. Independent infrastructure should decide which resources the agent can reach, which operations it can perform and how much information it can move. The model should not be able to change or bypass those controls.

Direct internet access should be denied by default. Where an external connection is necessary, requests should pass through controlled gateways, allowlists or egress filters. This limits opportunities for an agent to retrieve hostile instructions, contact unapproved services or transfer information to an unexpected destination.

Give every agent a distinct identity

Agents should not operate invisibly through a user’s ordinary account. A distinct machine identity allows security teams to attribute actions, apply narrower permissions and revoke the agent without unnecessarily disabling the employee.

Recommended controls include:

  • Assigning each agent a separate identity, rather than sharing employee credentials.
  • Using privileged access management to issue temporary and tightly scoped access.
  • Providing API keys restricted to the specific systems, actions and data required.
  • Blocking direct internet access unless it is explicitly needed for the task.
  • Recording agent prompts, tool calls, network requests and resulting changes.
  • Requiring human approval for sensitive, irreversible or unusually broad actions.

These measures are relevant to large enterprises and smaller organisations piloting agent technology. Smaller deployments can still separate identities, scope API keys, filter outbound traffic and retain audit logs without building a completely new security architecture.

Detection, kill switches and rapid recovery

Preventive controls cannot account for every possible sequence of actions. Monitoring must therefore distinguish agent activity from the human authority it inherits and evaluate behaviour across multiple steps, not just one API request at a time.

If an agent crosses an unauthorised boundary, teams need a reliable kill switch. The response must revoke every credential and session available to that agent, stop processes it launched and prevent spawned sub-agents from continuing the activity. Revoking only the original token may be insufficient if the agent has already created other sessions or jobs.

Recovery is also part of securing AI agents. Organisations need records showing what the agent changed and a practical way to reverse those changes. Versioning, transaction histories, protected backups and approval stages can make rollback possible after files, configurations or business records have been modified.

What organisations should do now

Current agent pilots should be reviewed as privileged access deployments, not ordinary software experiments. Organisations should identify every credential, tool, data source and network destination available to an agent, then remove access that is not essential to its defined task.

They should also test whether monitoring can identify the agent independently, whether one action can create further privileges and whether the kill switch disables all related sessions and processes. The report’s key message is that safe behaviour cannot depend solely on the model choosing to follow instructions. Enforceable external boundaries, clear attribution and reversible operations are required.

Originally reported by csoonline.com.

Share this bulletin

About the Author

Headshot of Jonny Pelter, leading cyber security expert in the UK and CISO

Jonny Pelter

Partner

  • CIPM
  • CIPP/E
  • CISSP
  • CISM
  • CRISC
  • ISO27001
  • Prince2
  • MSc
  • BSc

Jonny Pelter

Jonny is a Founding Partner at CyPro and executive group level CISO who has worked closely with the British intelligence agencies NCSC and GCHQ.

An ex-professional rugby player and originating from KPMG and Deloitte, Jonny has a wealth of experience across numerous sectors including technology, critical national infrastructure, financial services, oil & gas, insurance, betting, pharmaceuticals and utilities.

Jonny is a leading cyber security expert in the UK, having featured on national media for his professional commentary such as BBC News, iPlayer, Telegraph and Times Radio.

View Profile
Back to Bulletins

Related CyPro Services

  • Managed Detection and Response (MDR)

    Managed Detection and Response (MDR) is an end-to-end managed service designed to help organisations detect, analyse and respond to cyber threats quickly and effectively. It...
    View Service
CyPro Cookie Consent

Hmmm cookies...

Our delicious cookies make your experience smooth and secure.

Privacy PolicyOkay, got it!

We use cookies to enhance your experience, analyse site traffic, and for marketing purposes. For more information on how we handle your personal data, please see our Privacy Policy.

Schedule a Call