Every threat model for AI agents starts with the model. The part that holds the credentials, runs the commands, and decides what gets executed often gets waved through as plumbing.
The gear is the argument.
An AI agent investigating a failed deployment might read logs, inspect a repository and run tests, and with enough access it could just as easily change production settings or send internal files to an outside service. Which of those things it can actually do is not decided by the model. It is decided by the harness.
The harness is the software around the model that gives its output consequences. It assembles the context the model sees, exposes the tools the model can call, runs what the model asks to run, holds the credentials those tools need and keeps state between steps. Because an LLM on its own can only produce text, the harness is what connects that text to execution, and it coordinates the controls that constrain it. It also holds the tokens, cookies, filesystem handles and API keys the model never sees, which puts most of the privilege in an agentic system in its hands, shared with the operating system and the services it calls.
That layer is easy to overlook, because benchmarks and threat models tend to focus on the model and treat the harness as neutral wiring. Lasso Security tested that assumption from the attacker's side. They built an automated red-teaming agent, fixed its model, prompt, tools and targets, and ran it inside two off-the-shelf frameworks to see whether anything changed. For some models the framework changed how often attacks succeeded, and for one it changed which kind of attack the agent was good at. That is evidence about adversarial outcomes rather than a measurement of defensive containment, but it is enough to show that the wiring is not neutral.
This article is about securing that layer. Model-level defences such as adversarial training and input classifiers matter, and they appear where relevant, but the article treats them as signals the harness consumes rather than guarantees it can rely on. Its working principle is that the model proposes an action and the harness decides whether that action is permitted, then executes it within limits the model cannot alter. A convincing explanation from the model is an input to that decision, never a substitute for it.
The same scenario runs through every section. An engineer asks an agent to investigate a deployment failure and prepare a patch. To do that the agent needs the relevant logs, a checkout of the repository and somewhere to test its changes, and none of those imply permission to deploy the result, upload configuration files to a vendor or switch off monitoring. Keeping those distinctions intact from the first tool call to the final approval is the practical work of harness security.
Before the run does anything, the harness has to establish what it is allowed to do and keep that record somewhere the model cannot reach through conversation. That means binding the run to an authenticated user or service identity, to the organisation it belongs to and to a defined task, and then recording its operating scope, which is to say the resources it may access, the operations it may perform and when its authority expires. The model can help draft this scope, since it often understands the task well, but harness policy decides whether to accept the draft. A permission policy that a model generated is a proposal, not a grant.
For the deployment investigation, an initial scope might allow the agent to read selected logs, edit a dedicated branch and run tests in an isolated environment, with short-lived credentials limited to exactly those resources. Expiry alone is not enough, though. If the engineer cancels the run, the harness must be able to revoke its access immediately rather than waiting for a token to time out, because a cancelled run that still holds valid credentials is still a running agent.
With scope recorded, every tool request passes through a harness check against the current identity, the resource it names and the operation it wants. The receiving service should enforce access on its own side too, since the harness's view is not the only one that matters, but the harness check comes first and applies to every tool in the same way. It also has to check the right thing. Validating that a repository identifier is well formed says nothing about whether this agent may modify that repository, so the harness resolves the identifier to the actual resource and checks authorisation against what it finds.
These checks are only as good as their coverage. If one connector refuses an operation but another allows the same action without an equivalent check, the harness has a gap rather than a boundary, and the same is true of retries and recovery paths that bypass the normal route. Narrowly defined tools keep this manageable. A tool that can only create a patch on an approved branch is easy to reason about, whereas a general shell can express almost anything and needs correspondingly stronger containment, which a later section covers.
Scope also has a limit of its own, and it is worth being clear about it. It bounds where the agent can act, not whether what it does there is sound. A manipulated agent can introduce a vulnerability while staying entirely within its permitted branch and tools, and no permission check will notice, because nothing out of scope happened. Against that class of failure the harness can only ensure that the agent's output goes through the same code review, testing and independent validation as any other change, and that none of those steps can be skipped on the grounds that an agent produced it.
Scope controls what the agent may do. It says nothing about what the model can be persuaded to attempt, and the harness cannot assume the model will resist persuasion.
During the investigation the agent might find a troubleshooting page recommending that internal configuration files be uploaded to a diagnostic service. The advice sounds relevant and may be presented as a necessary step, but the page has no authority to approve that disclosure, and nothing about how it is phrased changes that. OpenAI describes a comparable attack in which email content claimed permission to retrieve employee information and submit it to a supposed compliance endpoint. What made it work was a plausible business process and an assertion of authority, not anything that looked like an attempt to override the agent's instructions.
Input scanners, classifiers and model training all help detect manipulation of this kind, and all of them can be wrong, particularly when the malicious instruction resembles legitimate task content. The harness should treat their verdicts as risk signals. A clean scan can inform a decision, but it must never expand permissions, so that even if the model accepts the vendor page's explanation, the tool layer still refuses the upload. The upload was never in scope, and no input check can put it there.
Provenance is where the harness's job becomes hard. The harness can label a document with its source as it enters context, but once the model has combined that document with logs and internal notes, there is no reliable way to tell which input shaped which part of the resulting summary. One conservative response is to apply every source's restrictions to the whole summary, which is safe but blocks harmless uses along with dangerous ones. Asking the model to explain where each part came from is not an alternative, because that explanation is just another claim from the component the harness is defending against.
Two harness architectures take provenance seriously. The dual-LLM pattern separates a privileged planner from a quarantined model that processes untrusted content without access to action tools, which shields the planner from direct exposure. The gap is that the quarantined model can still hand back a manipulated destination or file identifier, and the planner may act on it without knowing where it came from. CaMeL closes that gap by running the planner's program through an interpreter that tracks metadata on every value and enforces policies on how information reaches tools. It is a concrete enforcement mechanism, and its authors are candid about the costs, which include utility loss, the work of writing policies and side channels that remain. Its guarantees hold within a defined architecture and threat model, and they should not be read as proof that an unrestricted browser agent is secure.
Back in the investigation, all of this means the vendor's proposed upload destination stays an untrusted candidate no matter what the model concludes about it. Whether that destination and that data may be combined is a harness policy decision, and if the agent writes that "the engineer approved the upload", the harness consults its approval record rather than the sentence.
Policy decides what the agent may attempt. Containment limits what its processes can do regardless of what policy said, which makes it the harness's second line, and it has to hold even if the policy layer is bypassed or misconfigured.
The execution environment should hold the checkout, test fixtures and services the assignment requires and nothing more, with unrelated host directories and administrative interfaces out of reach. Limits on processing time, memory and storage catch runaway work. Every one of these restrictions has to extend to generated scripts and child processes, because a boundary that applies only to the first command is not a boundary.
Anthropic's sandboxing design is a worked example. It combines operating-system-enforced filesystem restrictions with network controls, and it adds a Git proxy that checks each scoped request before attaching credentials for the downstream service, so the agent can ask for an operation without ever holding the secret that performs it.
In our example, that pattern means the test runner does not inherit production credentials merely because they exist on the engineer's machine. A credential broker inside the harness can permit a push to the investigation branch while withholding any token that could touch other repositories. Doing so makes the broker the most sensitive component in the system, and the workload it serves must not be able to alter its configuration or read its credential store. Whatever protects the harness from the agent has to protect the broker most of all.
A branch-scoped push is only as narrow as what it sets in motion. In most repositories a push triggers something, whether a test workflow, a build, a preview deployment or a hook, and those processes run with their own credentials, which are often broader than the agent's. A modest repository operation can therefore start a far more privileged one. The threat model for the investigation has to include the workflows a push on that branch will trigger, what they can reach and which secrets they hold, and the harness should treat a push to a branch with production hooks as a different operation from a push to one without them.
Network restrictions are where containment gets uncomfortable. A hostname allowlist is easy to operate, but an approved host may contain attacker-controlled accounts or upload destinations, so allowing the host does not mean allowing every request to it. Checking the recipient and the operation gives finer control, at the price of service-specific knowledge and ongoing maintenance as those services change. General browsing is the hardest case of all, since a single page visit can trigger redirects and further requests the policy never saw.
The right boundary therefore depends on the assignment. An investigation may work well with a fixed set of documentation sites and package mirrors, while broader research may need a separate browsing environment with no production secrets in it at all. Restrictions will sometimes block legitimate work, a test that cannot run without a package install being the obvious case, so the harness needs an explicit exception path and a record of where the boundary gets in the way. What it must not do is relax the boundary automatically whenever the agent gets stuck, because a restriction that yields to inconvenience is not a restriction.
Everything so far assumes a single agent. In practice the harness may let the main agent hand log analysis to one worker and patch testing to another, and each handoff is a chance for scope to leak unless the harness carries it across.
Each worker needs only the access its part of the job requires, so the log analyst has no use for deployment credentials and the test worker has no use for the organisation's wider incident history. To enforce that, the harness limits each worker to the permissions the parent is allowed to delegate, narrows them further by the worker's assignment and by organisation policy, and records the parent run, permitted resources and expiry alongside the worker's access. A child cannot grant itself what its parent lacks, and cancellation propagates through the task tree, so stopping the main agent stops its workers too.
Worker output needs the same handling as any other input. A summary of hostile logs remains derived from hostile input even when another agent wrote it, so the harness carries that provenance across the boundary rather than letting it reset. If a worker reports that approval has been obtained, the main agent checks the harness's approval record and not the report. Adding a reviewing agent can improve judgement, but a second model does not by itself create an independent security boundary. That still has to come from enforcement outside the models.
Delegation also needs shared limits on worker count, recursion depth, spending and external requests. Set these explicitly for the workload rather than letting them emerge from the agents' behaviour, because a small investigation that can spawn workers without limit is no longer small.
The harness also inherits every trust relationship its integrations bring, and those relationships carry familiar application-security problems alongside prompt injection. The Model Context Protocol's security guidance covers confused-deputy attacks, token passthrough and server-side request forgery, all of which come down to the same requirement. A harness integrating such servers must verify the requesting identity, the intended audience of any credential and the destinations it may contact, because being authenticated to one interface does not confer access to every service behind it.
Third-party plugins, skills and tool definitions are changes to the harness itself and deserve the treatment given to any other software dependency. Review the code, pin and record versions, and assess what permissions an update requires. Their descriptions matter as much as their code, since descriptions shape how the model uses them. Above all, an agent must not be able to route around a refused operation by installing a more powerful extension, which means installation is a harness-level authority decision rather than something the workload performs for itself.
Tool responses need basic resource and format checks before they enter the model's context. Very large responses consume resources and crowd out useful information, so the harness should paginate where appropriate and make truncation visible, because an agent that does not know its evidence is incomplete will act as though it were complete. These measures improve reliability, but they say nothing about whether the content that remains is trustworthy.
Persistent memory extends the problem across time, and memory is harness state. AgentPoison demonstrates attacks through poisoned long-term memory or retrieval knowledge bases without any model retraining. Its findings concern particular experimental conditions, but they show why stored material needs a trust policy of its own.
In the investigation, a note that the vendor recommended an upload should remain a sourced, dated observation and never become a standing preference or permission. To keep it that way, the harness stores externally derived notes separately from authenticated authorisation records, retains version history and supports quarantining contaminated entries. The same separation protects the harness's own policy, plugin definitions and approval logic, which must stay out of the agent's write path, because a harness whose rules the workload can edit has no rules.
Human approval is a harness function too. It enters when the investigation produces a patch ready for production, because that is where the agent's authority ends and someone else's begins. The approval surface should show the reviewer the target environment, the exact commit or build, the proposed changes and their expected effects, and the recovery procedure where one exists, so that the reviewer sees the operation the harness will actually execute rather than the model's description of it.
Bind the decision to an authenticated person with authority over that environment, and bind it to the operation precisely. The approval record should carry an immutable artefact identifier such as a commit hash or image digest, the exact destination and every security-relevant parameter, and any mismatch at execution time should invalidate it unless the grant explicitly allowed that variation. Record the action, its expiry and the relevant policy version in protected storage. Rechecking immediately before execution is necessary but not sufficient, because the artefact or target can change between the check and the act, so where the platform allows, execute against the identifier that was approved rather than a mutable reference to it. A mutable reference can be pointed at a different artefact. A digest cannot.
Shared conversations sharpen this in two directions. Someone who can contribute debugging information may not be entitled to approve a deployment, so the harness tracks who authorised each consequential action instead of treating every participant as one undifferentiated user, and it keeps the approval channel for those actions outside any text the agent controls.
The same person may also not be entitled to see everything the agent found, and disclosure does not need a tool call. A final answer, a generated file, a rendered link that carries data in its URL, or a conversation shared onwards can all leak what the investigation uncovered, and OpenAI's account of prompt-injection attacks treats links and navigation as exactly this kind of channel. Each output surface should therefore be bound to an authorised audience. Who may approve an action and who may receive its results are separate questions, and the harness has to answer them separately.
There is a tension in all of this, because approval only works if it is rare enough to be considered. Anthropic identifies approval fatigue as a reason to move routine work inside a bounded sandbox, while noting that probabilistic approval controls retain a chance of missing unsafe actions. Reducing unnecessary prompts is the right instinct, provided automated approval is not mistaken for certainty.
The resolution is standing authorisation that is itself harness policy, with specified operations, resources, limits and duration and a record of who issued it. Reading the investigation logs fits comfortably inside such a grant. Deploying the patch does not, and that is the point, because review then stays focused on real changes in authority.
None of the above can be assumed to work, and harness tests have to be designed so that a well-behaved model does not mask a broken control. Test the whole investigation, including its awkward paths. Place malicious instructions in logs, documentation and worker summaries, check whether poisoned memory affects a later run, and exercise cancellation, expired credentials, changed deployment targets and attempts to switch connectors after a refusal.
Measure the model's behaviour separately from the system's outcome, because they fail differently and only the second is the harness's responsibility. A model might reject an attack in its response and then disclose the same information through a later tool call, and if the harness allowed that call, the harness failed. It might instead propose an unsafe action that the harness blocks, which is the harness working as designed, though it is worth knowing how often it has to. Swapping models while holding the harness constant, or deliberately using a compliant model that follows every injected instruction, helps show which controls hold on their own, but even a compliant model may never happen to try the bypass you care about.
So test the enforcement layer directly as well. Submit forbidden tool requests straight to the permission and execution layers without a model in the loop, replay an approved action against a changed target, cancel a run mid-operation, and call a second connector after the first refused. These tests make enforcement failures easier to reproduce without depending on the model to attempt a particular bypass. They complement the end-to-end runs rather than replacing them, since OpenAI's red-teaming work for Atlas describes attacks spanning many steps, which is exactly what isolated prompt tests miss.
Security results mean little without the other half of the ledger. Measure legitimate task completion, elapsed time, cost, approval frequency and incorrectly blocked actions alongside them, and state the threat model and coverage the tests represent, because passing a suite does not establish immunity and a change to the model, the tools or the permissions can invalidate last week's result.
Then prepare for the failures that get through. The harness's logs should connect the initiating identity to tool requests, policy decisions, approvals and resulting changes, without exposing secrets along the way, so that an investigator can reconstruct events without relying on the agent's account of them. Operators need harness-level ways to stop runs, revoke access, disable connectors and quarantine memory, and consequential retries need idempotency controls so that a timeout does not turn one authorised action into two. Backups can restore changed resources, but they cannot retrieve information already disclosed.
The investigation succeeds when the agent produces a tested patch within its scope and any production change receives the authorisation it requires. The harness decides whether the agent stays within that scope. It cannot decide whether what the agent did inside it was correct, which is why review and validation remain part of the system rather than something the agent's permissions replace. The question to ask before granting an agent more freedom follows from that division. If the model accepted a convincing false instruction, which harness control would still hold, what evidence supports that confidence, and which legitimate tasks does the control make harder? Those answers make the harness reviewable. The model's apparent confidence, and the fluency of its explanations, do not.