Book a consultation

Agent governance starts with the blast radius

A technical design for agent permissions, tool boundaries, prompt injection defenses, PII controls, approvals, and audit evidence.

Agent governance starts with the blast radius

An assistant that drafts a refund email can be wrong.

An agent that issues the refund can move money.

The model may be identical in both systems. The security boundary is not.

Governance has to execute inside the request path, not sit in a policy document. Audit trails, personal data checks, prompt versioning, and model change management all need runtime controls that can block or record an operation.

Layered agent blast radius showing least privilege, allowed data, approved tools, and staged autonomy
Autonomy is safe only inside explicit data, tool, and approval boundaries.
Risk to autonomy matrix mapping impact and reversibility to approval, controlled execution, recommendation, and automation
Impact and reversibility determine how much autonomy an action deserves.

Begin with a capability manifest

List what the agent can see and change before writing its system prompt.

agent: customer-refund-assistant
identity: service/refund-agent

read:
  - customer.profile.basic
  - orders.by_customer
  - refund_policy.current

propose:
  - refund.create

execute:
  - none

forbidden:
  - customer.payment_methods
  - employee_records
  - refund.approve
  - policy.write

This manifest gives security, engineering, and operations one artifact to review. It also exposes architectures where the agent runs under an administrator credential because that was easier during the PoC.

The tool registry should enforce the manifest. The prompt can describe the boundary, but it cannot implement it.

Least privilege applies at three layers

Agent permissions are often discussed as one setting. There are at least three separate controls.

The first is functionality. Which tools exist in the agent's context? A model cannot call a tool it never receives.

The second is tool authorization. Can this principal invoke this operation for this user, tenant, and purpose?

The third is downstream authorization. Does the CRM, database, or payment API enforce the same boundary when the tool calls it?

OWASP's Excessive Agency guidance recommends minimizing extensions, functions, permissions, and autonomy. It also recommends authorization inside downstream systems rather than trusting the LLM application to make the final access decision.

Use user-scoped credentials where the operation acts on behalf of a user. Use narrow service credentials for background jobs. Avoid one shared token that turns every agent request into the same privileged identity.

Separate read, propose, approve, and execute

A tool named manage_customer_account is too broad for a production agent.

Split capabilities by effect:

read_customer
read_orders
propose_address_change
approve_address_change
execute_address_change

The proposal is structured data, not an action disguised as a draft.

{
  "operation_id": "address-change-1842",
  "customer_id": "customer-772",
  "current_country": "PL",
  "proposed_country": "DE",
  "evidence": ["verified-user-request"],
  "risk_tier": "medium",
  "expires_at": "2026-09-02T14:10:00Z"
}

Approval signs this exact proposal. If arguments change after approval, the signature no longer matches and execution must stop.

This prevents a time-of-check to time-of-use problem where a reviewer approves one action and the agent modifies it before execution.

Treat retrieved content as untrusted input

Prompt injection is not limited to the user message. A web page, support ticket, PDF, or tool result can contain instructions aimed at the model.

Label content by trust zone:

system policy           trusted instructions
application state       trusted structured data
user request            untrusted content
retrieved documents     untrusted content
tool results            authenticated, not instruction-trusted

Authentication does not make content safe as an instruction. A valid support ticket can contain malicious text.

Keep policy and data in separate channels or structures supported by the model API. Tell the model which fields are evidence and which are instructions. More importantly, enforce sensitive actions outside the model.

Use allowlisted tools with typed arguments. Validate arguments against business rules. Recheck authorization at execution time. Limit network destinations and file paths. Assume an attacker may fully control the model's proposed tool call.

Make tool contracts narrow and typed

A dangerous tool accepts free-form commands.

def run_sql(query: str) -> str:
    ...

A safer tool exposes the operation the workflow needs.

class OrderLookup(BaseModel):
    order_id: str
    fields: list[Literal["status", "items", "delivery_date"]]

def get_order(request: OrderLookup, principal: Principal) -> OrderView:
    authorize(principal, "orders:read", request.order_id)
    ...

The second contract constrains syntax and semantics. It can apply row-level authorization and produce an audit event with a stable action name.

LangChain structured output can validate model responses against Pydantic, dataclass, TypedDict, or JSON Schema definitions. Validation improves the boundary between probabilistic output and deterministic application code. It does not replace authorization inside the tool.

Build approval policy from impact

Approving every tool call makes the system unusable. Approving none makes autonomy a default rather than a risk decision.

Classify operations using effect, reversibility, value, and audience.

read_public_policy:
  approval: never

send_internal_draft:
  approval: only_if_low_confidence

send_customer_email:
  approval: always_until_review_threshold_met

issue_refund:
  approval: always_above_25_usd

delete_customer_record:
  approval: always_two_person

The reviewer needs the proposed action, arguments, evidence, policy result, and expected effect. A generic "agent wants to call refund" dialog encourages blind approval.

LangChain human-in-the-loop middleware can pause selected tool calls and accept approve, edit, or reject decisions. LangGraph persists the interrupted state. Use a durable checkpointer and authenticate the reviewer outside the model conversation.

Edits need extra care. If a reviewer changes tool arguments, validate and authorize the edited action again.

Detect PII at ingress, context assembly, and egress

One PII filter at the final response misses several paths.

At ingress, classify user data before placing it in logs, traces, or third-party model calls.

During context assembly, enforce whether a source and field may enter this model request. Redact values the task does not need.

At egress, inspect generated text and structured output before returning it or passing it to another tool.

Tool results and traces need the same policy. An application can block PII from the final answer and still copy it into an observability vendor.

Use deterministic patterns and data classification for identifiers with known formats. Use named-entity recognition where formats vary. Treat model-based detection as another probabilistic signal, not the sole privacy control.

LangChain middleware includes PII detection controls, but the application still needs retention, access, and deletion policy for anything it records.

Audit actions, not private reasoning

An audit trail should answer:

  • who initiated the request
  • which principal the agent used
  • which version of the workflow ran
  • which data sources and policy versions were consulted
  • which tools were proposed and executed
  • which authorization decisions occurred
  • who approved an action
  • what changed in the downstream system

You do not need private chain-of-thought to answer these questions. Store observable transitions and evidence.

Tamper-evident storage matters for high-impact actions. Write audit events through a separate service identity. Restrict updates. Use retention rules that match the action class.

Article 12 of the EU AI Act requires defined high-risk systems to support automatic event logging. It does not apply that requirement to every chatbot. The broader engineering point is still useful. If an operation affects people or money, build logs that can reconstruct it before a compliance deadline forces a rushed implementation.

Test abuse as a workflow

A prompt-injection test that only asks the model to reveal its system prompt is too narrow.

Test complete attack paths:

malicious document
  ↓
retrieval
  ↓
model proposes forbidden tool
  ↓
tool registry rejects capability
  ↓
authorization denies action
  ↓
alert and trace label

The model may follow the malicious instruction. The system should still stop the effect.

Include cross-tenant retrieval, tool argument manipulation, encoded instructions, delayed attacks stored in memory, approval fatigue, and replayed operation IDs. Verify that the trace contains enough evidence to diagnose each test without storing the sensitive payload unnecessarily.

Governance needs runtime owners

Policies fail when nobody owns the control after launch.

Assign an owner for the workflow, tool registry, data sources, evaluator suite, and incident response. Define who can grant a capability, change an approval threshold, or accept a known risk. Record those changes like code deployments.

NIST's Generative AI Profile treats risk management as lifecycle work. That matches the engineering reality. Model behavior, permissions, data, and attackers change after release.

The next article covers two of the most behavior-changing dependencies: Prompts and models are production dependencies

Start with the work

Bring the workflow that needs attention.

Pick a time for a working session or send a short brief. Either way, we will come prepared to understand where the work gets stuck.

Talk through the work

Book a 20-minute consultation.

Bring the workflow that feels slow or fragile. We will determine whether it is a sensible candidate for an AI system.

Bartosz LuderaBartosz LuderaFounder, Harnessloop

Choose a time for a 20-minute consultation.

Send a workflow brief

Prefer to write it down?

Tell us where work waits, repeats, or falls through the cracks.