Oct 09 2026

The cheapest security control you’ll ever implement is the component you decided not to build

Less Engineering. More Effective AI Agents.

Why the simplest agent harness is also the most secure, most governable, and most auditable one…

Every component you add to an AI agent is a component an attacker can reach, an auditor must test, and your team must explain after an incident. That is the whole argument of this post, and it is why the most effective agents I review are also the plainest.

I see the same pattern in agent design reviews across B2B SaaS and financial services. Engineers who would reject a needless microservice in a web app will happily ship a five-agent pipeline, a vector database, and three MCP servers. The job underneath usually needed one model call inside a loop.

The demo looks impressive. The threat model does not. Each extra agent is another identity. Each extra tool is another path to a real-world effect. Each extra MCP server is another supplier in your chain. Each extra document in the context window is another place an injected instruction can hide.

Thesis: complexity in an agent harness is not a neutral engineering choice. It is attack surface, audit scope, and liability, accrued one convenient decision at a time. The cheapest security control you will ever implement is the component you decided not to build.

The golden rule, read as a security engineer

Faan Rossouw at AionSec recently published Stop Overengineering Your Agents, and it names the discipline cleanly. His golden rule of harness engineering: “Only ever make things as complex as they need to be.”

He builds it out of progressive disclosure, the context-engineering principle that a model should receive only what the current step requires and pull in more on demand. The idea comes from UX design, and engineers know it as lazy loading. Rossouw’s move is to apply the same shape, exactly what the job needs, nothing more, to every harness decision: tools, MCP, skills, orchestration, and RAG. The stopping point differs per layer. The rule is not “use less everywhere”. It is “match the need, then stop”.

He also cites Anthropic’s own guidance in Building Effective AI Agents: start with the simplest workable solution, add complexity only when it earns its place, and accept that the right answer is sometimes no agent at all. When the vendor selling model calls tells you that, believe it.

Rossouw frames the rule as an engineering principle. I want to add the governance reading, because it is the same rule wearing a different badge:

DisciplineName for the same idea
UX designProgressive disclosure
Software engineeringLazy loading, YAGNI
Context engineeringMinimal, just-in-time context
Security architectureLeast privilege, least functionality, attack surface reduction
AI governanceProportionality: controls scaled to intended use and impact
AuditScope minimisation: fewer components, fewer controls to evidence

The security community has preached least privilege for fifty years. Agent builders are rediscovering it from the performance side. Good: when the cost argument and the security argument point the same way, the right design stops needing a champion.

Complexity is attack surface, layer by layer

Every layer of a harness has a “just enough” point. Everything past it is risk you chose to carry with no business return. The table maps each layer to what overbuilding looks like, the threat it invites, and the matching category in the OWASP Top 10 for Agentic Applications (ASI) or the OWASP LLM Top 10.

Harness layerRight-sizedOverbuiltThreat it invitesOWASP mapping
ContextWhat this step needs, pulled just in timeEvery policy, ticket, and email pre-loaded “in case”More untrusted text in the window means more room for indirect prompt injection and sensitive data leakageASI01 Agent Goal Hijack; LLM01, LLM02
ToolsThe few tools this job requires, narrowly scopedA general execute_command or http_request tool “for flexibility”A hijacked agent inherits every capability you granted, whether or not the task needed itASI02 Tool Misuse; ASI05 Unexpected Code Execution
Identity and permissionsPer-task, short-lived credentialsOne broad service account shared across agentsPrivilege escalation and confused-deputy actionsASI03 Identity and Privilege Abuse
MCPA direct API call unless MCP gives a real advantageSeveral third-party MCP servers because “it’s the standard”Tool poisoning, rug-pull updates, and unvetted suppliers in the execution pathASI04 Agentic Supply Chain
SkillsAtomic, single-purpose instructionsSprawling do-everything documentsHidden instructions, unreviewable scope, and drift between what was approved and what runsASI04; ASI01
MemoryScoped, expiring, attributed entriesUnbounded persistent memory shared across sessions and tenantsPoisoning that survives the session, plus cross-tenant leakageASI06 Memory and Context Poisoning
OrchestrationOne agent, unless one demonstrably cannot do itPlanner, critic, router, and worker agents for a linear taskUntrusted messages between agents and errors that compound across hopsASI07 Inter-Agent Communication; ASI08 Cascading Failures
RAGOnly when knowledge is too large to inject directlyA vector store for a 40-page policy setEmbedding leakage, poisoned documents, and weak access control on the indexLLM08 Vector and Embedding Weaknesses
AutonomyHuman approval on consequential actionsFully autonomous loops for convenienceOver-trust of confident output and actions no one would have approvedASI09 Human-Agent Trust; ASI10 Rogue Agents

Read the “Overbuilt” column as a list of findings. I have written most of them in real assessment reports, and none of them was there because the business needed it. They were there because they were easy to add and nobody asked whether the job required them.

One compounding effect deserves its own line. Risks in a harness multiply rather than add. A broad tool is dangerous; a broad tool plus pre-loaded untrusted context is an exploit path; add a second agent that trusts the first agent’s output and the path now crosses a trust boundary nobody drew. Removing any one of those components breaks the chain.

Does this job need a model? The deterministic default is a security control

Rossouw pushes the rule down to the smallest unit: for each step, ask whether it needs a model at all. His answer is that code is the default, and the model earns a step only when the work requires interpretation, meaning ambiguity, judgment, or meaning.

His example is detecting beaconing in network telemetry. Whether a connection is regular enough to be a beacon is arithmetic, so code does it. Whether that regular connection is malicious or a software updater checking in is judgment, so the model does it. Same data, two jobs.

From a security and assurance view, that split matters more than the performance gain:

PropertyDeterministic codeModel step
Prompt injectionNot applicable; input cannot rewrite the logicEvery input is a potential instruction
TestingUnit tests prove behaviourEvaluations estimate behaviour at a pass rate
RepeatabilityIdentical on every runVaries with model version, prompt, and context
Audit evidenceCode, tests, and change recordsPrompts, evals, traces, and drift monitoring
Failure modeUsually loudCan be wrong in ways that look right
Cost per runEffectively zeroPaid on every call

This turns into a design rule I now use in reviews: controls belong in code; judgment belongs in the model. Authorisation checks, rate limits, allow-lists, value thresholds, data-classification gates, and approval routing are deterministic decisions. If your agent “decides” whether it is allowed to send a wire, export a dataset, or delete a record, you have placed a security control inside the one component that can be talked out of it. AionSec makes the same point in a companion piece, Your Prompt Is Not a Security Boundary.

In practice, most of a well-built agent is ordinary software. The model is a small, well-fenced judgment engine inside it.

Run every step through this gate. Only work that needs open-ended judgment reaches an agent, and even then the bounds live in code.

The smallest harness is the most auditable one

In my last post, The Harness Is the Contract, I argued that the useful question is not whether an agent can act. It is whether every consequential path was declared, bounded, and left evidence another party can verify. The golden rule is what makes that standard achievable at all.

Consider the work each property demands, and how it scales with harness size:

  • Declared. For every tool, you must write down its worst case if the agent is fully adversarial. Twelve tools means twelve worst cases, plus their combinations. Three tools is an afternoon. A general execute_command tool has no worst case you can write down, which is exactly the problem.
  • Bounded. Each path needs a deterministic limit outside the model: scope, rate, value, approval. Every extra agent, MCP server, and memory store is another place a bound has to be enforced and tested.
  • Verifiable. Each component must emit evidence a sceptical third party can rely on. A five-agent pipeline produces five interleaved traces and the hand-offs between them. A single agent with three tools produces one trace you can actually read.

Overengineering does not only add risk. It makes your governance claims harder to prove, and controls are rarely why organisations fail audits; evidence is. That held at Virtual Data Room, where we took the AI management system through ISO/IEC 42001 Stage 2 certification on the first attempt. The MCP Governance Standard I wrote there leaned on the same instinct: every integration had to justify itself before it was allowed into the execution path.

Here is how right-sizing lines up with the frameworks your customers and regulators will ask about. Treat it as a crosswalk, not a claim that one control satisfies another.

FrameworkWhere right-sizing shows upEvidence an assessor will ask for
ISO/IEC 42001:2023Clause 6 risk assessment and treatment; Annex A AI system lifecycle (A.6), intended use (A.9), and third-party relationships (A.10)Documented design rationale per component; supplier review for each MCP server and model provider
NIST AI RMF 1.0MAP: context and intended purpose; MANAGE: treat risks by removing capability, not only by adding controlsAgent inventory, tool-to-purpose mapping, risk treatment decisions
EU AI ActArt. 12 record-keeping, Art. 14 human oversight, Art. 15 robustness and cybersecurity, Art. 26 deployer obligationsLogs that reconstruct actions; oversight points on consequential steps
OWASP ASI Top 10ASI02, ASI03, ASI04, ASI08 shrink directly as tools, permissions, suppliers, and agents shrinkThreat model per declared path; test results per ASI category

The practical upshot for a CISO: the right-sizing review is not an engineering nicety you hope happens. It is a documented risk-treatment decision, and it belongs in your AI management system with an owner and a date.

What the rule does not cut: controls are not overengineering

There is a lazy reading of “less engineering” that I want to head off, because I have already heard it in client meetings: “Simplicity first, so we’ll skip the approval step and the logging for now.” That is not the golden rule. That is the golden rule used as an excuse.

The rule says match the need. For any agent that touches customer data, money, production systems, or regulated decisions, the need includes the safety layer. So spend your complexity budget deliberately:

Spend less onSpend enough on
Number of agentsPolicy enforcement outside the model
Number of tools and their breadthHuman approval on consequential, irreversible actions
Pre-loaded contextInput and output validation at trust boundaries
Third-party MCP serversSupplier vetting and version pinning for the ones you keep
Persistent memoryTamper-evident logs that reconstruct what the agent did
Clever orchestration patternsKill switch, rate limits, and spend caps

The left column is capability. The right column is assurance. Capability you do not need is risk; assurance you do need is not overhead. A simple agent with strong guardrails beats a sophisticated agent with weak ones every time, in production and in the audit.

One more caution: “simple” is a property of the system, not of the diagram. A single agent with one run_sql tool against the production database looks minimal on a whiteboard. It is not. Measure simplicity by the size of the blast radius, not by the number of boxes.

The Right-Sizing Review: twelve questions before you ship

Run this at design review and again before each release. Any “no” or “not sure” is either a component to remove or a risk to accept in writing.

Scope

  1. Could this be a single model call, or a workflow with fixed steps, instead of an agent?
  2. For each step, can you say why it needs judgment rather than code?

Context and memory

  1. Does each step receive only the data it needs, pulled at the time it needs it?
  2. Does persistent memory exist for a stated purpose, with scope, expiry, and attribution?

Tools and identity

  1. Can you name the worst case of every tool if the agent is fully hijacked, in one sentence?
  2. Is any tool general-purpose (shell, raw SQL, arbitrary HTTP) where a narrow function would do?
  3. Does the agent run on credentials scoped to this task, rather than a shared service account?

Integrations

  1. For each MCP server, what does it give you that a direct API call would not?
  2. Is every third-party server and model version pinned, reviewed, and owned?

Orchestration

  1. If there is more than one agent, what can the second one do that the first could not, and how is trust between them checked?

Assurance

  1. Are authorisation, limits, and approvals enforced in code outside the model?
  2. Could a sceptical outsider reconstruct what the agent did last Tuesday from your logs alone?

Twelve questions sounds like a lot. On a right-sized agent, it takes under an hour. On an overbuilt one, it takes a week, and that difference is the finding.

A 30-day right-sizing plan for agents already in production

Most teams are not designing from scratch. They are living with agents that grew. Here is how to shrink them safely.

  1. Week 1: Inventory. List every agent, tool, MCP server, memory store, vector index, and credential. Record the owner and the business purpose of each. Anything with no owner is your first finding.
  2. Week 2: Justify. Run the twelve questions against each agent. Tag every component keep, narrow, or remove. Expect a third of tools and most general-purpose tools to land in the last two buckets.
  3. Week 3: Cut and move controls. Remove unused tools and servers. Replace broad tools with narrow functions. Move any authorisation logic that lives in a prompt into code. Scope credentials per task.
  4. Week 4: Prove it. Re-test against the OWASP ASI categories that changed. Record each decision as a risk treatment in your AI management system, with evidence. Set the review to repeat at every major model or tool change.

The usual outcome is an agent that is cheaper to run, faster, more accurate, and has a threat model that fits on one page.

My perspective

The AI industry is about to learn the lesson the cloud industry learned a decade ago: most breaches will not come from exotic attacks on the model. They will come from capability nobody needed, granted to a component nobody reviewed, logged in a way nobody could reconstruct.

Rossouw is right that something about building with AI pulls disciplined engineers toward complexity. I would add that governance teams feel the same pull in reverse: piling on frameworks and policies instead of asking what the system actually does. The answer for both is the same rule. Build what the job needs. Govern what you built. Prove both.

Less engineering is not less rigour. It is rigour applied before the code exists, which is the cheapest place it will ever be.

Right-size your agents with DISC InfoSec

DISC InfoSec helps B2B SaaS and financial services teams build agents that are secure by design and certifiable on evidence. We led Virtual Data Room through ISO/IEC 42001 Stage 2 certification on the first attempt, and we bring that same practitioner lens to your agent harness.

Start wherever you are on the readiness ladder:

  1. Free 20-minute call. Walk through one agent and leave with the three components most likely to be overbuilt.
  2. Gap assessment. ISO 42001 gap assessment or ISO 27001 gap assessment.
  3. 7–10 day Quick-Start. An Agent Right-Sizing Review against OWASP ASI, ISO 42001, and NIST AI RMF, with a prioritised cut list and evidence pack.
  4. Implementation and certification support, or an ongoing vCISO / vCAIO retainer.

Book a call: calendly.com/hd-deurainfosec · Email: info@deurainfosec.com · Phone: (707) 998-5164 · deurainfosec.com

DiscInfosec – CISSP, CISM, ISO/IEC 42001 and 27001 Lead Implementer, is Principal Consultant at DISC InfoSec in Petaluma, California.

Publishing notes

FieldValue
Meta titleLess Engineering, More Effective AI Agents: Why Simple Harnesses Are Secure
Meta descriptionEvery tool, agent, and MCP server you add is attack surface and audit scope. A practitioner’s guide to right-sizing AI agents against OWASP ASI, ISO 42001, and NIST AI RMF.
Primary keywordAI agent security
Secondary keywordsagent harness, overengineering AI agents, least privilege AI agents, MCP security, OWASP agentic top 10, ISO 42001 AI agents
Slug/less-engineering-more-effective-ai-agents
Internal linksThe Harness Is the Contract; The Security Risks Hiding in AI Agent Memory

The cheapest security control you will ever implement is the component you decided not to build. Every tool you give an agent is a capability an attacker inherits. Every extra agent is another identity. Every MCP server is another supplier. Before you add one, ask the question Anthropic and AionSec both ask: does this job need it at all?

Sources

Tags: AI Agents, Less Engineering, Stop Overengineering

Leave a Reply

You must be logged in to post a comment. Login now.