Agent Harness: How We Run 40 AI Agents in Production
An agent harness is everything around the model: the code that decides what the model sees, which tools it may call, when it must stop, and what happens to what it proposes. We built and run an operations platform where around forty AI agents handle bookings, staff schedules, member records, events, invoices and communications. The model is the smallest part of it. The harness is where the engineering went, and it is why a person approves most of the changes these agents make. The product view covers what that feels like for staff; this is the AI agent harness behind it.
00 — what an agent harness is
The agent harness: seven parts around one model
A model on its own answers questions. A production agent needs a harness: a loop that plans and dispatches, a catalogue of agents, guardrails, a tool contract, a way to pause for a person, memory and limits, and a test bench. Each part below is one of those, as it exists in our platform.
01 — the loop
The harness loop: one supervisor graph, three stages
The core of the harness is a single LangGraph state graph. Agents are nodes inside it, not separate subgraphs. We started with per-agent subgraphs and consolidated, because one graph gives one place to interrupt, checkpoint and recover from errors.
A request moves through three stages. Decompose resolves who is asking (visitor, member or staff), ranks the most relevant agents, and asks the model for a typed task plan: sub-tasks, each with an operation verb, a domain and its dependencies. Execute starts each sub-task as soon as its dependencies finish, up to eight at a time; tasks that propose changes run one at a time. Synthesise turns the results into one answer. The model plans; it does not pick the next hop at every step.
02 — the catalogue
In the harness, an agent is a definition, not a subsystem
Each agent is a declarative definition in one registry: its domain, the operations it supports, the slots it extracts, its tools, whether it is read-only, which roles may use it, and its guidelines. The system prompt is generated from that definition, and a capability matrix keyed on (operation, domain) is generated from the whole registry. Routing a sub-task is a lookup in that matrix, with a semantic fallback, rather than another model call.
Of the 41 agents, 21 are query agents that only read and 20 are action agents that can propose changes. Adding a capability means adding or extending a definition; the graph, the gates and the approval flow come for free.
Two lessons came from treating agents as data. Role names written in two styles and compared case-sensitively silently hid 14 of 19 role-gated agents from the ranker: nothing failed, the assistant just got worse. And a write agent that omitted its read-only setting picked up the read-only default, slipping past the read-only gate, while its empty role list let any signed-in user reach it. A test now requires every write agent to declare its roles.
03 — the guardrails
Permissions are checked before an agent runs, not asked of it
Every sub-task passes four checks before its agent is called: a hard read-only gate for contexts that must not write, a role check against the agent's required roles, an entitlement check against the tenant's plan, and a capability gate that confirms the agent actually supports the operation. Anonymous visitors only reach an allowlist of four read-only agents.
The read-only gate has to be a hard check, not an instruction in a prompt: an approval interrupt fires after an agent has run, so by then it is too late to ask nicely. The clearest case is inbound customer email. The assistant drafts replies with full read access to the tenant's data, but runs with read-only forced on, because an email from outside can contain instructions aimed at the model. And the capability gate is rolling out in an observe mode that only logs; enforce is a configuration switch. Routing gaps show up in logs instead of as confident refusals.
04 — the tool contract
Tools declare whether they read, preview or write
Tools are plain functions in a shared library, wrapped per agent as closures that capture the tenant and user. The model never sees those IDs, so it cannot ask for another tenant's data even if prompted to. Each declared tool has one of three kinds.
This contract fixed a real class of bugs. Before it, which tools wrote anything was never declared, so the system sometimes reported a read as a completed action and occasionally wrote twice. Now a guard wraps every tool: it resolves names to IDs before refusing anything, keeps raw library errors out of the conversation, and flags any preview-write that succeeded without producing a preview.
Agents and the REST API also share one operations layer: a catalogue of verbs, each with its required roles and plan entitlement. Features built on it are shared by the dashboards and the agents, so those cannot drift apart. Our working rule: a feature without an operation is invisible to the assistant.
05 — the brake
Human in the loop: an approval is a paused graph, not a pending promise
When an action agent calls a preview-write tool, it records a pending change — old values, new values, the proposal — and the graph calls interrupt(). The interrupt streams to the interface as a server-sent event and the serverless invocation ends. Nothing is holding a connection open while a manager thinks about it.
Resuming is a new invocation that reloads the conversation's checkpoint from the database and continues the graph with the person's decision: approve, reject, revise (feedback goes back to the agent for a new preview), or approve with edits. Approval runs through one central apply function with a handler per operation, not a branch per agent. A compare-and-set moves the change from pending to executing, so two approvals cannot both apply it, and an unanswered change expires after an hour.
06 — memory and limits
A bounded harness: serverless, streaming and capped
Conversation state lives in LangGraph checkpoints in MongoDB, kept for 30 days. Long conversations are summarised past about twenty messages, keeping the latest ten verbatim. The agent service is a container-image Lambda that streams responses over server-sent events through a function URL, with a keep-warm ping. Models run on Bedrock: Claude Sonnet 4.6 by default, with Haiku 4.5 for lighter paths, and each tenant can override the model per agent.
Everything that can run away is capped: at most eight sub-tasks in parallel, four concurrent model calls, per-model rate limiters, two tool rounds for query agents, five loops for action agents and a recursion limit on the graph. Prompt caching covers the stable system prompt and the tool block.
Measure before you optimise: decomposition turned out to be the latency hot spot, at around 70% of end-to-end time — but 89% of that was embedding calls in the agent ranker, not the planning call to the model. Profiling, not intuition, found it.
07 — the test bench
Evaluating the agent harness on every change
There is no free-form long-term memory. Instead, a nightly job (switched on per deployment) can score each run, where an approved preview counts in its favour, and promote the best into a per-tenant set of worked examples that agents retrieve as few-shot guidance. The people approving changes are, in effect, teaching the system what good looks like for their business.
Testing works at two levels. A fake-model harness drives the real agent loops in unit tests. And an evaluation suite in CI scores every pull request that touches the agent code on plan schema, routing accuracy, dependencies, tool choice, whether a write correctly interrupts, preview structure, groundedness and PII, with a link to the experiment posted on the PR.
08 — the takeaway
Build the harness; the model is the easy part
Put the decisions that matter in code — routing, permissions, what counts as a write, how a change is applied — and let the model do what it is good at: understanding the request and preparing the work.
In production, an agent harness is a mix of model-driven agents, plain deterministic services (a nightly job that shares out oversubscribed requests uses ordinary code, no model at all), queue workers and human approval. The model plans and drafts. The graph, the gates, the typed tools and the people decide what actually happens.