Your AI agent can pass a staging benchmark, produce a plausible answer, and still take the wrong action in production. If that action is an email, a CRM update, a payment instruction, or a code change, discovering the mistake is only half the problem. You also need to stop the action path before queued work, retries, or delegated workers make the incident worse.
That requires one operating model for evaluation and runtime control. Evaluation tells you whether an agent is behaving acceptably. Runtime control decides whether it is still allowed to act. Enterprise agents need both, connected by the same identities, traces, policies, and incident evidence.
Treat evaluation and runtime control as one operating loop
Quality and authority are different questions. A strong evaluation score does not grant permanent permission, and revoking permission does not explain why the agent failed.
For every consequential action, your system should be able to answer four questions:
- What did the agent intend to do? Capture the proposed action, target, supporting inputs, and relevant model, prompt, policy, and tool versions.
- Was the proposal acceptable? Run deterministic and semantic evaluations appropriate to the action’s risk.
- Was the agent authorized at execution time? Check current deployment state, permissions, action limits, and human-approval requirements immediately before the side effect.
- What actually happened? Preserve the downstream service’s acceptance, rejection, completion, cancellation, or failure receipt.
This creates an action envelope around each side effect. The envelope should follow the work across queues, retries, tool calls, and delegated agents. A trace that ends when the model returns text is incomplete if a worker later turns that text into an external action.
| Control layer | Question | Useful evidence | Decision it supports |
|---|---|---|---|
| Deterministic evaluation | Does the proposal satisfy rules that can be checked exactly? | Schema validation, required fields, allowed values, thresholds, duplicate checks | Reject malformed or prohibited proposals |
| Semantic evaluation | Is the proposal correct and appropriate in context? | Human labels, calibrated judge scores, cited evidence, error-category results | Proceed, escalate, revise, or reject |
| Runtime authorization | May this deployment perform this action now? | Verified identity, stop state, policy version, approval status, current limits | Admit or deny the external attempt |
| Outcome reconciliation | What effect did the external system accept or complete? | Provider receipt, object version, delivery state, idempotency record | Close, cancel, compensate, or investigate |
The product requirement is not merely that the agent be accurate. It is that every important action be attributable, evaluable, revocable, and reconcilable. If any of those properties is missing, your incident response will depend on inference at exactly the moment you need evidence.
Build evals from observed failures, not benchmark theater
A generic success rate rarely tells you what to fix. Start with a concrete failure that a customer, operator, or reviewer can point to. Then turn that observation into an error category with an explicit test.
A seemingly narrow complaint about an AI-generated opportunity tree ultimately required four new evaluations and 16 experiment variants. The visible defect was a flat branch with too many sibling opportunities. The real work was determining whether the failure was structural, semantic, or produced earlier in the synthesis pipeline.
That is the pattern to apply to your own agent. Do not translate a complaint such as wrong account into improve accuracy. Preserve the failure’s shape: the agent confused entities with similar names, failed to use a disambiguating identifier, and then proposed an update against the wrong record. That statement tells you what data, assertion, and runtime policy you need.
- Write the failure in observable terms. State what the agent saw, decided, and attempted. Avoid diagnosing the prompt before you have isolated the failing stage.
- Locate the earliest detectable point. Ask whether the error first appeared during retrieval, entity resolution, planning, tool selection, argument construction, authorization, or execution.
- Create positive, negative, and near-miss cases. Include the original trace, similar cases that should fail, and difficult cases that must still pass. Near misses expose rules that are too broad.
- Add the cheapest reliable detector first. Use code assertions for exact properties. Reserve model-based judgment for meaning that cannot be reduced to a stable rule.
- Measure the error category separately. Keep high-consequence failures visible instead of hiding them inside one blended pass rate.
- Compare changes against a frozen regression set. A candidate prompt, model, or workflow must reduce the target failure without reopening failures you already controlled.
Use deterministic checks for shape and invariants
Code-based evaluations are appropriate when you can state the requirement without asking for an opinion. Examples include a valid schema, a required customer identifier, a maximum transaction amount, the absence of duplicate recipients, a citation for every extracted claim, or a limit on how many children a tree node may have before regrouping is required.
These checks are fast, repeatable, and easy to audit. They also have a hard boundary: they can confirm that an identifier exists, but not that it belongs to the intended company. Treating structural validity as semantic correctness is a common way for an agent to look reliable while making the wrong decision.
Calibrate model judges before using their scores
A model-based judge is useful when the failure depends on meaning: whether two items should be grouped, whether the evidence supports a conclusion, whether a response omitted a material constraint, or whether an escalation rationale is adequate. A second model call is not ground truth, however. It is another component that can misunderstand the task.
Build a calibration set that humans have reviewed against a written rubric. Run the judge without revealing the expected labels. Inspect false positives and false negatives separately, because the cost of each may differ. Revise the rubric and examples until the disagreements are understood, then freeze part of the set so future tuning cannot quietly redefine success.
Track judge performance by risk slice, not only in aggregate. A judge may perform acceptably on routine drafts while missing the few cases that change customer data. If people cannot agree on the label, mark the case as ambiguous and route comparable production cases to review. Do not force uncertainty into a pass or fail value simply to make the dashboard cleaner.
When you compare experiment variants, change one meaningful component at a time where practical. Record the model, prompt, retrieval configuration, tools, policy, and judge version. The winning candidate is not the one with the most attractive average. It is the one that meets the release gates for every critical category and makes an explicit tradeoff on the remaining errors.
Put the stop decision outside the agent and its workers
Stopping the visible worker is process control, not system control. A queued task may be waiting elsewhere. A delegated worker may have its own lifecycle. A retry may run after the original process exits. An external service may already have accepted a request.
The enforceable path should look like this:
Agent request → queue → worker → action gateway → external service
A separate control service records whether the deployment may act. The action gateway owns the downstream credentials and checks that control state immediately before admitting an outbound attempt. Workers can request an action, but they cannot bypass the gateway, mint a new identity, edit the stop record, or retrieve the gateway’s credentials.
This separation matters because task cancellation is cooperative in many runtimes. Temporal Activities, for example, can receive cancellation through heartbeats while the task implementation may accept or ignore it. Cancellation remains useful for releasing resources and ending unnecessary computation. It should not be the security boundary that protects an external system.
Your gateway needs a minimum control state that operators and services interpret consistently:
- Running: eligible actions may be admitted under the current policy.
- Paused: new side effects are denied while operators investigate; queued work may remain available for inspection.
- Revoked: new side effects are denied and the deployment must not resume without a deliberate reauthorization event.
The names can differ. The important part is that the state is durable, versioned, auditable, and evaluated outside the component you are trying to constrain. A stop flag stored only in worker memory disappears on restart. A flag that is cached longer than your stop objective cannot meet that objective.
Preserve identity through every handoff
Every unit of work should retain a verified deployment identity, run identity, parent-run identity, and action identity. Derive those values from trusted authentication and queue metadata. Do not trust an agent-authored field that says which deployment it belongs to.
A retry must request authorization again. Approval for one attempt should not become a reusable permission for later attempts. A child agent or restarted worker must inherit the constrained deployment identity rather than acquiring the default permissions of its new host process.
Inventory indirect routes, not just direct credentials
A worker may lack direct access to an email provider and still reach it through a shared connector. That connector is acting on the worker’s behalf. It must preserve the originating deployment identity and apply the same stop decision instead of substituting its own broader service identity.
Build the inventory from effects backward. For each external effect, identify every API, shared service, connector, scheduled job, database credential, administrative endpoint, and manual fallback that can produce it. If a worker can write directly to the case database, placing only the official case API behind a gateway does not create a real control boundary.
For consequential workflows, separate read, draft, approve, and commit permissions. An agent investigating an incident may need to read a case and draft a correction without retaining permission to commit the change. This lets you preserve useful diagnostic work while stopping irreversible or externally visible actions.
Define a stop guarantee, then try to break it
Stop the agent is too ambiguous to test. Convert it into a contract that names the affected deployment, the blocked action classes, the clock’s starting event, the propagation deadline, the treatment of already accepted work, and the behavior of unaffected deployments.
An illustrative bank scenario uses this contract shape: within 30 seconds of durably saving a stop request, every action gateway must refuse new email and case-update admissions from the deployment, including queued and background tasks. The 30-second value is illustrative, and the proposed architecture was not validated through a production experiment. Set your own objective from the consequence of the action and the measured behavior of your infrastructure; do not copy that number as a universal standard.
Start the clock at a machine-verifiable event, such as the durable control-service write. Detection time, operator decision time, and the time needed to reach the control service are separate intervals. Measuring them separately prevents a fast gateway from masking a slow incident response.
Also distinguish admission from completion:
- Not yet admitted: deny the attempt when it reaches the gateway, even if it was queued before the stop.
- Admitted but not completed: request cancellation only if the downstream provider supports it, then record whether cancellation succeeded.
- Already completed: do not describe the action as stopped. Use an approved compensating workflow, such as restoring the prior record version or contacting the recipient.
- Status unknown: reconcile using the provider’s receipt or idempotency record before retrying. A blind retry can duplicate the original effect.
An email accepted before the stop may still be delivered. A database write already committed will not disappear because a worker was terminated. Your operator interface should show this residual work plainly so a green stopped indicator is not mistaken for proof that nothing else can happen.
Run the failure exercise before production needs it
- Select one real side-effect path. Use a safe test tenant and a provider configuration that cannot affect customers, money, production data, or external recipients.
- Create work in every lifecycle state. Include a long-queued task, a delegated child, a scheduled retry, an in-flight gateway request, and an action the provider has already accepted.
- Request the stop. Record the durable-write timestamp and control-state version.
- Attack every route. Attempt the action through the normal worker, a restarted worker, a background worker, a shared connector, and any direct credential path found in the inventory.
- Verify enforcement. Confirm that each gateway denies new admissions within your stated objective and records the deployment, action, policy, state version, decision, and timestamp.
- Reconcile accepted work. Compare gateway records with downstream receipts. Identify actions that completed, were cancelled, remain pending, or have an unknown state.
- Restart the infrastructure and repeat. A restart must not clear the stop, refresh old work into a new identity, or reopen an indirect route.
Repeat the exercise when you add a tool, connector, queue, credential, or delegation mechanism. Those are changes to the control surface, even when the model and prompt stay the same.
Your operational dashboard should expose at least four measures: the proportion of known side-effect paths protected by a gateway, stop-propagation time, attempted actions denied after a stop, and external receipts that cannot be matched to an authorized action. Set alert levels from the consequences of each workflow. A customer-support draft and a payment instruction should not inherit the same failure policy by convenience.
Decide explicitly what happens when authorization infrastructure is unavailable. Denying writes protects integrity but can interrupt the business process; allowing them preserves availability but may defeat the control during an incident. For actions involving financial, legal, security, or regulatory exposure, document that choice with the accountable security, compliance, and business owners before launch.
Key takeaways
- Evaluate proposals and authorize actions as separate decisions. A quality score is not a permission token.
- Turn each observed failure into a named error category, a representative dataset, and the cheapest reliable detector.
- Use deterministic assertions for exact invariants and calibrated model judges for semantic questions. Do not treat judge output as ground truth.
- Keep downstream credentials in an action gateway that checks current control state immediately before every attempt.
- Carry verified deployment identity through queues, retries, restarts, connectors, and delegated agents.
- Define stopping in terms of denied admissions, not exited processes, and reconcile actions that an external service accepted before the stop.
Before your next agent release, choose one consequential action and run the full stop exercise. Queue it, delegate it, revoke the deployment, restart the worker, and let the work reach the gateway. If any route still succeeds, you do not yet have a system-level kill switch. You have found the next control to build and the next regression case to keep.
References
- Product Talk – 4 New Evals and 16 Experiment Variants to Fix 1 Customer Complaint
- The AI Runtime – How to Stop an AI Agent: AI System Design








