Your enterprise agent can sound certain and still be wrong for the business. It may have the policy but not the customer’s current product state, the account but not the user’s permissions, or the request but not the consequence of acting on it.
If you are deciding whether an agent is ready for production, stop treating context as text placed around a prompt. Design it as a governed product surface: the evidence the agent receives, the meaning attached to that evidence, the authority it can exercise, and the conditions under which it must stop.
Design context around the decision, not the model
A prompt tells an agent how to approach a task. Context tells it what is true for this case. A tool contract defines what the agent can change. Combining those concerns in a large prompt makes failures hard to locate: you cannot tell whether the model misunderstood an instruction, lacked a fact, or exceeded its authority.
Start with the decision the agent must make. “Help with this customer” is not a decision. “Determine whether this account is eligible for the requested feature and either draft the enablement action or explain what blocks it” is. The second formulation tells you which entities, policies, product signals, permissions, and outputs belong in the context package.
A useful context package answers these questions:
- Objective: What exact decision or action is expected, and what counts as completion?
- Actor: Who initiated the task, which tenant do they belong to, and what are they allowed to request?
- Subject: Which customer, account, request, product object, or workflow instance is being evaluated?
- Current state: What is true now in the systems that own the relevant records?
- Product behavior: What did the user actually encounter, attempt, complete, or fail to complete?
- Governing rules: Which policy, entitlement, routing rule, or approval requirement applies?
- Evidence quality: Where did each fact come from, when was it observed, and was it transformed?
- Action boundary: May the agent read, recommend, draft, approve, execute, or reverse the proposed change?
Customer profile data is not a substitute for product behavior. An account record can tell an agent what a customer bought while revealing nothing about what the customer experienced. Teams often omit this product context, leaving the agent to infer behavior from plan attributes, support language, or static documentation.
Do not respond by attaching the customer’s complete event history. Retrieve the behavior that can change the decision. For an adoption workflow, that may mean whether the feature was available, whether setup was started, which prerequisite remained incomplete, and whether the latest attempt succeeded. The context should contain decision evidence, not every fact the organization happens to possess.
Preserve the distinction between an observation, a derived value, and an inference. A recorded entitlement is an observation. A status calculated from documented rules is derived. A claim that the customer is “unlikely to adopt” is an inference. If those appear as equivalent fields, the agent can turn a prediction into a fact without noticing.
Absence also needs semantics. “No activation event returned” does not necessarily mean “the user never activated.” The event may be delayed, the query may cover the wrong identity, or instrumentation may not exist for that path. Encode unknown, unavailable, not applicable, and confirmed false as different states when they lead to different decisions.
Write a context contract before adding integrations
A context contract is a short specification for the facts an agent may receive and the rules governing those facts. It prevents a common implementation mistake: connecting every available system before anyone has agreed on what the agent actually needs.
Define the following for every decision-critical field:
- Business meaning: Describe what the field represents in language the workflow owner can verify.
- System of record: Name the authoritative system rather than accepting whichever connector responds first.
- Identity binding: Specify the tenant, account, user, and object keys needed to resolve the field safely.
- Allowed states: Include unknown, unavailable, conflicting, and not applicable where those states are possible.
- Freshness rule: State when the value becomes too old for this decision and what the agent must do then.
- Precedence rule: Define which record wins when systems disagree, or require escalation when neither can safely win.
- Sensitivity: Record who may retrieve the field, which purposes permit its use, and what must be redacted.
- Transformation: Document any filtering, aggregation, classification, or summarization applied before the value reaches the agent.
- Failure behavior: Decide whether the agent may proceed, ask for information, retry, propose an action, or stop.
The contract should also define the agent’s action budget. “Can update CRM” is too broad. The useful questions are which objects it may update, which fields are writable, what preconditions must hold, whether approval is required, and how the system detects that the intended change already happened.
For request triage, the decision might be: choose the correct queue and owner, or state why the request cannot yet be routed. Required context could include request type, required-field completeness, requester identity, current ownership, and routing policy. The agent might be allowed to add an internal classification and draft an assignment while being prohibited from closing the request or notifying an external customer.
This level of specificity changes architecture discussions. Instead of asking whether the agent needs access to the CRM, analytics platform, ticketing system, and knowledge base, you can ask which contract fields each system supplies and which decision would fail without them. If nobody can explain how a field could alter the output, leave it out until a real use case establishes the need.
Review the contract with product, operations, data, security, and the owner of the affected workflow. They do not need to agree on model behavior in the abstract. They do need to agree on authoritative records, conflict handling, access scope, approval boundaries, and stopping conditions. Those are business decisions disguised as implementation details.
Build a runtime that preserves identity, freshness, and authority
Resolve identity before retrieving meaning
Context assembly should follow a predictable sequence. Resolve the actor and tenant first. Resolve the subject through stable identifiers next. Only then retrieve records, product behavior, policy, and available actions. Similar names, email addresses, or semantic matches should not be allowed to establish tenant boundaries.
- Authenticate the actor and bind the request to a tenant.
- Resolve the customer, account, user, and workflow objects through approved identifiers.
- Fetch decision-critical state from the systems named in the context contract.
- Select the applicable policy or rule version and preserve its effective status.
- Filter tools and actions against the actor’s permissions and the agent’s action budget.
- Assemble a compact context manifest containing facts, provenance, timestamps, conflicts, and missing fields.
- Execute only the permitted workflow stage and record the resulting state.
Apply access filters before data reaches the model. A prompt instruction such as “do not reveal information from other accounts” is not tenant isolation. If an unauthorized record is included in model context, the retrieval layer has already crossed the access boundary, even if the final answer happens not to disclose it.
Conversation memory deserves the same treatment as any other data source. A previous turn may contain stale policy, an entity resolved under different permissions, or an inference that was never verified. Revalidate remembered identifiers and decision-critical facts instead of carrying them forward as trusted truth.
Make freshness and conflict visible to the agent
Freshness is not solved by retrieving documents again. The runtime needs to know whether a value is still valid for the decision. A policy without a version or effective status should not silently be treated as current. A cached entitlement should not override a newer authoritative record merely because it was easier to retrieve.
When systems disagree, do not flatten the records into a convenient summary. Pass the conflict, the provenance of each value, and the applicable precedence rule. If no approved precedence exists, the correct result is an escalation that names the conflict. Guessing is not graceful degradation.
Context compression must preserve decision boundaries. A summary should retain exceptions, unresolved contradictions, material state changes, and the origin of important claims. It can remove repetitive history and unrelated detail. It should not erase the reason an account is ineligible, the fact that an approval expired, or the distinction between customer-provided information and system-verified state.
Put human review at the consequential boundary
Human review is most valuable where uncertainty meets consequence. Read-only retrieval and evidence normalization usually need different controls from sending a customer message, changing an entitlement, committing funds, closing a case, or deleting a record.
A marketing operations team using Asana AI Teammates and Claude reduced a three-person request-triage workflow to a single QA review. The useful design lesson is not that review disappeared. The agent absorbed evidence gathering and routing work, while a person remained at a defined quality gate.
Do not make the reviewer reconstruct the case. An approval request should include the proposed decision, evidence used, missing or conflicting information, governing rule, intended action, and expected resulting state. Otherwise the human becomes the retrieval system, and the workflow keeps most of its original latency.
Writes need safeguards beyond approval. Require a current-state precondition before mutation, give the operation a stable identifier, record the result, and check the system state before retrying after a timeout. This prevents an uncertain response from turning into a duplicate side effect.
Evaluate the context system by breaking its assumptions
A polished answer on a complete case tells you little about production readiness. Context evaluations should deliberately remove, stale, contradict, misbind, and contaminate the evidence. The expected behavior is often to stop or escalate, not to produce the most fluent answer possible.
| Evaluation condition | Expected behavior | Failure it exposes |
|---|---|---|
| A required record is missing | Name the missing evidence and ask, retry, or stop according to the contract | Invented facts or unjustified completion |
| An obsolete policy is supplied | Reject it, retrieve the applicable version, or escalate if currency cannot be established | Stale-context acceptance |
| Authoritative systems disagree | Apply the approved precedence rule or surface the conflict without hiding it | Silent conflict resolution |
| A record belongs to another tenant | Exclude it before model processing and log the boundary violation | Identity misbinding or data leakage |
| Retrieved content contains an instruction to the agent | Treat the instruction as untrusted data unless it came through an authorized instruction channel | Authority confusion or prompt injection |
| The requested action exceeds the action budget | Refuse the write, offer a permitted proposal, or request the required approval | Authorization overreach |
| A tool times out after a write may have occurred | Inspect current state before retrying and preserve an auditable operation record | Duplicate or inconsistent side effects |
Add counterfactual tests alongside failure tests. Change a decision-critical fact and confirm that the outcome changes in the expected direction. Remove irrelevant context and confirm that the outcome remains stable. Update the governing policy and confirm that the old conclusion is not preserved merely because it appeared earlier in the conversation.
Measure the context system separately from the model’s prose. Useful dimensions include required-evidence coverage, provenance coverage, policy conformance, tenant isolation, action authorization, escalation quality, and consistency between the proposed action and resulting system state. Permission and policy failures should be release blockers; a favorable average cannot compensate for an agent crossing an access boundary.
Keep a context manifest for every consequential run. It should identify the actor, tenant, subject, record versions, retrieval times, applied rules, detected conflicts, permitted actions, tool calls, approvals, and final outcome. Govern the manifest under the same retention and access policies as the underlying data. Observability should not become an unreviewed copy of sensitive customer information.
Key takeaways
- Define the business decision before choosing context fields, integrations, or models.
- Keep instructions, case evidence, and action authority separate so failures can be diagnosed.
- Contract every critical field: meaning, source of record, identity binding, freshness, precedence, sensitivity, and failure behavior.
- Retrieve the smallest defensible evidence package, including relevant product behavior rather than relying only on account profiles and documentation.
- Enforce tenant and permission boundaries before information reaches the model.
- Place human approval at consequential actions and give reviewers the same evidence the agent used.
- Test missing, stale, conflicting, cross-tenant, and hostile context before enabling production writes.
Choose a bounded enterprise decision where the required evidence and permissible actions can be stated plainly. Write its context contract, build the context manifest, and run the failure cases before enabling writes. If the team cannot agree on provenance, freshness, precedence, and stopping conditions, the agent is not ready to act. A different model will not settle those decisions for you.
References
- Amplitude — I was the bottleneck
- Pendo — 3 types of context every AI agent needs, and the one everyone skips








