Your next customer may not visit your homepage, compare feature grids, or complete onboarding screen by screen. An agent may evaluate your product, create an account, consume a trial, call a tool, and initiate a purchase on the customer’s behalf. Inside your company, another set of agents may be building prototypes, qualifying demand, updating records, and coordinating operations.
That changes the product leadership question. You are no longer deciding where to add a chatbot. You are deciding whether nonhuman actors can complete valuable work without gaining ambiguous authority, consuming unbounded resources, or operating from stale business context. An agent-driven software business needs an operating model for action, trust, memory, and economics.
Key takeaways
- Design around a complete customer or business task, not an isolated AI interaction.
- Let agents handle interpretation and planning, but keep identity, permissions, data access, spending, and irreversible changes deterministic.
- Move abuse controls ahead of expensive work. Payment fraud detection is too late when the loss is compute already consumed.
- Store current state, decisions, and action history outside individual agent sessions.
- Measure the margin and quality of completed work. Token counts describe consumption, not value.
Start with an operating model, not an AI feature
A business becomes agent-driven when agents can move work from an accepted request to a verifiable outcome. Generating an answer is not enough. The agent needs a defined path to inspect state, select an action, use an approved tool, observe the result, and either finish or escalate.
This model has two sides. On the demand side, customer agents may discover, evaluate, buy, configure, and use your product. On the supply side, internal agents may build, sell, support, and operate it. The same control plane should govern both: identity, policy, budgets, data access, audit history, and revocation.
Do not begin by asking which departments should have agents. Choose one job whose successful completion is visible and whose failure is recoverable. A prototype-generation workflow is easier to govern than an agent that can deploy production code. A support agent that drafts a resolution is safer than one that can issue unrestricted refunds. Early autonomy should stop before the point where an error becomes expensive, external, or difficult to reverse.
Write an operating contract for the workflow before selecting a model. It should answer these questions:
- Trigger: What event is allowed to start the work?
- Inputs: Which records, instructions, and business facts may the agent read?
- Authority: Which tools may it call, what may it change, and what may it spend or disclose?
- Success: What observable state proves that the job is complete?
- Escalation: Which uncertainty, risk signal, or failed check transfers the work to a person?
- Evidence: What receipt, event, or evaluation must be stored after every consequential action?
- Memory: Which results become durable business state, and which remain temporary working context?
This contract forces an important distinction: autonomy is not one permission. An agent can be allowed to plan while being prohibited from executing. It can be allowed to execute a reversible change while being blocked from publishing it. It can have access to a budget without having permission to share confidential information. If those distinctions exist only in a prompt, they are preferences. If they exist in your identity, policy, and tool layers, they are controls.
Keep judgment agentic and controls deterministic
Language models are useful where the work is ambiguous: understanding intent, synthesizing context, comparing options, and constructing a plan. They are a poor place to enforce rules that must hold every time. Authentication, authorization, spending limits, tenant isolation, schema validation, and irreversible state changes should not depend on whether a model remembers an instruction.
| Concern | Safer default owner | What must be observable |
|---|---|---|
| Interpret a request | Agent | Intent, assumptions, and confidence |
| Choose a plan | Agent within policy | Selected steps and permitted tools |
| Authenticate a user or service | Deterministic identity system | Principal, tenant, and credential scope |
| Approve spending or data release | Deterministic policy engine | Rule evaluated, limit, and approver |
| Change business state | Typed, permissioned tool | Before state, requested change, and resulting state |
| Judge task quality | Evaluator plus human review where consequences warrant it | Acceptance criteria, evidence, and exception path |
Aha!’s app-building implementation illustrates the boundary. It kept authentication, SSO, database access, and email in deterministic components while separating AI generation into design-system, prototype, and backend phases. Its first containerized Ruby on Rails approach was also replaced with a single-instance, multi-tenant architecture using V8 isolates after proving too wasteful at scale. The broader lesson is not to copy that stack. It is to avoid asking a model to regenerate infrastructure and controls that your platform can supply reliably.
The same principle applies to the customer-facing surface. A conventional interface may become unexpectedly valuable when agents can operate it. A seven-year-old Stripe command-line interface saw usage rise sharply when agents discovered it. Agent readiness often comes from stable, composable actions rather than a new conversational screen.
For every action you expose, provide:
- A machine-readable name, purpose, input schema, and output schema.
- Explicit preconditions and permission requirements.
- An idempotency mechanism so retries do not duplicate payments, messages, records, or deployments.
- A preview or dry-run mode for consequential changes.
- Scoped credentials rather than a shared all-powerful token.
- Structured errors that tell the agent whether to correct, retry, wait, or escalate.
- A durable receipt containing the actor, policy decision, tool version, requested action, result, and cost.
- A compensating action or rollback route wherever the underlying change is reversible.
An API alone does not make a product agent-ready. The agent also needs to discover the action, understand its contract, obtain appropriately scoped authority, and determine whether the result satisfies the original request. Test that full path in a sandbox. A successful tool call with an unsuccessful customer outcome is still a failed workflow.
Trust and abuse determine how much demand you can serve
Traditional payment fraud systems evaluate a transaction. An agent-driven product may incur its loss much earlier. Cursor encountered accounts that could consume free credits, exhaust trials, or accumulate usage they would never pay for. By the time a conventional payment control had a transaction to inspect, the compute had already been spent.
This makes fraud management part of product design. If you add friction to every user, you suppress legitimate self-service adoption. If you allow expensive work before establishing trust, abusive accounts can turn growth into loss. Your trust system therefore defines which customers you can afford to serve and how quickly they can reach value.
Use a progressive trust ladder instead of a single approved-or-blocked decision:
- Unverified: Permit discovery, read-only access, sample data, and tightly bounded low-cost execution.
- Verified: Permit limited writes and paid usage within explicit account, task, and tool limits.
- Established: Increase concurrency or asynchronous capacity after successful payment and usage history.
- Privileged: Require a stronger role, explicit approval, or both for sensitive data, external commitments, and changes that are difficult to reverse.
The exact thresholds depend on your unit economics and risk. The structure should not. Every increase in autonomy needs a reason, a bounded scope, and a revocation path.
Keep three decisions separate:
- Budget authority: How much money or compute may this principal consume?
- Action authority: Is this principal allowed to perform this task on this object?
- Information authority: May the principal read or disclose the data needed to do it?
A wallet settles only the first question. Spending permission, task quality, and permission to share information are different judgments. If your system treats a valid payment method as universal authorization, a legitimate payer can still trigger an impermissible action or expose information. Put policy checks at the tool boundary, before the action and before the expensive model or infrastructure call where possible.
A practical control stack reserves budget before execution, limits cost at the account and task levels, rate-limits expensive operations, verifies identity progressively, and stops a run when its authorization changes. It should also separate a pause from a failure. A paused job can wait for approval without discarding its evidence or repeating completed actions.
Monitor the funnel as a distribution system, not merely a security queue. Track how many legitimate users are delayed or blocked, which controls prevent pre-payment loss, and how trust progression affects activation. A rule that catches abuse but makes the product unusable for good agents is not automatically a good rule.
Shared memory must know when it is no longer true
A stateless agent can complete a narrow task. A business process needs continuity across tasks, agents, people, and time. That continuity cannot live in a chat transcript or in an agent’s summary of its previous session. Those artifacts preserve language, but they do not reliably establish which fact is current.
The failure mode is subtle because the agent may reason correctly from obsolete context. In one self-documented agent-run consulting firm, five or six concurrent sessions drifted after positioning and customer-profile decisions changed. Older memory still carried retired language, a shared index rewrite dropped another session’s entry, and a dormant session later resumed from a profile that had been superseded twice. These were context and state failures rather than failures to reason from the supplied information.
Treat the model’s context window as a temporary workbench. Keep business memory in an external continuity layer with at least five distinct records:
- Canonical state: Current facts, their owners, versions, effective dates, and review or expiry conditions.
- Decision log: What was decided, why, by whom, what evidence supported it, and which condition should reopen it.
- Dependency map: Which offers, prompts, policies, documents, evaluations, or workflows depend on a fact or decision.
- Run packet: The current, task-specific context assembled when a job starts. Generate this from canonical state instead of trusting a prior session summary.
- Action ledger: Tool calls, policy decisions, costs, resulting state, quality evidence, and human interventions.
The safe write path is a state transition, not a free-form memory update:
- Read the current version and relevant dependencies.
- Propose a structured change with its reason and evidence.
- Validate permissions, schema, and the expected prior version.
- Commit the change atomically or reject it if the state has moved.
- Invalidate or queue updates for dependent artifacts.
- Append the action and result to the ledger.
Do not let an agent silently replace canonical memory because its new wording appears more complete. Let it propose the update. A deterministic service should check version, ownership, and permissions before committing it. This prevents one agent’s apparently harmless rewrite from erasing another workflow’s state.
Freshness also needs to be explicit. A customer segment, product position, policy, or price can be true when written and wrong later. Attach a review condition or expiry rule to facts that can age. When a decision changes, mark the previous state as historical and identify what must be regenerated. Retrieval quality cannot rescue a system that retrieves stale information with high confidence.
Price completed work and prove its economics
Agent usage produces plenty of measurable activity: tokens, calls, runs, steps, latency, and tool invocations. None of those metrics tells you whether the business created value. A useful economic model connects what a task costs with what the task earns. Counting tokens accurately is necessary for cost allocation, but it does not establish that the money was well spent.
Create a task-level economic ledger. For each completed or abandoned workflow, capture:
- The workflow, customer, trust tier, and model or route used.
- The intended outcome and the evidence that it was accepted, rejected, or reversed.
- Revenue attributable to the task or a defined internal value proxy.
- Model, tool, infrastructure, and third-party costs.
- Human review, correction, support, and rework.
- Abuse, refund, credit, and nonpayment loss.
- Latency, reliability, and the point where the workflow escalated or stopped.
The resulting contribution view is straightforward: task revenue or economic value, minus execution cost, human handling, and expected loss. Segment it by workflow rather than looking only at a company-wide average. A profitable support-resolution workflow can otherwise hide an unprofitable research or generation workflow, and a small abusive cohort can distort the economics of a healthy use case.
Keep three pricing concepts distinct. The cost metric tells you what the work consumes. The value metric tells you what outcome matters to the customer. The risk metric sets the limits under which you are willing to deliver it. Your external price may be a subscription, bounded consumption, a completed workflow, or an outcome-based charge, but it should not be a disguised copy of an internal token bill unless token consumption is genuinely the value the customer is buying.
Customer-side agents need similar discipline. Give them a maximum authorized spend and an approval rule, but do not expose the buyer’s maximum willingness to pay as an ordinary field in a seller-controlled workflow. A buying agent can communicate requirements, constraints, and the price it will currently accept without revealing the confidential ceiling that would weaken its negotiating position.
Before expanding an agent’s scope, run the workflow through this release checklist:
- Can the agent complete the full task in a sandbox, not merely call an individual tool?
- Does every consequential action have a scoped identity and an explicit policy decision?
- Will a retry duplicate an effect, or is the action idempotent?
- Can you revoke authority or stop spending while the workflow is running?
- Can you reconstruct what the agent knew, decided, called, changed, and cost?
- Does stale or conflicting state stop the workflow instead of being treated as current?
- Can an untrusted account consume expensive resources before payment or verification?
- Is there a defined human gate for sensitive data, external commitments, or irreversible changes?
- Can you connect the outcome to revenue, customer value, or a named internal result?
If you own this transition, choose one closed-loop workflow now. Write its operating contract, put deterministic controls around its tools, give it governed memory, and instrument its economics before adding more autonomy. You have an agent-driven business when an agent can finish valuable work and you can explain why each action was allowed, what it changed, what it cost, and how the company would contain or reverse a mistake.
References
- Nate Jones’s Substack – Stripe’s head of AI on the question every software business is about to face: can an agent buy from you?
- Product Talk – Creating Aha! Builder: Concept to Code, No Engineers Required
- Substack AI Topic – Agentic AI Case Study: A Firm’s Operations Run by Agents








