If your board asks how many AI agents the company can deploy, do not answer with an agent count. Ten agents attached to well-designed responsibilities may create more useful capacity than a thousand instances waiting for prompts, repeating failed work, or feeding a human review queue.
The decision you need to make is more precise: how many accepted outcomes can your agent system produce, at what total cost, with how much human involvement? Answer that question and you can budget, route models, redesign workflows, and scale without mistaking activity for operating leverage.
The constraint is usable work, not the number of agents
The compute buildout may eventually support hundreds of millions of AI agents, potentially more than the United States has workers. That is a supply estimate, not a demand forecast. It tells you little about how much economically valuable work your company can assign.
An agent is execution capacity attached to a workflow. That capacity becomes useful only when five things are available at the same time:
- Eligible demand: a queue of real work that meets clear entry criteria.
- Current context: the facts, decisions, constraints, and exceptions needed to do the work correctly.
- Permitted actions: access to the required tools, with authority limited to what the responsibility actually needs.
- A verifiable result: an output whose correctness or acceptability can be checked.
- An exception owner: a person or team responsible when the agent cannot proceed safely or confidently.
A useful planning model is:
Usable capacity = the smallest of eligible demand, context-ready work, permitted-action capacity, execution capacity, and verification capacity.
This is not an accounting identity. It is a constraint test. If reviewers can approve only a fraction of the output, buying more inference merely creates a larger review backlog. If the context is stale, additional concurrency produces mistakes faster. If permissions are missing, agents stop at the same handoff no matter how capable the model is.
The practical unit of deployment is therefore a responsibility, not an agent. A task is a single instruction. A responsibility has a trigger, a continuing objective, boundaries, and an escalation path. Early examples include an agent noticing a recurring calendar problem or revising downstream launch material when product scope changes. The important capability is not merely drafting a document. It is recognizing that a changed fact has created work across connected artifacts.
Before funding an agent, write a responsibility card with these fields:
- Trigger: What event or schedule starts the work?
- Inputs: Which systems and records are authoritative?
- Output: What concrete artifact, decision package, or system state must exist at the end?
- Definition of done: What must be true for the result to count?
- Allowed actions: What may the agent read, draft, change, send, or approve?
- Approval boundary: Which actions require a human before they affect a customer, commitment, account, or payment?
- Exception path: Who receives incomplete, conflicting, or high-risk cases?
- Expected demand: How often does the trigger occur, and is there enough recurring work to justify setup?
If you cannot fill in those fields, you do not yet have deployable demand. You have a promising demonstration looking for a job.
Measure capacity in accepted outcomes
Agent dashboards tend to emphasize calls, tokens, runtime, and tasks completed. Those are operating signals, but none proves that useful work reached the business. Your primary capacity measure should be the number of outcomes an accountable owner accepted during a defined operating period.
| Measure | How to define it | What it tells you |
|---|---|---|
| Eligible demand | Work items that satisfy the responsibility’s entry criteria | Whether more agent capacity would have anything valid to process |
| Initiated work | Eligible items the agent actually started | Whether execution, permissions, or scheduling constrained throughput |
| Autonomous completions | Items finished without human rescue during execution | How much of the workflow the agent can carry independently |
| First-pass accepted outcomes | Results approved without correction or rerun | The cleanest measure of effective output quality |
| Accepted outcomes after rework | Results approved only after correction, another attempt, or human completion | How much apparent throughput depends on hidden recovery work |
| Human touch time | Setup, context preparation, review, correction, and exception minutes | Whether automation is actually releasing human capacity |
| Post-acceptance defects | Problems discovered after a result was approved or acted upon | Whether the acceptance process is giving false confidence |
Two equations keep the dashboard honest:
Accepted throughput = accepted outcomes per operating period.
Human capacity released = baseline human time for the accepted work minus current setup, review, rework, and exception time.
The second calculation matters because a fast agent can still consume more human capacity than it releases. That happens when people reconstruct context before every run, inspect every detail afterward, or repair outputs outside the measured workflow. Those minutes belong in the denominator even when another department absorbs them.
Do not immediately translate released minutes into payroll savings. Time has economic value only when you can identify what happens to it. It may remove contractor spend, avoid an additional hire, shorten a revenue-producing process, or return capacity to a named backlog. If nothing changes after the time is released, record it as available capacity rather than realized cash.
Classify every unsuccessful item as well. A compact failure taxonomy makes investment decisions much easier:
- missing or stale context;
- ambiguous instructions or acceptance criteria;
- tool or integration failure;
- permission or policy block;
- reasoning or generation error;
- review rejection;
- no owner available for the exception.
Each category points to a different remedy. A better model may help with reasoning errors. It will not repair missing permissions, unclear ownership, or a source system nobody keeps current. Treating every failure as a model problem is one of the fastest ways to overspend.
Context preparation deserves its own measurement. When one person repeatedly explains project history, corrects assumptions, and supplies the exception that invalidates the obvious answer, that person’s time is part of the agent’s operating cost. Shared workspaces can let decisions, corrections, and contributions from multiple colleagues persist with the work. That can reduce repeated explanation, but only if someone owns freshness, access, and removal of obsolete information.
Calculate cost per accepted outcome, not cost per token
Model pricing matters, especially when a responsibility runs repeatedly. It is still only one line in the economics. The correct unit cost includes everything required to turn an eligible work item into an accepted result.
All-in cost per accepted outcome = total model, tool, infrastructure, context, review, rework, exception, and allocated fixed costs divided by accepted outcomes.
- Model execution: input, output, reasoning, retries, and any calls made by substeps.
- Tool execution: browsers, cloud computers, search, code environments, third-party APIs, and transaction fees.
- Context operations: retrieval, indexing, synchronization, storage, and human preparation of missing information.
- Human control: review, approval, escalation, correction, and incident handling.
- Fixed enablement: platform subscriptions, integration work, evaluations, security review, governance, and ongoing maintenance.
Use accepted outcomes as the divisor. Dividing spend by attempts rewards systems that retry frequently or produce outputs nobody uses.
Model selection can materially change this equation. OpenAI reports that GPT-6.1 Sol approaches Astra on selected evaluations at one-fifth of Astra’s standard input and output token prices. That is a vendor-reported result on selected evaluations, not evidence that the models are interchangeable on your workflow. The economic question is whether the lower-cost route meets your acceptance standard after retries, review, and failures are included.
A premium model can be cheaper per accepted outcome when it avoids repeated attempts or expensive human correction. A lower-cost model can be the right default when the work is repetitive, errors are easy to detect, and the result is reversible. Route by the cost of failure, not by an abstract ranking of intelligence.
| Work characteristic | Default route | Reason |
|---|---|---|
| Bounded, repetitive, easy to verify, and reversible | Lower-cost capable model | Verification controls downside while price supports higher volume |
| Ambiguous context or complex reasoning with a clear rubric | Stronger model, then evaluate the result | Higher execution cost may be offset by fewer retries and corrections |
| Customer commitments, payments, account changes, or other consequential actions | Appropriate model plus explicit human approval | Low token cost does not compensate for an uncontrolled external mistake |
| Novel work with no stable acceptance test | Shadow mode or assisted drafting | You need evidence about quality before granting action authority |
Use the same outcome definition when comparing routes. Otherwise, one model may receive routine cases while another receives exceptions, making the cost comparison meaningless. Capture attempt cost, acceptance, human minutes, and defects for each route on representative work.
Then attach a defensible value to an accepted outcome. Use a value finance and the operating owner can recognize:
- gross profit from an incremental completed transaction;
- external or contractor spend actually removed;
- the cost of capacity that no longer needs to be added;
- human capacity reassigned to a named, valuable queue;
- loss avoided through a defined risk-control method.
Avoid adding several versions of the same benefit. If released time already explains an avoided hire, counting both at full value double-counts the result.
For a recurring responsibility, the break-even model is straightforward:
Contribution per period = accepted outcomes multiplied by value per outcome, minus variable operating cost.
Payback period = setup and fixed enablement cost divided by contribution per period.
Run the calculation with expected and downside assumptions. Acceptance may fall when the work mix broadens. Review time may rise when volume arrives in bursts. Context maintenance may be negligible during a pilot and material across departments. If the economics work only when every assumption is favorable, you have not found scalable capacity yet.
Run a capacity pilot that answers a budget question
A strong pilot should tell you whether to add agent capacity, improve the workflow, upgrade the model, or stop. It should not end with a collection of impressive examples and no operating decision.
- Select one recurring responsibility. Choose work with a recognizable trigger, stable demand, accessible inputs, and an accountable outcome owner.
- Baseline the current process. Record demand, completion volume, cycle time, human touch time, error or rework patterns, and current cost using the same definition you will apply to the agent-assisted process.
- Write the acceptance rubric before running the agent. State what must be correct, what may vary, which evidence must accompany the result, and which failures are unacceptable.
- Assemble representative evaluation work. Include ordinary cases, known exceptions, conflicting inputs, missing information, and situations that should trigger escalation rather than confident completion.
- Start in shadow or draft mode. Let the agent produce results without taking consequential external actions. Record whether a reviewer would accept, correct, rerun, or reject each result.
- Compare model routes on comparable work. Measure cost per accepted result and human touch time, not just response quality or token price.
- Add bounded action authority. Permit reversible, observable actions first. Keep an approval gate for commitments or changes whose cost cannot be cheaply undone.
- Make a constraint-based decision. Add execution capacity only when valid demand is waiting and execution is the bottleneck. Fix context, tools, permissions, or review throughput when one of those is limiting the system.
Set the scale decision before the pilot begins. A responsibility is ready to expand when acceptance meets the business owner’s standard, human touch time is lower than the baseline, the all-in unit cost is below the value created, exceptions reach a named owner, and the review queue remains manageable under representative demand.
Redesign rather than scale when failures cluster around missing context, unclear definitions, or tool boundaries. Stop when nobody owns the outcome, correctness cannot be evaluated, consequential actions lack a dependable control, or the full cost exceeds the result’s value.
Concurrency is especially easy to misread. More simultaneous agents help only when execution throughput is the active constraint. If review is limiting throughput, additional concurrency increases work in progress and delays feedback. If context preparation is limiting throughput, it forces more people to package inputs. If demand is limiting throughput, it creates idle capacity. Buy or enable more of the constrained resource, not more of every resource.
Once the first responsibility works, manage the portfolio with a common scorecard. Rank candidate responsibilities by eligible demand, expected accepted throughput, value per accepted outcome, all-in unit cost, human touch time, failure consequence, and setup effort. This makes a small but profitable workflow visible beside a glamorous workflow with weak demand or expensive supervision.
Key takeaways
- An agent count measures supply. Accepted outcomes measure useful capacity.
- Your maximum throughput is set by the tightest constraint across demand, context, permissions, execution, and verification.
- Human setup, review, rework, and exception handling belong in both the capacity calculation and the unit cost.
- The cheapest model per token is not necessarily the cheapest model per accepted outcome.
- Model routing should reflect verifiability, reversibility, and the cost of failure.
- Scale a recurring responsibility only after its owner, acceptance standard, economics, and exception path are explicit.
At your next AI portfolio review, do not approve an unspecified number of agents. Approve a responsibility, its control boundary, its expected accepted throughput, and its unit economics. When the scorecard identifies a real execution bottleneck, add agent capacity there. When it identifies a context or review bottleneck, solve that instead. That discipline is what turns abundant AI execution into useful operating capacity.
References
- Epoch AI – Hundreds of millions of AI agents are coming. Is there work for them?
- Nate Jones’s Substack – DevDay recap: 3 of 20+ launches matter. Sol costs 1/5 of Astra.








