Your strongest AI operator can finish an implementation before the rest of the team has settled the decision behind it. Then the work sits: waiting for context, review, testing, approval, or release. You bought faster production, but customer-facing change still arrives at nearly the same pace.
That is not primarily an AI adoption problem. It is a flow problem. Your leadership job is to redesign the path from idea to accepted outcome so that faster local work does not create larger queues everywhere else.
Model capability is an input, not a productivity result
The model market is improving along several dimensions at once. OpenAI has moved newer coding, factual-accuracy, and computer-use capabilities into lower-cost models. Anthropic says its latest flagship handles typical workloads more than 30% faster and at 40% lower cost than its predecessor. Other releases emphasize longer autonomous tasks, self-checking, larger contexts, or faster inference. These are meaningful changes in what is technically and economically possible, although vendor performance claims still need validation in your own environment.
But a faster model does not automatically produce a faster team. Model performance affects the time required to generate an artifact. Team performance depends on what happens before and after that artifact appears.
- Local efficiency: How quickly and cheaply can a person and an agent produce a usable draft, analysis, design, test, or code change?
- Flow efficiency: How long does that work take to move from ready to accepted, integrated, and released?
- Outcome effectiveness: Did the released work change customer behavior, reduce operational effort, lower risk, or improve the business result it was meant to affect?
A dramatic improvement at the first layer can disappear at the second. The 10x individual versus 1.8x team contrast is useful as a management frame, not a universal benchmark. Your multiplier will be different. The durable point is that reducing one component of elapsed time cannot remove the time spent in every other component.
This distinction should change the target you set. I would not make model usage, generated output, or prompt volume the primary productivity goal. I would target the elapsed time to an accepted outcome, with quality and operating cost as guardrails.
Find the queue that absorbed the saved time
When implementation gets faster, the saved time rarely vanishes. It moves. Code waits for review. A prototype waits for a product decision. Research waits to be interpreted. A campaign waits for legal approval. An automated workflow waits for access to a production system. More work gets started because starting is now cheap, while the rate of acceptance remains unchanged.
Map one recurring value stream from the moment work becomes ready to the moment its intended user can rely on it. For each completed item, record both active work and waiting time. Use the timestamps your systems already produce where possible: ticket transitions, pull requests, test runs, approvals, deployments, publication, and customer enablement.
- Name the acceptance point. “Agent finished” is not acceptance. A reviewer approving the change, a customer using the capability, or an operations team relying on the output may be.
- Mark every handoff between creation and acceptance.
- Separate active work from waiting. A difficult review and a review that sat untouched are different problems.
- Identify the longest recurring wait, not the task that feels most sophisticated.
- Change that constraint before buying more generation capacity.
| What you observe | Likely constraint | Leadership response |
|---|---|---|
| Prototypes are ready while requirements keep changing | Decision ownership | Name the person who can settle the unresolved question and record the decision the team needs. |
| Pull-request volume rises but merge time does not improve | Review capacity or change size | Route changes by risk, automate repeatable checks, and require reviewable units rather than larger batches. |
| Agents repeatedly fail on environment-specific details | Missing access, context, or feedback | Give the workflow a safe way to inspect the real environment and verify its assumptions. |
| More work starts while completed work stays flat | Excess work in progress | Stop launching additional parallel tasks and finish the oldest blocked items first. |
| The same expert rescues every failed run | Hidden support dependency | Turn recurring fixes into checks, examples, or documented recovery paths with a separate owner. |
| Delivery accelerates but defects and reversals increase | Weak verification | Narrow the agent’s authority and add reproducible tests at the point where failures can still be contained. |
The important shift is from asking, “How can the agent generate more?” to asking, “What prevents generated work from becoming trusted work?” The answer might sit in engineering, but it may also be a product decision, an approval policy, an overloaded domain expert, or an unclear definition of done.
Standardize the handoff, not everyone’s working style
The usual response to uneven AI performance is to ask the strongest operator to document every prompt, tool, and habit. That can help beginners, but it confuses personal technique with organizational infrastructure. Models change. Tools change. Effective operators adapt their methods to the task. A mandatory prompt stack can freeze yesterday’s practice while imposing friction on the people already learning fastest.
Standardize what crosses a team boundary instead. Every consequential agent-assisted deliverable should arrive with an artifact contract: the minimum information another person needs to evaluate, reproduce, continue, or reverse the work.
- Intent: the problem being solved and the acceptance criteria.
- Inputs: the relevant files, data, environment, constraints, and permissions.
- Assumptions: what the agent treated as true, including anything it could not verify.
- Output: the actual change, analysis, design, or recommendation rather than a transcript of the process.
- Evidence: reproducible tests, evaluation results, source links, or another check appropriate to the work.
- Risk: the areas most likely to fail and the expected effect if they do.
- Continuation state: what is complete, what remains open, and how another person can resume.
- Accountability: the human owner who reviewed the result and can answer for the decision.
This contract makes different tools and working styles interoperable. One engineer can use several agents in parallel. Another can work sequentially with a single assistant. A product manager can use a long-context model to inspect discovery material. None of them must expose identical internal methods if their outputs meet the same acceptance standard.
The strength of the evidence should match the consequence of an error. A reversible internal draft may need only a clear owner and a quick review. A production change needs repeatable technical checks. A workflow that can alter permissions, delete data, commit money, or communicate with customers needs an explicit human checkpoint and a recovery path. Do not grant unattended authority merely because an agent passed a local test.
This is also where many approval processes should be redesigned. If an old review existed to catch formatting, syntax, or another mechanically testable condition, automate that check and reserve human attention for intent, trade-offs, and risk. If a review protects a real control, preserve the control but give the reviewer structured evidence instead of a raw transcript and a large generated artifact.
Scale the practice without consuming your strongest operator
Your most effective AI user is evidence that a better operating method exists. That does not make this person the default trainer, support desk, prompt librarian, workflow maintainer, and reviewer for everyone else. Assigning all of those jobs can erase the capacity gain you hoped to spread.
Extract the reusable parts of the practice while protecting the operator’s time. Focus on stable interfaces and failure prevention, not a complete imitation of the person’s setup.
- Observe a real workflow. Capture where the operator branches work, what context is prepared, how results are checked, and which failures require judgement.
- Convert recurring judgement into a reusable aid where possible. This may be a check runner, evaluation set, example repository, acceptance template, decision record, or recovery procedure.
- Keep irreducible judgement human. If the work depends on product taste, architecture, policy, or risk appetite, name the accountable decision-maker instead of hiding the decision inside a prompt.
- Transfer ownership. The original operator can help create the first reusable asset, but a manager, platform owner, or enablement owner should accept responsibility for maintenance and support.
- Retire what is not used. Templates and checks are products. If nobody relies on one, maintaining it is overhead rather than leverage.
Charge the full support cost to the rollout. If several people appear faster only because the strongest operator spends more time rescuing their work, the organization has borrowed productivity from its scarcest contributor. The team-level result may be flat even while individual dashboards look better.
Communities of practice can spread discoveries, but they should not become mandatory demonstrations of personal workflow. Bring concrete failures, checks, and completed artifacts. A passing test that everyone can run is more scalable than a presentation about how one person phrases prompts.
Run a two-week trial around one complete workflow
A broad rollout makes attribution almost impossible. Start with a recurring workflow that has a visible acceptance point and enough activity to expose its queues. A two-week workflow trial is long enough to reveal handoff problems without pretending to establish a permanent productivity benchmark.
Record the existing workflow before changing it, then use the same definitions during the trial. Track five measures:
- End-to-end lead time: elapsed time from ready to accepted. This is the number local speed must eventually improve.
- Accepted throughput and unit cost: work accepted per period, alongside model fees and other variable costs. Generated volume does not count until it passes the agreed acceptance point.
- Handoff wait: time spent waiting for review, a decision, access, testing, approval, or release. This reveals where the saved production time moved.
- Quality and rework: material revisions, reopened work, failed evaluations, defects, rollbacks, or other corrections relevant to the workflow.
- Expert tax: interruption and maintenance time contributed by the strongest operators. This catches gains that were borrowed rather than created.
Read the measures as a system. If task execution gets faster but end-to-end lead time does not, work on the largest handoff wait. If accepted throughput rises while rework rises, strengthen verification or narrow the scope of autonomy. If the trial succeeds only through repeated expert intervention, improve the reusable checks and ownership model before expanding it. If accepted throughput improves while quality, cost, and expert tax remain controlled, move to an adjacent workflow with similar inputs and risks.
Do not use a short trial to claim a durable business outcome that takes longer to appear. Its purpose is to test the operating mechanism: whether agent-assisted work can move through the team with less waiting and without shifting hidden costs onto reviewers, maintainers, or customers.
Key takeaways
- A better model expands what your team can attempt; it does not automatically shorten the path to customer value.
- Measure from ready to accepted, then find the recurring queue that absorbed the saved execution time.
- Standardize intent, evidence, continuation state, and accountability rather than forcing everyone to use the same prompts.
- Turn repeated support into checks and reusable assets, with maintenance owned beyond the fastest operator.
- Scale only when accepted throughput improves without unacceptable rework, cost, or expert interruption.
At your next planning review, do not begin with “How do we get everyone onto the newest model?” Choose one recurring workflow, name its acceptance point, expose the waiting time, and remove its largest constraint. When the next model improvement arrives, your operating system will be ready to convert more of that capability into finished work.
References
- Nate Jones’s Substack — Executive Briefing: Your Best Engineer Got Ten Times Faster. Your Team Got 1.8x. Here’s How to Close the Gap.
- Substack AI Topic — GPT 6 Sol, Gemini 3.8 TTS, Grok 4.7, Opus 5.5, MiMo 2.6, WorldCrafter: AI News








