If AI can handle a meaningful share of your team’s work, you face a deceptively simple decision: reduce headcount, retrain people, or redesign the role. Starting with headcount is usually a mistake. It treats a job as one indivisible unit and ignores where the work, risk, and value actually sit.
The more useful response is to separate tasks from jobs, production from learning, and assistance from accountability. That gives you a practical way to adopt AI without preserving obsolete workflows, weakening your team’s judgment, or automating decisions you still need people to own.
Stop planning at the level of jobs
Andrew Ng’s useful provocation is that AI could perform 30-40% of many jobs while making the remaining human contribution more valuable. The percentage is a directional estimate, not a universal benchmark. A product manager, engineer, support specialist, and recruiter will have different task mixes, tools, constraints, and costs of error.
The important idea is economic complementarity. If AI makes one input cheaper, the scarce input paired with it can become more valuable. Cheap first drafts increase the value of editorial judgment. Faster code generation raises the importance of architecture, testing, security, and problem selection. Abundant analysis makes deciding which evidence deserves attention more important, not less.
That doesn’t guarantee every position will remain intact. Partial automation can still reduce staffing when demand is fixed, tasks are highly standardized, or a company chooses cost reduction over expansion. It does mean that a flat claim such as AI will replace product managers or AI will not affect product managers is too crude to guide a real operating decision.
Audit the work your team actually performs, not the responsibilities preserved in an old job description:
- Collect the recurring artifacts produced during an ordinary work cycle: briefs, queries, designs, tickets, forecasts, interview notes, code, decisions, and customer communications.
- Break each artifact into tasks small enough to assign separately. Researching a market, choosing evidence, drafting a narrative, challenging assumptions, and approving a recommendation are different tasks even when they end up in the same document.
- Describe the model’s possible role with a precise verb: retrieve, classify, summarize, generate, compare, recommend, or act. The verb exposes the level of authority you are considering.
- Name the human complement. It may be selecting context, defining acceptance criteria, interpreting ambiguity, checking evidence, resolving a trade-off, earning trust, or accepting accountability.
- Prioritize work that is frequent, reviewable, and reversible. Treat opaque or high-consequence decisions as a different class of problem, even when a model appears capable of producing an answer.
A simple decision table helps prevent an impressive demonstration from turning into a careless rollout.
| Work pattern | Useful AI role | Human complement | Leadership decision |
|---|---|---|---|
| Research and synthesis | Retrieve, cluster, and draft | Choose credible evidence and detect missing context | Augment with explicit citation and review requirements |
| Repeatable production | Generate code, copy, queries, or variations | Define tests and approve release quality | Automate only where acceptance can be checked reliably |
| Ambiguous prioritization | Surface options and challenge assumptions | Make trade-offs and own consequences | Use AI as decision support, not decision authority |
| Relationship work | Prepare context and capture follow-up | Build trust, negotiate, and read the room | Keep the interaction human-led |
The output of this audit isn’t a list of jobs to remove. It is a map of where capacity can expand, where review work will grow, which skills become bottlenecks, and which decisions must retain a named human owner.
Redesign the workflow, not just the prompt
Most weak AI adoption leaves the operating system untouched. Employees receive access to a chat interface, attend a demonstration, and are told to find use cases. The result is scattered experimentation: some people save time, some create extra review work, and nobody can explain whether the organization is becoming more capable.
A prompt is only one component of a reliable workflow. Before scaling a use case, define the full path from input to accepted outcome:
- Input contract: Specify the context the model needs, where that context may come from, and what data must not enter the system.
- Output contract: Define the required structure, evidence, assumptions, and uncertainty. A polished paragraph is not a useful output specification.
- Acceptance test: Decide how a person or system will determine whether the result is correct enough to use. If correctness can’t be checked, the workflow is not ready for unattended automation.
- Escalation trigger: Identify the conditions that require human review, such as conflicting evidence, missing context, sensitive data, or an action that is difficult to reverse.
- Accountable owner: Name the person responsible for the outcome. The model can perform a task, but it cannot carry organizational accountability.
- Failure record: Capture recurring error patterns, not just successful examples. Those failures should shape prompts, retrieval, evaluations, permissions, and training.
This changes how you evaluate an AI pilot. Time saved matters, but it is incomplete. You also need to notice where time moved. Faster generation may create a review bottleneck. More variations may slow approval. A summary tool may reduce reading while increasing the chance that a weak source influences a decision.
Separate assistance from agency as well. A system that drafts a customer reply creates a different risk from one that sends it. A system that proposes a database query differs from one that runs the query against production data. A system that recommends a roadmap change differs from one that alters priorities in the system of record. Each additional permission changes the failure surface and should require stronger evaluation, observability, and rollback controls.
The goal isn’t maximum automation. It is a better allocation of attention. Let the model absorb work that is costly because it is repetitive or voluminous. Preserve human attention for consequential trade-offs, novel failure modes, and situations where context exists outside the model’s reach.
Protect learning from the convenience of correct answers
AI can improve the artifact in front of you while weakening the capability behind it. Ng’s warning is specifically about common usage: better homework output can coexist with worse retention when the model supplies the reasoning a learner should be practising. That distinction matters well beyond school.
At work, the same failure appears when an engineer ships code they can’t later explain, a product manager presents analysis they couldn’t reconstruct, or a leader accepts a strategy whose assumptions were generated but never examined. The immediate task is complete. The next task becomes harder because no durable mental model was formed.
The problem isn’t that an LLM explains too much. It is that fluent answers make recognition feel like understanding. You can follow a solution and still be unable to produce, adapt, or challenge it without the conversation in front of you.
Use separate modes for production and learning. In production mode, optimizing for speed is reasonable when the task is understood, the output is verifiable, and the consequences are controlled. In learning mode, useful friction is part of the design.
- Attempt before asking. Write your own hypothesis, outline, query, or solution first. This reveals the gap you actually need help with.
- Ask for critique before replacement. Have the model identify flaws, counterexamples, or missing assumptions without immediately supplying a finished answer.
- Interrogate the mechanism. Ask why the solution works, where it fails, what constraints matter, and what would change the recommendation.
- Reconstruct without the transcript. Close the conversation and explain the idea or rebuild the result in your own words. If you can’t, the model completed the task but you did not learn it.
- Transfer the idea. Apply the underlying principle to a different case. Transfer is a stronger signal of understanding than reproducing the original output.
- Record the durable lesson. Save the rule, decision, failure mode, or example in a place you can retrieve independently of the chat history.
Leaders should build this distinction into enablement. A tool demonstration teaches interface mechanics. It doesn’t prove that a team can frame a problem, verify an answer, or recover when the model fails. Training should therefore include explanation, reconstruction, and error diagnosis alongside prompt examples.
This is especially important for junior employees. If AI removes every first-draft task, it may also remove the practice through which they acquire judgment. Don’t preserve manual work merely because it is traditional. Do decide which experiences develop the skills your future reviewers and decision-makers will need, then make those experiences explicit.
Hire and manage for the new bottleneck
Once generation becomes cheaper, volume stops being a strong proxy for contribution. The person who produces the most documents, tickets, or code may simply be creating the largest review burden. Productive AI use shows up in accepted outcomes, stronger decisions, shorter feedback loops, and fewer preventable failures.
The sharper employment pressure is therefore not simply human versus model. It is often a competition between people who can redesign their work around AI and people who preserve a manual workflow. The concern that education is still preparing people for 2022-style work instead of the capabilities needed for 2028 should change both hiring and internal development.
For AI-exposed roles, assess whether a candidate can:
- turn an ambiguous goal into a clear task and acceptance criteria;
- assemble relevant context without exposing data the tool is not allowed to process;
- choose when to use AI, when to use a deterministic tool, and when to work unaided;
- verify claims against evidence instead of trusting fluent output;
- notice failure patterns and improve the workflow rather than retrying prompts indefinitely;
- explain the final reasoning in their own words; and
- retain ownership when an AI-assisted recommendation is wrong.
A practical work sample can reveal these abilities without pretending AI doesn’t exist. Ask the candidate to frame the problem and define what a good answer would require. Then allow AI-assisted execution and observe what context they supply, what they challenge, and how they check the result. Finally, remove the tool and ask them to explain the recommendation, its weakest assumption, and the evidence that would change their mind.
This structure distinguishes tool fluency from borrowed competence. Banning AI throughout the exercise hides how the candidate will actually work. Allowing invisible AI use hides whether they understand what was produced.
Apply the same logic to performance management. Don’t reward prompt activity or raw output volume. Look for better problem selection, sound verification, reusable workflows, documented failures, responsible handling of context, and evidence that saved effort was reinvested in customer understanding or decision quality. AI literacy is not the ability to make a model respond. It is the ability to produce dependable outcomes through an AI-enabled system.
Lead between panic and complacency
Public AI debate often asks leaders to choose between two moods: catastrophe or inevitability. Neither is an operating model. Claims that some AI fear campaigns benefit incumbents through regulatory capture raise a legitimate incentive question, but they do not prove the motive behind every safety proposal. Likewise, an exaggerated analogy doesn’t make concrete risks disappear.
Translate every broad claim, optimistic or alarming, into a risk statement your team can examine:
- Harm: What specific bad outcome could occur?
- Pathway: How would this product, model, user, and permission set produce it?
- Exposure: Who or what would be affected, and can the action be reversed?
- Evidence: Is the concern based on observed failures, a credible mechanism, a stress test, or speculation?
- Control: Can access restrictions, retrieval boundaries, human approval, evaluation, monitoring, or rollback reduce the risk?
- Residual decision: After controls, who decides whether the remaining risk is acceptable?
This keeps policy arguments separate from product controls. You don’t need to settle the future of artificial general intelligence before deciding whether customer data may enter an external model, whether generated code needs security review, or whether an agent may take an irreversible action. If your data policy prohibits information from leaving a controlled environment, don’t paste that information into an external model. Use an approved architecture or keep the task human-operated until one exists.
The same discipline prevents fear from becoming paralysis. A hypothetical systemic concern should not automatically block a reversible, observable workflow with low-consequence data. Evaluate the actual deployment, define the control, and collect evidence. Governance should help the organization make differentiated decisions, not turn every use case into the same yes or no.
Key takeaways
- A job is a bundle of tasks. Map the bundle before making hiring or headcount decisions.
- Partial automation increases the importance of complementary skills such as judgment, verification, context selection, and accountability.
- A reliable AI workflow needs an input contract, an output contract, an acceptance test, an escalation path, and a named owner.
- Production and learning require different behavior. Fast answers can improve today’s artifact while weakening tomorrow’s capability.
- Hire for problem framing, verification, failure diagnosis, and ownership, not prompt fluency alone.
- Turn sweeping AI claims into specific harms, pathways, evidence, controls, and accountable decisions.
At your next operating review, choose a role close to an important bottleneck and map its real tasks. Assign the model a precise role, define the human complement, and write the acceptance and escalation rules before adding another tool. Then inspect not only what became faster, but what your people stopped practising. That is where the future of work becomes a management decision rather than a prediction.
References
- The AI Corner — Andrew Ng: AI Already Does 30-40% of Your Job
- The AI Corner — The Man Who Taught 8 Million People AI Just Said AI Is Terrible for Learning








