You pay for several capable AI tools, but muscle memory still sends almost every task to the same one. The alternative appears only when your default disappoints. It gets a colder prompt, less context, and less opportunity to revise. That isn’t a useful comparison. It is a biased audition.
You don’t need an elaborate router to fix this. You need a small operating system for your own work: assign models clear jobs, keep durable context outside their chats, use a consistent handoff packet, and preserve the corrections that would otherwise disappear into conversation history.
Route the state of the work, not its subject
Labels such as writing model, coding model, or research model are too broad. A strategy memo can move through ambiguous exploration, evidence gathering, structured drafting, adversarial review, and mechanical formatting. Those stages do not require the same kind of intelligence.
Start by separating four kinds of work:
- Explore: The problem is still unsettled. Ask a model to expose assumptions, generate materially different approaches, and identify decisions you have not made. Do not ask for polished prose yet.
- Produce: The objective, constraints, and acceptance criteria are clear. Give a model a bounded artifact to create, such as a decision memo, analysis query, prototype, or interview rubric.
- Challenge: A plausible answer exists, but you need to find what it missed. Give another model the original brief as well as the candidate output, then ask it to locate unsupported claims, contradictions, unhandled cases, and weak evidence.
- Transform: The judgment has already happened. Use a suitable lower-cost or faster model for formatting, classification, extraction, conversion, or other repeatable operations with stable instructions.
The challenger does not have to produce the artifact you keep. A rejected prototype can still be valuable if it reveals choices you did not know you needed to specify. In product work, that might be a missing permission state, an undefined audience, an overlooked exception, or a trade-off hidden inside a vague requirement.
| Work state | Model’s job | What you provide | What good looks like |
|---|---|---|---|
| Ambiguous | Explorer | Problem, evidence, constraints, unknowns | Important decisions and distinct options become visible |
| Defined | Producer | Brief, acceptance criteria, output contract | A usable artifact that satisfies the stated checks |
| Plausible but consequential | Challenger | Original brief, evidence, candidate artifact | Specific defects, missing evidence, and testable objections |
| Routine | Transformer | Stable instructions, examples, required format | Consistent output with little review burden |
Keep model names out of this table. Roles should survive product launches, pricing changes, usage limits, and model upgrades. Your explorer this month may become your producer later. The job definition stays stable even when the cast changes.
Compare models on time to accepted work
First-response quality is only part of performance. A model that responds sooner can create more room for feedback, correction, testing, and another build before your deadline. In a real application test, a faster initial build received more revision cycles and became the version that was kept. That does not prove it was universally better. It shows that latency changed the amount of learning available to the workflow.
Measure time to accepted output, not time to first output. The distinction matters whenever the task is iterative: prototyping, executive writing, data analysis, coding, research synthesis, or product-specification work.
Your default model also enjoys a hidden advantage. It has received more of your context, more examples of your preferences, and more chances to recover from mistakes. You have learned how to phrase requests in a form it handles well. A rival often receives the same prompt style without the same accumulated practice, creating an unequal comparison between the familiar model and the challenger.
Use this audition protocol when a task is important enough to influence your routing rules:
- Choose a representative task: Use real work whose quality you can judge. A synthetic puzzle may reveal capability, but it will not tell you whether a model fits your operating environment.
- Freeze the brief: Give each candidate the same objective, constraints, evidence, source files, output format, and definition of done.
- Equalize access: Keep tool permissions and relevant context comparable. If a candidate cannot use a required tool, record that as a workflow limitation rather than quietly changing the task.
- Allow comparable recovery: Give each candidate the same type of feedback and a fair chance to revise. Do not compare a polished incumbent result with a challenger’s first draft.
- Record useful surprises: Credit a model for identifying an unasked question or a hidden decision, even if you keep another model’s final artifact.
- Score the finished workflow: Track acceptance-criteria coverage, severity of errors, manual editing, elapsed time, review burden, and ease of recovery after feedback.
Do not collapse those observations into a universal winner. Produce a narrow routing decision: use this model for unsettled argumentation, that model for bounded execution, and another option for routine transformation. A model can be excellent at finding the problem and mediocre at packaging the answer. That still makes it valuable.
For adversarial review, resist the vague instruction to critique this. Give the challenger the original acceptance criteria and ask it to classify its findings: missing requirement, unsupported claim, internal contradiction, edge case, or presentation issue. Require it to point to the relevant passage and propose a verification step. This turns disagreement into work you can inspect.
Keep context in a portable workspace
A multi-model workflow breaks down when the only complete version of your project lives inside a long chat. Every handoff then becomes an exercise in reconstructing what happened, and your best model appears better partly because it owns the history.
Move durable context into ordinary files that you control. A practical project workspace can contain:
brief.mdfor the objective, audience, constraints, non-goals, and definition of done.decisions.mdfor settled decisions, rationale, unresolved questions, and assumptions that still need evidence.sources/for approved documents, notes, transcripts, and the extracted text actually available to the model.skills/for reusable procedures such as writing an executive memo, reviewing a product requirement, or checking an analysis.corrections.mdfor recurring failures and the behavior you want instead.outputs/for artifacts worth preserving independently of any chat session.
This does not require a specialized memory platform. One workable setup keeps reusable knowledge and skills in an Obsidian vault shared across Claude Code and Codex, making the workspace accessible from different devices. Plain Markdown files are valuable precisely because they do not depend on one model’s private memory.
Portable does not mean sending the entire workspace with every request. More context can bury the current decision, introduce stale instructions, and expose information the task does not need. Select the smallest packet that preserves the relevant state.
Use a handoff packet instead of a conversation dump
A clean handoff should answer these questions:
- Goal: What decision or artifact is needed, and who will use it?
- Current state: What has already been attempted, and what remains unresolved?
- Evidence: Which files and observations may support the answer?
- Settled decisions: What must not be reopened without new evidence?
- Open questions: Where is judgment still required?
- Constraints: What legal, privacy, technical, brand, timing, or organizational boundaries apply?
- Output contract: What structure and level of detail should the result have?
- Verification: How will you decide whether the result is acceptable?
Ask the producing model to return an updated artifact plus any new assumptions it introduced. Save the artifact, update the decision record, and give the challenger that same packet. This prevents important state from being trapped inside an invisible chain of reasoning or a vendor-specific conversation.
Debug the representation before rewriting the prompt
Document work needs an additional check. A PDF can look correct to you while its extracted text has separated a table value from its heading or moved a footnote away from the sentence it qualifies. When an answer grounded in a document is wrong, inspect the extracted text before changing the prompt.
Try to answer the disputed question using only the representation the model received. If the necessary relationship is missing or ambiguous there, repair the extraction or provide a clearer structured version. Retrieval cannot find a relationship that parsing discarded, and repeated prompting cannot reliably reconstruct information that never reached the model.
Convert repeated corrections into durable skills
A correction buried in a chat improves one output. A correction stored as a reusable instruction improves the workflow. This is where personal AI use begins to compound.
When the same kind of correction recurs, record it in a model-neutral form:
- Trigger: Describe when the instruction applies.
- Observable failure: State what the bad result looks like without guessing why the model produced it.
- Required behavior: Explain the replacement behavior in direct, testable language.
- Example: Preserve a compact before-and-after example when wording or structure matters.
- Check: Add a way for the model or reviewer to verify compliance.
For example, do not save a preference as make executive writing better. Save the behavior: lead with the requested decision, separate evidence from recommendation, surface unresolved risks, and remove background that does not change the decision. The latter can be applied and checked by any capable model.
Do not turn every isolated disappointment into a permanent rule. A correction may be specific to a malformed input, a particular task, or a temporary model behavior. Promote it only when it expresses a preference or quality check you expect to remain useful. The durable pattern is to turn recurring corrections into reusable skills rather than losing them in old conversations.
Spend judgment where it changes the outcome
The strongest model does not need to touch every stage. Once the judgment has been resolved and the instructions are stable, a cheaper option can be appropriate for the clearly specified parts of a long job. The important boundary is not glamorous work versus menial work. It is unresolved judgment versus repeatable execution.
Before assigning a lower-cost or faster model, ask:
- Can success be described with observable checks?
- Are the inputs complete and consistently structured?
- Can a bad result be detected before it affects a customer, candidate, employee, or executive decision?
- Would reviewing the output take less effort than producing it directly?
- Can the task be retried without losing data or creating an external commitment?
Judge cost at the point of acceptance. Include model usage, elapsed time, human review, rework, and the consequence of an undetected error. A cheap call that creates expensive review is not a cheap workflow. Nonurgent, well-bounded work can also be scheduled around usage constraints; a portable agent setup can move scheduled tasks to an always-available machine and run deferred work later.
Set the data boundary before adding the challenger
Using another model creates another path through which your information may travel. That matters for roadmap details, customer records, employee information, candidate materials, source code, contracts, and internal financial data. A fair model comparison never justifies copying sensitive material into an unapproved system.
Record which tools are approved for each data class, what may be redacted, which integrations may retrieve internal files, and which outputs require human approval before an external action. Apply the same boundary to every route. If the challenger cannot receive the necessary evidence safely, test it on a sanitized equivalent or exclude it from that job.
Apply the system to your next consequential task
You can build the workflow while doing real work. Start with a task you already need to complete:
- Write an acceptance card containing the audience, requested decision or artifact, constraints, and observable checks.
- Mark the current work state as exploration, production, challenge, or transformation.
- Assemble a handoff packet from the relevant files instead of pasting an entire chat history.
- Assign a primary model based on the job. Save its output outside the conversation.
- Use a challenger only where ambiguity, consequence, or uncertainty makes another perspective valuable. Give it the original brief and evidence, not merely the primary model’s answer.
- Record what required manual correction, what useful surprise appeared, how much review was needed, and whether feedback produced a reliable recovery.
- Update the routing rule or reusable skill only when the result gives you a clear reason to do so.
What this looks like in product leadership
- Strategy memo: Use an explorer to surface strategic assumptions and alternative frames. Let a producer draft from the decisions you accept. Give a challenger the evidence and ask it to locate claims that would fail under executive scrutiny.
- Customer research: Verify transcript or document extraction before synthesis. Use a producer to organize evidence, then ask a challenger to search for disconfirming observations and places where the conclusion exceeds the notes.
- Product specification: Use exploration to expose undefined states and edge cases. Move to production only after the acceptance criteria are explicit. Use challenge mode to test permissions, failure handling, and interactions with existing behavior.
- Hiring: Create the structured rubric from approved job criteria before evaluating material. Use a challenger to identify inconsistent reasoning or unsupported inference, but keep accountability for candidate decisions with the human decision-makers.
- Internal AI transformation: Treat the portable workspace, approved data routes, reusable skills, and acceptance checks as operating infrastructure. Adoption without these artifacts produces isolated chats, not organizational capability.
Key takeaways
- Route models by the state of the work: explore, produce, challenge, or transform.
- Evaluate time to accepted output, including revision speed and review burden.
- Give competing models comparable context, permissions, criteria, and recovery opportunities.
- Keep briefs, decisions, evidence, outputs, and reusable skills outside model-specific chats.
- Inspect document extraction when grounded answers fail; prompting cannot repair missing structure.
- Use stronger judgment selectively, move stable procedures to suitable lower-cost models, and keep sensitive data inside approved routes.
For your next live task, pause before opening the default chat. Write the acceptance card, name the work state, and choose the model for that job. Save the result and the correction outside the conversation. The goal is not to use more models. It is to make each model earn a defined place while your context and judgment continue to compound independently of all of them.
References
- Nate Jones’s Substack – You pay for two frontier models and route almost everything to one
- Learn AI Together – LAI #142: My AI Setup in 2026








