Author: Shivam Tiwari

  • How to Prove the ROI of an AI Product Before You Scale It

    How to Prove the ROI of an AI Product Before You Scale It

    Your AI product is getting used. The demos land well, task completion is improving, and internal enthusiasm is high. Then the CFO asks a harder question: what changed in the business because this product exists?

    You cannot answer that question with prompt volume, response quality, adoption, or tickets touched. You need a measurement system that separates activity from incremental value, counts the full operating cost, and makes risk visible before a rollout gets larger. Here is how to build one.

    Start with the decision your ROI model must support

    ROI is not a retrospective slide assembled after launch. It is a decision rule. Before development begins, decide what evidence would justify launching, scaling, redesigning, rolling back, or retiring the capability.

    That distinction changes the conversation. Instead of asking whether the agent is accurate enough or popular enough, you ask whether a measurable change in customer behavior produces a measurable business result without crossing an unacceptable risk threshold.

    Build a driver tree with four levels:

    1. Company outcome: revenue growth, lower cost to serve, or reduced business risk.
    2. Customer outcome: the user completes a valuable job, reaches value sooner, or resolves a problem without unnecessary effort.
    3. Product behavior: the AI capability changes conversion, expansion, self-service completion, containment, handle time, or escalation.
    4. Controllable lever: the team changes the workflow, model behavior, conversation design, human review, or product guidance.

    The chain matters because a model metric is rarely a business metric. Better answer quality may improve task completion, which may improve trial-to-paid conversion. The ROI case depends on the full chain, not the first link.

    Value pathBusiness outcomeLeading evidenceGuardrails
    RevenueHigher conversion, average order value, or expansionTime-to-first-value and self-service completionErrors, complaints, and policy violations
    CostLower cost to serveContainment, deflection, and reduced handle timeEscalations, false resolution, and downstream customer harm
    RiskLower frequency or impact of harmful failuresHuman-review events and detected violationsFalse positives, false negatives, hallucinations, and security breaches

    Choose one primary value path for the investment case. Revenue, cost, and risk can all appear on the scorecard, but declaring all three as primary makes it too easy to rescue a weak result with whichever metric moved after launch.

    A support agent, for example, may appear successful because it contains more conversations. But containment is only valuable if customers actually resolve their problems. A conversation that never reaches a human can reduce measured support volume while increasing complaints or churn risk. This is why revenue, cost, and risk measures must be evaluated together.

    Write the measurement contract before you build the dashboard

    A measurement contract is a short agreement among product, data, finance, and the operational team affected by the AI workflow. It prevents the definitions, cost boundaries, and success thresholds from changing after results arrive.

    Your contract should answer these questions:

    • Who is eligible? Define the users, accounts, tasks, channels, and exclusions. Do not mix workflows with materially different economics.
    • What is the intervention? Name the AI capability and the version being evaluated. A model, prompt, retrieval pipeline, policy, or escalation change can alter the result.
    • What is the primary outcome? Select the business metric that determines whether the hypothesis passed.
    • What are the leading indicators? Use measures such as time-to-first-value, containment, and self-service completion to diagnose movement before lagging results mature.
    • What are the guardrails? Predefine acceptable limits for errors, hallucinations, false positives, false negatives, escalations, complaints, security events, and policy violations.
    • What is the baseline? Freeze the comparison period or control group before exposing the eligible population to the capability.
    • How will incrementality be proven? Specify the experiment, holdout, assignment unit, and minimum detectable effect.
    • What costs count? Agree on model or API consumption, labeling, evaluation, human review, and ongoing oversight before calculating value.
    • What action follows each result? Record the thresholds for launch, scale, redesign, rollback, and retirement.

    The contract should distinguish an outcome OKR from an output OKR. Shipping the agent, generating responses, and increasing feature use are outputs. Improving conversion, lowering verified cost to serve, or reducing harmful failures are outcomes. Outputs can explain what happened, but they cannot establish value on their own.

    Instrument the complete journey, not just the conversation

    An AI log tells you what the model did. An ROI dataset must also tell you what the user did next.

    Connect the journey from eligibility to business outcome:

    1. The user or account became eligible for the capability.
    2. The AI experience was offered, viewed, and engaged.
    3. A task was attempted, completed, abandoned, or repeated.
    4. A response was accepted, corrected, regenerated, or sent for human review.
    5. The interaction was contained, escalated, or handed to another workflow.
    6. The downstream conversion, expansion, support, retention, or complaint event occurred.
    7. The associated model cost, labeling work, and human-oversight cost were recorded.

    Carry a stable user or account identifier, experiment assignment, agent version, and journey identifier across those events. Without that connective tissue, the team may have an impressive agent dashboard and no defensible way to attribute a business outcome to the experience.

    Use behavioral analytics and session replay to understand why a metric moved. Use journey mapping and retention analysis to locate the friction worth solving in the first place. Product tours and in-app guidance can then help eligible users reach a validated workflow. This creates a closed loop from journey friction to experiment and measurable outcome, instead of a collection of disconnected AI metrics.

    Calculate economic value without turning activity into savings

    Start with net business value:

    Net business value = incremental revenue + cost avoided – total operating cost – quantified risk loss

    If finance requires an ROI percentage, divide net business value by the agreed investment base. Keep both the numerator and denominator visible. A percentage without its cost boundary is easy to inflate and hard to audit.

    Count only incremental revenue

    Do not credit the AI product with every transaction it touched. Credit it with the difference between the exposed population and the valid control or holdout.

    A practical revenue calculation is:

    Incremental revenue = eligible volume x measured outcome lift x value per additional outcome

    The measured outcome might be trial-to-paid conversion, self-service upsell, average order value, or expansion. Use the same eligibility definition, attribution window, and revenue treatment for the intervention and control. If the AI experience merely appears somewhere in a successful journey, that is influenced revenue, not proof of incremental revenue.

    Separate capacity from cashable savings

    Cost claims require more care than a deflection count. A contained interaction may create capacity without reducing expenditure. That capacity can still be valuable, but it should not be presented as cash savings unless spending actually changes.

    • Capacity created: employees have time available for other work, but the existing cost base remains.
    • Variable cost avoided: the company no longer incurs a cost that would have grown with each additional interaction.
    • Cashable savings: an approved budget, vendor charge, or staffing requirement is actually reduced.

    Report these separately. Otherwise, the same saved minute can be counted once as employee capacity and again as reduced spend.

    Validate that a deflected task was resolved, not abandoned or displaced to another channel. Then calculate avoided cost from the incremental lift in verified resolution, not the total number of conversations the agent handled.

    Include the operating costs that make the agent dependable

    Model or API cost is only one part of the investment. Include labeling, evaluation, human review, and operational oversight. If a safer workflow requires more review, that review is part of the product’s economics, not an external inconvenience to exclude from the model.

    Segment cost by agent, workflow, and outcome. Cost per response is useful for infrastructure management, but cost per verified successful outcome is the better economic unit. A cheap response that triggers retries, escalations, or corrections may be more expensive than a higher-cost response that completes the job.

    Do not bury risk inside an average ROI number

    Risk adjustment should make uncertainty visible, not create false precision. Use three layers:

    1. Hard guardrails: security and policy conditions that trigger containment or rollback regardless of financial upside.
    2. Observed risk indicators: error, hallucination, escalation, complaint, false-positive, and false-negative rates tracked by workflow and cohort.
    3. Financial adjustment: expected loss deducted from net value only when the probability and impact assumptions are credible enough for finance and risk owners to accept.

    Do not let a low-frequency, high-consequence failure disappear inside a high average success rate. If the downside cannot be defensibly monetized, keep it as an explicit decision constraint rather than assigning it a convenient dollar value.

    Prove incrementality before claiming impact

    The strongest ROI calculation still fails if the attribution is weak. A before-and-after improvement may come from seasonality, pricing, traffic quality, a support policy change, or another product release. The AI capability needs a counterfactual: what would have happened to comparable eligible users without it?

    Use an A/B test or holdout whenever the product and risk profile allow it. Make these choices before launch:

    • Assignment unit: Randomize at the level where the outcome occurs. If expansion is measured per account, account-level assignment can prevent users in the same customer organization from receiving conflicting experiences.
    • Primary outcome: Pick the metric that determines success and keep diagnostic metrics secondary.
    • Minimum detectable effect: Precompute the smallest lift worth detecting based on the baseline, available population, and business value. If the experiment cannot detect a decision-relevant change, extending the metric list will not fix it.
    • Guardrails: Test quality, escalation, complaints, security, and policy outcomes alongside the primary metric.
    • Analysis population: For a product-level ROI claim, analyze eligible users according to their assigned experience. Looking only at people who voluntarily used the agent introduces selection bias.
    • Measurement horizon: Keep the holdout long enough to observe the outcome named in the contract. Leading indicators can guide iteration, but they should not be substituted for retention, churn, Net Recurring Revenue, or other lagging outcomes.

    If randomization is not practical, use a fixed holdout or a frozen comparison period and document the limitations. A weaker design can still inform a decision, but the ROI claim should carry less confidence. Do not quietly promote correlation to causation because the rollout has executive attention.

    Interpret the result as a system. Suppose self-service completion rises but the business outcome does not. The agent may be solving a low-value task, attracting users who would have converted anyway, or shifting effort to a later step. If conversion improves while complaints or policy violations cross the guardrail, the value hypothesis may be valid but the implementation is not ready to scale.

    This is eval-driven development applied to product economics: define acceptable behavior and business success, measure both under controlled conditions, diagnose the failures, and repeat the test after a meaningful change.

    Turn ROI into a portfolio operating system

    A one-time business case goes stale as models, prompts, traffic, user behavior, and operating costs change. Maintain an Agent Analytics view for every production capability.

    Each agent scorecard should show:

    • The primary business outcome and current experiment result.
    • Leading journey metrics from eligibility through verified completion.
    • Revenue contribution, cost avoided, and total operating cost using the agreed definitions.
    • Quality and risk guardrails, including escalations and human-review events.
    • Performance by relevant customer, task, and journey cohort.
    • The agent, model, policy, and workflow version associated with the result.
    • The current decision status: exploring, launching, scaling, redesigning, contained, or retiring.

    Use the dashboard to make portfolio decisions, not merely to report trends:

    • Scale when the primary outcome clears the precommitted threshold, guardrails hold, net value is positive, and the result remains credible across the cohorts that matter.
    • Redesign when leading indicators improve but the business outcome does not, or when human review and escalation erase the economic gain.
    • Contain or roll back when a hard security, policy, or customer-harm threshold is breached, even if average financial performance is positive.
    • Retire when controlled measurement shows no decision-relevant incrementality or when dependable operation costs more than the value created.

    Review operational signals with frontline teams because they can explain patterns hidden by aggregate metrics. Review portfolio value in QBRs with product, data, finance, and risk owners so investment follows evidence rather than novelty.

    Only accelerate adoption after the workflow has demonstrated unit value. In-app guides, product tours, and lifecycle nudges can bring more eligible users into a validated flow. Measure whether those interventions increase the business outcome, not merely clicks or agent sessions. Scaling exposure to an unproven workflow scales its cost and risk as readily as its potential benefit.

    Key takeaways

    • Treat ROI as a precommitted decision rule for launch, scale, redesign, rollback, or retirement.
    • Connect model behavior to customer behavior and then to revenue, cost, or risk through a driver tree.
    • Freeze the baseline, cost boundary, guardrails, attribution method, and success thresholds before results arrive.
    • Credit only incremental revenue and verified avoided cost. Keep created capacity separate from cashable savings.
    • Include model consumption, labeling, evaluation, human review, and oversight in the operating cost.
    • Use controlled experiments or holdouts, with a decision-relevant minimum detectable effect, to separate causal impact from correlation.
    • Keep severe risk conditions as explicit constraints when they cannot be responsibly converted into a financial estimate.
    • Scale adoption only after the AI workflow has shown positive unit value under acceptable risk.

    Pick one high-friction customer journey and complete its measurement contract before the next roadmap review. If the team cannot name the baseline, control, primary outcome, cost boundary, guardrails, and decision thresholds, the capability is still an exploration. Label it honestly, instrument it properly, and earn the right to make an ROI claim.

    References

  • How to Deploy an Operator AI Agent in Customer Operations

    How to Deploy an Operator AI Agent in Customer Operations

    Your support team probably does not need another chatbot that summarizes a ticket on command. It needs help with the operational work surrounding every ticket: finding why escalations changed, keeping knowledge accurate, correcting broken automations, coordinating incident communication, and showing human reps what deserves attention next.

    An operator AI agent can take on that work, but only if you design it as an operating system for customer operations rather than a conversational layer over support APIs. The useful version closes the loop from signal to diagnosis to tested change. The dangerous version produces plausible commentary and receives permission to act before it has earned trust.

    Define the job as a closed loop, not a chat box

    A customer-facing AI agent handles an individual customer’s request. An operator agent works on the system around those requests: conversations, help content, automation configuration, performance data, incident workflows, and the human queue.

    That distinction changes the product requirement. The agent is not complete when it answers a question such as why escalations increased. It is complete when it can investigate the increase, identify a supported cause, determine which operational object needs attention, prepare a change, test that change where possible, and route it to the right person for approval.

    1. Observe: Detect a question, anomaly, scheduled task, failed conversation, release brief, or incident.
    2. Diagnose: Select the relevant metrics and attributes, inspect representative conversations, and separate recurring patterns from isolated cases.
    3. Locate the control point: Determine whether the problem sits in knowledge, guidance, a procedure, a data connector, an automation rule, or a human workflow.
    4. Propose: Produce a concrete artifact such as an article diff, configuration change, procedure, incident audience, or prioritized queue.
    5. Verify: Run a simulation or another appropriate check and expose failures, edge cases, and remaining uncertainty.
    6. Act and learn: Apply an approved change, record what happened, and monitor the affected outcome for regression.

    Consider the prompt, Why did escalations rise last week? A reporting copilot returns a chart. A useful operator identifies which escalation definition applies, segments the change, reads relevant conversations, finds the repeated cause, checks whether the corresponding help content or automation is deficient, and prepares the smallest defensible correction. That progression from an operational question to an actionable proposal is already possible across analysis, knowledge maintenance, automation building, and human support workflows.

    Write the acceptance criteria around that complete handoff. Require the evidence used, the proposed artifact, the scope of impact, the verification result, the named reviewer, and any action the agent is forbidden to take. If the output still leaves an operations manager rebuilding the context manually, you have a chat assistant, not an operator.

    Build reliability below the model and price that work honestly

    A foundation model with API access can make a persuasive prototype. It can query ticket data, summarize conversations, and write a report that appears coherent. The hard part begins when different workspaces use different fields, configurations, workflows, permissions, and definitions of success.

    The model should not have to rediscover your operating rules on every run. Encode those rules in purpose-built tools and reusable skills. A tool performs one bounded operation, such as retrieving a conversation, searching knowledge, or running a defined report. A skill coordinates several tools to complete a business job, such as debugging a failed resolution or rolling a policy change through the help center.

    Operator’s production architecture is described as having more than 50 tools and 10 multi-step skills. Those counts are not targets to copy. They illustrate how quickly the hidden surface area grows once an agent must do dependable operational work instead of demonstrating a few API calls.

    System layerJob it must performFailure you should test forControl to add
    Semantic retrievalFind content by meaning, not only exact wordsIrrelevant or incomplete evidence produces a confident diagnosisEvaluate retrieval against real support questions and known content gaps
    Attribute awarenessKnow which metrics, fields, and custom attributes are populated and meaningfulThe agent invents a pattern from sparse or unused fieldsExpose field definitions, coverage, allowed joins, and missing-data signals
    Atomic toolsPerform narrow reads or writes predictablyA broad API wrapper allows an unintended query or changeUse typed inputs, constrained scopes, explicit permissions, and structured results
    Domain skillsChain tools according to a repeatable customer-operations methodThe same request follows a different process on each runDefine required steps, exit conditions, evidence, and escalation paths
    Review interfaceTurn reasoning into charts, diffs, tests, and proposalsA reviewer approves a wall of prose without understanding the changeRender the decision in the format appropriate to the object being changed

    Semantic retrieval and attribute awareness deserve particular attention. Retrieval grounds the agent in the content that can actually answer the question. Attribute awareness stops it from treating every available field as equally meaningful. A custom field that exists but is almost never populated should not become the foundation of an operational recommendation.

    Give every tool a contract before the model can call it:

    • The business purpose and the questions it is allowed to answer.
    • The read and write permissions it requires.
    • The preconditions that must be true before it runs.
    • The evidence and identifiers it must return.
    • Its behavior when data is missing, ambiguous, stale, or inconsistent.
    • The audit event, approval requirement, and rollback path for a write.

    Evaluate build versus buy beyond the demonstration

    A proof of concept establishes that a model can produce a plausible answer with your data. It does not establish that the answer is grounded, that the proposed action is safe, or that the system will behave consistently as configurations change.

    For a build decision, include retrieval tuning, permission design, tenant isolation, tool maintenance, skill development, evaluation data, observability, proposal interfaces, audit history, rollback behavior, and on-call ownership. Also ask who will update the agent when a support object, metric definition, product policy, or API changes. If these responsibilities do not have durable owners, the internal agent will age like any other unsupported operations system.

    For a buy decision, ask the vendor to demonstrate your difficult cases rather than its preferred prompts. Use a conversation with conflicting evidence, an unused custom attribute, an outdated localized article, a misconfigured rule, and a proposed write with a wide blast radius. Inspect the evidence, tool trace, permissions, diff, test result, and audit record. The quality of the generated prose is one of the least informative parts of that evaluation.

    Put a proposal boundary around every material action

    Moving from analysis to live changes is a different class of production problem. A wrong summary wastes time. A wrong configuration can degrade customer outcomes across every conversation that matches it. An incorrect outbound message cannot be recalled after customers have read it.

    I would give the agent autonomy according to consequence, not according to how confident its language sounds:

    1. Read: Search content, inspect conversations, calculate approved metrics, and assemble evidence. Run these tasks autonomously within access controls and log every operation.
    2. Recommend: Explain a root cause or rank an opportunity. Attach the underlying conversations, segments, fields, and assumptions so a person can challenge the conclusion.
    3. Prepare: Draft an article, procedure, rule, connector configuration, customer response, or queue. Save it as a proposal with no production effect.
    4. Change: Publish, configure, send, or otherwise alter the live operation only after the required reviewer sees the exact scope and explicitly approves it.

    A proposal is a structured change object, not a paragraph asking for trust. Production-grade operator systems can present reviewable diffs before applying changes, allowing the reviewer to accept, reject, or refine the work. The same principle should govern any operator implementation.

    Your review screen should answer six questions without forcing the approver into another tool:

    • What object will change?
    • What exact fields, passages, rules, or recipients are affected?
    • What evidence connects the observed problem to this change?
    • What test ran, and which cases failed or remained untested?
    • Who must approve, and which permission will execute the action?
    • How can the change be reversed, and what cannot be reversed?

    Customer outreach needs the strictest treatment because sending is effectively irreversible. Do not approve a batch from a conversational summary that hides the audience. The safe alternative is a preview containing the resolved customer list, inclusion logic, exclusions, exact message variants, delivery channel, and approver. Start by allowing the agent to prepare that package while a person performs the send.

    Simulation also needs a visible place in the proposal. If the agent modifies an automation procedure, show which representative conversations were tested, the expected outcome for each, the observed outcome, and why any mismatch occurred. An overall pass label is not enough to reveal an important edge case.

    Human approval is not a permanent substitute for system quality. If reviewers routinely accept proposals without inspecting them, the control has become ceremonial. Track corrections, rejections, rollbacks, and the evidence reviewers open. Use those signals to improve the relevant retrieval rule, tool, skill, or interface.

    Roll out workflows in increasing order of consequence

    Choose the first workflow by its operating characteristics. A strong starting candidate recurs frequently, consumes expert attention, has accessible evidence, produces a clear artifact, and has a named reviewer. It should also allow the agent to be useful before it receives broad write permission.

    A practical rollout sequence looks like this:

    1. Recurring operations analyst. Give the agent one standing question, such as what changed in escalations or automation performance. Define the metric, comparison period, relevant segments, evidence requirements, and report destination. Require links to representative conversations and allow the conclusion that no action is warranted. Compare its reasoning with an experienced operator’s review until the failure modes are understood.
    2. Knowledge steward. Feed it a release brief or policy change. Ask it to find affected help content, identify missing coverage, and prepare article diffs in the required voice and format. Include localized variants where they exist. The reviewer should validate product behavior, instructions, links, policy language, and whether the proposed set of pages is complete before publishing.
    3. Automation maintainer. Start with known failed conversations. Ask the agent to distinguish a content gap from a rule, procedure, guidance, or connector problem; prepare the smallest correction; define triggers and edge cases; and simulate the result. Do not grant live configuration access until the tool trace and tests make the diagnosis reproducible.
    4. Human-operations coordinator. Use the agent to assemble an incident audience, draft targeted responses, prepare coaching evidence, or prioritize a rep’s queue. These workflows can save substantial coordination time, but they touch customer communication and employee decisions. Begin in preparation mode, expose the selection logic, and expand autonomy only after identity, permission, review, and audit controls have been exercised.

    This sequence is a risk ordering, not a universal maturity model. A read-only weekly analysis is easier to inspect and reverse than an outbound incident campaign. A knowledge proposal has a reviewable artifact. A live automation change affects future conversations, while customer communication may create an immediate and irreversible consequence. Move forward when the evidence and controls for the next class of action are ready, not merely because the previous feature launched.

    Measure the completed loop, not chat activity

    Prompt counts and conversation volume tell you that people opened the product. They do not tell you that customer operations improved. Build the scorecard around the operational loop:

    • Diagnostic quality: Whether the proposed root cause survives expert review, whether its evidence supports the conclusion, and how often factual correction is required.
    • Operational throughput: Time from a detected signal to a reviewed proposal and from an approved proposal to a verified change.
    • Artifact quality: Acceptance, revision, rejection, and rollback patterns for knowledge, automation, configuration, and communication proposals.
    • Customer outcome: Resolution, escalation, repeat contact, and sentiment for the affected topic after the change, interpreted alongside volume and case mix.
    • Safety: Permission denials, attempted out-of-scope actions, failed simulations, unauthorized writes, rollbacks, and missing audit events.
    • Human leverage: Expert time spent collecting evidence, recreating context, drafting the artifact, and reviewing the final proposal.

    Do not make automation rate the only goal. A higher rate can coexist with poor resolutions or avoidable escalations. Treat it as one diagnostic measure and pair it with customer outcomes, correction rates, and topic-level regressions.

    Create an evaluation set from real operating conditions: known content gaps, misconfigured rules, legitimate escalations, sparse attributes, conflicting evidence, localized content, and incidents with precise audience criteria. Give each case an expected outcome, required evidence, allowed tools, and forbidden action. Re-run the set when the model, retrieval system, tool, skill, permissions, or support configuration changes.

    Scheduled work is where the leverage begins to compound. An operator can run recurring analysis and deliver the resulting report without waiting for a manager to remember the question. Keep an owner on every scheduled job, however. That owner should know where failures appear, when the task last completed, which data it used, and how to pause it.

    Key takeaways

    • An operator agent improves the system around customer conversations; it is not simply another customer-facing bot.
    • The product boundary should cover observation, diagnosis, proposal, verification, approval, action, and monitoring.
    • Reliable behavior comes from grounded retrieval, attribute awareness, bounded tools, encoded domain skills, and structured review surfaces.
    • Grant autonomy by consequence: broad freedom to inspect approved data, tighter controls to prepare changes, and explicit approval for production writes.
    • Roll out recurring analysis before knowledge changes, automation configuration, and customer communication unless your own risk profile clearly supports another order.
    • Measure supported diagnoses, accepted artifacts, customer outcomes, human time, and safety events rather than prompt volume alone.

    Your next step is to choose one recurring operational question and write down the evidence it requires, the artifact a good answer should produce, the person who will review it, and the actions the agent must not take. Once that loop works reliably, add one downstream proposal. That is a much stronger foundation for an operator agent than beginning with an open-ended prompt and a broad API key.

    References

  • Our Operating Model Is the Product—Why We Built Product Partners to Accelerate Outcomes

    Our Operating Model Is the Product—Why We Built Product Partners to Accelerate Outcomes

    I’ve learned that customers don’t just buy features—they buy the way we discover, decide, build, ship, and support. In other words, the operating model is the product. That realization has shaped how my team and I at HighLevel translate product strategy into tangible, repeatable outcomes that show up in quality, reliability, onboarding, and consultative support every single day.

    We created Product Partners to codify that operating model and scale it with discipline. It’s a blueprint and operating rhythm that unifies product strategy with go-to-market strategy, customer success, and solutions engineering—so empowered product teams can move faster without sacrificing clarity, governance, or customer trust.

    First, we anchored on continuous discovery. Product trios work shoulder-to-shoulder with customer-facing teams to run customer interviews, journey mapping, and A/B testing, then validate insights with session replay and behavioral analytics. We use driver trees and opportunity solution trees to connect problems to outcomes, ensuring prioritization is evidence-based and aligned to product-market fit—not just output.

    Second, we elevated delivery excellence. Our practices emphasize CI/CD, feature flags, observability, SRE-informed incident management, and DORA metrics to shorten feedback loops while raising the bar on stability. Privacy-by-design, data governance, and regulatory compliance are built into our workflows, and we make deliberate build vs buy decisions to protect platform scalability and long-term velocity.

    Third, we integrated go-to-market alignment from day one. Solutions engineering and customer success shape requirements early, so launches include in-app guides, product tours, onboarding paths, and consultative support that accelerate user activation. We tie outcomes vs output OKRs to stakeholder management rituals, ensuring sales-led and product-led growth motions reinforce each other instead of competing for focus.

    Finally, we closed the loop with a unified analytics platform. Activation, retention analysis, and Net Recurring Revenue (NRR) sit alongside qualitative signals from customer interviews and support. This single source of truth helps us refine product positioning, sharpen value propositions, and improve roadmapping and sprint planning with clear, testable hypotheses.

    What does this mean for our partners and customers? Faster time-to-value, fewer handoffs, clearer expectations, and a shared lens on the metrics that matter. Product Partners isn’t a side program; it’s how we operationalize trust—through transparency, consistent rituals, and a bias toward learning that compounds.

    If this resonates, you’ll feel it in how we discover, build, and support together. I’ll continue to share our playbooks—covering continuous discovery, onboarding, and outcome-based planning—so we can keep raising the standard for product management leadership and product-led growth, one operating rhythm at a time.


    Inspired by this post on Product School.


    Book a consult png image
  • AI-Enabled Enzymatic Recycling: A Product Leader’s Playbook

    AI-Enabled Enzymatic Recycling: A Product Leader’s Playbook

    You have an AI-enabled materials proposal in front of you, a promising set of enzyme candidates, and a difficult decision: fund another round of discovery or start building toward industrial scale. The candidate sequences may be impressive, but they are not yet the product.

    Your decision should turn on whether the full system can repeatedly transform a defined waste stream into usable monomers at an economically viable cost. That framing connects model performance, laboratory evidence, process engineering, and commercial reality before an exciting demonstration becomes a stranded pilot.

    Define the product around recovered monomers

    Only 10% of the plastic manufactured gets recycled. That ceiling is not merely a sorting or consumer-behavior problem. Traditional recycling commonly shortens polymer chains instead of restoring their original molecular building blocks, so the resulting material can lose quality and move toward downcycling.

    Enzymatic recycling changes the intended output. An engineered enzyme can deconstruct a polymer into its original monomers, which can then become inputs for new, high-quality plastic. The difference is fundamental: the product is not processed waste or a smaller plastic fragment. It is recovered molecular feedstock.

    This distinction gives you a better product boundary. A generated protein sequence is a feature. An enzyme that shows activity in one assay is a technical result. The product is a repeatable monomer-recovery system with a defined input, output, operating envelope, and cost structure.

    Before approving a roadmap, require the team to define five contracts:

    • Input contract: Which polymer, packaging format, mixture, and contamination profile will the process accept? “Mixed plastic” is not a specification. Name the included materials and the variation the system must tolerate.
    • Transformation contract: Which polymer bonds must the enzyme break, and what conversion and selectivity must the reaction demonstrate?
    • Output contract: Which monomers will be recovered, what downstream use must they support, and how will the team determine that the output is suitable for that use?
    • Operating contract: What reaction conditions, throughput, energy consumption, and process controls must hold outside a small laboratory assay?
    • Economic contract: Which cost per ton must the integrated process approach, and which assumptions currently separate measured economics from projected economics?

    Selectivity is especially important. An enzyme can target a particular plastic within a mixed waste stream, potentially reducing the need to treat every input as chemically identical. But selectivity does not make an undefined waste stream manageable. The process still needs to know which target material is present, whether the enzyme can reach it, and how the desired products will be recovered.

    Write the product brief in one sentence: For this defined feedstock, transform this polymer into these monomers, within this operating envelope, output specification, and cost boundary. If a number is unknown, leave a visible blank and assign an experiment to fill it. Do not hide the uncertainty inside a broad ambition such as “make plastic circular.”

    Build the AI as a closed learning system

    AI changes the economics of searching enzyme-design space. Protein language models can generate candidates, multi-step agents can coordinate specialized tasks, and computational evaluations can eliminate weak options before scarce laboratory capacity is used. Advances in protein structure prediction have expanded what can be explored, but prediction does not remove the need for physical validation.

    The useful architecture is therefore not a model that emits sequences. It is a closed loop in which every physical result makes the next design round better. Rhea’s Factory combines protein language models, an agentic pipeline, domain constraints, and proprietary wet-lab feedback. The product lesson is broader than any one implementation: generation, evaluation, experimentation, and learning need to operate as one traceable system.

    1. Encode the objective. Convert the product contract into machine-readable constraints: target polymer, desired products, acceptable operating conditions, and the metrics that will decide whether a candidate advances.
    2. Generate candidates. Explore multiple plausible designs rather than optimizing immediately around the first promising family.
    3. Apply computational gates. Reject candidates that violate explicit constraints, preserve the reasons for rejection, and rank the remaining candidates for laboratory use.
    4. Run controlled wet-lab experiments. Test candidates under recorded conditions and capture successes, failures, and inconclusive results.
    5. Update domain predictions. Use the measured outcomes to improve ranking and candidate selection for the next round.
    6. Feed process evidence back into discovery. When a candidate struggles under reactor or feedstock conditions, turn that failure into a new design constraint instead of treating it as a separate engineering problem.

    Agentic AI is valuable here because the workflow is multi-step, not because an agent should make every decision autonomously. At each handoff, define the required input, expected output, validator, and failure behavior. A generation step should not advance an incomplete candidate. A computational score should not be presented as a laboratory observation. A promising assay should not silently become a scale claim.

    Exploration also needs an explicit lane. Higher model-sampling temperatures can produce more unusual enzyme candidates and reach beyond the safest local variations. Controlled model “hallucination” can be useful during candidate exploration when downstream guardrails prevent novelty from being mistaken for evidence.

    Separate the candidate portfolio into three buckets: improvements near known winners, adjacent designs that test a clear hypothesis, and high-variance exploration. Give each bucket a deliberate laboratory budget. Raise sampling temperature only in the exploratory lane, and never allow generated assay values, reaction outcomes, or scale results into the measured-data record.

    The durable advantage sits in the feedback data. In a narrow, high-signal domain, even hundreds of relevant proprietary laboratory observations can support a useful domain prediction model. That is not a general claim that small datasets are always sufficient. It means contextual quality can matter more than indiscriminate volume when the problem, assay, and outcomes are tightly defined.

    For every experiment, preserve enough context to make the result reusable:

    • The enzyme identity, sequence version, and design lineage.
    • The target polymer, material format, mixture, and relevant contamination profile.
    • The assay and protocol version used for the test.
    • The reaction conditions and duration.
    • The measured conversion, selectivity, yield, and uncertainty available from the experiment.
    • The full result, including failure, no-result, and inconclusive outcomes.
    • The relationship between the candidate, computational evaluations, physical test, and model or data release.

    A spreadsheet of winning sequences is not a data moat. A traceable record of why candidates were proposed, how they were tested, what failed, and how each result changed the next decision can become one.

    Use stage gates that end in physical evidence

    AI product teams often gravitate toward a model leaderboard because it creates a clean sense of progress. Enzymatic recycling does not have one adequate master score. A candidate can look structurally plausible and fail in the lab. It can perform in a controlled assay and miss the required throughput. It can convert the polymer and still lose economically once the rest of the process is counted.

    Use a hierarchy of evidence that moves from design compliance to laboratory performance, operating fit, and scale economics:

    GateDecision questionRequired evidenceRed flag
    Design complianceDoes the candidate satisfy the stated target and pipeline constraints?Deterministic checks, recorded constraint evaluations, and candidate provenanceA candidate advances mainly because it appears novel
    Wet-lab performanceDoes the enzyme convert the target with the required selectivity under defined conditions?Repeatable measured observations, including negative and inconclusive runsOnly the best run is retained or shared
    Operating fitDoes useful performance hold within the intended controlled, low-temperature process and throughput requirements?Process measurements tied to reaction conditions, conversion, yield, throughput, and energy useActivity is reported without the process context needed to interpret it
    Scale economicsCan the integrated system move toward cost parity with inexpensive oil-based plastic?A cost and energy model tied to measured inputs, with assumptions and sensitivities exposedCommercial viability is inferred from enzyme activity alone

    Set pass, hold, and stop conditions before seeing the result. Otherwise, an interesting candidate will repeatedly earn one more experiment while the commercial requirement drifts. Relative improvement is useful for learning, but an enzyme that is twice as good as an unusable baseline may still be unusable. Every relative metric should sit beside the absolute requirement it is meant to approach.

    Keep conversion, selectivity, yield, throughput, and energy per ton separate. Combining them too early into a single score can conceal the actual tradeoff. A team should be able to show why it is advancing a faster candidate with lower selectivity, or a more selective candidate with a different operating burden, without claiming that the candidates are equivalent.

    Three common metric substitutions deserve direct scrutiny:

    • Low reaction temperature is not automatically low total energy. Count the energy demands of the complete process rather than the enzyme reaction in isolation.
    • Polymer conversion is not automatically usable monomer recovery. Measure whether the desired output can be recovered to the specification required downstream.
    • Bench performance is not automatically scaled performance. Treat increasing process scale as a new evidence gate, not a routine deployment step.

    My rule is simple: model output can earn laboratory time; only measured process evidence can earn scale capital.

    Plan the roadmap backward from cost parity

    The commercial benchmark is unforgiving. Enzymatic recycling ultimately has to compete with inexpensive oil-based plastic production. A greener reaction that cannot approach a viable delivered cost will remain dependent on special conditions rather than becoming a broadly adopted circular process.

    Build the economic model while discovery is still underway. At minimum, separate these cost lines:

    • Feedstock acquisition, sorting, and rejected material.
    • Preparation required before the enzyme can act on the target polymer.
    • Enzyme production, delivery, useful lifetime, and replacement.
    • Reactor capacity, reaction time, process control, and energy.
    • Monomer recovery and purification.
    • Waste handling, downtime, and variability in plant utilization.

    Do not wait for perfect values. Use ranges, label each input as measured or assumed, and run sensitivity analysis. The purpose is to identify which uncertain variable can kill the business case. If enzyme lifetime dominates cost, another candidate-generation run may be rational. If purification dominates, generating thousands of additional sequences may be a distraction from the real constraint.

    Pair every scientific milestone with an industrial question:

    • Discovery gate: Is activity and selectivity reproducible enough to justify process work?
    • Process gate: Does the candidate perform inside the intended operating envelope rather than only under a convenient assay condition?
    • Feedstock gate: Does performance survive representative material formats and mixtures, including difficult packaging such as clamshells?
    • Demonstration gate: Can the system sustain the required material flow, output quality, and energy profile at a scale that tests the major engineering assumptions?
    • Commercial gate: Does the cost case remain credible when feedstock composition, utilization, throughput, and other sensitive inputs move away from the preferred case?

    A planned 5,000-ton demonstration plant in California illustrates why demonstration capacity belongs on the product roadmap. A plant is not simply a larger laboratory. It tests whether biology, equipment, controls, feedstock variability, and recovery operations behave as an integrated product.

    Before committing meaningful scale capital, ask six kill questions:

    1. Which assumption has the largest effect on delivered cost per ton?
    2. Which inputs are measured, and which still come from a design estimate?
    3. At what physical scale was each important input measured?
    4. What fails first when the feedstock mix changes?
    5. If enzyme performance improves as planned, which downstream step becomes the bottleneck?
    6. Which observed result will stop, narrow, or materially redesign the program?

    Expansion into additional plastics should follow the same discipline. Enzyme selectivity creates a plausible path toward enzyme blends for mixed streams, and new plastic types and mixed-plastic blends remain important development directions. Treat each added polymer as a new product vertical with its own input contract, assays, process interactions, recovery requirements, and economics. A new enzyme is not automatically a low-cost extension of the first process.

    Key takeaways for your next roadmap review

    • Define success as repeatable recovery of specified monomers, not the generation of novel enzyme sequences.
    • Run discovery as a closed loop connecting product constraints, AI generation, computational gates, wet-lab measurements, and process feedback.
    • Treat proprietary experimental context—including failures—as the data asset; candidate count alone is not a defensible moat.
    • Use separate gates for design compliance, laboratory performance, operating fit, and scale economics.
    • Work backward from cost parity and direct the next experiment toward the assumption that most threatens the integrated business case.

    For your next review, ask the team to bring one page containing the input and output contracts, a diagram of the learning loop, the current stage-gate thresholds, the experimental data schema, and a cost sensitivity model with measured and assumed inputs clearly separated. Every roadmap item should change one of those artifacts or produce evidence for a named decision.

    If the team cannot fill those fields yet, that is the immediate product work. The first defensible milestone is one traceable loop from a defined industrial problem through candidate generation, laboratory measurement, and an updated cost model. Repeat that loop with increasing realism before increasing capital exposure. That is how you determine whether programmable biology is becoming an industrial recycling product rather than remaining an impressive AI demonstration.

    References

  • Figma’s Executive Scaling Playbook for IPO Readiness

    Figma’s Executive Scaling Playbook for IPO Readiness

    If your company may pursue an IPO, the tempting move is to wait until the filing window is visible, then recruit executives who have done it before. That sequence solves for credentials. It does not necessarily build the decision system the business will need.

    Figma took a different path. An operator who joined when the company had about 30 people and was not yet charging for its product grew into the CFO role, while the company adopted public-company habits years before its 2025 IPO. The useful lesson for you is not simply to promote insiders. It is to develop executive judgment, operating cadence, and economic instrumentation as one connected system.

    Key takeaways: the playbook on one page

    • Expand a leader’s decision scope before expanding the title. Look for progression from doing the work, to framing the questions, to allocating resources, to improving decisions across the company.
    • Install public-company behaviors before the transaction demands them. Figma began operating this way three years before its IPO, using quarterly rhythms, tighter controls, a close that could withstand scrutiny, and a coherent forward-looking narrative.
    • Treat product, finance, and go-to-market as joint owners of the economic model. They need the same driver tree, definitions, telemetry, and assumptions before they debate pricing or investment.
    • Manage AI investment as a portfolio of explicit bets. Usage, customer value, cost-to-serve, decision triggers, and risks should be visible even when the underlying economics are changing quickly.
    • Use leadership transitions to redraw decision rights. Replacing a departing executive without reconsidering the operating model preserves yesterday’s bottlenecks.

    Scale executive judgment before you scale titles

    Praveer Melwani joined Figma in 2017 as its first business operations and finance hire. The company was still around 30 people and had not begun charging for the product. He became CFO in 2022 and helped lead the company through its IPO in 2025.

    The important pattern is the sequence of work. Early responsibilities included building driver trees, challenging go-to-market assumptions, and establishing the mechanics of board management. Later responsibilities moved toward defining the questions the company needed to answer, directing capital, and shaping the operating cadence. The role grew because the decisions grew.

    You can use that sequence as an executive-readiness ladder. It is more informative than tenure or the seniority of a candidate’s last title.

    Executive modeWork that demonstrates readinessFailure signal to watch
    OperatorBuilds the model, tests assumptions, and makes the basic process reliable.Produces accurate work but cannot explain which decision it should change.
    Question-setterIdentifies the uncertainty that matters, frames options, and defines success.Waits for the founder or another executive to determine what deserves attention.
    AllocatorConnects product evidence, financial constraints, and strategic upside to resource choices.Treats the budget as a fixed entitlement rather than a set of revisable bets.
    System leaderImproves the cadence, decision rights, narrative, and judgment of the wider team.Remains the indispensable reviewer for every important decision.

    Do not promote someone merely because they are excellent in the first row. Give them work from the next row and observe what happens. Ask a strong operator to frame an ambiguous company problem, recommend where resources should move, document the trade-offs, and run the decision through the relevant functions. You are testing whether the person can create clarity beyond the boundaries of the original role.

    I use a similar first-principles test when evaluating a prospective VP, especially in a function the founder does not know deeply:

    • Can the candidate map how the business creates and captures value?
    • Can they define success metrics and show where those metrics could mislead the team?
    • Can they explain a meaningful trade-off in plain language?
    • Can they describe the team and decision system they would build, rather than only the work they would personally perform?
    • Can they teach the executive team something useful in 30 minutes?

    Run this test on your actual business context, not a generic case interview. A candidate who asks sharper questions, exposes a hidden assumption, and improves the decision has demonstrated more than someone who recites the standard playbook from a previous employer. Prior experience still matters, but learning velocity and expanding scope deserve more weight than familiarity alone.

    Make IPO readiness a company cadence, not a finance workstream

    Figma began behaving like a public company three years before its IPO. That is not a universal countdown for every company. The more important point is the order of operations: the habits came before the event that would test them.

    Late preparation forces teams to create controls, reconcile definitions, improve forecasting, and construct a credible narrative while the stakes are already high. Early preparation turns the same work into ordinary management. It also reveals weak ownership and unreliable data while the company still has room to correct them.

    • Quarterly operating rhythm: Review changes in the business drivers, the assumptions behind the forecast, the resulting resource choices, and the risks that could alter the plan. A performance presentation without a decision is reporting, not an operating review.
    • Close and controls: Make ownership, evidence, access, and material judgments explicit. The goal is not bureaucracy for its own sake. It is to produce numbers that leaders can use without reopening the entire chain of custody every time.
    • Forward-looking narrative: Connect past performance to the decisions now being made. Explain what changed, why management believes it changed, what will be done next, and what evidence would invalidate that view.
    • Decision record: Preserve the assumptions, alternatives, owner, and follow-up trigger behind a material choice. This prevents the company from rewriting the reasoning after the outcome is known.

    This discipline can accelerate decisions because product, finance, and go-to-market stop renegotiating the basic facts in every meeting. Product brings evidence about behavior and roadmap alternatives. Finance brings the model, sensitivities, and constraints. Go-to-market brings customer context, commercial implications, and execution dependencies. The executive owner makes the cross-company choice and records what would cause it to change.

    A useful quarterly decision packet should answer the following questions:

    <!– wp:list {
  • No More Accidental Agents: How We Engineered Global Agent’s Helpful, Curious Personality

    No More Accidental Agents: How We Engineered Global Agent’s Helpful, Curious Personality

    Most teams ship AI agent personalities by accident—emergent quirks, brittle prompts, and uneven behavior. We refused to let that happen. From day one, we treated personality as a first-class product surface, one that should be designed, instrumented, and iterated with the same rigor as any core capability.

    Learn how we designed Global Agent’s personality and fine-tuned its inquisitiveness and helpfulness using Agent Analytics.

    In my role leading product at HighLevel, Inc., I framed our approach around agentic AI and conversation design: personality is not “flavor text”; it is the control system for how an agent interprets context, asks questions, and decides when to act. Our product strategy prioritized clarity, empathy, and consistency—so the agent would be curious enough to resolve ambiguity without becoming interrogatory, and helpful enough to move work forward without overstepping.

    We made that intent measurable. Using behavioral analytics, we defined operational signals such as clarification-question rate, resolution-path efficiency, and escalation quality. We combined eval-driven development with targeted A/B testing to compare prompt patterns and tool strategies, ensuring each change had a clear hypothesis and measurable outcome.

    To calibrate inquisitiveness, we mapped decision points where the agent should ask follow-ups versus proceed autonomously. Prompt engineering codified those thresholds, while a retrieval-first pipeline reduced unnecessary questions by improving context completeness up front. When the agent did ask, we constrained tone and cadence to keep queries concise, respectful, and progress-oriented.

    To enhance helpfulness, we prioritized precise action-taking and unambiguous guidance. Context window management preserved relevant facts without diluting intent, and guardrails aligned with AI risk management principles ensured the agent stayed within policy, privacy, and compliance boundaries. The result was an assistant that resolved more tasks end-to-end, with fewer stalls and clearer handoffs when human help was warranted.

    Agent Analytics became our nervous system. We instrumented every dialog turn to attribute outcomes to design choices, then used driver trees to connect micro-behaviors to macro results like time-to-resolution and customer satisfaction. This closed-loop view let us ship confidently, knowing which levers improved helpfulness, which sharpened curiosity, and which merely added noise.

    Process mattered as much as tooling. Product trios ran continuous discovery with customers to surface edge cases—ambiguous intents, multi-intent turns, and sensitive scenarios—while our engineering partners operationalized experiments with clean rollback paths. We favored small, testable changes over sweeping rewrites, building momentum and trust with each iteration.

    The payoff is a personality that feels consistent across use cases: curious when clarity is missing, decisive when action is obvious, and transparent when limits are reached. Users experience fewer dead ends, faster resolutions, and a brand voice that shows up the same way every time—because it was defined, measured, and improved on purpose.

    If you’re building agentic AI, don’t leave personality to chance. Treat it like a product: set clear outcomes, instrument deeply with Agent Analytics, and iterate with eval-driven development and A/B testing. That’s how curiosity becomes a feature, helpfulness becomes a habit, and your agent becomes reliably, intentionally excellent.


    Inspired by this post on Amplitude – Best Practices.


    Book a consult png image
  • Knowledge Management for AI Sales Agents: A Practical System

    Knowledge Management for AI Sales Agents: A Practical System

    Your AI sales agent answers the pricing question, then recommends the wrong plan. It identifies a promising buyer, then sends the conversation to the wrong queue. If the underlying facts, decision rules, or routing policy are missing, another prompt adjustment cannot fix the problem.

    You need a knowledge operating system, not a larger folder of sales collateral. The goal is to give the agent the smallest reliable path from a buyer’s question to an accurate answer, an appropriate recommendation, useful qualification, and the correct next step.

    Key takeaways

    • Separate product facts, decision guidance, and execution policy. Each solves a different part of the sales conversation.
    • Turn long documents into focused, approved knowledge records with an owner, scope, effective date, and explicit boundaries.
    • Launch a complete sales motion for a narrow set of buyer intents before trying to document everything.
    • Test recommendations, qualification, routing, and escalation behavior, not just whether the agent can repeat a correct sentence.
    • Convert unanswered, incorrect, and disengaged conversations into a managed improvement queue.

    Design for answering, recommending, qualifying, and routing

    A conventional knowledge base helps someone find information. An AI sales agent has a harder job: it must interpret a buyer’s situation and decide what to do with the information it retrieves.

    A language model does not inherently know your current plans, qualification criteria, commercial boundaries, or customer-specific use cases. That context is unique to your business and must be made explicit. Fluency cannot compensate for a missing policy.

    I find it useful to divide sales knowledge into three layers:

    Knowledge layerWhat it containsWhat the agent should do with itTypical failure when it is missing
    Product factsPricing, plan structure, capabilities, limitations, availability, and supported use casesGive a direct, accurate answerThe agent guesses, gives a vague response, or repeats obsolete information
    Decision guidancePlan-fit logic, relevant constraints, case studies, approved comparisons, and the context behind product factsExplain which option fits and whyThe answer is technically correct but does not help the buyer decide
    Execution policyQualification questions, required fields, routing conditions, escalation rules, and actions the agent may takeAdvance the conversation within defined authorityThe agent collects irrelevant details, makes an unsupported commitment, or routes the buyer incorrectly

    This distinction exposes why uploading product pages is not enough. Public pages and product documentation are useful starting points, but a capable inbound motion also needs FAQs, pricing explanations, case studies, competitive material, qualification criteria, and internal sales guidance.

    Facts answer, “What does the product do?” Decision guidance answers, “Is this appropriate for my situation?” Execution policy answers, “What should happen next?” Audit your knowledge against all three questions.

    Set a hard boundary around commercial exceptions. The agent should not infer an unlisted discount, invent a contractual commitment, or turn an internal hypothesis into a buyer-facing claim. It should state the approved terms, gather the information required by policy, and route the exception to an authorized person. A plausible but unauthorized promise can create financial and legal exposure.

    Turn scattered documents into governed knowledge

    Build sales-ready knowledge records

    A long document can be correct and still be poor input for an agent. Pricing may be buried below an obsolete introduction. A feature table may omit the condition that changes plan fit. A battlecard may combine approved facts with a rep’s unverified notes.

    Convert those documents into focused records. Each record should cover one buyer intent or one tightly related decision. Use a consistent template:

    • Buyer intent: The question or decision this record addresses, including common alternative phrasings.
    • Approved answer: The direct response the agent may give without qualification.
    • Decision context: Why the fact matters and when it changes the recommendation.
    • Constraints and exceptions: What the answer does not cover, including conditions that require clarification.
    • Next question or action: The appropriate follow-up, qualification step, route, or escalation.
    • Scope: The plans, markets, customer types, channels, or agents allowed to use the record.
    • Evidence location: The canonical product, pricing, or policy record from which the answer was derived.
    • Owner and approver: The people accountable for accuracy and authorization.
    • Lifecycle metadata: Effective date, review status, and whether the record replaces an earlier version.

    For example, a record about plan fit should not stop after naming a plan. It should state the relevant requirement, identify the condition that changes the answer, give the agent an approved follow-up question, and define where to route a buyer whose situation falls outside the standard policy. The recommendation then becomes reproducible rather than improvised.

    Keep buyer language in the record. Prospects rarely use your internal taxonomy, and the same intent may appear as a product question, an outcome question, or a comparison. Alternative phrasing helps the retrieval layer recognize that these expressions belong to the same approved answer.

    Create an authority hierarchy

    A centralized repository is valuable only if the agent can distinguish current authority from historical residue. Define the hierarchy before connecting more content:

    1. Designate one canonical record for each product fact, commercial rule, or routing policy.
    2. Make approved sales explanations point back to that record rather than becoming independent versions of the truth.
    3. Treat scripts and examples as phrasing aids unless they are explicitly approved to carry facts.
    4. Keep drafts, call notes, chat fragments, and retired material outside the agent’s usable knowledge until they are reviewed.

    Do not let the agent reconcile conflicting records by choosing the newest upload or blending the language. When two approved items disagree, the safe behavior is to withhold the disputed claim, follow the defined escalation path, and send the conflict to its owner.

    Ownership should also be specific. A knowledge owner maintains the record. A domain approver authorizes sensitive claims. An operations owner monitors how the agent uses the record in conversations. One person may hold more than one role, but every role needs a name rather than a department-shaped placeholder.

    Target knowledge by audience and action

    Internal knowledge is not automatically buyer-facing knowledge. A qualification score may guide routing without being disclosed. A battlecard may help frame an approved comparison without exposing internal commentary. A security question may require an authorized answer rather than the agent’s summary of a sales note.

    Mark each record as buyer-answerable, decision-only, action-only, or restricted. Then expose only the appropriate material to each agent, channel, market, and sales motion. This kind of centralization and content targeting reduces duplication while keeping internal policy separate from the words a prospect sees.

    Launch the smallest complete sales motion

    Trying to clean every sales document before launch creates a long project with no conversational evidence. Launching with disconnected FAQs creates a different failure: the agent answers isolated questions but cannot move the buyer forward.

    The better unit of scope is a complete motion for a bounded set of intents. For each selected intent, the agent needs an answer, the relevant fit logic, the next qualification question, a route or resolution, and an escalation path.

    Prioritize by demand and consequence

    Start with questions that appear repeatedly, delay buyers, consume sales time, signal meaningful intent, or cause material damage when answered incorrectly. Pricing, plan differences, core capabilities, common use cases, qualification, and routing are natural candidates when they dominate your actual inbound conversations.

    Two simple calculations help quantify repetitive work: team time reclaimed = average response composition time x question frequency, while buyer wait avoided = number of prospects asking x average response time. These calculations are useful for prioritizing knowledge work, but neither should be presented as revenue without downstream evidence.

    Add consequence to the ranking. A frequent low-risk question may save time, but an infrequent error involving price, eligibility, security, or a contractual promise may deserve earlier treatment. Frequency tells you where the volume is. Consequence tells you where control matters.

    Your initial release is ready when the selected motion has:

    • Approved answers for the recurring and commercially important questions in scope.
    • Plan-fit or use-case guidance where the buyer needs a recommendation rather than a fact.
    • Explicit qualification fields and follow-up questions.
    • Routing rules for the standard paths.
    • A visible no-answer and human-escalation path.
    • Owners and lifecycle metadata for every active record.
    • A representative evaluation set based on real buyer language.

    Test the conversation, not the sentence

    A retrieval test can show that the right paragraph was found. It cannot show that the agent asked the necessary follow-up, respected a restriction, or routed the lead correctly. Evaluate the complete interaction.

    Your test set should include direct questions, paraphrases, multi-part questions, ambiguous requests, outdated assumptions, missing qualification details, requests for exceptions, and scenarios that should be escalated. For each case, check:

    • Is every factual claim aligned with approved knowledge?
    • Does the response answer the buyer’s actual question before adding detail?
    • Does the recommendation apply the right conditions rather than matching a keyword?
    • Does the agent ask only for information required by the qualification policy?
    • Does it avoid unsupported commitments and internal-only language?
    • Does the final route or escalation match the execution rule?

    Record pass or fail at the behavior level and attach a reason to every failure. Do not allow a strong average score to conceal a severe pricing or policy error. High-consequence failures should block that behavior from release until the knowledge or policy is corrected.

    Use a controlled launch with an obvious human path. The point is to begin collecting real conversational evidence early, not to claim autonomy before the boundaries are reliable. Fast deployment and continuous iteration work when the feedback loop is designed before traffic arrives.

    Run the knowledge flywheel from real conversations

    Once the agent is live, conversation failures become your most useful knowledge backlog. Do not place every poor result under a generic label such as bad answer. Classify the mechanism so the right owner can fix it.

    • Coverage gap: No approved record addresses the buyer’s intent.
    • Retrieval failure: The right knowledge exists, but the wrong record was selected or the right one was missed.
    • Freshness failure: The agent used information that should have been replaced or retired.
    • Guidance failure: The fact was correct, but the recommendation ignored relevant context.
    • Qualification failure: The agent skipped a required question, collected unnecessary information, or misread the answer.
    • Routing failure: The collected information was correct, but the next action did not follow policy.
    • Boundary failure: The agent disclosed restricted material or made an unauthorized claim.
    • Conversation failure: The content was accurate, but the response was unclear, repetitive, or poorly sequenced.

    Turn each confirmed failure into a work item containing the conversation, intent, root cause, affected knowledge record, accountable owner, proposed change, and regression test. A change is not complete when the wording is edited. It is complete when the test passes, the approved version is published, and conflicting text is retired.

    Track a small set of operational signals that lead to decisions:

    SignalWhat it tells youWhat to do with it
    CoverageWhich eligible buyer intents have an approved answer and action pathAdd knowledge where demand and consequence justify it
    Reviewed correctnessWhether sampled claims match the approved recordRepair facts, retrieval, or response generation
    Knowledge conflictsWhere active records disagree or overlap ambiguouslyResolve authority and retire obsolete material
    Qualification completionWhether required information was collected for eligible conversationsImprove questions, field definitions, or sequencing
    Routing complianceWhether the next action matched the approved ruleCorrect policy logic or integrations
    Buyer progressionWhether the conversation reached its intended next stepInspect guidance and friction, then validate changes against downstream outcomes
    Content healthWhich active records lack an owner, approval, scope, or lifecycle statusRepair governance before stale content becomes a live failure

    Review these signals by intent and sales path. A global average can look acceptable while one plan, market, or routing branch fails repeatedly. Conversion can be a useful downstream outcome, but it is not proof that a knowledge change caused the result. Traffic mix, offer changes, seasonality, and human follow-up can also move it. Use controlled comparisons where practical and pair outcome data with conversation-level review.

    Reserve a recurring weekly block for unanswered questions, disengaged prospects, high-consequence errors, and unresolved conflicts. Process pricing, product, and policy changes as immediate knowledge events rather than waiting for the review block. This is where knowledge management becomes an operating responsibility instead of a cleanup project.

    The compounding effect comes from the loop: approved knowledge improves conversations; conversations reveal gaps; each repaired gap becomes a reusable capability. That is why small, well-chosen content improvements can have effects beyond the original conversation.

    Start with your recent inbound conversations. Choose the most repeated unanswered question, the most consequential incorrect answer, and one routing failure. Convert each into an owned knowledge record, add a regression test, and release only the behavior that passes. That small loop is the foundation of a sales agent you can trust with progressively more of the funnel.

    References

  • I Pointed a “Ralph Wiggum” AI Loop at My Product for a Week—The Data That Stopped Chaos

    I Pointed a “Ralph Wiggum” AI Loop at My Product for a Week—The Data That Stopped Chaos

    I spent a week pointing a "Ralph Wiggum loop" at my product to see how far an agentic AI could take pragmatic, everyday improvements without human micromanagement. It was equal parts exhilarating and nerve-wracking. The short version: the loop moved fast and broke assumptions, but Amplitude analytics kept it from going off the rails—and turned chaos into controlled acceleration.

    By "Ralph Wiggum loop," I mean a deliberately naive, endlessly curious cycle: try something small, ship it behind a flag, watch the data, then try again. It is the product equivalent of a fearless intern who experiments constantly. That energy is invaluable for discovery, but it absolutely demands strong guardrails and a clear definition of success.

    Before I started, I framed the outcomes I cared about: user activation within the first session, reduction in time-to-value, and early retention indicators. I set baselines and a minimum detectable effect (MDE) for A/B testing so the loop could distinguish noise from signal. I also documented a driver tree of behaviors we wanted to influence and ensured every event was cleanly instrumented in Amplitude analytics to support reliable behavioral analytics.

    The guardrails mattered most. I put every change behind feature flags with instant rollback. I defined "off the rails" conditions upfront, including regression thresholds for activation and retention analysis, and enabled anomaly detection to surface unexpected spikes or drops. Session replay was ready to diagnose confusion fast, and I kept a daily evaluation cadence so the loop never ran unattended for long.

    Day by day, the loop proposed micro-experiments: onboarding copy variants, tooltip timing, in-app guide sequencing, and subtle changes to progressive disclosure. Each iteration shipped behind a flag to a small cohort. I watched leading indicators in real time, then zoomed out to cohort views to guard against short-term gains that might erode longer-term value. When something looked promising, we expanded exposure methodically; when something looked risky, we paused immediately.

    We had a pivotal moment where the loop suggested a bolder call-to-action that spiked activation. On the surface, it looked like a win. Amplitude cohorts told a fuller story: downstream engagement softened, and anomaly detection flagged a pattern that hinted at premature conversion rather than genuine intent. A quick rollback through feature flags saved the week—and reminded me why eval-driven development should be the default for agentic AI workflows.

    The most surprising part was how quickly the loop unlocked small compounding gains once the measurement scaffolding was in place. With a unified analytics platform and crisp guardrails, the system became a safe sandbox where the AI could explore aggressively while we stayed anchored to outcomes. The combination of behavioral analytics, A/B testing discipline, and daily human review turned raw speed into durable learning.

    My takeaways are direct. Agentic AI can accelerate discovery, but only if you define stop conditions and wire strict feedback loops into your stack. Measurement is product strategy here—without it, you get noisy activity instead of progress. Invest in instrumentation first, treat feature flags as non-negotiable, and let anomaly detection and session replay be your early warning system. Most of all, tie every experiment to activation, engagement, or retention, not vanity metrics.

    If you’re considering your own week with a "Ralph Wiggum loop," start painfully small, constrain the blast radius, and insist on decision-quality data. Do that, and you’ll turn a chaotic agent into a compounding engine for product discovery—one that moves fast, learns faster, and stays on track.


    Inspired by this post on Amplitude – Perspectives.


    Book a consult png image
  • From Vision to Execution: Building Agentic, Data‑Driven Products with Real‑World Rigor

    From Vision to Execution: Building Agentic, Data‑Driven Products with Real‑World Rigor

    When I consider where product development is headed, one statement captures the mandate perfectly: "Eric Carlson is a Principal AI Engineer helping to shape and build Amplitude's next generation vision of of agentic and data driven product development." That vision resonates deeply with how I lead teams—anchoring strategy in behavioral analytics while enabling agentic AI to act on insights with speed, safety, and measurable impact.

    Translating that vision into execution starts with clarity of outcomes. I frame driver trees that connect customer value to leading indicators—activation, engagement depth, and retention—then instrument product telemetry with Amplitude analytics and behavioral analytics to surface the moments that matter. From there, we operationalize learning with A/B testing and feature flags, ensuring each hypothesis gets a fair, observable run and that we can safely ramp what works.

    Agentic AI changes the operating model. Instead of static dashboards, we design autonomous workflows that observe signals, reason over context, and take action—grounded in a retrieval-first pipeline and governed by eval-driven development. For product managers, this demands fluency with LLMs for product managers and practical prompt engineering, plus rigorous AI Strategy around data governance, privacy-by-design, and risk scoring so agents remain trustworthy under real-world conditions.

    Cross-functional cadence is everything. I partner closely with Principal AI Engineers and product trios to blend continuous discovery with execution: rapid user interviews to reveal intent, opportunity solution trees to prioritize, and outcomes vs output OKRs to align incentives. The result is a system where insights are unified, decisions are explainable, and agents improve through tight feedback loops across analytics, experimentation, and production telemetry.

    If you’re building toward an agentic, data-driven future, invest in a unified analytics platform, shorten the path from signal to action, and measure learning velocity as carefully as feature delivery. With the right foundations, agentic AI becomes more than a feature—it becomes a force multiplier for product strategy, customer value, and sustainable growth.


    Inspired by this post on Amplitude – Perspectives.


    Book a consult png image
  • From Prototype to Production: How I Built Reliable AI-Generated Opportunity Solution Trees

    From Prototype to Production: How I Built Reliable AI-Generated Opportunity Solution Trees

    I just wrapped an all-out engineering sprint. That still sounds odd coming from me, because while I’ve written code on and off for years, I don’t self-identify as an engineer. I’m a product manager who used to be a designer. It’s been a long time since I wrote code for a living.

    But AI has expanded what’s just now possible—for our products, and for us. It’s pushed me to do more than I imagined. In that spirit, I want to share a recent engineering story. It includes technical details, and a year ago I couldn’t have done any of it. I learned it with the help of AI, and my aim is to show what’s now within reach.

    I’ve been building two services with a partner at Vistaly: AI-generated interview snapshots and AI-generated opportunity solution trees. We put out a call for alpha partners, received over 100 applicants, and selected eight design partners to start.

    Opportunity Solution Tree diagram with a blue Desired Outcome branching to green Opportunity nodes, yellow Solution nodes, and orange Assumption Tests for product discovery and AI workflows.
    A clear, color‑coded map from desired outcome to opportunities, solutions, and assumption tests—showing how to structure discovery work and prompt AI to generate, compare, and validate product ideas.

    Each team uploaded three customer interviews. I identified the key moments and opportunities and then generated an opportunity solution tree from those snapshots. I provide the AI services; Vistaly is building the UI and workflows around them.

    Early feedback was strong. Teams immediately asked to upload more interviews—exactly the kind of demand signal you hope to see—so we got to work making that possible.

    Dark interface screenshot of an opportunity solution tree with colored cards and dotted connectors, showing merged, moved, and evidence-added Opportunity notes about onboarding, support, and bot readiness.
    Go behind the scenes as AI turns raw feedback into a clear Opportunity Solution Tree. Linked cards reveal user needs—onboarding, support offload, and bot-readiness signals—so product teams can spot priorities and next steps at a glance.

    Updating an opportunity solution tree with new interview content is far harder than generating a new tree from scratch. I initially underestimated the complexity. Our goal wasn’t to produce a tree and declare it truth. We wanted teams to engage, correct, and collaborate with the AI—scaffolding cross-interview synthesis instead of doing it for them.

    To support that, we needed a way to communicate precisely how a tree would change after new interviews were added. We took inspiration from git diff and set out to build the equivalent for opportunity solution trees—step-by-step change sets that explain each proposed modification.

    Diagram of an opportunity solution tree with an Outcome node pointing to Opportunity A and Opportunity B; B branches to child opportunities and shows source evidence, labeled “Updates Can't Result in Data Loss.”
    A clear visual of AI‑generated opportunity solution trees: outcomes feed opportunities that branch into sub‑opportunities, while evidence is preserved. The structure ensures updates stay traceable and never cause data loss.

    That decision was right, but the lift was larger than I expected. It wasn’t enough to generate an updated tree; I also had to provide a clear, ordered walkthrough of what changed and why.

    I often see the same pattern with AI: it’s easy to get to an impressive prototype, but much harder to reach a production-grade product. That was exactly my experience here. My service actually comprised two sub-services: generating a new tree from scratch and updating an existing tree with new interviews. The first worked well in alpha; the second had to be built before anyone could add a fourth interview.

    Opportunity Solution Tree diagram: teal Outcome links to Opportunities A and B; Opportunities C and D branch under B; right panel lists the change set steps for adding nodes.
    Explore how an outcome expands into an Opportunity Solution Tree: Opportunities A and B stem from the goal, with C and D nested under B, while a concise change set tracks every node added along the way.

    On the surface, these services look similar. In reality, updates must preserve existing structure unless new evidence requires a change. You have to account for compound operations—merges, splits, deletes—while guaranteeing no data loss. Every node has source opportunities (supporting evidence from interviews) and children (tree sub-opportunities), and neither can be dropped.

    In classic AI fashion, I got a reasonable version working in a few days and shipped it to our design partners. One team quickly hit our beta limits and asked to convert to a paid subscription so they could keep going. They showed a willingness to pay, converted, and started uploading aggressively.

    Diagram of an Opportunity Solution Tree showing how parent 'Opportunity A' with children x, y, z is split into 'Opportunity A' and 'Opportunity B' to reassign evidence and connections.
    Watch an Opportunity Solution Tree evolve: the original parent A with x, y, z branches is split into A and B, shifting evidence while preserving links—mirroring how AI refines scope and structure in discovery.

    At the 14th, 15th, and 16th uploads, the cracks appeared. We saw odd behavior in some trees. The Vistaly team noticed that the change sets—the step-by-step instructions emitted by my service—didn’t always reconstruct the final tree my service also emitted. We needed those steps to match exactly, so teams could review and accept, modify, or reject each change with confidence.

    They flagged the issue the day I was flying to New Orleans for Jazz Fest. In hindsight, I’m glad I didn’t grasp the scope of what awaited me. I had roughly 80% of the work still to do to make tree updates rock solid. At least I got to enjoy the music first.

    Flowchart merging two opportunity solution trees: Opportunity B with children y and z, and Opportunity C with t, u, v, consolidated into one tree led by Opportunity C connected to five child opportunity nodes.
    From fragments to focus: this diagram shows how Opportunities B and C are merged into a single Opportunity Solution Tree, removing duplicates and unifying context so AI can rank and explore five related opportunities with clarity.

    Back home, I started diagnosing. My service was a pipeline: several LLM-driven steps followed by deterministic code to compare trees and produce change sets. As I dug in, I realized that approach was flawed. Tree diffs, unlike linear document diffs, are ambiguous.

    In a document, if I add a sentence, the diff shows an addition. If I delete a paragraph and rewrite it, the diff shows a removal and an addition. Simple. But trees are different. Suppose I split opportunity A into A and B, and later merge B with C. The split can disappear from the final diff.

    Diagram of an opportunity solution tree labeled 'Input Tree' showing an Outcome node branching to Opportunity A and C, each with child nodes x-z and t-v, with arrows indicating hierarchy.
    Peek inside our process: a simple opportunity solution tree maps an outcome to prioritized opportunities A and C with downstream options x-z and t-v. A clear snapshot of how AI organizes product discovery.

    When the model splits an opportunity, it must distribute A’s source opportunities and children between A and B. For instance, if A has source opportunities 1, 2, 3 and children x, y, z, after the split A might keep 1, 2, and x, while B takes 3, y, and z.

    Now suppose the model merges B into C. If C originally had source opportunities 4 and 5 and children t, u, v, then after the merge C now has source opportunities 3, 4, 5 and children t, u, v, y, z. When you compare the original and final trees, it looks like A somehow donated some evidence and children directly to C. The split and merge that explain why are invisible to a naive diff.

    Opportunity Solution Tree diagram titled Output Tree: a blue Outcome node branches to green Opportunity A and Opportunity C, which expand to nodes x-v with arrows; Product Talk badge.
    See how an AI-generated Opportunity Solution Tree unfolds: one Outcome flows to Opportunities A and C, then into options x–v. Clean colors and arrows reveal the hierarchy from goal to opportunities at a glance.

    That was the core insight: we didn’t just need to show what changed—we needed to show why it changed. I had to reconstruct each move step-by-step. That meant getting the model to show its work, which opened a new can of worms.

    I refactored my prompts so the model produced both the final output and the exact change set it used to get there. The action language was explicit: add, delete, reframe, merge, split, and so on. Crucially, I asked the model to describe its moves in user-meaningful terms—“split A into A and B, then merge B into C”—not as opaque reassignments of sources and children.

    Diagram of an AI-generated Opportunity Solution Tree: blue Outcome node with children Opportunity A and Opportunity B; B branches to Opportunity C and D. A right-hand list shows the change set for each step.
    Watch an opportunity solution tree take shape: start with the outcome, add opportunities A and B, then extend B to C and D. The paired change set makes every edit transparent—ideal for AI-assisted product discovery.

    For each LLM step, the model now emitted its recommendation and the corresponding change set. This helped, but it wasn’t perfect. After extensive testing and error analysis, two classes of errors emerged: (1) the model attempted an invalid move, and (2) the change set didn’t actually generate the recommendation.

    Category 1 felt like designing a game while the model played it creatively. For example, what happens when the model tries to merge a parent with a child? If opportunity A has children B, C, and D and the model merges A with B, the merge is directional. If the instruction is “keep A, delete B,” that works—the parent absorbs the child. But if the instruction is “keep B, delete A,” then C and D become orphans. These puzzles were solvable and even fun.

    Diagram of Opportunity Solution Tree merge rules: merging node B into parent A is allowed, while merging A into B is not because it would orphan opportunities B, C, and D.
    Visual explainer from Product Talk on AI-generated Opportunity Solution Trees. It contrasts an allowed merge (B into A) with a not-allowed merge (A into B) that leaves child opportunities orphaned, guiding safe hierarchy edits.

    Category 2 was harder. Despite prompt iterations, I could only push the discrepancy rate down to about 1 in 40 instances. With 10–20 LLM calls per run, that meant roughly half of all runs still failed. Not acceptable for production. I hit a wall. A paying customer was waiting, and more design partners were queued up.

    Next, I tried to correct the model’s mistakes with deterministic code. I had promised that my change sets would generate the output tree, so I wrote verifiers: detect conflicts (e.g., delete a node, then try to use it later), guard against data loss, prevent orphaned nodes, and more. Detection was straightforward; correction was not. Fixing issues required guessing the model’s intent. If the sequence said “delete A, then merge A with B,” should I remove A entirely or salvage A’s sources and children by merging into B? There were dozens of such cases with no unambiguous answer.

    Workflow diagram titled 'My Simple Repair Loop' showing an iterative validation cycle: Generate the Change Set → Run the validation tool → Check Result, with branches to retry on failure or exit on pass.
    A step-by-step loop shows how changes are validated: generate a change set, run a validation tool, review the result, then repeat on failure and exit on pass—mirroring iterative work behind AI-built Opportunity Solution Trees.

    After 11 straight days of deep work—including weekends—I was exhausted. I dislike hustle culture; this isn’t how I design my life. But I was stuck, and then I had an insight.

    On a walk with my husband (also an engineer), I realized I could have the LLM repair its own mistakes. My data contract with Vistaly requires that the change set must generate the output tree. I had already built robust validation code. I knew exactly when a change set failed—and why. No amount of prompt tuning alone was fixing it. So I turned the validator into a tool for the model and created a simple agentic loop.

    The loop works like this: the model proposes a change set, calls the validation tool, and gets back a pass/fail plus specific feedback. If it fails, the model uses those instructions to repair the change set and calls the tool again. Iterate until success or a max number of turns.

    I prototyped in Node.js with a single model call, a verifier pass, and a repair attempt. At first, the loop didn’t converge—it just accumulated compute. I experimented with how to communicate errors, how much context to include, and how to sequence feedback. Eventually, it clicked: the model began fixing its own mistakes and typically returned a valid change set in one or two repairs. It was, in practice, eval-driven development applied to LLM outputs.

    I had already built an agent loop utility for another AI workflow, so I productionized quickly: model call, optional tool invocation, tool result returned to the model, repeat until the validator signals success or the loop times out. I integrated the new loop into the pipeline and shipped the revamped service to Vistaly on Monday at noon. They’re integrating now, and it will be in the hands of our design partners shortly. I’m relieved—and ready for a day off.

    Reflecting on the last two weeks, a few things stand out. First, I shed limiting beliefs about being an engineer. To make this reliable, I had to solve legitimately hard problems, and that feels good.

    Second, this was genuinely fun. Designing the action set and watching the model push those boundaries was like working through elegant puzzles. Models are incredibly creative, and harnessing that creativity with the right constraints is deeply satisfying.

    Third, I learned when I can and can’t trust Claude to write code for me. Since Opus 4.6 came out, I gave Claude a much longer leash. After the past two weeks, Claude is back on a short leash. I found a lot of gaps in my implementation in areas where I simply trusted that Claude got it right, when in fact it didn’t. If you don’t have the right infrastructure—planning, testing, code review—this can be disastrous. I’ll be investing more here and sharing what I learn.

    Finally, if this work had been spread over two months, it would have been thoroughly enjoyable. I’m discovering how much I like being an AI engineer. It feels like a new chapter where I can combine opportunity solution trees with modern AI engineering—and deliver real value to product teams doing continuous discovery.

    I’m excited to share more of what we’re building with Vistaly and to onboard more design partners soon. If you’re interested, get on the waiting list. And if you’ve been hesitant to stretch beyond your current skill set, I hope this story nudges you to take the first small step toward what’s just now possible.


    Inspired by this post on Product Talk.


    Book a consult png image
  • AI-Assisted Product Strategy: A Practical Operating System

    AI-Assisted Product Strategy: A Practical Operating System

    You can get an AI model to produce a roadmap in minutes. That is precisely the problem. A polished roadmap can hide weak evidence, unresolved trade-offs, and a strategy that never made a real choice.

    The useful question is not whether AI can do product management work. It is where AI should accelerate the path from evidence to decision, where human judgment must remain explicit, and how you will know the resulting strategy is working. The operating system below gives you that separation.

    Key takeaways

    • Give AI a defined role in the decision process. It can extract, organize, challenge, and draft; the product leader still owns choices, trade-offs, and commitments.
    • Build a strategy chain from customer problem to business result before asking AI for initiatives. Otherwise, the model will fill strategic gaps with plausible language.
    • Ground every workflow in canonical product context, and require every important claim to point back to evidence.
    • Use AI to shorten discovery synthesis, not to turn a limited set of interviews or support conversations into false market certainty.
    • Carry the same strategic hypothesis through the roadmap, experiment, launch, and learning review. Changing the success definition between those stages makes measurement meaningless.

    Start with decision architecture, not a better prompt

    Most weak AI-assisted strategy work begins with an underspecified request: analyze this feedback, prioritize these ideas, or build a roadmap. The model responds by making silent assumptions about the customer, the business objective, and the meaning of priority. Its output may read well while answering a question nobody deliberately chose.

    Write a decision brief before opening the model. This is not a conventional product requirements document. It is a compact contract defining the decision AI is helping you make.

    • Decision: State the choice in one sentence. For example, decide which onboarding opportunity deserves discovery capacity in the next planning cycle.
    • Target customer and context: Name the segment, job, and situation. Feedback from an administrator configuring an account should not be blended with feedback from an end user completing a daily task.
    • Desired outcome: Identify the customer behavior you want to change and the business result it is expected to influence.
    • Evidence in scope: List the interviews, behavioral data, support conversations, journey maps, and prior experiments the model may use.
    • Constraints: Include privacy requirements, technical dependencies, commercial commitments, capacity limits, and non-goals.
    • Decision owner: Name the person accountable for accepting the trade-off. An AI-generated recommendation does not distribute accountability.

    Build a strategy chain the model can inspect

    Your strategy should form a traceable chain:

    1. Choose the customer and job that matter.
    2. Define the value proposition, including what must match the market and what should be meaningfully different.
    3. Name the customer outcome and business outcome.
    4. Break that outcome into drivers the product can influence.
    5. Select an opportunity supported by evidence.
    6. Form a testable product bet.
    7. Decide what evidence would justify continuing, changing, or stopping.

    A driver tree makes this chain concrete. It creates a visible connection between roadmap work and measures such as activation, retention, expansion, and Net Recurring Revenue. AI is useful here as a critic. Ask it to identify unsupported jumps, duplicated drivers, initiatives disguised as outcomes, and metrics the proposed product change cannot plausibly affect.

    Keep outputs and outcomes separate. Shipping an AI onboarding assistant is an output. Changing a defined activation behavior for a defined customer segment is an outcome. The model can help rewrite output-oriented objectives, but it cannot choose a credible target without baseline data, business context, and an accountable owner.

    Force a distinction between fact, inference, and assumption

    Require the model to label every material statement as one of three things:

    • Observed: Directly supported by a supplied interview, event, support conversation, or experiment.
    • Inferred: A reasonable interpretation that combines observations but is not explicitly stated by the customer or proven by the data.
    • Assumed: Necessary for the recommendation to work but not yet supported by the supplied evidence.

    This simple classification prevents an attractive narrative from laundering assumptions into facts. It also improves discovery planning: the most consequential assumption with the weakest evidence becomes a candidate for the next test.

    A useful instruction is: Use only the supplied material. For every recommendation, show the observations that support it, the inference connecting those observations to the recommendation, the assumptions that remain, and the evidence that could disprove it. If support is missing, say that it is missing.

    Build a controlled workflow from context to decision record

    AI assistance becomes reliable when it is a workflow rather than a chat session. A chat encourages improvisation: context changes, instructions disappear, and nobody can reconstruct why an answer looked different the next time. A workflow gives each pass a defined input, output, and approval gate.

    Ground the model in canonical product context

    Start with a retrieval-first set of canonical documents. At minimum, that context should include the current vision, product strategy, target segments, value proposition, OKRs, metric definitions, analytics dashboards, relevant discovery evidence, decision history, and definition-of-done checks.

    Canonical does not mean comprehensive. More context can make conflicts harder to notice. Give each item an owner, a freshness indicator, and an authority level. If an old positioning document conflicts with the approved strategy, the workflow should identify the conflict rather than silently averaging the two.

    Include exclusions as well. Tell the model which documents are historical, which metrics are deprecated, which segments are out of scope, and which proposals have already been rejected. Without those boundaries, previously abandoned ideas can return as apparently new recommendations.

    Separate extraction, synthesis, challenge, and approval

    1. Extract: Pull observations, customer language, events, metrics, decisions, and unresolved questions from the supplied material. Preserve links to the original evidence.
    2. Synthesize: Group related observations and propose opportunity statements. Keep contradictory evidence visible.
    3. Challenge: Look for alternative explanations, missing segments, weak causal claims, metric gaming, dependencies, and reasons the recommendation could fail.
    4. Decide: Have the accountable product leader and relevant partners accept, modify, or reject the recommendation. Record the trade-off explicitly.
    5. Publish: Store the decision, evidence, owner, expected outcome, guardrails, and next review trigger in the system the team already uses.

    Do not combine these passes into one request for a final answer. Extraction should not quietly prioritize. Synthesis should not hide inconvenient evidence. A challenge pass should test a proposed direction without changing the original evidence set. The human approval gate should be visible, not implied by the fact that somebody copied the output into a roadmap.

    Raw interviews, support threads, CRM records, and analytics exports can contain personal or confidential data. Do not paste them into an unapproved model. Minimize the data, remove identifiers that are not needed for the decision, use the governed environment approved by your organization, and retain only what the workflow requires. Privacy-by-design belongs at intake because redacting an output does not undo an inappropriate disclosure in the input.

    For recurring workflows, add acceptance criteria and evaluation cases. A discovery synthesis evaluation might check whether every theme retains evidence links, whether contradictions survive summarization, and whether unsupported market-size claims are rejected. A strategy evaluation might check whether every initiative maps to an outcome driver and whether an output has been mislabeled as an objective. Re-run those checks when the model, prompt, context set, or output schema changes.

    Use AI in discovery without laundering uncertainty

    Discovery generates exactly the kind of material language models handle well: interview transcripts, support conversations, journey notes, behavioral patterns, and open-ended hypotheses. AI can reduce the time between collecting this material and discussing it. It cannot make a biased sample representative or turn a correlation into a cause.

    Run synthesis as part of a weekly learning cadence that combines customer evidence with journey and behavioral analysis. Waiting for a large quarterly research readout increases the distance between observation and decision. Treating every new conversation as a roadmap mandate creates the opposite problem. A regular review gives the team a stable point at which evidence can accumulate, conflict, and change an existing belief.

    A cluster is a lead, not a finding

    Theme clustering is useful for navigation. It is not proof of importance. A frequent topic in support data may reflect product friction, a noisy customer segment, a documentation gap, or a recent incident. The model sees only the supplied dataset, not the market outside it.

    Require each proposed opportunity to include:

    • The affected segment and the context in which the problem occurs.
    • The job the customer is trying to complete.
    • Links to supporting observations, including direct customer language where it preserves important nuance.
    • The observed count within the supplied dataset, clearly distinguished from prevalence in the customer base or market.
    • Behavioral evidence that supports or challenges the qualitative pattern.
    • The outcome driver the opportunity could influence.
    • Contradictory evidence and plausible alternative explanations.
    • The unanswered question that creates the greatest decision risk.
    • The next piece of evidence that would materially change the decision.

    Then place the opportunity in an opportunity solution tree. Keep the opportunity separate from candidate solutions. If the branch says customers need an AI assistant, it has already collapsed a customer problem into a preferred implementation. Rewrite it in terms of the customer’s obstacle or desired progress, then generate multiple ways to address it.

    At the weekly review, ask four practical questions: What did the team observe? Which belief changed? Which important assumption remains weakly supported? What evidence should be collected next? AI can prepare the evidence packet and show deltas from the prior review. The product trio should decide what the evidence means and whether it changes the opportunity being pursued.

    Connect roadmap, experiment, launch, and learning

    A strategy loses integrity when each delivery stage invents its own explanation. The roadmap promises one outcome, the experiment measures another, the launch emphasizes a feature, and the retrospective celebrates shipping. AI can help maintain the thread, but only if the same hypothesis and metric definitions travel with the work.

    Decision layerUseful AI assistanceRequired human judgmentArtifact to preserve
    StrategyCheck the chain from customer value to business result and expose unsupported jumpsChoose the segment, differentiation, outcome, and trade-offsStrategy brief and driver tree
    DiscoveryExtract observations, cluster themes, retain contradictions, and draft opportunitiesInterpret evidence and choose the next uncertainty to reduceEvidence-linked opportunity record
    RoadmapMap candidate initiatives to drivers, surface dependencies, and prepare option comparisonsAllocate capacity and accept opportunity costPrioritization decision record
    ExperimentDraft hypotheses, instrumentation, guardrails, edge cases, and analysis checksApprove the test design, statistical assumptions, and decision ruleExperiment brief
    LaunchAdapt release notes, in-product guidance, support material, and segment messagingApprove claims, rollout risk, positioning, and readinessLaunch plan and approved message set
    LearningSummarize funnels, cohorts, retention patterns, qualitative feedback, and anomaliesDecide whether to continue, revise, expand, or stopLearning review and updated decision

    Make the roadmap show its reasoning

    Ask AI to produce roadmap options, not a single supposedly objective ranking. Each option should show the outcome driver it targets, evidence strength, important dependencies, unresolved risk, stakeholder impact, and the work displaced by choosing it. A priority score can organize inputs, but it cannot resolve a strategic disagreement about which customer or outcome matters most.

    Every roadmap item should answer: Why this customer problem, why now, what behavior should change, which business result should follow, and what observation would make the team reconsider? If the answer is merely that customers requested it or a competitor has it, the strategy is incomplete.

    Make experiments decision-ready before they run

    An AI-drafted experiment brief should contain a falsifiable hypothesis, eligible population, primary metric, guardrail metrics, instrumentation plan, exposure logic, expected mechanism, known confounders, and decision rule. For A/B testing, define the minimum detectable effect before interpreting results. The value must be tied to a practically meaningful change and checked against baseline behavior and available traffic; a model cannot infer those constraints from a feature description.

    Instrumentation deserves its own review. Specify the event, properties, eligibility conditions, trigger, and expected sequence in the funnel. Use behavioral analytics to check that exposure and activation are measured consistently across variants. Feature flags can separate deployment from release, support a controlled ramp, and limit exposure while the team checks behavior.

    For an AI-powered product experience, add eval-driven checks alongside product metrics. Define the behavior the model should exhibit, edge cases it must handle, unacceptable outputs, privacy constraints, and regression cases. Product success cannot compensate for a model behavior that violates an explicit safety or trust requirement.

    Keep launch language tied to the original value proposition

    AI can adapt UX copy, product tours, tooltips, release notes, in-app guides, and support macros for different segments. Give every channel the same approved value proposition, capability boundaries, terminology, and claims. Otherwise, speed creates message drift: the release note promises an outcome the interface does not support, while the support macro describes a different workflow again.

    After release, bring the original decision brief into the learning review. Examine the target cohort, funnel behavior, activation, retention, qualitative feedback, and guardrails. Do not ask only whether the feature was adopted. Ask whether the intended customer behavior changed, whether the assumed mechanism appears credible, and whether the business outcome remains a reasonable consequence.

    Scale the workflow only when another person can audit it

    Before expanding AI assistance across the product organization, hand one completed decision package to a colleague who was not part of the workflow. They should be able to identify the governing strategy, trace each important claim to evidence, see which assumptions remain open, understand the trade-off, and find the metric that will trigger the next decision.

    If they cannot, do not solve the problem with a longer prompt. Repair the missing artifact, unclear ownership, broken evidence link, or inconsistent metric definition. That is where strategic reliability lives.

    Start with one decision entering your next weekly discovery review. Build its evidence set, label observations and assumptions, run separate synthesis and challenge passes, and publish the human decision with its reversal signal. Once that chain survives review, reuse the workflow. The goal is not more AI-generated product work. It is a shorter, more inspectable path from customer evidence to a measurable strategic choice.

    References

  • What the Intercom-to-Fin Rebrand Teaches Product Leaders

    What the Intercom-to-Fin Rebrand Teaches Product Leaders

    If you are deciding whether an AI product should become your company name, you probably do not have a naming problem. You have a portfolio commitment problem. The rename will make your bet visible, but it will also force you to explain what existing customers still own, what will keep improving, and what now defines the company’s future.

    The Intercom-to-Fin move offers a clean way to think about that decision. The company is now named Fin, while Intercom remains its customer service software platform; Intercom 2 has also launched as a complete rebuild with continued investment behind it. The growth brand moves up to the corporate level without erasing the durable product brand beneath it. That is the strategic work of this rebrand.

    The decisive choice is what you do not rename

    The most important word in this rebrand is not Fin. It is “remains.” Intercom remains a product, a customer commitment, and a place where the company can keep creating value. Fin becomes the corporate identity and the clearest expression of the next growth thesis.

    Changing the company name while retaining the established product name is not an incomplete rebrand. It is deliberate brand architecture. The two names answer different customer questions:

    • The company brand answers: What future is this organization building toward?
    • The product brand answers: What can I buy, operate, renew, and rely on today?
    • The category brand answers: What new capability should I understand, budget for, and compare with alternatives?

    Those answers do not always belong under one name. Forcing them together can make the new strategy sound smaller than it is or make the established product appear to be on its way out. Keeping Intercom as the platform avoids turning corporate ambition into accidental product deprecation.

    Before approving a similar rename, write a transition contract. This is not a legal document. It is a short internal statement that every product, sales, marketing, support, recruiting, finance, and communications leader can use without improvising. It should answer:

    1. Exactly which entity is being renamed?
    2. Which products keep their current names?
    3. What changes for an existing customer because of the rename?
    4. What explicitly does not change?
    5. Where will investment increase, continue, or decline?
    6. How should someone describe the relationship between the company and each product?

    If the answers vary by executive, your organization is not ready to communicate the rename. Customers will encounter every inconsistency as a separate strategic story.

    A new category needs a clean place in the buyer’s mind

    Established brands are efficient because buyers use them as shorthand. The same shorthand becomes restrictive when a company wants to define a substantially different category. People do not continuously reassess every vendor from first principles. They attach new information to what they already believe.

    That is why a legacy name can create friction even when it has strong awareness and customer trust. The problem is not that buyers dislike the old brand. The problem is that they already know where to file it. Every pitch for the new category begins with a correction: the company you associate with one product is now asking you to understand it as something else.

    Fin had time to develop as a distinct service-agent identity before becoming the company name. The business introduced Fin three years before the corporate rename and deliberately led with that name while keeping Intercom in the background. That sequence matters. It allowed the category proposition to earn meaning before the corporate identity was placed behind it.

    You should look for the same underlying evidence before elevating a product brand:

    • Prospects ask for the new product or category by name instead of treating it as another feature of the established platform.
    • The product has a distinct job, competitive set, buying conversation, and roadmap.
    • Your largest resource-allocation decisions increasingly revolve around the new category.
    • The existing company name repeatedly requires explanation before buyers understand the new proposition.
    • The legacy product can remain a coherent, investable business under its own name.
    • Leadership is willing to keep prioritizing the category when it competes with comfortable, near-term work elsewhere in the portfolio.

    Wait if the new product still depends almost entirely on legacy demand, if “AI” is the only thing making it sound like a new category, or if leaders cannot explain the future of the existing portfolio. A corporate rename should settle a strategic truth that is already visible in the product and resource decisions. It cannot manufacture that truth.

    Test the strategy before you test the name

    Name preference is the least important question at the start. A memorable name cannot rescue an unstable thesis, and a room full of favorable reactions cannot prove that the proposed architecture makes sense. Test the decisions the name is meant to encode.

    Strategic permanence

    Ask whether the new identity can survive normal product evolution. A company named after a feature will eventually outgrow its name. A company named for a durable category, customer outcome, or long-term platform has more room to expand.

    Pressure-test the choice against plausible roadmap changes. If the current interface changes, the underlying models improve, or the product expands into adjacent workflows, does the name still represent the company? If one disappointing planning cycle would make leadership retreat to the old story, the corporate rename is premature.

    Customer comprehension

    Do not ask customers whether the new brand “makes sense.” That question invites politeness. Show them the proposed naming hierarchy without an explanation and ask them to describe:

    • What the company does.
    • What they can buy.
    • What happened to the existing product.
    • Which name they expect to see in the application, documentation, support experience, and commercial relationship.
    • Whether the new offering feels like a feature, a product, a platform, or a category.

    The vocabulary in their answers matters more than a preference score. If customers merge the company and product into one ambiguous object, the hierarchy needs work. If established customers assume their product is being replaced, the continuity story is too weak. If prospects still describe the company only through the old category, the new position has not yet become legible.

    Portfolio durability

    Every product affected by the rename needs a stated fate: promoted, retained, integrated, or retired. Silence creates its own answer, and customers usually interpret it as declining commitment.

    The Intercom-to-Fin architecture avoids that ambiguity. The corporate brand follows the AI growth engine, while the established platform receives a rebuilt product and continued investment. You can apply the same discipline by requiring a roadmap, owner, customer promise, and success measure for every brand that survives the transition.

    Operating commitment

    A company name is a resource-allocation claim. Check whether hiring plans, executive attention, roadmap capacity, sales enablement, partner priorities, and operating metrics already support the future implied by the name.

    This is where weak rebrands reveal themselves. The homepage changes, but planning continues to favor the old center of gravity. Sales compensation rewards the previous motion. Product teams keep describing the AI offer as an add-on. Recruiting language promises one future while internal goals fund another. If those contradictions remain, the market will believe the operating behavior rather than the new identity.

    Turn the rebrand into an operating model

    A corporate rename touches more than brand assets. It changes the nouns people use to make product, commercial, and technical decisions. Treat it as a cross-functional migration with a defined architecture, owners, dependencies, and observable failure modes.

    Before launch, remove internal ambiguity

    Start with an inventory of named objects. Separate the corporate brand, legal entity, product names, application name, AI agent, domains, documentation, status pages, integrations, partner listings, support channels, and customer-facing team names. They may not all change together, and some should not change at all.

    Create a controlled vocabulary for each object. Record the approved name, a plain-language definition, the transition phrase, phrases to avoid, and the person responsible for exceptions. Then apply it to roadmap documents, release notes, sales materials, onboarding, job descriptions, support macros, analytics labels, and executive reporting. This prevents each function from inventing a slightly different portfolio.

    Keep the public brand change separate from legal and payment instructions. A new display name does not automatically mean that the contracting entity, tax information, or bank details changed. Telling customers to update those records without confirmation can create payment failures, procurement delays, and fraud risk. Legal and finance owners should identify any real operational changes and communicate them through established, verifiable channels.

    Build the customer FAQ from actual consequences, not brand language. Cover logins, existing contracts, invoices, data handling, support access, integrations, domains, saved links, product roadmaps, and administrative work. For every item, say whether action is required. “No action required” is useful only when you have verified it across the relevant systems.

    At launch, separate ambition from continuity

    Lead with the scope of the change. Say which name belongs to the company, which belongs to the existing product, and how the new category fits. Then explain why the corporate identity is changing. Follow that with a precise account of what existing customers should expect.

    Do not rely on “nothing changes” as reassurance. It is usually too broad to be credible, especially when a new product strategy and increased investment are central to the story. Name the stable elements instead: the product that remains, the workflows that continue, the commitments that persist, and any interfaces or commercial records that stay the same.

    Use the same architecture everywhere a customer can encounter the company. A clear launch page cannot compensate for an application header, help center, invoice, partner marketplace entry, or sales deck that implies a different relationship. Transitional wording can help connect the names, but it should have an exit condition rather than becoming permanent clutter.

    After launch, measure the translation tax

    Launch reach tells you that people saw the rename. It does not tell you that they understood it. Establish a pre-launch baseline where possible, then monitor evidence of confusion:

    • Support conversations asking whether the existing product is being discontinued or replaced.
    • Sales calls in which representatives must repeatedly correct the company-product relationship.
    • Documentation searches that mix old and new names in ways your information architecture does not handle.
    • Broken redirects, failed bookmarks, authentication problems, or integration errors caused by changed domains or labels.
    • Procurement and accounts-payable questions about the company name, contracting entity, or invoice sender.
    • Prospect descriptions of the category after encountering the new positioning.
    • Retention, adoption, and expansion for the established product, tracked separately from awareness of the new corporate brand.

    Review the language in those interactions, not just their volume. The words customers use will show whether the new mental model has formed. Retire transition copy only when support, sales, search, and customer interviews indicate that people can move between the names without assistance.

    Key takeaways for your own portfolio decision

    • The Fin corporate name expresses the future growth bet; retaining Intercom protects a valuable product identity and signals continued commitment.
    • A corporate rename is a brand-architecture and resource-allocation decision, not a cosmetic marketing project.
    • Elevate a product name only when the category, roadmap, buying conversation, and operating priorities already support it.
    • Tell customers exactly what is renamed, what remains, what changes, and whether they need to act.
    • Validate comprehension with unscripted customer explanations, not name-preference questions.
    • Measure confusion across support, sales, documentation, procurement, integrations, and product health after launch.

    If this decision is in front of you, bring a one-page transition contract to your next portfolio review. Ask product, sales, support, legal, finance, and recruiting to describe the company and its products using the same nouns. If they cannot, keep working on the architecture. If they can, and your resource allocation already matches the story, the rename can do its real job: make the strategy easier for the market to understand.

    References