Month: November 2025

  • UX Product Management Career Playbook: Build Proof, Not Polish

    UX Product Management Career Playbook: Build Proof, Not Polish

    You are probably not wondering whether UX matters. You are trying to decide whether to move closer to design, how to make that move without becoming a second designer, and what evidence will convince a hiring manager that you can own the work.

    The answer is not another UX certificate or a more polished portfolio. You need proof that you can connect customer friction to a product decision, shape an experience with design and engineering, and measure whether the resulting behavior creates business value. This playbook shows you how to build that proof.

    Decide whether you want the work, not just the title

    A UX product manager owns the customer experience end to end while steering toward measurable outcomes. That does not mean producing every wireframe, conducting every research session, or making every interface decision. It means remaining accountable for the connection between a user’s problem, the experience the team ships, and the behavior that follows.

    The distinction matters because the role sits in an overlap, not in a gap. A designer should not need a product manager to practice design. A product team does need someone who can turn customer evidence into a prioritized problem, make trade-offs explicit, and keep discovery connected to delivery.

    Role emphasisPrimary questionStrong evidence
    Product designHow should this experience work for the user?Research synthesis, flows, interaction decisions, usability findings, and design-system judgment
    Product managementWhich problem should the team solve, for whom, and why now?Prioritization, value proposition, outcome definition, trade-offs, and business impact
    UX-oriented product managementWhich experience change will help a defined user reach value, and how will the team know?Customer evidence, experience strategy, cross-functional decisions, instrumentation, and behavioral outcomes

    You are likely suited to the overlap if you want to do all of the following:

    • Investigate why users struggle before debating what the team should build.
    • Move comfortably between a journey-level problem and a specific piece of microcopy.
    • Accept accountability for an outcome even though design, engineering, marketing, support, and the user all affect it.
    • Use qualitative evidence to explain behavior and quantitative evidence to establish its scale.
    • Partner closely with a designer without treating collaboration as permission to direct every screen.

    If those are not the decisions you want to own, do not force a title change. A product manager can deepen UX judgment without becoming a UX product manager, and a designer can develop product sense without leaving design. Choose the work you want to be accountable for.

    Build the three capabilities around one real user problem

    The fastest way to look shallow is to collect disconnected skills: a research course, an analytics dashboard, a prototype, and a prioritization framework that never touch the same decision. Build customer insight, product strategy, and experience design around one observable problem instead.

    Onboarding is a useful practice field because it exposes the whole system. You must identify the user’s intended value, find where progress breaks, decide what not to explain yet, shape guidance, and measure whether people reach a meaningful action. If onboarding is not relevant to your product, choose a core workflow with a clear start, a meaningful completion event, and visible friction.

    Customer insight: explain the friction before proposing a fix

    Start with a defined segment and a job the user is trying to complete. Then combine behavioral evidence with direct customer evidence. Funnel data can show where people leave; interviews, support conversations, and usability observation can help explain why.

    Create a compact evidence packet containing:

    • The target segment and the situation that brings the user into the experience.
    • The job the user believes they are completing, stated in the user’s terms.
    • The current critical path from entry to value.
    • Observed drop-off, delay, confusion, or repeated support demand.
    • Direct evidence behind the suspected cause, separated from your interpretation.
    • Assumptions that remain untested.

    That last distinction is career evidence. A strong UX product manager can say, “Users leave at this step” as an observation, “They may not understand the permission request” as a hypothesis, and “Changing the explanation should improve completion” as a testable prediction. Blending those statements into one confident story makes weak discovery look stronger than it is.

    Product strategy: turn the insight into a choice

    Customer pain is not automatically a priority. Connect it to a value proposition and an outcome. A useful framing is: “For this segment, improve this meaningful behavior by removing this verified barrier, because the behavior is part of reaching product value.”

    Now compare problem-level alternatives. The team might remove a step, change its sequence, defer a decision through progressive disclosure, clarify the value with UX writing, or provide contextual guidance. Do not jump from “users are confused” to “build a product tour.” A tour, an in-app guide, and a tooltip are interventions, not strategies. Each is appropriate only when it addresses the cause of the friction.

    Record what you will not pursue and why. This is where prioritization becomes visible. A hiring manager learns more from a rejected alternative with a sound trade-off than from a long feature list with no decision logic.

    Experience design: make the hypothesis concrete enough to test

    Work with design and engineering to turn the chosen problem into a testable flow. Trace the happy path, but also inspect empty states, errors, permission requests, loading behavior, recovery paths, and the moment when the user must make a consequential choice.

    Treat language as product behavior. A vague button label, an unexplained requirement, or a tooltip shown without context can create the same friction as a poor interaction. Good UX writing tells the user what will happen, why an input is needed, and how to recover when something goes wrong.

    Your artifact does not need visual polish. It needs enough fidelity to expose assumptions. Annotate the flow with the user question each step must answer, the behavior you expect, and the event required to measure it. That turns a prototype into a decision instrument rather than a gallery piece.

    Use activation as a diagnostic system, not a vanity metric

    Activation is a strong practice area because it forces you to define what “reaching value” means. It can also mislead you. Account creation, a completed tour, or a clicked button is not necessarily activation. The event should represent meaningful progress toward the reason the user adopted the product.

    Use this sequence for an activation project:

    1. Choose the segment. Different users may enter with different jobs, permissions, data, or expectations. Do not let an overall average hide a segment-specific failure.
    2. Define the value event. Name the behavior that indicates the user has experienced a meaningful part of the product’s promise. Explain why it matters rather than selecting the easiest event to count.
    3. Map the critical path. Identify the necessary steps between entry and value. Separate required complexity from friction the product has introduced.
    4. Locate the barrier. Combine funnel behavior with usability observation, customer language, and support evidence. A drop-off identifies a location, not a cause.
    5. Write the hypothesis. State the segment, barrier, intervention, expected behavioral change, and reason the change should occur.
    6. Define the read before launch. Specify the primary outcome, relevant guardrails, instrumentation, segments, and the decision you will make under each plausible result.

    Your tooling might include Amplitude, Pendo, or Intercom for funnels, product behavior, experiments, and customer signals. The brand matters less than the discipline: events must represent the intended behavior, properties must support the relevant segmentation, and exposure to an experiment must be distinguishable from eligibility for it.

    If you run an A/B test, set the minimum detectable effect before interpreting the result. Without an explicit MDE, an inconclusive read is easy to recast as success or failure after the fact. The purpose is not to make experimentation look scientific. It is to decide what size of change would matter and whether the test can detect it.

    Read activation alongside time-to-value and adoption of the core capability. Then inspect retention rather than assuming an early lift created durable value. If activation improves while retention does not, you may have accelerated an action without improving the underlying experience. If usability feedback improves but the behavioral metric does not, the altered friction may not have been the limiting factor. Both outcomes are useful when they lead to a sharper next decision.

    A practical experiment brief should answer these questions before delivery begins:

    • Which user segment is eligible?
    • What verified barrier are you addressing?
    • Which behavior should change, and why?
    • What is the smallest experience change that can test the causal assumption?
    • What is the primary outcome, and what must not degrade?
    • Which events and properties are required?
    • What MDE makes the test worthwhile?
    • What decision follows a positive, negative, mixed, or inconclusive result?

    This is how you keep discovery attached to delivery. A sprint should carry a learning goal or an outcome, not merely a collection of screens to complete.

    Build a portfolio that exposes your decisions

    A UX product management portfolio is not a design portfolio with extra charts. Its job is to make your reasoning inspectable. A reviewer should be able to see what you knew, what you assumed, which choices were available, why you selected one, and how evidence changed the next decision.

    Structure each case study as a decision journal:

    1. Context: Identify the segment, user job, product state, business relevance, and constraints.
    2. Problem evidence: Show the qualitative and quantitative signals. Distinguish observations from interpretations.
    3. Outcome: Define the behavior the team intended to change. Explain why it represented customer and business value.
    4. Alternatives: Present the credible options, including a smaller intervention and the option to do nothing.
    5. Decision: Explain the trade-off, who contributed, and which uncertainty the team accepted.
    6. Validation: Describe the prototype, usability work, production experiment, instrumentation, or retention analysis used.
    7. Result and next move: Report what the evidence justified. If it was ambiguous, explain what remained unresolved and what you changed next.

    Include screens only when they help the reader understand a decision. An annotated flow showing where a hypothesis enters the experience is more valuable than a polished sequence with no explanation. Likewise, a metric screenshot is not evidence of impact unless you define the segment, behavior, comparison, and decision attached to it.

    If the work was exploratory or self-directed, label it clearly. Do not imply that a concept shipped, that users were interviewed, or that business impact occurred when it did not. You can still demonstrate strong judgment by showing how you would instrument the experience, which assumptions require validation, and what evidence would cause you to stop.

    Your starting discipline determines which gaps the portfolio must close:

    • If you are a designer: make prioritization, value proposition, business trade-offs, outcome definition, and sequencing visible. Do not let the quality of the screens carry the case.
    • If you are a product manager: make the research plan, critical path, journey decisions, usability evidence, UX writing, and interaction trade-offs visible. Do not reduce UX to a feature requirement handed to design.

    Prepare interview stories around consequential decisions, not project tours. Start with the tension. Name the alternatives. Explain the riskiest assumption and how you tested it. Then state what you decided and what the evidence changed. This gives the interviewer material to assess your judgment under uncertainty.

    A strong resume bullet follows the same logic: “Changed [behavior] for [segment] through [experience decision], using [evidence or method], which informed [product or business decision].” Replace every bracket with facts you can defend. If you cannot name the behavior or the decision, the bullet is probably describing output.

    Lead the product trio without taking over another craft

    Your career will stall if UX fluency turns into design control. The useful version of the role creates a tighter product trio: product keeps the segment, problem, priority, and outcome visible; design leads the coherence and usability of the experience; engineering brings feasibility, system constraints, delivery insight, and instrumentation into the decision early. Important choices are shaped together.

    Use a lightweight operating loop:

    • Before planning: align on the user problem, current evidence, target behavior, unresolved assumptions, and the next learning goal.
    • During discovery: pair customer evidence with prototypes and technical investigation. Involve engineering before the team commits to a flow whose cost or constraints are unknown.
    • During delivery: preserve the hypothesis in the acceptance criteria and instrumentation. Do not let the ticket retain the interface while losing the reason for it.
    • After release: review behavior and customer signals together. Decide whether to continue, adjust, investigate, or stop.

    Tailor the decision narrative to the audience. Executives need the trade-off, business consequence, evidence strength, and decision required. Engineers need constraints, sequencing, edge cases, event definitions, and the reason behind the behavior. Designers need the user job, journey context, friction evidence, and experience assumptions. Other stakeholders need to know what changed, why it changed, how success will be judged, and which new evidence could alter the plan.

    A reusable update can stay simple: “For [segment], we are trying to change [behavior] because [evidence] indicates [barrier]. We chose [intervention] over [alternative] because [trade-off]. We will judge it through [outcome and guardrail]. The next decision occurs when [evidence condition].” That format reduces status theater because it keeps the decision and its evidence in view.

    Key takeaways

    • A UX product manager connects customer insight, experience decisions, and measurable product outcomes; the role is not a substitute for product design.
    • Build customer insight, product strategy, and experience design around the same real problem so your skills form a coherent body of evidence.
    • Use activation to diagnose the path to value, but verify downstream adoption and retention before claiming durable impact.
    • Define segments, events, guardrails, MDE, and decision rules before reading an experiment.
    • Make your portfolio a decision journal that includes constraints, alternatives, ambiguous evidence, and rejected ideas.
    • Demonstrate leadership by improving the product trio’s decisions, not by absorbing the responsibilities of design or engineering.

    Choose one experience in your current product and build the full evidence chain: segment, problem, critical path, hypothesis, experience change, instrumentation, outcome, and next decision. When you can show that chain clearly, you are no longer asking a hiring manager to infer your UX product judgment. You are giving them proof.

    References

  • AI-First Customer Support for Sustainable Ecommerce Growth

    AI-First Customer Support for Sustainable Ecommerce Growth

    Your ecommerce support queue is growing, but cutting ticket volume is not the real decision in front of you. The harder question is which customer outcomes you can let AI own – from order questions to address changes and refunds – without creating a faster path to a wrong answer or action.

    AI-first support earns its place when it completes customer work safely, gives human agents the full context when it cannot, and produces evidence you can use to improve the buying and ownership experience. Growth does not mean forcing a sale into every conversation. It means removing avoidable friction before purchase, resolving post-purchase problems well, and turning repeated support demand into better product and operational decisions.

    Define the unit of automation as a resolved customer job

    A message is not a resolution. An answer is not always a resolution either. If a customer asks to cancel an order, sending the cancellation policy may be factually correct while leaving the actual job unfinished.

    For an AI agent to resolve that request, it must verify the customer and order, check whether cancellation is allowed, execute the permitted action, confirm the exact outcome, and recognize when an exception requires a person. This distinction matters because a deflected conversation can still represent an unresolved customer and a second contact waiting to happen.

    Start by separating support demand into four kinds of work:

    • Informational work: order status, delivery information, return-policy questions, and other requests that can be completed with a grounded answer.
    • Bounded transactional work: changing an eligible shipping address, cancelling an order, issuing an allowed refund, or performing another action with clear rules and permissions.
    • Advisory work: helping a shopper find a suitable product using current catalog data and the constraints the shopper has provided.
    • Judgment-heavy work: policy exceptions, ambiguous intent, conflicting account data, unusual financial consequences, or emotionally sensitive cases where discretion matters.

    Use a workflow map like this before choosing what to automate:

    Customer jobAI needsEvidence of completionWhen AI must stop
    Get current order informationVerified identity, correct storefront, and current order dataThe requested state is returned from the commerce systemIdentity, store, or order data is missing or inconsistent
    Change a shipping addressAn eligible order, editable fields, an authorized tool, and customer confirmationThe commerce platform accepts the new value and returns the updated orderThe order has progressed too far, the address is ambiguous, or the tool fails
    Cancel or refund an orderPolicy rules, order state, transaction permissions, and explicit confirmationThe platform confirms the exact cancellation or refund that occurredThe request is an exception, the amount is unclear, or execution is incomplete
    Choose a productCurrent catalog data and relevant shopper constraintsThe shopper receives grounded options or a clean route to human adviceRequired constraints are unknown or the catalog cannot support the recommendation

    For example, a Shopify support integration can distinguish between retrieving order information and executing actions such as address edits, cancellations, refunds, and duplicate-order workflows. That separation is the architectural principle to preserve: knowing something about an order is not the same as having permission to change it.

    Prioritize each workflow using three factors: how much customer demand it represents, how ready the required data and tools are, and how costly a wrong outcome would be. High frequency alone is a poor selection rule. A common request with unreliable data will produce common failures, while a lower-volume workflow with clear rules may be the better place to prove the operating model.

    Build shared context, bounded actions, and deliberate handoffs

    Treating AI as infrastructure and assigning clear ownership of its performance changes the design question. You are no longer adding a writing assistant to an inbox. You are creating a customer-facing system that reads business state, applies policy, calls tools, and hands work to people.

    The minimum useful context for ecommerce support usually includes verified customer identity, storefront, order and customer records, applicable policies, product or catalog information, conversation history, and the current state of any attempted workflow. Multi-store merchants need the store identifier to travel with the conversation. A valid order number in the wrong storefront is still the wrong context.

    Data architecture deserves the same attention as the model. Capabilities such as multi-store handling, synchronized custom fields, updated data mappings, and EU workspace support illustrate the practical requirements. If the AI cannot determine which record is authoritative, it should expose the conflict and stop. It should never manufacture the missing state.

    Give every action an explicit contract

    A prompt is not an adequate control for a transactional workflow. Every tool the AI can call should have an action contract that defines:

    • Preconditions: what must be true before the action is available.
    • Required inputs: which values must come from verified commerce data and which may come from the customer.
    • Permissions: which customers, agents, stores, order states, and transaction types are eligible.
    • Confirmation: the exact order, field, amount, or consequence the customer must approve.
    • Execution response: a structured success or failure state returned by the commerce platform, not a guess based on generated text.
    • Duplicate-submission protection: how the system prevents the same action from being executed twice.
    • Failure behavior: whether to retry, stop, reverse a reversible step, or hand the case to a person.
    • Audit data: what action was requested, which policy was applied, what the tool returned, and what the customer was told.

    Separate permissions by consequence. Reading authenticated order status is different from drafting a proposed change. Drafting is different from executing a reversible update. A cancellation or refund carries financial and customer-trust consequences, so it needs stricter eligibility checks, explicit confirmation, and a reliable human path for exceptions. Customer confirmation does not compensate for an ineligible order or an unreliable tool.

    The integration method does not remove these obligations. Whether a tool is exposed through a native connector, an internal API, or Model Context Protocol, the AI still needs a constrained schema, narrow permissions, deterministic validation, and an unambiguous result.

    Make escalation a designed path, not a failure bucket

    AI-first does not mean AI-only. Humans should enter when judgment adds value or when a control condition is triggered. Define those conditions before launch rather than expecting the model to improvise them.

    Escalate when identity cannot be verified, records conflict, a policy exception is requested, a consequential action falls outside permission, a tool returns an incomplete result, the customer disputes an executed action, or the customer asks for a person. A model confidence score is not enough unless you have calibrated it against the actual intents and failure costs in your environment.

    The human receiving the conversation should get a compact handoff package containing:

    • The customer’s current request and the reason for escalation.
    • The verified customer, storefront, and order identifiers.
    • A short summary of facts already established.
    • Every action attempted and the exact tool result.
    • The unresolved decision or exception.
    • Anything already promised to the customer.

    The customer should not have to reconstruct the case. When the AI has enough context to recognize that it cannot finish, passing that context forward is part of the resolution experience.

    Measure verified outcomes, system reliability, and growth impact

    Deflection is an activity measure. It tells you a human did not enter the conversation, but it does not prove the customer received the right answer, the requested action succeeded, or the issue stayed resolved. An AI-first operating model should instead emphasize resolution, impact, and system reliability.

    Define a successful automated resolution before you build a dashboard. A practical definition is: the AI correctly understood an eligible request, delivered the correct answer or completed the authorized action, communicated the outcome accurately, and did not create an avoidable repeat contact within a fixed follow-up window. Choose the window for your business and apply it consistently.

    Report coverage and success separately. A strong success rate on a very narrow set of conversations can look impressive while leaving most customer demand untouched. A broad coverage rate can hide weak execution. At minimum, track these metric layers:

    • Eligibility and coverage: the share of total conversations that match a workflow AI is allowed to handle, followed by the share it actually attempts.
    • Resolution quality: verified correctness by intent, policy adherence, repeat contact, customer dispute, and the rate of unnecessary escalation.
    • Action reliability: successful tool execution, rejected actions, duplicate attempts, incomplete results, and wrong or unauthorized changes.
    • Handoff quality: whether the right cases escalate, whether the context package is complete, and whether customers must repeat information.
    • Customer experience: time to the completed outcome and satisfaction segmented by intent and resolution path.
    • Business impact: cost per verified resolution, pre-purchase assisted conversion where attribution is credible, and downstream retention or repeat-purchase signals.

    Do not present an association as growth causation. Customers who contact support may already differ from those who do not. Use controlled experiments where they are practical, compare like-for-like intent cohorts, and treat retention as a downstream signal unless the measurement design supports a stronger claim.

    Ownership matters as much as measurement. Assign someone to own AI support as a product surface, someone to govern knowledge and policy, someone to own commerce integrations and permissions, and someone to review quality and customer harm. These are responsibilities, not mandatory job titles. A smaller organization may place several with one person, but none should be left implicit.

    During a live rollout, I would review every failed or disputed write action and sample successful actions across each active intent every operating day. Once the important failure modes are understood and performance is stable, intent-level review can move to a weekly cadence. Scope changes should still happen through an explicit release decision, not because the queue happens to be busy.

    Roll out one dependable resolution lane at a time

    The safest path to meaningful automation is not a site-wide chatbot launch. It is a sequence of narrow resolution lanes, each with grounded data, an evaluation set, clear permissions, a human fallback, and a rollback path.

    1. Establish the baseline. Group current conversations by customer intent and record volume, time to outcome, repeat contact, escalation, and the systems or policies each intent depends on.
    2. Select a narrow first lane. Favor a request with clear rules, reliable data, and low action reversibility. Authenticated order information is often a better proving ground than refunds, but your own data readiness should decide.
    3. Create an evaluation set from real, appropriately handled conversations. Include ordinary cases as well as missing orders, stale data, multi-store ambiguity, policy exceptions, tool errors, changed customer intent, and explicit requests for a person.
    4. Write expected outcomes before testing. For every case, specify whether AI should answer, act, ask for missing information, or escalate. Classify unauthorized disclosure, wrong transactional action, and missed consequential escalation as critical failures that an overall average cannot hide.
    5. Observe before granting broad action permissions. If your platform supports a draft or shadow mode, compare proposed behavior with the expected outcomes. Then launch to a limited storefront, channel, workflow, or customer cohort with active monitoring.
    6. Add one write action at a time. Confirm the action contract, permissions, confirmation language, duplicate protection, audit trail, human fallback, and rollback mechanism before expanding eligibility.
    7. Protect peak periods. Do not introduce a consequential workflow immediately before your highest-demand period unless it has already passed realistic evaluation and the operating team can disable it quickly. Keep staffing and fallback capacity based on verified workload movement, not projected deflection.

    This expansion model creates a compounding loop. Every failed or repeated conversation should produce a specific improvement task: repair missing knowledge, correct a data mapping, clarify a policy, tighten an action permission, improve the handoff, or send a recurring upstream problem to product, merchandising, fulfillment, or operations. The value is not only that AI absorbs work. It is that support demand becomes structured evidence about where ecommerce growth is leaking.

    Continue expanding only when a lane remains dependable under real conditions. Tight merchant feedback loops and peak-season planning are especially important as the agent moves from answering questions to taking actions. Pause when unresolved contacts or ambiguous cases rise. Roll back immediately when the system performs an unauthorized or incorrect consequential action.

    Key takeaways

    • Optimize for completed customer jobs, not avoided human conversations.
    • Separate information retrieval from transactional authority, and give every action a testable contract.
    • Make verified identity, storefront, order state, policy, and tool state part of the shared context.
    • Design human escalation before launch so judgment-heavy cases arrive with their context intact.
    • Report eligibility, coverage, resolution quality, action harm, and business impact separately.
    • Expand through evaluated resolution lanes with explicit release, monitoring, and rollback decisions.

    Your next move is concrete: choose one customer job, write down its required data, allowed actions, stop conditions, success evidence, and human fallback. If you cannot make those five elements explicit, the workflow is not ready for autonomous resolution. If you can, you have the first building block of an AI-first support system that can grow without asking customers to absorb the risk.

    References

  • From KPIs to Comebacks: How I Lead Through Setbacks with Curiosity, Care, and Discovery

    From KPIs to Comebacks: How I Lead Through Setbacks with Curiosity, Care, and Discovery

    Setbacks are the tax we pay for doing meaningful product work. As a VP of Product Management, I’ve learned that what separates resilient teams from the rest isn’t a lack of failures—it’s how we metabolize them. This episode of All Things Product with Teresa Torres and Petra Wille is a powerful reminder that recovery, reflection, and rigorous product discovery are as essential as speed and execution.

    Listen to this episode on: Spotify https://open.spotify.com/episode/10LYRya7boYJBHTYBnE79E?ref=producttalk.org | Apple Podcasts https://podcasts.apple.com/kh/podcast/dealing-with-setbacks/id1794203808?i=1000737190520&ref=producttalk.org

    What struck me most is how Teresa shares a deeply personal story about her long recovery from an injury—and how that journey mirrors the nonlinear reality of product development. In product, just like in healing, progress is rarely a straight line. We have surges, stalls, and moments that feel like reversals. Yet with the right mindset and rituals, we still move forward.

    Professionally, we all face moments when your product fails to move a single KPI, when a launch falls flat, or when you just feel stuck. I’ve been there—in quarterly reviews, post-launch standups, and board prep. The instinct is to sprint straight into solutions. The wiser move is to respond with curiosity, emotional honesty, and resilience, then re-engage our discovery habits with intention.

    If you’re a PM, designer, or researcher, consider this an invitation to rebalance. Recovery and reflection are just as important as velocity and success. That’s not soft talk—it’s how empowered product teams build durable performance without burning out.

    On the emotional reality of setbacks, I’ve learned to normalize naming the loss. We put immense pressure on ourselves, and it’s okay (and necessary) to grieve product failures. When we acknowledge the disappointment, we regain the ability to observe clearly—and to learn.

    Leaders play a crucial role here. I create space for teams to recover before jumping into post-mortems. We don’t whiteboard over feelings; we schedule time for decompression, then conduct a crisp, blameless review. That sequencing transforms the quality of insights and strengthens psychological safety.

    Another lesson that resonates is the danger of tying performance too tightly to outcomes. Outcomes matter, but they are lagging indicators influenced by many externalities. I evaluate performance on behaviors: clarity of problem framing, rigor in discovery, quality of decision-making, and stakeholder alignment. This aligns with outcomes vs output OKRs and keeps us focused on controllable excellence.

    How do we build resilience? Continuous discovery builds resilience by normalizing failure. When we test assumptions routinely with customers and data, we turn large, risky bets into a series of small, learnable steps. Teams recover faster because failure becomes feedback—frequent, cheap, and informative.

    For perspective, I often use the 10–10–10 framework (from Decisive by Chip & Dan Heath). I ask: How will this setback feel in 10 minutes, 10 months, and 10 years? The answers de-escalate urgency, expand our time horizon, and produce better, calmer decisions.

    Here are the key takeaways I’m carrying forward. Setbacks are not just inevitable—they’re part of doing meaningful product work. Giving teams time and space to process failure builds long-term resilience. Mourning losses is just as important as celebrating wins.

    Healthy discovery cultures embrace reflection, psychological safety, and emotional honesty. And most importantly, staying consistent with discovery habits helps teams recover faster and learn more deeply.

    Notable moments that stood out for me include: [00:02:00] Teresa shares the story of her injury and what it’s taught her about patience and setbacks. The parallel to product cadence is both humbling and motivating.

    [00:10:00] Petra talks about a team whose carefully planned launch didn’t move a single KPI. I’ve led similar debriefs; when we anchor on customer insight gaps rather than blame, the next iteration improves dramatically.

    [00:20:00] Discussion on allowing space for grief and frustration after failure. In my teams, we time-box “emotional processing” before we enter analysis mode—it humanizes the work and sharpens the learning.

    [00:30:00] Why organizations must decouple performance reviews from short-term outcomes. I align evaluations to strategy execution quality, hypothesis discipline, and cross-functional collaboration.

    [00:40:00] How continuous discovery can help teams normalize—and even learn to appreciate—setbacks. When discovery is weekly, momentum becomes self-healing.

    If you want to dig deeper, here are useful links from the episode. Follow Teresa Torres: https://ProductTalk.org

    Follow Petra Wille: https://Petra-Wille.com

    Mentioned in the episode: Decisive by Chip & Dan Heath — The 10–10–10 framework for perspective in decision-making https://heathbrothers.com/books/decisive/?ref=producttalk.org

    Teresa Torres’ Continuous Discovery Habits — Building resilience through ongoing discovery practices. https://www.amazon.com/Continuous-Discovery-Habits-Discover-Products/dp/1736633309?dchild=1&keywords=continuous+discovery+habits&qid=1621385051&sr=8-2&linkCode=sl1&tag=teresatorres-20&linkId=34bc439ac78da06e1398f7bf069b219e&language=en_US&ref_=as_li_ss_tl&ref=producttalk.org

    Join the Conversation: Have thoughts on this episode? Leave a comment below. I’d love to hear how you create space for recovery while sustaining product velocity.

    Full Transcript: Full transcripts are only available for paid subscribers.


    Inspired by this post on Product Talk.


    Book a consult png image
  • PendomoniumX London: An Operating Model for AI Products

    PendomoniumX London: An Operating Model for AI Products

    If your AI portfolio has plenty of prototypes but little habitual use, the gap is probably not access to better models. It is operating design. A team can ship an impressive assistant and still fail because it chose a weak workflow, buried the feature, measured clicks instead of changed behavior, or treated trust as a post-launch review.

    At PendomoniumX London, more than 350 software leaders gathered around AI transformation and product innovation. The useful signal for product leaders was the move from broad enthusiasm to execution: clearer customer problems, measurable adoption, faster learning, and explicit governance. You can turn that signal into an operating model for your own AI roadmap.

    Transform a customer workflow, not a feature list

    An AI feature generates, summarizes, classifies, recommends, or takes an action. An AI product transformation changes how a person completes a meaningful job. The distinction matters because customers do not adopt model capabilities in isolation. They adopt a faster, easier, or more reliable way to get something done.

    Starting with the model usually produces a familiar failure mode: the team finds technically plausible places to insert AI, ships several disconnected experiences, and then struggles to explain why customers should change their behavior. Starting with the workflow forces the team to identify the user, the moment of friction, the desired behavior, and the evidence that would justify further investment.

    I would not approve an AI roadmap item until the team can complete this sentence:

    For a specific user completing a specific workflow, the product will use AI to remove a named source of effort or uncertainty, leading to an observable behavior change and a defined customer or business outcome, within explicit trust boundaries.

    Build the statement in this order:

    1. Describe the current workflow. Write the steps a customer takes now, including any handoffs, repeated decisions, manual checks, or places where work is abandoned.
    2. Isolate one consequential friction point. Avoid vague problems such as “the workflow is inefficient.” Name the decision, delay, rework, or uncertainty that prevents progress.
    3. Define the assistance. State whether AI will draft, recommend, retrieve, classify, predict, or act. These modes create different expectations and require different controls.
    4. Name the behavior that should change. Examples include completing a setup step, accepting or editing a recommendation, resolving a case, or returning to use the capability again.
    5. Connect the behavior to an outcome. A click is not an outcome. Faster time-to-value, lower abandonment, greater task completion, and sustained use are closer to the value you need to establish.
    6. Write the boundary before the prototype. Specify what data the system may use, what the user must verify, when a human remains responsible, and what happens when the system cannot produce an acceptable result.

    This framing also gives you a useful way to reduce an overcrowded AI roadmap. Reject ideas that cannot name a recurring workflow, an observable behavior, and a credible path to customer value. A clever demonstration without those elements is an experiment, not yet a product commitment.

    Run one evidence loop from discovery through go-to-market

    AI work becomes slow when discovery, delivery, analytics, and go-to-market operate as separate projects. Research identifies one problem, engineering explores another, marketing promises a broad capability, and analytics arrives after launch. Each function can appear busy while the product accumulates uncertainty.

    The better unit of management is one evidence loop:

    1. Discovery identifies the costly moment. Combine customer interviews with behavioral data. Interviews explain the user’s reasoning and workarounds; analytics shows where the behavior occurs, which segments encounter it, and whether the problem is frequent enough to matter.
    2. Prioritization exposes the assumptions. Compare bets using problem severity, workflow frequency, data readiness, trust burden, reach, and speed of learning. Do not hide weak evidence behind a single calculated score. Record why each factor received its assessment.
    3. Sprint planning targets uncertainty. A prototype should answer a specific question: whether customers want assistance at this moment, whether the available context supports an acceptable output, or whether users understand how to review the result. Building the full workflow before answering the riskiest question creates expensive evidence.
    4. Go-to-market explains the changed job. Lead with what the customer can now accomplish. “AI-powered” describes an implementation choice; it does not tell a customer when to use the capability, what input it needs, or what outcome to expect.
    5. Post-launch behavior changes the roadmap. Compare actual use with the original baseline and bet statement. Look at starts, completions, acceptance or editing of outputs, abandonment, repeated use, and downstream outcomes. Feed those observations into the next discovery decision.

    A lightweight decision log keeps this loop honest. For every AI bet, record the customer problem, riskiest assumption, evidence collected, decision made, owner, and next review condition. The log prevents a prototype from quietly becoming a permanent commitment simply because significant effort has already been spent.

    A prototype that misses the mark can still be valuable if it retires uncertainty. If customers do not recognize the problem, stop. If they value the workflow but distrust the output, change the interaction or control model. If the output is useful but discovery is weak, address distribution and onboarding. Those are different diagnoses, so they should not all produce the same response of adding more features.

    Make adoption part of the product itself

    Launching an AI capability does not teach customers when to trust it, what information to provide, or how it fits into an existing routine. That education is part of the experience, especially when the product asks someone to replace a familiar manual process with a probabilistic system.

    Examples at PendomoniumX paired Pendo’s in-app guides and product tours with behavioral analytics to improve activation and reduce friction around important onboarding moments. The transferable lesson is not to add a tour to every AI release. It is to place guidance at the moment of intent and measure whether it helps the customer reach value.

    Instrument the adoption path before you publish the guidance:

    • Eligible: the right user reaches the relevant workflow and has permission to use the AI capability.
    • Exposed: the user can see the entry point or receives contextual guidance.
    • Started: the user initiates the AI-assisted action.
    • Delivered: the system returns an output or completes the requested action.
    • Evaluated: the user accepts, edits, rejects, retries, or reverses the result.
    • Completed: the user finishes the larger workflow in which the AI action sits.
    • Repeated: the user chooses the capability again when the relevant need returns.

    This sequence prevents a common measurement mistake. A guide view shows exposure, not activation. A button click shows curiosity, not value. Even a generated output may not matter if the user discards it or fails to complete the surrounding task. Define activation at the first point where the customer receives meaningful value, then monitor whether that behavior repeats.

    Keep the guidance proportional to the decision:

    • Use a short contextual prompt when the customer only needs to notice a new action.
    • Use a tooltip when the customer needs one local explanation, such as what information the model will use.
    • Use a multi-step tour only when the workflow itself spans multiple unfamiliar steps.
    • Show an example input when output quality depends heavily on how the request is framed.
    • Explain review and fallback behavior next to the action, not in a distant help page.
    • Let experienced users dismiss education that no longer helps them.

    If traffic and risk permit a controlled experiment, compare eligible guided and unguided cohorts on workflow completion and repeated use. If you cannot create a credible control group, use a documented baseline and staged rollout. In either case, do not claim that guidance caused adoption merely because guide views and feature use rose at the same time.

    Make trust boundaries and decision rights explicit

    Trust is not a legal checklist appended to an otherwise finished AI experience. It affects what the system may do, what the interface must explain, which events need monitoring, and whether the customer remains in control. Deferring these decisions creates rework because the team may later need to change data flows, permissions, interaction design, or the scope of automation.

    For each workflow, answer these questions in language the product team can implement:

    • What customer, account, or third-party data may enter the system?
    • What context is necessary, and what data should be excluded even if it could improve the output?
    • What is retained, for what purpose, and who can access it?
    • Which outputs are suggestions, and which can cause an action in the customer’s environment?
    • What must the user review or confirm before an action becomes consequential?
    • How does the experience communicate uncertainty, missing context, or inability to complete the task?
    • What fallback lets the customer continue when the AI path fails?
    • Which signals trigger investigation, rollback, or a narrower release?
    • Who owns customer feedback, incidents, and changes to the evaluation criteria?

    When personal data, sensitive customer information, or regulated decisions are involved, bring privacy, security, and legal reviewers into discovery. The safe alternative to making assumptions is to narrow the data and action scope until the appropriate review is complete.

    Governance must be matched by clear decision rights. An empowered product team is not an ungoverned team. It is a team that knows which decisions it can make, the evidence expected, and the boundary at which another owner must participate.

    A practical division is to distinguish three layers:

    • Team-owned decisions: workflow design, contextual education, experiments within approved boundaries, evaluation cases, and roadmap changes supported by product evidence.
    • Cross-functional review: new data access, material changes to retention, model-provider changes, higher-impact automation, and controls that affect security, privacy, support, or compliance.
    • Leadership decisions: risk tolerance, strategic investment across portfolios, shared platform choices, and conflicts that cannot be resolved within the product outcome.

    Write these rights into the AI bet rather than relying on organizational memory. Also define the conditions for continuing, reworking, pausing, or stopping the work. The exact thresholds should come from your baseline and risk context, but the decisions should exist before launch. Otherwise, encouraging signals will be celebrated while contradictory evidence is explained away.

    Key takeaways

    • Frame every AI investment around a recurring customer workflow, not a model capability.
    • Require a bet statement that connects assistance, behavior change, customer value, and trust boundaries.
    • Use one evidence loop across discovery, prioritization, sprint planning, go-to-market, and post-launch learning.
    • Measure the full adoption path from eligibility to repeated use; guide views and feature clicks are intermediate signals.
    • Treat in-app education as contextual product design, not a substitute for a clear value proposition.
    • Set data boundaries, human-review points, fallback behavior, decision rights, and stop conditions before broad release.

    In your next planning cycle, choose one live AI initiative and rewrite it as a workflow bet. Add its behavioral baseline, activation event, trust boundary, decision owner, and stop condition. Then instrument the path before expanding the feature set. If the team cannot agree on those elements, the roadmap item is not ready. If it can, AI has started to become a managed product capability rather than a collection of prototypes.

    References

  • Brand Visibility in AI Answer Engines: A Product Playbook

    Brand Visibility in AI Answer Engines: A Product Playbook

    If your CEO asks why an AI answer names a competitor but leaves out your brand, the tempting response is to publish more pages or look for a ChatGPT optimization trick. That treats the symptom. The real question is whether the answer engine can confidently connect your brand to the user’s decision, verify the connection, and explain it accurately.

    Treat AI visibility as a product system. You can improve its inputs, test its outputs, and assign owners to its failure modes. You cannot guarantee a mention, but you can increase the probability of an accurate inclusion by building a clear public identity, credible evidence, reliable retrieval, and useful actions.

    Define the decision you want to be present for

    Brand visibility is too vague to manage. Visibility for what? A category definition, a shortlist, an integration question, a troubleshooting task, and a product comparison are different jobs. Each requires different evidence.

    Start with an intent map. Use the customer journey, support conversations, sales objections, onboarding friction, and product analytics to identify the decisions that matter. Then connect each decision to the artifact an answer engine would need.

    User jobTypical questionArtifact to publishDesired answer behavior
    Understand the categoryWhat problem does this category solve?Category explainer and glossaryRecognize the brand’s category and relevant use cases
    Evaluate optionsWhich product fits this workflow or constraint?Use-case page, comparison, and evidenceInclude the brand when it genuinely fits and state the tradeoffs
    Get startedHow do I reach the first useful outcome?Quick-start documentationReturn accurate prerequisites and steps
    IntegrateDoes this product connect to another system?Integration page and API documentationDescribe compatibility, setup, and limitations correctly
    Resolve a problemWhy is this workflow failing?Troubleshooting documentationRetrieve a grounded diagnosis and resolution path
    Check current statusIs this feature available, and what changed?Changelog and release notesUse current product facts instead of stale descriptions

    For each row, define when your brand is actually eligible. A weak objective says, ‘The brand should appear.’ A useful objective says, ‘The brand is relevant when the user needs this capability, works under these constraints, and can verify these claims.’

    That distinction protects the program from vanity metrics. Your product should not appear in every answer. It should appear in the answers where it can help, in the correct category, with an honest account of its strengths and limits. My rule is simple: a mention that misclassifies the product is a failure, even if the brand name is present.

    Prioritize prompt families using product judgment. Start where a better answer could affect a meaningful buying, activation, integration, or support decision. Within that set, look for the largest evidence gap: an important question for which your current public material is missing, contradictory, gated, or stale. That gives you a defensible backlog rather than an open-ended demand for more content.

    Build a canonical brand record before producing more content

    An answer engine has a harder job when your homepage describes one category, your documentation uses another product name, a partner directory lists an old capability, and a comparison page makes a broader claim than the evidence supports. Publishing another page adds volume without resolving the identity problem.

    Create an internal brand fact record that becomes the contract for every public property. It should contain:

    • The official organization, product, and feature names, including approved abbreviations.
    • The primary category and a plain-language description of what the product does.
    • The users, jobs, and constraints for which the product is relevant.
    • The capabilities and integrations that can be stated publicly.
    • The limitations or eligibility conditions that materially change a recommendation.
    • The evidence behind important claims, such as documentation, case studies, API references, or release notes.
    • An owner and review trigger for every fact that can change.

    Use this record to audit the homepage, product pages, documentation, API references, GitHub repositories, partner listings, review profiles, and conference descriptions. Do not force identical prose everywhere. Do keep the underlying identity, category, capability, and product status consistent.

    Your site architecture should make that identity easy to follow. Connect category explainers to use-case pages, use-case pages to product documentation, documentation to integrations and troubleshooting, and changing capabilities to release notes. The links should reflect a real path from understanding to evaluation to action.

    Then inspect the technical path an unauthenticated visitor can use. The essentials are concrete:

    • Put foundational product facts in semantic HTML rather than only inside images, videos, or interfaces that require a login.
    • Keep robots.txt and XML sitemaps friendly to public product and documentation pages.
    • Use canonical tags to concentrate signals when similar pages exist.
    • Apply schema.org types such as Organization, Product, HowTo, and FAQPage only where the visible content supports them.
    • Use descriptive headings and rich alt text so page meaning is not dependent on presentation.
    • Keep public pages fast enough to retrieve reliably.
    • Leave foundational documentation open when there is no business, privacy, or security reason to gate it.

    Do not loosen access controls in the name of visibility. Public product facts, help content, and approved evidence belong in the retrievable footprint. Customer data, internal plans, private support records, and administrative documentation do not. The right fix for a gated public fact is a safe public page, not broader access to a private system.

    Write pages that answer prompts without requiring guesswork

    Traditional marketing pages often ask the visitor to infer the product’s category, audience, and value from slogans. An answer engine needs explicit relationships. It should be able to identify what the product is, who it is for, what task it performs, what conditions apply, and where the supporting evidence lives.

    Use a predictable page contract

    Write as if you are teaching a capable assistant that lacks your internal context. A useful page contract contains:

    • A short opening that directly answers the page’s primary question.
    • A clear definition of the product, feature, workflow, or integration.
    • Prerequisites and eligibility conditions before the instructions begin.
    • Steps or decision criteria in the order the user needs them.
    • Limitations, tradeoffs, and unsupported cases near the claim they qualify.
    • Links to evidence and deeper documentation.
    • A visible path to the next task, such as setup, troubleshooting, or an API operation.

    Define acronyms where they first appear. Use descriptive headings rather than clever labels. Add concise question-and-answer sections when they match real prompts. Repeat canonical facts consistently, but do not bury the useful answer under repeated positioning language.

    Match the artifact to the intent

    A single generic landing page cannot cover the full journey. Build the artifact that makes the intended answer defensible:

    • Category explainers should define the problem, the common workflow, the relevant buyer, and the boundaries of the category.
    • Use-case pages should connect a specific user job to product capabilities and show the conditions under which the fit holds.
    • Comparison pages should state points of parity, meaningful differences, user fit, limitations, and migration considerations without turning every dimension into a victory claim.
    • Quick starts should identify prerequisites, the setup sequence, the first observable success, and common failure paths.
    • Integration pages should state supported objects or workflows, authentication requirements, data direction, limitations, and links to the relevant API or setup instructions.
    • Troubleshooting pages should connect symptoms to likely causes, corrective steps, and a way to verify that the fix worked.
    • Release notes and changelogs should make changing availability, behavior, and terminology explicit.

    Comparison content deserves particular care because it directly affects product positioning. Do not hide obvious points of parity or invent distinctions that a buyer cannot verify. Explain where the alternatives differ, who benefits from each difference, and when the distinction should change the decision. Honest limits make the rest of the page more credible.

    Maintain a claim ledger behind these pages. Record the exact claim, its evidence, the public locations where it appears, its owner, and the event that should trigger review. A product rename, integration change, policy update, or feature release should update the ledger and the affected pages together. This is how content operations become part of product operations.

    Layer authority, live retrieval, and useful actions

    AI visibility can happen at different layers. Treating them as one channel makes diagnosis difficult:

    1. Public-footprint visibility comes from a clear, consistent body of information that helps an engine recognize the brand and its category.
    2. Retrieval visibility happens when the engine or an attached workflow fetches current material during the conversation.
    3. Action visibility happens when a connector or tool lets the user complete a task through the assistant.

    The public footprint needs distribution as well as first-party content. Keep product facts consistent across documentation, API references, GitHub repositories, partner directories, reputable media, conference material, and legitimate third-party reviews. Pursue inclusion in structured knowledge bases such as Wikidata only when the brand meets the relevant eligibility requirements.

    Do not manufacture authority through fabricated claims, fake reviews, or spammy link schemes. Those tactics create contradictions and reputational risk. The durable strategy is to be verifiably useful on the surfaces where practitioners already look for answers.

    Live retrieval becomes important when an answer depends on current documentation, account context, or a changing product state. A retrieval-first pipeline should fetch the relevant material before the response is generated. Its quality depends on more than adding documents to an index.

    • Chunk documentation around a coherent task or concept rather than breaking related instructions apart.
    • Carry the heading and parent context with each chunk so a retrieved paragraph retains its meaning.
    • Add metadata for product, feature, version or status, intent, update state, and access permissions.
    • Prefer canonical documentation when duplicate explanations compete.
    • Return citations or document identifiers that allow the answer to be checked.
    • Test retrieval against the same prompt families used for visibility measurement.

    A ChatGPT connector or CustomGPT workflow adds the action layer. Publish a high-quality OpenAPI specification, keep each action narrowly scoped, and describe its inputs, permissions, output, and failure conditions clearly. The assistant should be able to choose the correct operation without guessing between overlapping tools.

    Privacy-by-design belongs in the architecture, not in a warning added after launch. Enforce the user’s permissions before retrieval, preserve tenant boundaries, minimize the data passed into the model context, and keep secrets out of indexed content. If an action changes data or creates an external consequence, use clear confirmation and guardrails appropriate to that action.

    A connector does not replace the public footprint. It improves accuracy and task completion for users who can access it. Public explanations still establish category relevance, authority, and discoverability before the user invokes a tool.

    Measure visibility as a product system, not a screenshot

    A favorable answer copied into a presentation is not a measurement system. Answer behavior can vary with wording, context, model configuration, accessible material, and tool availability. Build a stable panel of priority prompts and track its outputs over time.

    Each prompt in the panel should have an intent identifier, target user, task, wording, expected eligibility condition, claims that must be correct, and an artifact owner. Include natural variants across category discovery, evaluation, setup, integration, and troubleshooting. Preserve the panel long enough to compare changes instead of rewriting it after every result.

    Score more than whether the name appeared:

    • Eligible mention rate: how often the brand appears when the predefined fit conditions are present.
    • Grounded citation rate: how often the answer points to appropriate first-party or credible third-party evidence.
    • Factual accuracy: whether the answer passes a predefined set of product facts.
    • Positioning accuracy: whether the brand is placed in the right category, use case, and competitive context.
    • Freshness: whether changing capabilities and product status match the canonical record.
    • Retrieval success: whether the workflow returns the document needed for the task.
    • Action completion: whether an enabled connector completes the intended task under the correct permissions.

    Share of voice can help, but only within eligible prompts. A rising mention rate paired with falling accuracy is not progress. Nor is a citation useful when it points to an outdated page.

    Use the failure pattern to choose the next intervention:

    • If the brand is absent across an entire intent family, inspect coverage, category clarity, and external authority.
    • If it appears under the wrong category, reconcile names and definitions across the canonical record and public properties.
    • If it appears without evidence, strengthen the relevant artifact and its links to documentation or proof.
    • If the facts are stale, repair canonical pages, release notes, metadata, and duplicate content.
    • If retrieval returns the wrong page, adjust chunking, metadata, canonical preference, and evaluation queries.
    • If the answer is correct but the action fails, inspect the OpenAPI description, authentication, permissions, inputs, and error handling.

    Test changes with the same discipline used for a product experiment. State the hypothesis before shipping. Freeze the evaluation rubric. Capture a baseline, compare the candidate under the same conditions, and use repeated samples rather than interpreting one convenient response. Use an A/B design only where exposure can be isolated; otherwise label the result as a before-and-after observation and avoid claiming causality.

    Set the minimum detectable effect before reviewing the outcome. In this context, it is the smallest improvement large enough to justify a decision. That prevents a tiny movement in a noisy prompt panel from becoming a success story merely because the team wants the release to work.

    Assign ownership by failure class. Product marketing can own canonical positioning, documentation can own instructional accuracy, the web team can own crawlability and structured markup, engineering can own retrieval and connectors, and product or analytics can own the evaluation panel. A shared dashboard is useful only when each red metric has a named route to action.

    Key takeaways

    • Optimize for eligibility in a real user decision, not for raw brand-name frequency.
    • Establish one canonical brand fact record before adding more public content.
    • Publish answer-shaped artifacts for category, comparison, setup, integration, troubleshooting, and product-change intents.
    • Combine a trustworthy public footprint with live retrieval and carefully scoped actions.
    • Measure mentions, citations, accuracy, freshness, retrieval, and task completion separately.
    • Tie every content or technical change to a hypothesis, a stable prompt panel, and a minimum detectable effect.

    Start with the prompt family closest to a real buying, activation, integration, or support decision. Capture the baseline answer, identify the smallest missing or unreliable artifact, fix it, and rerun the same evaluation. Expand to adjacent intents only after the first one produces consistently accurate, well-grounded answers.

    The goal is not to make an assistant say your name. It is to make your brand a defensible inclusion for the right question, supported by current evidence and a working next step.

    References

  • How I Use ChatGPT to Supercharge PM: Smart Workflows, Killer Prompts, and Real-World Wins

    How I Use ChatGPT to Supercharge PM: Smart Workflows, Killer Prompts, and Real-World Wins

    Every week, I lean on ChatGPT to cut through noise, reduce rework, and move faster with more confidence. It’s not a silver bullet, but it has become an unfair advantage in my day-to-day leadership of product strategy, discovery, and delivery. Unlock workflows, prompts, and real PM tips showing how ChatGPT quietly reshapes product management behind the scenes.

    Here’s my stance: ChatGPT doesn’t replace product judgment. It amplifies it. Used well, it accelerates product discovery, clarifies roadmaps, sharpens positioning, and strengthens stakeholder management. Used poorly, it creates noise and risk. What follows are the specific workflows and prompts that reliably save me hours while protecting quality and trust.

    Discovery and research are where I see the biggest upside. I use ChatGPT to draft interview guides, transform raw notes into theme clusters, and generate “Jobs to Be Done” problem statements—then I validate them with customers. I anonymize inputs to protect privacy and follow privacy-by-design and data governance commitments; AI risk management matters more than ever when we’re handling real user data.

    When I move from insight to definition, ChatGPT helps me spin up crisp PRDs and user stories. I provide context about our users, constraints, and success metrics and ask for structured outputs: goals, non-goals, acceptance criteria, and risks. This keeps our product trios aligned and focused on outcomes vs output OKRs, not just shipping features.

    For competitive analysis and positioning, I feed in public information and ask for points of parity, points of differentiation, and potential messaging angles. I treat the output as a starting point for my value proposition and battlecards—not the final word. It’s a fast way to surface hypotheses and pressure-test our product-led growth narrative.

    Roadmapping and sprint planning also benefit. I use ChatGPT to map dependencies, draft milestone narratives, and transform epics into well-formed backlogs. When we align quarterly plans, I ask for risk scenarios and contingency options so we can make trade-offs explicit before we commit.

    On analytics and experiments, ChatGPT is my drafting partner. It helps me define A/B testing plans, clarify the minimum detectable effect (MDE), and outline instrumentation requirements. I still verify numbers in our analytics stack, but the scaffolding is done in minutes, not hours—freeing me to focus on retention analysis and activation levers.

    Stakeholder communication is where the time savings compound. I use ChatGPT to produce executive summaries, QBRs vs OKRs comparisons, and board-ready narratives that highlight outcomes, risks, and next steps. It’s a powerful way to stay crisp and consistent across leadership updates without losing the nuance that matters.

    Prompt patterns make or break results. I keep four rules: set the role, provide rich context, define constraints, and specify the output format. For example: “You are a senior PM advisor. Context: [user, market, problem]. Constraints: [privacy, timeline, budget]. Output: PRD with goals, acceptance criteria, and risks.” With larger inputs, I use context window management by chunking content and asking for summaries before synthesis.

    For internal knowledge, I lean on a retrieval-first pipeline. Instead of pasting long docs, I reference curated, approved sources so answers track to current reality. CustomGPT workflows and a simple ChatGPT connector help with governance: they increase speed while reducing the chance of hallucinations and stale information.

    Guardrails are non-negotiable. We never paste sensitive data into prompts; we redact PII, spot-check against source-of-truth systems, and red-team important outputs. AI risk management isn’t just a checkbox—it’s how we maintain trust while scaling productivity with gen ai.

    Finally, enablement turns personal productivity into team capability. I run short playbooks for empowered product teams: discovery synthesis, PRD drafting, roadmap storytelling, and stakeholder-ready updates. The result is higher-quality thinking, faster cycles, and fewer meetings to align on the essentials.

    ChatGPT for product managers isn’t hype; it’s a practical edge when you apply discipline. Start with one workflow that drains your time, add a prompt template, and measure the outcome. In a week, you’ll have proof. In a quarter, you’ll have a new operating system for how your team learns, decides, and ships.


    Inspired by this post on Product School.


    Book a consult png image
  • Evidence-Based Product Marketing: From Claims to Behavior

    Evidence-Based Product Marketing: From Claims to Behavior

    Your campaign can beat its click target and still fail. If the message attracts people who never reach value, the dashboard is reporting distribution, not evidence that the promise worked.

    The practical fix is to connect each important product marketing claim to an expected customer response, an observable product behavior, and a business decision. That chain gives you something stronger than a collection of campaign metrics: it tells you what to scale, what to revise, and what to stop.

    Start with the decision, not the dashboard

    Evidence-based product marketing does not mean attaching a metric to every asset. It means deciding what must be true for a claim to deserve more investment, then collecting evidence capable of answering that question.

    Begin by naming the decision in plain language. Most product marketing work needs to answer one of four questions:

    • Clarify: Do the intended customers recognize themselves, understand the problem, and repeat the outcome accurately?
    • Launch: Does the message motivate the right people to take the next meaningful step?
    • Scale: Does the campaign create incremental activation or qualified demand without damaging the customer experience?
    • Standardize: Does the promise continue to hold after acquisition, through early value, retention, and commercial outcomes?

    Those decisions require different evidence. Customer interviews can reveal whether the language is clear. Funnel data can show whether exposed customers behave differently. A controlled experiment can isolate the effect of a headline or narrative. Retention and revenue can show whether the acquired behavior was durable. No single metric answers all four questions.

    I find it useful to write the evidence chain before discussing creative execution:

    1. Claim: What outcome are you promising?
    2. Interpretation: What should the intended customer understand or believe?
    3. Immediate action: What is the next meaningful behavior if the message resonates?
    4. Product consequence: Which first-value or activation milestone should improve?
    5. Durable consequence: What should happen to early engagement, retention, or revenue?
    6. Decision: What will you do if the evidence supports, weakens, or contradicts the claim?

    Consider a hypothetical claim that customers can reach first value with less setup. The predicted consequence is not merely a higher click-through rate. Eligible customers should complete the relevant onboarding milestone more often or reach it sooner. If more people start but activation does not improve, the message may be generating curiosity, setting the wrong expectation, or attracting the wrong audience. The evidence should lead you to revise the claim or targeting, not celebrate the larger top of funnel.

    For category education or an unfamiliar product, immediate purchase may be the wrong primary outcome. You still need a defined next behavior, such as exploring the relevant use case, beginning an evaluation, or returning for deeper consideration. The point is not to force every campaign into a purchase funnel. It is to stop treating attention as self-validating.

    Turn positioning into a testable claim card

    Positioning becomes useful when it can survive contact with customers and product data. A strong positioning foundation makes explicit who the product serves, which urgent problem it owns, the category customers recognize, the outcome it promises, its points of parity, its differentiation, and the proof behind the promise.

    Put those elements into a one-page claim card. This is the contract between product marketing, product management, analytics, sales, and the product experience:

    Claim-card fieldQuestion it must answerWhat to record
    Audience and contextExactly who should recognize this problem?The narrowest viable segment, situation, and trigger
    ProblemWhat costly or frustrating job needs to be solved?Customer language, not an internal feature description
    CategoryWhat familiar frame helps the buyer understand the product?The recognized category and likely comparison set
    Outcome claimWhat changes for the customer?One outcome stated without feature soup
    Points of parityWhich table-stakes expectations must be met?The capabilities buyers reasonably assume
    DifferentiationWhy choose this over the primary alternative?Two or three defensible distinctions, not a feature inventory
    Current proofWhy should the buyer believe the promise?Relevant results, usage, social proof, or integrations that actually exist
    Behavioral predictionWhat should a persuaded customer do next?A named event, milestone, or qualified sales action
    Disconfirming signalWhat result would force a revision?A failure condition decided before launch

    The last two rows change positioning from an assertion into a hypothesis. They also expose weak claims early. If nobody can name the behavior that should change, the claim is probably too abstract. If nobody can describe a result that would disconfirm it, the team is preparing to rationalize any outcome.

    For a hypothetical workflow product, a claim card might predict that a simpler setup promise will increase completion of the first workflow and shorten time to activation. The test should also protect early feature engagement and retention. If trial starts rise while first-workflow completion stays flat, the message has increased acquisition without delivering better customer progress. That is evidence against scaling the current version, even if the campaign dashboard looks healthy.

    You can produce a first claim card in a focused 30-minute working session: spend five minutes on the target and problem, five on the category, ten on the outcome plus parity and differentiation, five on available proof, and five defining a customer-language check and a controlled message test. Keep the result to one page. Its job is to drive a decision, not become another positioning deck.

    Do not merge language evidence with performance evidence. When customers repeat your value proposition accurately, you have evidence of comprehension. When their behavior changes, you have evidence of consequence. When a controlled comparison isolates the message as the cause, you have causal evidence. Each answers a different question.

    Instrument the path from exposure to durable value

    A claim cannot be evaluated if campaign exposure and product behavior live in disconnected systems. Before launch, define the path you need to observe and make sure the identifiers survive every handoff.

    At minimum, campaign and product events need stable properties that identify the message and its context. Useful fields include campaign_id, creative_theme, entry_channel, audience_mood, and landing_variant. Use only properties your team can define and populate reliably. A sophisticated taxonomy filled with ambiguous or missing values creates false precision.

    Map the journey in the order the customer experiences it:

    1. Qualified exposure: The intended message and variant were actually delivered to an eligible person.
    2. Meaningful entry: The person took the next action implied by the campaign rather than producing a passive page view.
    3. First value: The person reached the earliest product moment that demonstrates the promised outcome.
    4. Activation: The person completed the behavior or set of behaviors associated with becoming a viable user.
    5. Early depth: The activated person used the relevant capability beyond the minimum milestone.
    6. Retention: The person returned and repeated a valuable behavior in the time window appropriate to the product.
    7. Commercial outcome: The journey produced qualified pipeline, conversion, revenue, or expansion where those outcomes apply.

    Your activation definition must belong to the product, not the campaign. A landing-page scroll is not activation simply because it is easy to measure. Choose a milestone that represents real progress toward value, document its event logic, and use the same definition in the campaign analysis, product dashboard, and decision log.

    Audit the measurement path before spending heavily on distribution:

    • Confirm that event names and triggers have one documented meaning.
    • Verify that the assigned creative and landing variants are preserved after the first session.
    • Test the transition from an anonymous visitor to a known account or user.
    • Check that campaign and product timestamps use a consistent interpretation.
    • Make sure CRM integration carries the identifiers needed to connect marketing exposure with qualified sales outcomes.
    • Document exclusions such as employees, test accounts, bots, duplicate events, and ineligible users.
    • Inspect missing-property rates and unexpected values before trusting segment comparisons.

    Do this with test records that you can trace from the first campaign event to the final system. A dashboard rendering successfully does not prove that identity resolution, variant assignment, or CRM handoffs are correct.

    Once the data is trustworthy, cohort customers by creative theme, channel, audience, or landing variant. That analysis can reveal whether one narrative is associated with faster activation or stronger retention. It does not, by itself, establish that the narrative caused the difference. Channels often reach different people, and audiences can arrive with different levels of intent. Use cohort analysis to find patterns and controlled experiments to test causal claims.

    Match the strength of the evidence to the claim

    Evidence is not a binary label. A customer interview, a funnel comparison, and a randomized experiment can all be useful, but they support different statements. The language in your readout should reflect that difference.

    • Customer-language evidence supports statements about relevance, comprehension, vocabulary, and objections. It helps you learn why a claim makes sense or fails to land.
    • Observed behavioral evidence supports statements about association. It can show that a campaign cohort activated or retained differently, but other differences between the cohorts may explain the result.
    • Experimental evidence supports an incremental claim when assignment, exposure, measurement, and analysis are sound. It helps isolate the effect of a narrative, headline, or creative treatment.
    • Durability evidence supports the commercial importance of a result. It tests whether an early lift reaches activation, retention, and revenue instead of ending with a shallow conversion.

    That distinction prevents a common reporting error: using a strong verb with weak evidence. Say that a theme was associated with higher activation when you observed cohorts. Say that it caused an incremental change only when the design supports that conclusion. If the evidence is directional, label it directional.

    Write the test brief before launching the variant

    A useful A/B test brief should fit on one page and contain the following:

    1. Hypothesis: For a named audience, changing one defined message should change one expected behavior because of a stated reason.
    2. Eligibility and exposure: Specify who enters the test and what counts as seeing the treatment.
    3. Assignment unit: Decide whether assignment happens at the user, account, or another appropriate level, then keep that assignment stable.
    4. Primary metric: Choose the single outcome that answers the decision question. Supporting metrics can diagnose the mechanism, but they should not compete for the verdict.
    5. Business threshold: State the smallest improvement that would justify implementation or further investment.
    6. Minimum detectable effect: Size the test around an explicit MDE so you know which effects the design can and cannot resolve.
    7. Guardrails: Protect the experience with relevant checks such as activation, retention, or NPS. Match the guardrail to the test horizon; some retention and sentiment outcomes need a later read.
    8. Segments: Predefine any audience cuts that could change the decision. Treat unplanned segment findings as hypotheses for another test.
    9. Decision rule: Write what you will do if the primary metric improves, remains unresolved, or moves against the claim.

    The business threshold and MDE are related, but they are not automatically the same. The first asks which effect is worth acting on. The second describes which effect the planned test is equipped to detect. If the design can detect only effects much larger than the improvement you care about, the test cannot settle the decision. Change the design, gather more eligible traffic, or narrow the claim instead of treating an inconclusive result as proof of no effect.

    Low-volume teams still need discipline. When a well-powered test is not practical, use session quality, content depth, return visits, and other directional signals to understand the path, then combine them with customer language and sales objections. Keep the conclusion modest. Directional evidence can justify another iteration; it should not be rewritten as causal proof.

    Also look beyond a positive average. A message may improve trial starts while reducing activation, attract one segment while confusing another, or pull forward behavior that would have happened anyway. The primary metric gives you a verdict on the declared hypothesis. Guardrails and predefined segments tell you whether acting on that verdict is responsible.

    Make the evidence change what the team does

    Measurement creates value only when it changes positioning, distribution, onboarding, the roadmap, or sales execution. That requires one operating cadence and one record of the decision.

    Carry the same promise through the surfaces that customers encounter. The category and value proposition should remain coherent across campaigns, pricing, product tours, onboarding guidance, CRM notes, and sales collateral. Consistency does not mean repeating identical copy. It means the product experience delivers the outcome that marketing introduced.

    Use a shared dashboard or notebook, annotate launches and instrumentation changes, and review the evidence with product and go-to-market partners on a weekly cadence. A useful review answers six questions:

    1. Which claim and audience are under review?
    2. Was exposure delivered as intended, and is the measurement path healthy?
    3. What happened to the declared primary metric?
    4. What happened to activation, retention, experience, and commercial guardrails that are mature enough to read?
    5. Which result is causal, associated, directional, or still unresolved?
    6. What decision follows, who owns it, and when will the next evidence arrive?

    Record the answer in an evidence ledger rather than leaving it in a meeting. For every important claim, capture its audience, product version or context, evidence type, primary result, guardrails, known limitations, status, decision, owner, and review date. Useful statuses include untested, directional, supported in a defined context, contradicted, and stale.

    The context matters. A message supported for one audience, channel, or product experience has not been validated everywhere. Product changes can also make old proof stale. Reopen the claim when the promised workflow changes, the target segment expands, or a new channel reaches customers with materially different intent.

    This operating model also sharpens accountability. Product marketing owns the clarity and integrity of the claim. Product management connects it to value and activation. Analytics protects definitions and interpretation. Sales contributes objection patterns and qualified outcomes. Customer success contributes evidence about expectation gaps and durable value. The exact ownership can vary, but the claim, metric, and decision cannot be ownerless.

    Keep campaign output separate from customer outcomes. Shipping a landing page, launching a narrative, or producing enablement is work completed. Activation, retention, qualified demand, and revenue are outcomes. Reviewing outcomes rather than celebrating output makes it harder for an attractive campaign to survive after the customer evidence turns against it.

    Key takeaways

    • Start with the product marketing decision, then choose the evidence capable of supporting it.
    • Convert positioning into a claim card with an audience, outcome, proof, behavioral prediction, and disconfirming signal.
    • Instrument the complete path from qualified exposure through first value, activation, retention, and commercial outcomes.
    • Treat customer language, observed behavior, experiments, and durability as different forms of evidence.
    • Define the primary metric, MDE, guardrails, segments, and decision rule before reading test results.
    • Keep an evidence ledger so supported claims are reused, contradicted claims are retired, and old proof does not quietly become permanent truth.

    Before your next campaign, take its strongest claim and complete one claim card. Confirm that the campaign identifier reaches the activation event, name one primary metric and one guardrail, and write the decision rule before launch. If you cannot trace the promise to customer value, fix that measurement path before buying more attention.

    References

  • Taming 1,000+ Vendor Emails: How Xelix’s AI Helpdesk Delivers Fast, Confident Answers

    Taming 1,000+ Vendor Emails: How Xelix’s AI Helpdesk Delivers Fast, Confident Answers

    Chaos in vendor communications is a problem I see across finance operations: sprawling accounts payable inboxes, slow response times, and missed context. That’s why this build caught my attention—not just because it’s GenAI, but because it’s a disciplined product strategy that converts email overload into measurable outcomes.

    Accounts payable inboxes can see 1,000+ vendor emails a day. Xelix’s new Helpdesk turns that chaos into structured tickets, enriched with ERP data, and pre-drafted replies—complete with confidence scores.

    I dug into the end-to-end approach with the team—Claire Smid — AI Engineer, Xelix; Emilija Gransaull — Back-End Tech Lead, Xelix; Talal A. — Product Manager, Xelix—focusing on how they scoped the problem, iterated fast, and de-risked AI in production.

    Their product thesis is refreshingly pragmatic. They prototyped with “daily slices” (Carpaccio-style) and built a retrieval-first pipeline that matches vendors, links invoices, and drafts accurate responses—before a human ever clicks “send.” That framing matters: enrichment and matching take center stage, with the model amplifying precision instead of improvising.

    We unpacked the tricky bits that make or break an AI helpdesk at scale: vendor identity matching, Outlook threading, UX pivots from “inbox clone” to ticket-first views, and the metrics that prove real impact (handling time, stickiness, auto-closed spam). The pipeline architecture and email processing choices were grounded in operational realities, not just AI aspirations.

    Several takeaways are worth pinning to any AI product roadmap. “Start narrow to win: pick high-volume, high-cost requests (invoice status & reminders).” “Enrichment > magic: accurate replies come from great retrieval/matching, not just a bigger LLM.” “Design for adoption: familiar inbox view helps onboarding, but a ticket-first UI unlocks AI features.” These are the kinds of decisions that drive adoption, trust, and ROI.

    Data enrichment challenges dominated early learning curves: stitching ERP context into tickets, handling vendor identification at scale, managing email thread continuity, and calibrating response generation for accuracy. On the generation side, the team emphasized precision over verbosity—clean responses that reflect system-of-record truth—then instrumented the experience to “Evaluate System Performance” with production-grade telemetry.

    Trust was treated as a product feature. “Measure outcomes, not vibes: track ‘messages sent from Helpdesk’, % auto-resolved.” And critically, “Confidence builds trust: show match quality and response confidence so humans know when to edit.” By surfacing match quality and confidence scores, they shortened coaching loops and made human-in-the-loop supervision feel natural, not burdensome.

    What’s next is equally compelling: “targeted generation, multiple specialized responders, and more agentic routing.” That direction aligns with agentic AI patterns I recommend for operations-heavy workflows—route first, retrieve deeply, then generate with intent. It’s a scalable path from assistive AI to autonomous resolution while maintaining governance and auditability.

    If you want a quick map of the journey, the conversation flowed from 0:00 Meet the Team: Claire, Emilija, and Talal, 00:36 Introduction to Xelix and Its Products, 01:08 Understanding Accounts Payable Teams, 01:37 Help Desk Product Overview, 03:11 Challenges Faced by Accounts Payable Teams, 04:03 AI Integration in Help Desk, 05:47 Automating Reconciliation Requests, 07:45 Development Methodology: Carpaccio, 09:11 Prototyping and Beta Testing, 12:00 Manual Tagging and Data Collection, 16:39 Focusing on High-Impact Use Cases, 18:55 User Experience and Interface Design, 24:56 Pipeline Architecture and Email Processing, 28:21 Data Enrichment Challenges, 29:04 Handling Vendor Identification, 33:33 Email Thread Management, 36:15 Generating Accurate Responses, 40:48 Evaluating System Performance, 49:20 Future Developments and Goals.

    My takeaway for product leaders: when the domain is high-volume and rules-heavy (like AP), retrieval-first beats model-first. Start with the narrowest, costliest intents; prove lift with “messages sent from Helpdesk” and “% auto-resolved”; then graduate UX from familiar to AI-native (ticket-first) once trust is earned. That’s how you turn vendor chaos into answers—reliably, scalably, and fast.


    Inspired by this post on Product Talk.


    Book a consult png image
  • Inside-Out vs Outside-In: How I Balance Both to Build Products Users Love—and CFOs Trust

    Inside-Out vs Outside-In: How I Balance Both to Build Products Users Love—and CFOs Trust

    Inside-out or outside-in thinking? I choose both. The strongest product strategies fuse a bold internal vision with relentless customer evidence, creating a flywheel that lifts adoption, engagement, and revenue while reducing risk.

    When I lead with inside-out thinking, I articulate a clear product thesis, technical roadmap, and platform leverage. This is where we define points of parity and differentiation, sharpen our value proposition, and ensure our architecture scales. It’s disciplined, outcomes-first, and anchored in product positioning—not output checklists.

    Outside-in thinking ensures that vision stays honest. I listen to customers, analyze friction in onboarding, instrument user activation, and study retention analysis to validate whether our promises translate into real user value. This is where product discovery, A/B testing, and in-app signals tell me what’s working, what needs refinement, and what we should stop doing.

    In practice, I operationalize this balance through Software Experience Management. “Increase revenue, cut costs, and reduce risk with Pendo’s Software Experience Management platform. Optimize the entire software experience to drive adoption and improve engagement.” That promise captures the core of how I align strategy with reality inside the product, not just around it.

    Concretely, I combine product analytics with in-app guides and product tours to accelerate onboarding and improve user activation. I run targeted experiments to de-risk decisions, and I iterate quickly based on what users actually do—not just what they say. The result is a product-led growth engine that compounds over time.

    This approach also builds trust with finance and go-to-market partners. Inside-out clarity gives us confident, sequenced bets; outside-in data provides proof that those bets pay off. When engagement expands and adoption climbs, the business case writes itself.

    If you’re deciding where to start, begin with three moves: define activation events aligned to your value proposition, instrument the experience end-to-end, and ship one high-impact in-app guide to remove a known onboarding blocker. Then measure, learn, and iterate—quickly.

    The truth is, great products emerge when conviction meets evidence. Inside-out sets the vision. Outside-in earns the right to scale it.


    Inspired by this post on Pendo – Perspectives.


    Book a consult png image
  • AI Won’t Replace Engineers—Engineers Using AI Will: A Practical Playbook for Your Next Move

    AI Won’t Replace Engineers—Engineers Using AI Will: A Practical Playbook for Your Next Move

    Will AI replace software engineers or reshape their roles? Explore risks, opportunities, and alternative career paths in tech.

    I’m often asked whether AI will make software engineers obsolete. My short answer: AI is already automating tasks, not eliminating the role. The engineers who learn to orchestrate models, systems, and stakeholders will create more value—not less. The real shift is from keystrokes to judgment, from writing code to designing socio-technical systems that deliver outcomes.

    Today’s gen ai assistants—think Claude Code and ChatGPT connector—excel at unit test scaffolding, boilerplate generation, refactoring, docstrings, and code search. When integrated into CI/CD, they can open draft pull requests, annotate diffs, and propose fixes. This lifts developer productivity and frees time for higher-leverage work: problem framing, architecture decisions, and customer discovery.

    What changes in the role? We spend more cycles on product discovery, privacy-by-design, and AI Strategy, and fewer on repetitive implementation. We design agentic AI workflows that combine retrieval, tools, and guardrails; we evaluate trade-offs that blend performance, cost, and safety; and we partner with empowered product teams to ship the smallest valuable slice, learn, and iterate.

    Measure what matters. If AI is working, DORA metrics should improve: higher deployment frequency, shorter lead time for changes, stable change failure rate, and faster MTTR. Pair that with outcomes vs output OKRs to avoid gaming the system—shaving seconds off a build is meaningless if it doesn’t move activation, retention, or revenue. A unified analytics platform can help connect engineering signals to business impact.

    Risk is real—and manageable. AI risk management and data governance are now core competencies, not afterthoughts. Protect IP with robust access controls, context window management, and red-teaming. In production, instrument threat detection and response to catch prompt injection, data leakage, and model drift. Treat this like any other reliability discipline alongside SRE.

    If parts of coding get automated, where can great engineers thrive? Several high-impact paths are emerging: platform engineering for LLMs (tooling, evals, observability), SRE for AI-infused systems, developer evangelism and education, product management for AI-native experiences, security engineering focused on model and data threats, and forward deployed engineers who pair with customers to solve messy, real-world problems.

    How to upskill fast: build an AI product toolbox and ship small. Prototype gen ai features end-to-end—retrieval, function calling, human-in-the-loop QA—and connect them to your CRM integration or support stack. Use A/B testing with a clear minimum detectable effect (MDE) to validate impact. Leverage CustomGPT workflows for internal enablement and in-app guides or product tours to onboard users safely.

    Here’s a pragmatic 90-day plan. Week 0–2: audit your top 10 engineering tasks by time spent; identify 3 that are ripe for AI augmentation. Week 3–6: pilot inside CI/CD with explicit guardrails; track DORA metrics and developer sentiment. Week 7–10: productionize the wins; document runbooks; add incident management paths. Week 11–12: share learnings with product trios, refine your value proposition, and set next-quarter OKRs.

    AI won’t replace software engineers; engineers who master AI will outpace those who don’t. If we embrace the shift—toward systems thinking, responsible governance, and customer outcomes—we’ll build better products faster and open new, rewarding career paths. The opportunity is here and compounding.


    Inspired by this post on Product School.


    Book a consult png image
  • A Quality System for Trustworthy AI-Assisted UX Research

    A Quality System for Trustworthy AI-Assisted UX Research

    Your AI-generated synthesis can be polished, plausible, and wrong. The dangerous failures are rarely obvious fabrications. They are quieter: a biased sample becomes a universal claim, a participant’s opinion becomes a product need, or a tidy theme loses the contradiction that should have changed the roadmap.

    If you are deciding whether to trust AI-assisted UX research, do not judge the fluency of the summary. Judge the evidence chain behind it. You need to see how a product decision connects to the participants recruited, the questions asked, the underlying observations, the analytical interpretation, and the behavioral data used to check it.

    Key takeaways

    • Research quality is mostly determined before an AI tool sees a transcript. Start with the decision, learning question, and hypothesis.
    • Use AI to accelerate transcription, extraction, tagging, clustering, and contradiction searches. Keep interpretation, confidence, and product judgment under human control.
    • Require every theme to retain its participant coverage, supporting evidence, counterexamples, and unresolved uncertainty.
    • Pair qualitative findings with funnels, cohorts, session evidence, and CRM data when those signals are relevant. Neither qualitative nor quantitative evidence should carry the decision alone.
    • Finish with an atomic insight and a recorded choice. A summary that does not change a decision, test, or learning priority is not finished research.

    Define quality at the decision boundary

    Many teams begin AI-assisted research by asking which model should summarize their transcripts. That is too late in the process. The first quality control is the decision the research must inform.

    Strong discovery begins with a decision statement, an explicit learning goal, and a hypothesis the team is willing to falsify. Without those constraints, an AI system can generate an impressive taxonomy of themes while leaving the actual product question untouched.

    Before recruiting participants or writing prompts, create a short research contract:

    • Decision: Name the choice that is genuinely open. Examples include whether to pursue an opportunity, which problem to solve first, or whether a proposed workflow deserves further testing.
    • Decision condition: State what you would need to learn to proceed, pause, narrow the audience, or reject the current direction.
    • Learning question: Ask about the behavior, context, constraint, or unmet need that makes the decision uncertain.
    • Hypothesis: Write the current belief in a form that evidence could disprove. If every possible interview result would support it, it is not a useful hypothesis.
    • Relevant population: Specify whose behavior matters to this decision and which segments could experience the problem differently.
    • Evidence plan: Identify what interviews can reveal and which behavioral or operational signals could challenge the interpretation.
    • Data boundary: Decide what the AI tool is allowed to receive, what must be removed, and who may review the resulting artifacts.

    This contract changes how you evaluate the output. You are no longer asking whether the summary sounds reasonable. You are asking whether the evidence changes a named choice under stated conditions.

    My standard is simple: a decision-grade insight must survive a skeptical review without relying on the model’s authority. A reviewer should be able to inspect the underlying evidence, see which participants and segments it covers, understand the interpretation applied to it, and identify what remains unknown.

    Keep one distinction visible throughout the work:

    • Observation: What the participant did, described, showed, or failed to complete.
    • Interpretation: What that behavior may mean about a goal, anxiety, constraint, or job.
    • Implication: What the product team may choose to change, test, or leave alone.

    AI can help produce all three, but it should never blur them into a single sentence. Once an inference is written as if it were an observed fact, the rest of the synthesis becomes difficult to audit.

    Protect the signal before AI touches it

    An LLM cannot repair a convenient sample or a leading interview guide. It can only reorganize the resulting bias, often in language that makes the bias look more certain.

    Recruit for the decision, not for convenience

    If you interview only power users, you risk treating advanced workflows as mainstream needs. If you interview only vocal detractors, the roadmap can become a queue of complaints. A more useful recruiting frame includes new users, churned users, people who evaluated but did not convert, and adjacent personas where the decision calls for them.

    Build a participant matrix before outreach. Use rows for the segments that could materially change the decision and columns for relevant states, such as adoption stage, conversion outcome, or workflow maturity. The matrix is not a quota formula. It is a visibility tool. It should make overrepresented groups and missing perspectives obvious.

    Carry that segment metadata into synthesis. A theme that appears among established customers should not silently become a claim about evaluators. When a segment is absent, write that limitation into the insight rather than hiding it in an appendix.

    Ask for behavior before interpretation

    Questions about whether someone likes an idea invite speculation, politeness, and solution theater. Ask about the last relevant event instead. Have the participant reconstruct what triggered it, what they tried, where they hesitated, who else became involved, what workaround they used, and what happened next.

    Neutral, behavior-first questions become stronger when participants can support the account with artifacts such as screenshots or workflow examples. The artifact does not automatically prove the interpretation, but it helps distinguish remembered behavior from a general opinion.

    Pilot the guide with the product trio. Remove product terminology that telegraphs the preferred answer. Check whether each question could produce evidence against the working hypothesis. If the guide repeatedly asks participants to react to your solution, it is a concept evaluation guide, not an open discovery guide. Label it accordingly.

    Set privacy boundaries before uploading transcripts

    Consent to an interview does not automatically settle how AI will be used in transcription, analysis, storage, or sharing. Tell participants how their material will be handled, follow your organization’s data governance requirements, and remove identifiers that are not needed for the decision.

    Do not place sensitive participant data into an unapproved prompt workflow. If the tool’s handling, retention, or access controls have not been approved, keep raw transcripts out of it and work with appropriately de-identified material in an authorized environment. The downside is not merely a poor synthesis; it is unnecessary exposure of participant and customer information.

    De-identification should not erase the context required for analysis. Preserve non-identifying segment labels, workflow stage, and participant codes when they are relevant. The goal is to minimize sensitive data while retaining enough context to audit coverage and interpretation.

    Make AI produce an auditable synthesis

    The most reliable workflow separates extraction from clustering and clustering from judgment. Asking for findings, recommendations, sentiment, and a roadmap in one prompt encourages the model to fill gaps and compress uncertainty.

    1. Prepare the evidence set. Preserve the original transcript or recording, assign a participant code, attach relevant segment metadata, and remove unnecessary identifiers. Do not let an AI-generated summary replace the underlying material.
    2. Extract participant-level observations. Ask the model to work through each participant separately. Capture the behavior or event, its context, the supporting excerpt or evidence location, and any missing information. Do not ask for themes yet.
    3. Review the extraction. Check whether the observation is grounded in the transcript and whether the model has converted an opinion into behavior or inferred a motive the participant did not provide.
    4. Cluster reviewed observations. Group similar evidence only after the participant-level pass. Require each cluster to retain the contributing participant codes, segment coverage, supporting evidence, and meaningful variations.
    5. Search for contradictions. Ask which observations do not fit the cluster, which participants experienced the situation differently, and which alternative explanations remain plausible. Do not treat dissent as noise merely because it makes the summary less tidy.
    6. Draft atomic insights. Turn a defensible pattern into a small evidence packet containing the finding, evidence, coverage, contradictions, confidence rationale, product implication, and unresolved question.
    7. Triangulate relevant claims. Compare the qualitative interpretation with funnels, cohorts, session evidence, in-product paths, or CRM data when those systems contain a useful signal.
    8. Conduct the decision review. A person accountable for the product choice inspects the evidence chain, challenges the interpretation, and records what the team will do or learn next.

    You can make the separation explicit with narrowly scoped prompts.

    Extraction prompt: Use only the supplied transcript. For each relevant event, return the participant code, observed or reported behavior, context, supporting excerpt, evidence location, and uncertainty. Do not merge participants, infer motives, or recommend a solution. Flag information that is missing.

    Clustering prompt: Use only the reviewed observations. Group evidence by shared behavior and context. For every cluster, retain participant codes, represented segments, supporting observations, material variations, counterexamples, and plausible alternative explanations. Do not use repetition in the transcript as a substitute for participant coverage.

    Challenge prompt: Review the proposed themes as a skeptical researcher. Identify unsupported generalizations, segment differences that were flattened, interpretations written as observations, contradictory evidence, and claims that cannot be traced to the supplied material. Do not invent missing evidence.

    Prompt design helps, but it does not replace review. Keep the prompt, relevant tool or model information, input scope, and human corrections with the research artifact. If the synthesis later changes, you should be able to determine whether the cause was new evidence, a different analytical instruction, or a human judgment.

    AI is well suited to accelerating transcription, tagging, theme clustering, Jobs to Be Done extraction, and searches for hesitation or sentiment. Treat the latter outputs as interpretations to validate, not measurements generated by an objective instrument. A sentiment label is useful only when a reviewer can return to the behavior and language that produced it.

    Validate the insight, then record the decision

    A good synthesis review is not a copy-edit. It is an attempt to break the claim before the claim influences a roadmap.

    Run a quality review against the evidence chain

    • Traceability: Can a reviewer move from the insight to the contributing participants and the exact supporting material?
    • Coverage: Does the claim name the segments represented, and does it disclose relevant segments that are missing?
    • Construct validity: Is the finding about the behavior the study intended to understand, or has a nearby opinion been used as a proxy?
    • Separation: Are observation, interpretation, and product implication visibly distinct?
    • Contradiction: Does the artifact preserve disconfirming cases and material variations instead of forcing consensus?
    • Triangulation: Where behavioral data is relevant, does it support, narrow, or challenge the qualitative account?
    • Decision relevance: Does the finding change a live choice, a test, or the next learning priority?

    Do not outsource confidence to the model. A confident tone is a language property, not an evidence assessment. Record confidence as a human rationale based on the clarity of the underlying behavior, the relevance and coverage of participants, consistency and counterexamples, and any corroborating behavioral evidence.

    Quantitative and qualitative signals answer different parts of the question. Funnels, cohorts, and retention analysis can show where behavior changes or where people leave. Interviews and artifacts can expose the goals, anxieties, organizational constraints, and workarounds behind that behavior. Pairing those signals is how a team moves from observing what happened to developing a testable account of why.

    When the signals disagree, do not average them into a vague conclusion. Check whether the interview sample represents the population in the analytics, whether the event instrumentation reflects the behavior being discussed, whether segments have been combined, and whether the evidence refers to the same stage of the journey. A contradiction is often the next research question.

    Use an atomic insight format

    A reusable insight should be small enough to inspect and complete enough to guide a choice. Use this structure:

    • Decision: The product choice this evidence informs.
    • Finding: The observed behavioral pattern and the context in which it occurs.
    • Evidence: Participant codes, excerpts or artifact locations, and any relevant behavioral signal.
    • Coverage: The represented segments and known gaps.
    • Interpretation: The best current explanation, clearly labeled as an inference.
    • Contradictions: Cases or data that weaken, narrow, or complicate the interpretation.
    • Confidence: A short rationale grounded in evidence quality, coverage, consistency, and triangulation.
    • Product implication: The opportunity, risk, constraint, or tradeoff the team should consider.
    • Disposition: Act, test further, monitor, or take no action.
    • Next unknown: The uncertainty most likely to change the decision.

    Useful insight records also prevent familiar synthesis mistakes. Replace a broad label such as onboarding friction with the specific behavior, actor, context, and consequence. Do not let a memorable quotation stand in for a pattern. Do not describe a participant’s requested feature as the underlying need. Do not convert an AI-generated cluster into a roadmap item until the evidence packet survives review.

    Bring the atomic insights to a decision review with the product trio. Record the choice, its rationale, what the team is deliberately not doing, and the evidence that could reopen the decision. Connect the chosen action to an outcome or learning objective rather than treating delivery of a feature as proof that the research was correct.

    For your next study, start with one live decision and run the evidence through this chain. If a theme cannot be traced, mark it as a hypothesis. If participant coverage is lopsided, narrow the claim. If qualitative and behavioral evidence conflict, investigate the conflict before committing the roadmap. That is how AI becomes a fast, inspectable research assistant instead of an unaccountable author of customer truth.

    References

  • How to Evaluate AI Voice Support in Real-World Conditions

    How to Evaluate AI Voice Support in Real-World Conditions

    You have a shortlist of AI voice support products, a polished recording, and a decision that could affect thousands of customer conversations. The hard question is not whether an agent can sound convincing during one ideal call. It is whether the system stays useful when a caller interrupts, corrects themselves, asks an ambiguous question, waits on a backend system, or needs a human.

    You can answer that question before a broad rollout. The method is to test complete support outcomes, introduce controlled complications, score failures separately from conversational polish, and use the result to define a limited production pilot.

    Evaluate the support outcome, not the performance

    A natural voice can create an impression of competence before the agent has done anything useful. Pleasant pacing, expressive speech, and a quick opening matter, but they cannot compensate for retrieving the wrong account, misunderstanding the request, or claiming that an action succeeded when it did not.

    Treat the unit of evaluation as a completed support job. Depending on the intent, that job may require the agent to identify the caller, understand the request, retrieve the right information, explain the answer, perform an authorized action, confirm the resulting state, and send a follow-up or transfer the conversation. If you score only the spoken answer, you leave most of the product untested.

    One live Fin Voice call illustrated this end-to-end standard in about 90 seconds: the agent verified identity, retrieved account information, managed an interruption, presented options, completed a workflow, and sent a follow-up email. That sequence is a useful model for constructing a test. It is not, by itself, proof of reliability across other calls.

    Before anyone places a test call, write an outcome contract for each scenario:

    • Caller goal: What is the person trying to accomplish?
    • Starting state: What customer, account, order, subscription, or case data exists before the call?
    • Available evidence: Which knowledge, policies, and records may the agent use?
    • Permitted actions: What may the agent change, create, send, cancel, or escalate?
    • Required clarification: Which missing or conflicting facts must be resolved before an answer or action?
    • Completion evidence: What observable state proves that the request was resolved?
    • Unacceptable outcome: What error would make the call a failure even if the conversation sounded good?

    This contract prevents a common scoring mistake: confusing non-transfer with resolution. A call can remain inside the AI channel and still leave the customer with a wrong answer, an incomplete action, or no idea what happens next. Conversely, an intentional transfer can be the correct resolution when the agent reaches a policy, permission, or confidence boundary.

    Build scenarios around the ways real calls become difficult

    Start with support intents your operation actually receives. Prioritize intents that are frequent, expensive to handle, important to customer trust, or dependent on multiple systems. Do not begin with trivia questions that merely demonstrate broad language-model knowledge. You are evaluating support execution.

    For every core intent, create a straightforward case and several controlled variants. Keep the customer objective constant while changing one condition at a time. That makes a failure diagnosable instead of merely disappointing.

    A practical scenario matrix

    • Clean path: The caller gives the relevant facts in a clear order. This establishes whether the basic workflow works at all.
    • Missing information: Omit a detail the agent needs. Check whether it asks a focused question instead of guessing or restarting the intake.
    • Ambiguous intent: Use wording that could map to two support issues. The agent should disambiguate before retrieving data or taking action.
    • Mid-call correction: Let the caller change an account detail, date, product, or preferred option. Check whether the corrected fact replaces the old one throughout the workflow.
    • Interruption: Speak while the agent is answering. Observe whether it stops cleanly, understands the new input, and continues from the right point.
    • Backend delay: Introduce a slow retrieval or action. Evaluate how the agent manages the wait and whether it distinguishes a pending operation from a completed one.
    • Backend failure: Make a required system unavailable or return an error. The agent should not fabricate a result or promise completion it cannot verify.
    • Policy boundary: Ask for something the agent is not allowed to do. Test the explanation, alternatives, and escalation path.
    • Human request: Ask directly for a person. Verify that the agent follows the configured policy without turning the handoff into an argument.
    • Listening conditions: If your deployment must support different languages, accents, devices, or noisy environments, test each condition explicitly rather than treating one clear studio call as representative.

    Give testers the goal, account state, and one complication. Do not script every sentence. A fully written dialogue tests whether the agent can follow the dialogue you anticipated; a goal-based scenario tests whether it can manage the conversation the caller actually creates.

    Keep a few variants undisclosed until the live session. This is not a trick. It prevents the evaluation from becoming a memorized path while still keeping every test fair and reproducible. Record the exact variant afterward so another evaluator can run it again.

    Run the call through the systems you expect to deploy

    An unedited live call is more informative than a produced recording, but live alone is not enough. A live test can still use ideal data, a simplified integration, a practiced caller, and a workflow that avoids the hard parts of your environment.

    Ask to run the scenario through a path that resembles the intended deployment:

    1. Place a normal phone call through the proposed telephony route. If production will use call forwarding, test the forwarding path rather than a direct internal endpoint.
    2. Use a safe test account containing representative records, permissions, and history.
    3. Require the agent to retrieve data from the backend system that will be authoritative in production.
    4. Introduce the chosen interruption, correction, ambiguity, delay, or error during the live conversation.
    5. Require a real test action where it is safe to do so, not a verbal description of what the agent would have done.
    6. Inspect the backend state after the call. Confirm that the correct record changed once, with the expected values.
    7. Verify every promised follow-up, case creation, notification, or handoff outside the voice channel.
    8. Retain the recording, transcript, timestamps, tool activity, and final system state for scoring.

    This is especially important when an agent can take consequential actions. A fluent confirmation is not evidence that the action happened. The system of record is the evidence.

    Repeat important scenarios with different wording and a different caller. One successful run demonstrates that the capability can work. Repeated variants reveal whether the capability depends on a narrow phrase, a rehearsed cadence, or an unusually forgiving path.

    Key takeaways

    • Score complete resolution, including backend state and follow-up, rather than voice quality alone.
    • Change one condition at a time so you can identify why a call failed.
    • Test interruptions, corrections, ambiguity, system delays, system errors, and escalation.
    • Measure different kinds of waiting separately; a lookup pause and a turn-detection problem are not the same defect.
    • Treat a successful demo as evidence for a pilot, not permission for an unrestricted rollout.

    Score conversation, reasoning, and operational closure separately

    A single overall rating hides the information you need to make a product decision. The call may sound awkward but reach the correct outcome, or sound excellent while making a dangerous mistake. Separate the evaluation into three layers.

    LayerWhat to inspectEvidence of a passTypical failure
    Conversation mechanicsTurn detection, interruption handling, pacing, response length, and intelligibilityThe caller can speak naturally, correct the agent, and follow the response without fighting for the floorThe agent talks over the caller, leaves confusing silence, or delivers answers too long to retain by ear
    Decision qualityIntent recognition, clarification, use of account context, policy application, and answer accuracyThe agent asks only for missing information, uses the correct evidence, and avoids unsupported conclusionsThe agent guesses, asks redundant questions, ignores a correction, or applies the wrong policy
    Operational closureIdentity checks, tool calls, state changes, confirmation, follow-up, and escalationThe verified backend state matches the caller’s request and the agent’s final explanationThe agent claims success without a completed action, changes the wrong record, duplicates work, or drops context during handoff

    Use a simple 0-2 score for each criterion: 0 for failed or unsupported, 1 for completed with material caller effort or recovery, and 2 for correct and usable. The scale is deliberately small. Evaluators can usually distinguish failure, friction, and success more consistently than they can defend the difference between seven and eight on a ten-point scale.

    Do not average away critical errors. A wrong account action, failed identity control, fabricated completion, or forbidden disclosure should remain visible as a release blocker even if many low-risk calls receive high scores. Record both the criterion scores and the count of critical failures.

    Break latency into moments the caller can feel

    Latency is not one number. Capture at least three moments: the time the agent takes to recognize that the caller has finished, the time it spends reasoning or waiting for a system, and the time needed to begin and complete the spoken response.

    • End-of-turn delay: A long delay after every caller turn makes the exchange feel unresponsive and can encourage both sides to start speaking at once.
    • Reasoning or retrieval delay: A pause can be appropriate when the agent is checking account data or invoking a backend workflow. Brief pauses were audible during live subscription and backend checks, which is more informative than editing those waits out.
    • Response delivery: A fast start does not help if the answer becomes a long monologue. Voice responses need structure and pacing that work for listening, not merely text that sounds acceptable when read.

    Ask what is happening during a pause. If the system is doing useful work, the next statement should reflect that work and the action log should verify it. If the pause is long enough to make a caller wonder whether the call has dropped, the experience needs an appropriate progress cue. If the agent answers instantly but guesses, speed is concealing a quality problem.

    Review individual timings as well as an average. A generally responsive agent with occasional severe stalls creates a different operational problem from one that is consistently a little slow. Your test recordings and timestamps should make both patterns visible without inventing a universal pass threshold that ignores the complexity of the workflow.

    Make recovery and escalation part of the product test

    The strongest voice experiences are not the ones that never encounter confusion. They are the ones that recover without making the caller restart. Recovery is therefore a capability to test, not an embarrassing exception to hide.

    Interrupt the agent in the middle of an answer. Correct a fact it has already used. Add a second request after the first appears resolved. Say that an explanation was unclear. Ask for a human. These moves reveal whether the agent maintains conversational state or merely produces plausible turns one at a time.

    During recovery, look for specific behavior:

    • It stops speaking promptly when the caller takes the turn.
    • It identifies what changed instead of repeating the whole interaction.
    • It replaces corrected information rather than carrying both versions forward.
    • It asks a narrow clarification when the next action is uncertain.
    • It does not claim to understand when the transcript or subsequent action shows otherwise.
    • It preserves verified context and the reason for contact when a human takes over.
    • It tells the caller what will happen next instead of ending on an internal routing label.

    Tone belongs in this test, but not as a beauty contest between synthetic voices. Evaluate whether pacing, brevity, acknowledgement, and word choice suit the moment. A caller correcting a billing detail needs a clear acknowledgement and an accurate update, not theatrical empathy. A caller who sounds uncertain may need a shorter explanation and a confirming question. Tone is the behavior of the conversation, not just the timbre selected in a settings menu.

    Escalation should also count as a valid outcome when it is timely and informed. Define which conditions require a handoff, which allow one, and what context must travel with it. Then test the handoff from the caller’s side. If the customer reaches a person but has to repeat identity, intent, and every attempted step, the routing technically worked while the support experience failed.

    Turn the evaluation into a controlled pilot decision

    A strong live evaluation earns the right to run a pilot. It does not justify sending every eligible call to the agent. Production introduces variation in callers, data quality, traffic, integrations, and issue combinations that a demonstration cannot reproduce fully.

    I would require five gates before approving even a limited external pilot:

    1. Capability gate: Every must-have intent has completed its end-to-end workflow, including at least one controlled complication.
    2. Critical-risk gate: No unresolved failure can expose the wrong account, bypass a required check, perform an unauthorized action, or report a false completion.
    3. Conversation gate: The agent can handle interruptions, corrections, clarification, and explicit human requests without trapping the caller in a loop.
    4. Operations gate: Your team can configure terminology, guidance, escalation behavior, greetings, voice, and deployment controls for the intended support environment.
    5. Learning gate: Owners can inspect recordings, transcripts, tool activity, outcomes, and failures, then change the knowledge, workflow, policy, or conversation design responsible.

    Start the pilot with a reversible slice of traffic and a clear human fallback. Select intents whose correct outcome can be verified in your systems. Define who reviews failed and escalated calls, who can pause the rollout, and who owns each class of fix. An answer-quality issue, a telephony issue, and a backend integration issue require different owners even when the caller experiences all three as one bad call.

    Expand only when observed calls meet the outcome contracts you wrote before the demo. If the definition of success keeps changing after failures appear, the evaluation is no longer protecting the decision.

    For your next vendor session, replace “show me your best call” with a scenario pack, a test account, and a request to inspect the final system state. You will learn more from one imperfect call that recovers correctly than from a flawless recording that never had to recover at all.

    References