Author: Shivam Tiwari

  • Evidence-Based Product Marketing: From Claims to Behavior

    Evidence-Based Product Marketing: From Claims to Behavior

    Your campaign can beat its click target and still fail. If the message attracts people who never reach value, the dashboard is reporting distribution, not evidence that the promise worked.

    The practical fix is to connect each important product marketing claim to an expected customer response, an observable product behavior, and a business decision. That chain gives you something stronger than a collection of campaign metrics: it tells you what to scale, what to revise, and what to stop.

    Start with the decision, not the dashboard

    Evidence-based product marketing does not mean attaching a metric to every asset. It means deciding what must be true for a claim to deserve more investment, then collecting evidence capable of answering that question.

    Begin by naming the decision in plain language. Most product marketing work needs to answer one of four questions:

    • Clarify: Do the intended customers recognize themselves, understand the problem, and repeat the outcome accurately?
    • Launch: Does the message motivate the right people to take the next meaningful step?
    • Scale: Does the campaign create incremental activation or qualified demand without damaging the customer experience?
    • Standardize: Does the promise continue to hold after acquisition, through early value, retention, and commercial outcomes?

    Those decisions require different evidence. Customer interviews can reveal whether the language is clear. Funnel data can show whether exposed customers behave differently. A controlled experiment can isolate the effect of a headline or narrative. Retention and revenue can show whether the acquired behavior was durable. No single metric answers all four questions.

    I find it useful to write the evidence chain before discussing creative execution:

    1. Claim: What outcome are you promising?
    2. Interpretation: What should the intended customer understand or believe?
    3. Immediate action: What is the next meaningful behavior if the message resonates?
    4. Product consequence: Which first-value or activation milestone should improve?
    5. Durable consequence: What should happen to early engagement, retention, or revenue?
    6. Decision: What will you do if the evidence supports, weakens, or contradicts the claim?

    Consider a hypothetical claim that customers can reach first value with less setup. The predicted consequence is not merely a higher click-through rate. Eligible customers should complete the relevant onboarding milestone more often or reach it sooner. If more people start but activation does not improve, the message may be generating curiosity, setting the wrong expectation, or attracting the wrong audience. The evidence should lead you to revise the claim or targeting, not celebrate the larger top of funnel.

    For category education or an unfamiliar product, immediate purchase may be the wrong primary outcome. You still need a defined next behavior, such as exploring the relevant use case, beginning an evaluation, or returning for deeper consideration. The point is not to force every campaign into a purchase funnel. It is to stop treating attention as self-validating.

    Turn positioning into a testable claim card

    Positioning becomes useful when it can survive contact with customers and product data. A strong positioning foundation makes explicit who the product serves, which urgent problem it owns, the category customers recognize, the outcome it promises, its points of parity, its differentiation, and the proof behind the promise.

    Put those elements into a one-page claim card. This is the contract between product marketing, product management, analytics, sales, and the product experience:

    Claim-card fieldQuestion it must answerWhat to record
    Audience and contextExactly who should recognize this problem?The narrowest viable segment, situation, and trigger
    ProblemWhat costly or frustrating job needs to be solved?Customer language, not an internal feature description
    CategoryWhat familiar frame helps the buyer understand the product?The recognized category and likely comparison set
    Outcome claimWhat changes for the customer?One outcome stated without feature soup
    Points of parityWhich table-stakes expectations must be met?The capabilities buyers reasonably assume
    DifferentiationWhy choose this over the primary alternative?Two or three defensible distinctions, not a feature inventory
    Current proofWhy should the buyer believe the promise?Relevant results, usage, social proof, or integrations that actually exist
    Behavioral predictionWhat should a persuaded customer do next?A named event, milestone, or qualified sales action
    Disconfirming signalWhat result would force a revision?A failure condition decided before launch

    The last two rows change positioning from an assertion into a hypothesis. They also expose weak claims early. If nobody can name the behavior that should change, the claim is probably too abstract. If nobody can describe a result that would disconfirm it, the team is preparing to rationalize any outcome.

    For a hypothetical workflow product, a claim card might predict that a simpler setup promise will increase completion of the first workflow and shorten time to activation. The test should also protect early feature engagement and retention. If trial starts rise while first-workflow completion stays flat, the message has increased acquisition without delivering better customer progress. That is evidence against scaling the current version, even if the campaign dashboard looks healthy.

    You can produce a first claim card in a focused 30-minute working session: spend five minutes on the target and problem, five on the category, ten on the outcome plus parity and differentiation, five on available proof, and five defining a customer-language check and a controlled message test. Keep the result to one page. Its job is to drive a decision, not become another positioning deck.

    Do not merge language evidence with performance evidence. When customers repeat your value proposition accurately, you have evidence of comprehension. When their behavior changes, you have evidence of consequence. When a controlled comparison isolates the message as the cause, you have causal evidence. Each answers a different question.

    Instrument the path from exposure to durable value

    A claim cannot be evaluated if campaign exposure and product behavior live in disconnected systems. Before launch, define the path you need to observe and make sure the identifiers survive every handoff.

    At minimum, campaign and product events need stable properties that identify the message and its context. Useful fields include campaign_id, creative_theme, entry_channel, audience_mood, and landing_variant. Use only properties your team can define and populate reliably. A sophisticated taxonomy filled with ambiguous or missing values creates false precision.

    Map the journey in the order the customer experiences it:

    1. Qualified exposure: The intended message and variant were actually delivered to an eligible person.
    2. Meaningful entry: The person took the next action implied by the campaign rather than producing a passive page view.
    3. First value: The person reached the earliest product moment that demonstrates the promised outcome.
    4. Activation: The person completed the behavior or set of behaviors associated with becoming a viable user.
    5. Early depth: The activated person used the relevant capability beyond the minimum milestone.
    6. Retention: The person returned and repeated a valuable behavior in the time window appropriate to the product.
    7. Commercial outcome: The journey produced qualified pipeline, conversion, revenue, or expansion where those outcomes apply.

    Your activation definition must belong to the product, not the campaign. A landing-page scroll is not activation simply because it is easy to measure. Choose a milestone that represents real progress toward value, document its event logic, and use the same definition in the campaign analysis, product dashboard, and decision log.

    Audit the measurement path before spending heavily on distribution:

    • Confirm that event names and triggers have one documented meaning.
    • Verify that the assigned creative and landing variants are preserved after the first session.
    • Test the transition from an anonymous visitor to a known account or user.
    • Check that campaign and product timestamps use a consistent interpretation.
    • Make sure CRM integration carries the identifiers needed to connect marketing exposure with qualified sales outcomes.
    • Document exclusions such as employees, test accounts, bots, duplicate events, and ineligible users.
    • Inspect missing-property rates and unexpected values before trusting segment comparisons.

    Do this with test records that you can trace from the first campaign event to the final system. A dashboard rendering successfully does not prove that identity resolution, variant assignment, or CRM handoffs are correct.

    Once the data is trustworthy, cohort customers by creative theme, channel, audience, or landing variant. That analysis can reveal whether one narrative is associated with faster activation or stronger retention. It does not, by itself, establish that the narrative caused the difference. Channels often reach different people, and audiences can arrive with different levels of intent. Use cohort analysis to find patterns and controlled experiments to test causal claims.

    Match the strength of the evidence to the claim

    Evidence is not a binary label. A customer interview, a funnel comparison, and a randomized experiment can all be useful, but they support different statements. The language in your readout should reflect that difference.

    • Customer-language evidence supports statements about relevance, comprehension, vocabulary, and objections. It helps you learn why a claim makes sense or fails to land.
    • Observed behavioral evidence supports statements about association. It can show that a campaign cohort activated or retained differently, but other differences between the cohorts may explain the result.
    • Experimental evidence supports an incremental claim when assignment, exposure, measurement, and analysis are sound. It helps isolate the effect of a narrative, headline, or creative treatment.
    • Durability evidence supports the commercial importance of a result. It tests whether an early lift reaches activation, retention, and revenue instead of ending with a shallow conversion.

    That distinction prevents a common reporting error: using a strong verb with weak evidence. Say that a theme was associated with higher activation when you observed cohorts. Say that it caused an incremental change only when the design supports that conclusion. If the evidence is directional, label it directional.

    Write the test brief before launching the variant

    A useful A/B test brief should fit on one page and contain the following:

    1. Hypothesis: For a named audience, changing one defined message should change one expected behavior because of a stated reason.
    2. Eligibility and exposure: Specify who enters the test and what counts as seeing the treatment.
    3. Assignment unit: Decide whether assignment happens at the user, account, or another appropriate level, then keep that assignment stable.
    4. Primary metric: Choose the single outcome that answers the decision question. Supporting metrics can diagnose the mechanism, but they should not compete for the verdict.
    5. Business threshold: State the smallest improvement that would justify implementation or further investment.
    6. Minimum detectable effect: Size the test around an explicit MDE so you know which effects the design can and cannot resolve.
    7. Guardrails: Protect the experience with relevant checks such as activation, retention, or NPS. Match the guardrail to the test horizon; some retention and sentiment outcomes need a later read.
    8. Segments: Predefine any audience cuts that could change the decision. Treat unplanned segment findings as hypotheses for another test.
    9. Decision rule: Write what you will do if the primary metric improves, remains unresolved, or moves against the claim.

    The business threshold and MDE are related, but they are not automatically the same. The first asks which effect is worth acting on. The second describes which effect the planned test is equipped to detect. If the design can detect only effects much larger than the improvement you care about, the test cannot settle the decision. Change the design, gather more eligible traffic, or narrow the claim instead of treating an inconclusive result as proof of no effect.

    Low-volume teams still need discipline. When a well-powered test is not practical, use session quality, content depth, return visits, and other directional signals to understand the path, then combine them with customer language and sales objections. Keep the conclusion modest. Directional evidence can justify another iteration; it should not be rewritten as causal proof.

    Also look beyond a positive average. A message may improve trial starts while reducing activation, attract one segment while confusing another, or pull forward behavior that would have happened anyway. The primary metric gives you a verdict on the declared hypothesis. Guardrails and predefined segments tell you whether acting on that verdict is responsible.

    Make the evidence change what the team does

    Measurement creates value only when it changes positioning, distribution, onboarding, the roadmap, or sales execution. That requires one operating cadence and one record of the decision.

    Carry the same promise through the surfaces that customers encounter. The category and value proposition should remain coherent across campaigns, pricing, product tours, onboarding guidance, CRM notes, and sales collateral. Consistency does not mean repeating identical copy. It means the product experience delivers the outcome that marketing introduced.

    Use a shared dashboard or notebook, annotate launches and instrumentation changes, and review the evidence with product and go-to-market partners on a weekly cadence. A useful review answers six questions:

    1. Which claim and audience are under review?
    2. Was exposure delivered as intended, and is the measurement path healthy?
    3. What happened to the declared primary metric?
    4. What happened to activation, retention, experience, and commercial guardrails that are mature enough to read?
    5. Which result is causal, associated, directional, or still unresolved?
    6. What decision follows, who owns it, and when will the next evidence arrive?

    Record the answer in an evidence ledger rather than leaving it in a meeting. For every important claim, capture its audience, product version or context, evidence type, primary result, guardrails, known limitations, status, decision, owner, and review date. Useful statuses include untested, directional, supported in a defined context, contradicted, and stale.

    The context matters. A message supported for one audience, channel, or product experience has not been validated everywhere. Product changes can also make old proof stale. Reopen the claim when the promised workflow changes, the target segment expands, or a new channel reaches customers with materially different intent.

    This operating model also sharpens accountability. Product marketing owns the clarity and integrity of the claim. Product management connects it to value and activation. Analytics protects definitions and interpretation. Sales contributes objection patterns and qualified outcomes. Customer success contributes evidence about expectation gaps and durable value. The exact ownership can vary, but the claim, metric, and decision cannot be ownerless.

    Keep campaign output separate from customer outcomes. Shipping a landing page, launching a narrative, or producing enablement is work completed. Activation, retention, qualified demand, and revenue are outcomes. Reviewing outcomes rather than celebrating output makes it harder for an attractive campaign to survive after the customer evidence turns against it.

    Key takeaways

    • Start with the product marketing decision, then choose the evidence capable of supporting it.
    • Convert positioning into a claim card with an audience, outcome, proof, behavioral prediction, and disconfirming signal.
    • Instrument the complete path from qualified exposure through first value, activation, retention, and commercial outcomes.
    • Treat customer language, observed behavior, experiments, and durability as different forms of evidence.
    • Define the primary metric, MDE, guardrails, segments, and decision rule before reading test results.
    • Keep an evidence ledger so supported claims are reused, contradicted claims are retired, and old proof does not quietly become permanent truth.

    Before your next campaign, take its strongest claim and complete one claim card. Confirm that the campaign identifier reaches the activation event, name one primary metric and one guardrail, and write the decision rule before launch. If you cannot trace the promise to customer value, fix that measurement path before buying more attention.

    References

  • Taming 1,000+ Vendor Emails: How Xelix’s AI Helpdesk Delivers Fast, Confident Answers

    Taming 1,000+ Vendor Emails: How Xelix’s AI Helpdesk Delivers Fast, Confident Answers

    Chaos in vendor communications is a problem I see across finance operations: sprawling accounts payable inboxes, slow response times, and missed context. That’s why this build caught my attention—not just because it’s GenAI, but because it’s a disciplined product strategy that converts email overload into measurable outcomes.

    Accounts payable inboxes can see 1,000+ vendor emails a day. Xelix’s new Helpdesk turns that chaos into structured tickets, enriched with ERP data, and pre-drafted replies—complete with confidence scores.

    I dug into the end-to-end approach with the team—Claire Smid — AI Engineer, Xelix; Emilija Gransaull — Back-End Tech Lead, Xelix; Talal A. — Product Manager, Xelix—focusing on how they scoped the problem, iterated fast, and de-risked AI in production.

    Their product thesis is refreshingly pragmatic. They prototyped with “daily slices” (Carpaccio-style) and built a retrieval-first pipeline that matches vendors, links invoices, and drafts accurate responses—before a human ever clicks “send.” That framing matters: enrichment and matching take center stage, with the model amplifying precision instead of improvising.

    We unpacked the tricky bits that make or break an AI helpdesk at scale: vendor identity matching, Outlook threading, UX pivots from “inbox clone” to ticket-first views, and the metrics that prove real impact (handling time, stickiness, auto-closed spam). The pipeline architecture and email processing choices were grounded in operational realities, not just AI aspirations.

    Several takeaways are worth pinning to any AI product roadmap. “Start narrow to win: pick high-volume, high-cost requests (invoice status & reminders).” “Enrichment > magic: accurate replies come from great retrieval/matching, not just a bigger LLM.” “Design for adoption: familiar inbox view helps onboarding, but a ticket-first UI unlocks AI features.” These are the kinds of decisions that drive adoption, trust, and ROI.

    Data enrichment challenges dominated early learning curves: stitching ERP context into tickets, handling vendor identification at scale, managing email thread continuity, and calibrating response generation for accuracy. On the generation side, the team emphasized precision over verbosity—clean responses that reflect system-of-record truth—then instrumented the experience to “Evaluate System Performance” with production-grade telemetry.

    Trust was treated as a product feature. “Measure outcomes, not vibes: track ‘messages sent from Helpdesk’, % auto-resolved.” And critically, “Confidence builds trust: show match quality and response confidence so humans know when to edit.” By surfacing match quality and confidence scores, they shortened coaching loops and made human-in-the-loop supervision feel natural, not burdensome.

    What’s next is equally compelling: “targeted generation, multiple specialized responders, and more agentic routing.” That direction aligns with agentic AI patterns I recommend for operations-heavy workflows—route first, retrieve deeply, then generate with intent. It’s a scalable path from assistive AI to autonomous resolution while maintaining governance and auditability.

    If you want a quick map of the journey, the conversation flowed from 0:00 Meet the Team: Claire, Emilija, and Talal, 00:36 Introduction to Xelix and Its Products, 01:08 Understanding Accounts Payable Teams, 01:37 Help Desk Product Overview, 03:11 Challenges Faced by Accounts Payable Teams, 04:03 AI Integration in Help Desk, 05:47 Automating Reconciliation Requests, 07:45 Development Methodology: Carpaccio, 09:11 Prototyping and Beta Testing, 12:00 Manual Tagging and Data Collection, 16:39 Focusing on High-Impact Use Cases, 18:55 User Experience and Interface Design, 24:56 Pipeline Architecture and Email Processing, 28:21 Data Enrichment Challenges, 29:04 Handling Vendor Identification, 33:33 Email Thread Management, 36:15 Generating Accurate Responses, 40:48 Evaluating System Performance, 49:20 Future Developments and Goals.

    My takeaway for product leaders: when the domain is high-volume and rules-heavy (like AP), retrieval-first beats model-first. Start with the narrowest, costliest intents; prove lift with “messages sent from Helpdesk” and “% auto-resolved”; then graduate UX from familiar to AI-native (ticket-first) once trust is earned. That’s how you turn vendor chaos into answers—reliably, scalably, and fast.


    Inspired by this post on Product Talk.


    Book a consult png image
  • Inside-Out vs Outside-In: How I Balance Both to Build Products Users Love—and CFOs Trust

    Inside-Out vs Outside-In: How I Balance Both to Build Products Users Love—and CFOs Trust

    Inside-out or outside-in thinking? I choose both. The strongest product strategies fuse a bold internal vision with relentless customer evidence, creating a flywheel that lifts adoption, engagement, and revenue while reducing risk.

    When I lead with inside-out thinking, I articulate a clear product thesis, technical roadmap, and platform leverage. This is where we define points of parity and differentiation, sharpen our value proposition, and ensure our architecture scales. It’s disciplined, outcomes-first, and anchored in product positioning—not output checklists.

    Outside-in thinking ensures that vision stays honest. I listen to customers, analyze friction in onboarding, instrument user activation, and study retention analysis to validate whether our promises translate into real user value. This is where product discovery, A/B testing, and in-app signals tell me what’s working, what needs refinement, and what we should stop doing.

    In practice, I operationalize this balance through Software Experience Management. “Increase revenue, cut costs, and reduce risk with Pendo’s Software Experience Management platform. Optimize the entire software experience to drive adoption and improve engagement.” That promise captures the core of how I align strategy with reality inside the product, not just around it.

    Concretely, I combine product analytics with in-app guides and product tours to accelerate onboarding and improve user activation. I run targeted experiments to de-risk decisions, and I iterate quickly based on what users actually do—not just what they say. The result is a product-led growth engine that compounds over time.

    This approach also builds trust with finance and go-to-market partners. Inside-out clarity gives us confident, sequenced bets; outside-in data provides proof that those bets pay off. When engagement expands and adoption climbs, the business case writes itself.

    If you’re deciding where to start, begin with three moves: define activation events aligned to your value proposition, instrument the experience end-to-end, and ship one high-impact in-app guide to remove a known onboarding blocker. Then measure, learn, and iterate—quickly.

    The truth is, great products emerge when conviction meets evidence. Inside-out sets the vision. Outside-in earns the right to scale it.


    Inspired by this post on Pendo – Perspectives.


    Book a consult png image
  • AI Won’t Replace Engineers—Engineers Using AI Will: A Practical Playbook for Your Next Move

    AI Won’t Replace Engineers—Engineers Using AI Will: A Practical Playbook for Your Next Move

    Will AI replace software engineers or reshape their roles? Explore risks, opportunities, and alternative career paths in tech.

    I’m often asked whether AI will make software engineers obsolete. My short answer: AI is already automating tasks, not eliminating the role. The engineers who learn to orchestrate models, systems, and stakeholders will create more value—not less. The real shift is from keystrokes to judgment, from writing code to designing socio-technical systems that deliver outcomes.

    Today’s gen ai assistants—think Claude Code and ChatGPT connector—excel at unit test scaffolding, boilerplate generation, refactoring, docstrings, and code search. When integrated into CI/CD, they can open draft pull requests, annotate diffs, and propose fixes. This lifts developer productivity and frees time for higher-leverage work: problem framing, architecture decisions, and customer discovery.

    What changes in the role? We spend more cycles on product discovery, privacy-by-design, and AI Strategy, and fewer on repetitive implementation. We design agentic AI workflows that combine retrieval, tools, and guardrails; we evaluate trade-offs that blend performance, cost, and safety; and we partner with empowered product teams to ship the smallest valuable slice, learn, and iterate.

    Measure what matters. If AI is working, DORA metrics should improve: higher deployment frequency, shorter lead time for changes, stable change failure rate, and faster MTTR. Pair that with outcomes vs output OKRs to avoid gaming the system—shaving seconds off a build is meaningless if it doesn’t move activation, retention, or revenue. A unified analytics platform can help connect engineering signals to business impact.

    Risk is real—and manageable. AI risk management and data governance are now core competencies, not afterthoughts. Protect IP with robust access controls, context window management, and red-teaming. In production, instrument threat detection and response to catch prompt injection, data leakage, and model drift. Treat this like any other reliability discipline alongside SRE.

    If parts of coding get automated, where can great engineers thrive? Several high-impact paths are emerging: platform engineering for LLMs (tooling, evals, observability), SRE for AI-infused systems, developer evangelism and education, product management for AI-native experiences, security engineering focused on model and data threats, and forward deployed engineers who pair with customers to solve messy, real-world problems.

    How to upskill fast: build an AI product toolbox and ship small. Prototype gen ai features end-to-end—retrieval, function calling, human-in-the-loop QA—and connect them to your CRM integration or support stack. Use A/B testing with a clear minimum detectable effect (MDE) to validate impact. Leverage CustomGPT workflows for internal enablement and in-app guides or product tours to onboard users safely.

    Here’s a pragmatic 90-day plan. Week 0–2: audit your top 10 engineering tasks by time spent; identify 3 that are ripe for AI augmentation. Week 3–6: pilot inside CI/CD with explicit guardrails; track DORA metrics and developer sentiment. Week 7–10: productionize the wins; document runbooks; add incident management paths. Week 11–12: share learnings with product trios, refine your value proposition, and set next-quarter OKRs.

    AI won’t replace software engineers; engineers who master AI will outpace those who don’t. If we embrace the shift—toward systems thinking, responsible governance, and customer outcomes—we’ll build better products faster and open new, rewarding career paths. The opportunity is here and compounding.


    Inspired by this post on Product School.


    Book a consult png image
  • A Quality System for Trustworthy AI-Assisted UX Research

    A Quality System for Trustworthy AI-Assisted UX Research

    Your AI-generated synthesis can be polished, plausible, and wrong. The dangerous failures are rarely obvious fabrications. They are quieter: a biased sample becomes a universal claim, a participant’s opinion becomes a product need, or a tidy theme loses the contradiction that should have changed the roadmap.

    If you are deciding whether to trust AI-assisted UX research, do not judge the fluency of the summary. Judge the evidence chain behind it. You need to see how a product decision connects to the participants recruited, the questions asked, the underlying observations, the analytical interpretation, and the behavioral data used to check it.

    Key takeaways

    • Research quality is mostly determined before an AI tool sees a transcript. Start with the decision, learning question, and hypothesis.
    • Use AI to accelerate transcription, extraction, tagging, clustering, and contradiction searches. Keep interpretation, confidence, and product judgment under human control.
    • Require every theme to retain its participant coverage, supporting evidence, counterexamples, and unresolved uncertainty.
    • Pair qualitative findings with funnels, cohorts, session evidence, and CRM data when those signals are relevant. Neither qualitative nor quantitative evidence should carry the decision alone.
    • Finish with an atomic insight and a recorded choice. A summary that does not change a decision, test, or learning priority is not finished research.

    Define quality at the decision boundary

    Many teams begin AI-assisted research by asking which model should summarize their transcripts. That is too late in the process. The first quality control is the decision the research must inform.

    Strong discovery begins with a decision statement, an explicit learning goal, and a hypothesis the team is willing to falsify. Without those constraints, an AI system can generate an impressive taxonomy of themes while leaving the actual product question untouched.

    Before recruiting participants or writing prompts, create a short research contract:

    • Decision: Name the choice that is genuinely open. Examples include whether to pursue an opportunity, which problem to solve first, or whether a proposed workflow deserves further testing.
    • Decision condition: State what you would need to learn to proceed, pause, narrow the audience, or reject the current direction.
    • Learning question: Ask about the behavior, context, constraint, or unmet need that makes the decision uncertain.
    • Hypothesis: Write the current belief in a form that evidence could disprove. If every possible interview result would support it, it is not a useful hypothesis.
    • Relevant population: Specify whose behavior matters to this decision and which segments could experience the problem differently.
    • Evidence plan: Identify what interviews can reveal and which behavioral or operational signals could challenge the interpretation.
    • Data boundary: Decide what the AI tool is allowed to receive, what must be removed, and who may review the resulting artifacts.

    This contract changes how you evaluate the output. You are no longer asking whether the summary sounds reasonable. You are asking whether the evidence changes a named choice under stated conditions.

    My standard is simple: a decision-grade insight must survive a skeptical review without relying on the model’s authority. A reviewer should be able to inspect the underlying evidence, see which participants and segments it covers, understand the interpretation applied to it, and identify what remains unknown.

    Keep one distinction visible throughout the work:

    • Observation: What the participant did, described, showed, or failed to complete.
    • Interpretation: What that behavior may mean about a goal, anxiety, constraint, or job.
    • Implication: What the product team may choose to change, test, or leave alone.

    AI can help produce all three, but it should never blur them into a single sentence. Once an inference is written as if it were an observed fact, the rest of the synthesis becomes difficult to audit.

    Protect the signal before AI touches it

    An LLM cannot repair a convenient sample or a leading interview guide. It can only reorganize the resulting bias, often in language that makes the bias look more certain.

    Recruit for the decision, not for convenience

    If you interview only power users, you risk treating advanced workflows as mainstream needs. If you interview only vocal detractors, the roadmap can become a queue of complaints. A more useful recruiting frame includes new users, churned users, people who evaluated but did not convert, and adjacent personas where the decision calls for them.

    Build a participant matrix before outreach. Use rows for the segments that could materially change the decision and columns for relevant states, such as adoption stage, conversion outcome, or workflow maturity. The matrix is not a quota formula. It is a visibility tool. It should make overrepresented groups and missing perspectives obvious.

    Carry that segment metadata into synthesis. A theme that appears among established customers should not silently become a claim about evaluators. When a segment is absent, write that limitation into the insight rather than hiding it in an appendix.

    Ask for behavior before interpretation

    Questions about whether someone likes an idea invite speculation, politeness, and solution theater. Ask about the last relevant event instead. Have the participant reconstruct what triggered it, what they tried, where they hesitated, who else became involved, what workaround they used, and what happened next.

    Neutral, behavior-first questions become stronger when participants can support the account with artifacts such as screenshots or workflow examples. The artifact does not automatically prove the interpretation, but it helps distinguish remembered behavior from a general opinion.

    Pilot the guide with the product trio. Remove product terminology that telegraphs the preferred answer. Check whether each question could produce evidence against the working hypothesis. If the guide repeatedly asks participants to react to your solution, it is a concept evaluation guide, not an open discovery guide. Label it accordingly.

    Set privacy boundaries before uploading transcripts

    Consent to an interview does not automatically settle how AI will be used in transcription, analysis, storage, or sharing. Tell participants how their material will be handled, follow your organization’s data governance requirements, and remove identifiers that are not needed for the decision.

    Do not place sensitive participant data into an unapproved prompt workflow. If the tool’s handling, retention, or access controls have not been approved, keep raw transcripts out of it and work with appropriately de-identified material in an authorized environment. The downside is not merely a poor synthesis; it is unnecessary exposure of participant and customer information.

    De-identification should not erase the context required for analysis. Preserve non-identifying segment labels, workflow stage, and participant codes when they are relevant. The goal is to minimize sensitive data while retaining enough context to audit coverage and interpretation.

    Make AI produce an auditable synthesis

    The most reliable workflow separates extraction from clustering and clustering from judgment. Asking for findings, recommendations, sentiment, and a roadmap in one prompt encourages the model to fill gaps and compress uncertainty.

    1. Prepare the evidence set. Preserve the original transcript or recording, assign a participant code, attach relevant segment metadata, and remove unnecessary identifiers. Do not let an AI-generated summary replace the underlying material.
    2. Extract participant-level observations. Ask the model to work through each participant separately. Capture the behavior or event, its context, the supporting excerpt or evidence location, and any missing information. Do not ask for themes yet.
    3. Review the extraction. Check whether the observation is grounded in the transcript and whether the model has converted an opinion into behavior or inferred a motive the participant did not provide.
    4. Cluster reviewed observations. Group similar evidence only after the participant-level pass. Require each cluster to retain the contributing participant codes, segment coverage, supporting evidence, and meaningful variations.
    5. Search for contradictions. Ask which observations do not fit the cluster, which participants experienced the situation differently, and which alternative explanations remain plausible. Do not treat dissent as noise merely because it makes the summary less tidy.
    6. Draft atomic insights. Turn a defensible pattern into a small evidence packet containing the finding, evidence, coverage, contradictions, confidence rationale, product implication, and unresolved question.
    7. Triangulate relevant claims. Compare the qualitative interpretation with funnels, cohorts, session evidence, in-product paths, or CRM data when those systems contain a useful signal.
    8. Conduct the decision review. A person accountable for the product choice inspects the evidence chain, challenges the interpretation, and records what the team will do or learn next.

    You can make the separation explicit with narrowly scoped prompts.

    Extraction prompt: Use only the supplied transcript. For each relevant event, return the participant code, observed or reported behavior, context, supporting excerpt, evidence location, and uncertainty. Do not merge participants, infer motives, or recommend a solution. Flag information that is missing.

    Clustering prompt: Use only the reviewed observations. Group evidence by shared behavior and context. For every cluster, retain participant codes, represented segments, supporting observations, material variations, counterexamples, and plausible alternative explanations. Do not use repetition in the transcript as a substitute for participant coverage.

    Challenge prompt: Review the proposed themes as a skeptical researcher. Identify unsupported generalizations, segment differences that were flattened, interpretations written as observations, contradictory evidence, and claims that cannot be traced to the supplied material. Do not invent missing evidence.

    Prompt design helps, but it does not replace review. Keep the prompt, relevant tool or model information, input scope, and human corrections with the research artifact. If the synthesis later changes, you should be able to determine whether the cause was new evidence, a different analytical instruction, or a human judgment.

    AI is well suited to accelerating transcription, tagging, theme clustering, Jobs to Be Done extraction, and searches for hesitation or sentiment. Treat the latter outputs as interpretations to validate, not measurements generated by an objective instrument. A sentiment label is useful only when a reviewer can return to the behavior and language that produced it.

    Validate the insight, then record the decision

    A good synthesis review is not a copy-edit. It is an attempt to break the claim before the claim influences a roadmap.

    Run a quality review against the evidence chain

    • Traceability: Can a reviewer move from the insight to the contributing participants and the exact supporting material?
    • Coverage: Does the claim name the segments represented, and does it disclose relevant segments that are missing?
    • Construct validity: Is the finding about the behavior the study intended to understand, or has a nearby opinion been used as a proxy?
    • Separation: Are observation, interpretation, and product implication visibly distinct?
    • Contradiction: Does the artifact preserve disconfirming cases and material variations instead of forcing consensus?
    • Triangulation: Where behavioral data is relevant, does it support, narrow, or challenge the qualitative account?
    • Decision relevance: Does the finding change a live choice, a test, or the next learning priority?

    Do not outsource confidence to the model. A confident tone is a language property, not an evidence assessment. Record confidence as a human rationale based on the clarity of the underlying behavior, the relevance and coverage of participants, consistency and counterexamples, and any corroborating behavioral evidence.

    Quantitative and qualitative signals answer different parts of the question. Funnels, cohorts, and retention analysis can show where behavior changes or where people leave. Interviews and artifacts can expose the goals, anxieties, organizational constraints, and workarounds behind that behavior. Pairing those signals is how a team moves from observing what happened to developing a testable account of why.

    When the signals disagree, do not average them into a vague conclusion. Check whether the interview sample represents the population in the analytics, whether the event instrumentation reflects the behavior being discussed, whether segments have been combined, and whether the evidence refers to the same stage of the journey. A contradiction is often the next research question.

    Use an atomic insight format

    A reusable insight should be small enough to inspect and complete enough to guide a choice. Use this structure:

    • Decision: The product choice this evidence informs.
    • Finding: The observed behavioral pattern and the context in which it occurs.
    • Evidence: Participant codes, excerpts or artifact locations, and any relevant behavioral signal.
    • Coverage: The represented segments and known gaps.
    • Interpretation: The best current explanation, clearly labeled as an inference.
    • Contradictions: Cases or data that weaken, narrow, or complicate the interpretation.
    • Confidence: A short rationale grounded in evidence quality, coverage, consistency, and triangulation.
    • Product implication: The opportunity, risk, constraint, or tradeoff the team should consider.
    • Disposition: Act, test further, monitor, or take no action.
    • Next unknown: The uncertainty most likely to change the decision.

    Useful insight records also prevent familiar synthesis mistakes. Replace a broad label such as onboarding friction with the specific behavior, actor, context, and consequence. Do not let a memorable quotation stand in for a pattern. Do not describe a participant’s requested feature as the underlying need. Do not convert an AI-generated cluster into a roadmap item until the evidence packet survives review.

    Bring the atomic insights to a decision review with the product trio. Record the choice, its rationale, what the team is deliberately not doing, and the evidence that could reopen the decision. Connect the chosen action to an outcome or learning objective rather than treating delivery of a feature as proof that the research was correct.

    For your next study, start with one live decision and run the evidence through this chain. If a theme cannot be traced, mark it as a hypothesis. If participant coverage is lopsided, narrow the claim. If qualitative and behavioral evidence conflict, investigate the conflict before committing the roadmap. That is how AI becomes a fast, inspectable research assistant instead of an unaccountable author of customer truth.

    References

  • How to Evaluate AI Voice Support in Real-World Conditions

    How to Evaluate AI Voice Support in Real-World Conditions

    You have a shortlist of AI voice support products, a polished recording, and a decision that could affect thousands of customer conversations. The hard question is not whether an agent can sound convincing during one ideal call. It is whether the system stays useful when a caller interrupts, corrects themselves, asks an ambiguous question, waits on a backend system, or needs a human.

    You can answer that question before a broad rollout. The method is to test complete support outcomes, introduce controlled complications, score failures separately from conversational polish, and use the result to define a limited production pilot.

    Evaluate the support outcome, not the performance

    A natural voice can create an impression of competence before the agent has done anything useful. Pleasant pacing, expressive speech, and a quick opening matter, but they cannot compensate for retrieving the wrong account, misunderstanding the request, or claiming that an action succeeded when it did not.

    Treat the unit of evaluation as a completed support job. Depending on the intent, that job may require the agent to identify the caller, understand the request, retrieve the right information, explain the answer, perform an authorized action, confirm the resulting state, and send a follow-up or transfer the conversation. If you score only the spoken answer, you leave most of the product untested.

    One live Fin Voice call illustrated this end-to-end standard in about 90 seconds: the agent verified identity, retrieved account information, managed an interruption, presented options, completed a workflow, and sent a follow-up email. That sequence is a useful model for constructing a test. It is not, by itself, proof of reliability across other calls.

    Before anyone places a test call, write an outcome contract for each scenario:

    • Caller goal: What is the person trying to accomplish?
    • Starting state: What customer, account, order, subscription, or case data exists before the call?
    • Available evidence: Which knowledge, policies, and records may the agent use?
    • Permitted actions: What may the agent change, create, send, cancel, or escalate?
    • Required clarification: Which missing or conflicting facts must be resolved before an answer or action?
    • Completion evidence: What observable state proves that the request was resolved?
    • Unacceptable outcome: What error would make the call a failure even if the conversation sounded good?

    This contract prevents a common scoring mistake: confusing non-transfer with resolution. A call can remain inside the AI channel and still leave the customer with a wrong answer, an incomplete action, or no idea what happens next. Conversely, an intentional transfer can be the correct resolution when the agent reaches a policy, permission, or confidence boundary.

    Build scenarios around the ways real calls become difficult

    Start with support intents your operation actually receives. Prioritize intents that are frequent, expensive to handle, important to customer trust, or dependent on multiple systems. Do not begin with trivia questions that merely demonstrate broad language-model knowledge. You are evaluating support execution.

    For every core intent, create a straightforward case and several controlled variants. Keep the customer objective constant while changing one condition at a time. That makes a failure diagnosable instead of merely disappointing.

    A practical scenario matrix

    • Clean path: The caller gives the relevant facts in a clear order. This establishes whether the basic workflow works at all.
    • Missing information: Omit a detail the agent needs. Check whether it asks a focused question instead of guessing or restarting the intake.
    • Ambiguous intent: Use wording that could map to two support issues. The agent should disambiguate before retrieving data or taking action.
    • Mid-call correction: Let the caller change an account detail, date, product, or preferred option. Check whether the corrected fact replaces the old one throughout the workflow.
    • Interruption: Speak while the agent is answering. Observe whether it stops cleanly, understands the new input, and continues from the right point.
    • Backend delay: Introduce a slow retrieval or action. Evaluate how the agent manages the wait and whether it distinguishes a pending operation from a completed one.
    • Backend failure: Make a required system unavailable or return an error. The agent should not fabricate a result or promise completion it cannot verify.
    • Policy boundary: Ask for something the agent is not allowed to do. Test the explanation, alternatives, and escalation path.
    • Human request: Ask directly for a person. Verify that the agent follows the configured policy without turning the handoff into an argument.
    • Listening conditions: If your deployment must support different languages, accents, devices, or noisy environments, test each condition explicitly rather than treating one clear studio call as representative.

    Give testers the goal, account state, and one complication. Do not script every sentence. A fully written dialogue tests whether the agent can follow the dialogue you anticipated; a goal-based scenario tests whether it can manage the conversation the caller actually creates.

    Keep a few variants undisclosed until the live session. This is not a trick. It prevents the evaluation from becoming a memorized path while still keeping every test fair and reproducible. Record the exact variant afterward so another evaluator can run it again.

    Run the call through the systems you expect to deploy

    An unedited live call is more informative than a produced recording, but live alone is not enough. A live test can still use ideal data, a simplified integration, a practiced caller, and a workflow that avoids the hard parts of your environment.

    Ask to run the scenario through a path that resembles the intended deployment:

    1. Place a normal phone call through the proposed telephony route. If production will use call forwarding, test the forwarding path rather than a direct internal endpoint.
    2. Use a safe test account containing representative records, permissions, and history.
    3. Require the agent to retrieve data from the backend system that will be authoritative in production.
    4. Introduce the chosen interruption, correction, ambiguity, delay, or error during the live conversation.
    5. Require a real test action where it is safe to do so, not a verbal description of what the agent would have done.
    6. Inspect the backend state after the call. Confirm that the correct record changed once, with the expected values.
    7. Verify every promised follow-up, case creation, notification, or handoff outside the voice channel.
    8. Retain the recording, transcript, timestamps, tool activity, and final system state for scoring.

    This is especially important when an agent can take consequential actions. A fluent confirmation is not evidence that the action happened. The system of record is the evidence.

    Repeat important scenarios with different wording and a different caller. One successful run demonstrates that the capability can work. Repeated variants reveal whether the capability depends on a narrow phrase, a rehearsed cadence, or an unusually forgiving path.

    Key takeaways

    • Score complete resolution, including backend state and follow-up, rather than voice quality alone.
    • Change one condition at a time so you can identify why a call failed.
    • Test interruptions, corrections, ambiguity, system delays, system errors, and escalation.
    • Measure different kinds of waiting separately; a lookup pause and a turn-detection problem are not the same defect.
    • Treat a successful demo as evidence for a pilot, not permission for an unrestricted rollout.

    Score conversation, reasoning, and operational closure separately

    A single overall rating hides the information you need to make a product decision. The call may sound awkward but reach the correct outcome, or sound excellent while making a dangerous mistake. Separate the evaluation into three layers.

    LayerWhat to inspectEvidence of a passTypical failure
    Conversation mechanicsTurn detection, interruption handling, pacing, response length, and intelligibilityThe caller can speak naturally, correct the agent, and follow the response without fighting for the floorThe agent talks over the caller, leaves confusing silence, or delivers answers too long to retain by ear
    Decision qualityIntent recognition, clarification, use of account context, policy application, and answer accuracyThe agent asks only for missing information, uses the correct evidence, and avoids unsupported conclusionsThe agent guesses, asks redundant questions, ignores a correction, or applies the wrong policy
    Operational closureIdentity checks, tool calls, state changes, confirmation, follow-up, and escalationThe verified backend state matches the caller’s request and the agent’s final explanationThe agent claims success without a completed action, changes the wrong record, duplicates work, or drops context during handoff

    Use a simple 0-2 score for each criterion: 0 for failed or unsupported, 1 for completed with material caller effort or recovery, and 2 for correct and usable. The scale is deliberately small. Evaluators can usually distinguish failure, friction, and success more consistently than they can defend the difference between seven and eight on a ten-point scale.

    Do not average away critical errors. A wrong account action, failed identity control, fabricated completion, or forbidden disclosure should remain visible as a release blocker even if many low-risk calls receive high scores. Record both the criterion scores and the count of critical failures.

    Break latency into moments the caller can feel

    Latency is not one number. Capture at least three moments: the time the agent takes to recognize that the caller has finished, the time it spends reasoning or waiting for a system, and the time needed to begin and complete the spoken response.

    • End-of-turn delay: A long delay after every caller turn makes the exchange feel unresponsive and can encourage both sides to start speaking at once.
    • Reasoning or retrieval delay: A pause can be appropriate when the agent is checking account data or invoking a backend workflow. Brief pauses were audible during live subscription and backend checks, which is more informative than editing those waits out.
    • Response delivery: A fast start does not help if the answer becomes a long monologue. Voice responses need structure and pacing that work for listening, not merely text that sounds acceptable when read.

    Ask what is happening during a pause. If the system is doing useful work, the next statement should reflect that work and the action log should verify it. If the pause is long enough to make a caller wonder whether the call has dropped, the experience needs an appropriate progress cue. If the agent answers instantly but guesses, speed is concealing a quality problem.

    Review individual timings as well as an average. A generally responsive agent with occasional severe stalls creates a different operational problem from one that is consistently a little slow. Your test recordings and timestamps should make both patterns visible without inventing a universal pass threshold that ignores the complexity of the workflow.

    Make recovery and escalation part of the product test

    The strongest voice experiences are not the ones that never encounter confusion. They are the ones that recover without making the caller restart. Recovery is therefore a capability to test, not an embarrassing exception to hide.

    Interrupt the agent in the middle of an answer. Correct a fact it has already used. Add a second request after the first appears resolved. Say that an explanation was unclear. Ask for a human. These moves reveal whether the agent maintains conversational state or merely produces plausible turns one at a time.

    During recovery, look for specific behavior:

    • It stops speaking promptly when the caller takes the turn.
    • It identifies what changed instead of repeating the whole interaction.
    • It replaces corrected information rather than carrying both versions forward.
    • It asks a narrow clarification when the next action is uncertain.
    • It does not claim to understand when the transcript or subsequent action shows otherwise.
    • It preserves verified context and the reason for contact when a human takes over.
    • It tells the caller what will happen next instead of ending on an internal routing label.

    Tone belongs in this test, but not as a beauty contest between synthetic voices. Evaluate whether pacing, brevity, acknowledgement, and word choice suit the moment. A caller correcting a billing detail needs a clear acknowledgement and an accurate update, not theatrical empathy. A caller who sounds uncertain may need a shorter explanation and a confirming question. Tone is the behavior of the conversation, not just the timbre selected in a settings menu.

    Escalation should also count as a valid outcome when it is timely and informed. Define which conditions require a handoff, which allow one, and what context must travel with it. Then test the handoff from the caller’s side. If the customer reaches a person but has to repeat identity, intent, and every attempted step, the routing technically worked while the support experience failed.

    Turn the evaluation into a controlled pilot decision

    A strong live evaluation earns the right to run a pilot. It does not justify sending every eligible call to the agent. Production introduces variation in callers, data quality, traffic, integrations, and issue combinations that a demonstration cannot reproduce fully.

    I would require five gates before approving even a limited external pilot:

    1. Capability gate: Every must-have intent has completed its end-to-end workflow, including at least one controlled complication.
    2. Critical-risk gate: No unresolved failure can expose the wrong account, bypass a required check, perform an unauthorized action, or report a false completion.
    3. Conversation gate: The agent can handle interruptions, corrections, clarification, and explicit human requests without trapping the caller in a loop.
    4. Operations gate: Your team can configure terminology, guidance, escalation behavior, greetings, voice, and deployment controls for the intended support environment.
    5. Learning gate: Owners can inspect recordings, transcripts, tool activity, outcomes, and failures, then change the knowledge, workflow, policy, or conversation design responsible.

    Start the pilot with a reversible slice of traffic and a clear human fallback. Select intents whose correct outcome can be verified in your systems. Define who reviews failed and escalated calls, who can pause the rollout, and who owns each class of fix. An answer-quality issue, a telephony issue, and a backend integration issue require different owners even when the caller experiences all three as one bad call.

    Expand only when observed calls meet the outcome contracts you wrote before the demo. If the definition of success keeps changing after failures appear, the evaluation is no longer protecting the decision.

    For your next vendor session, replace “show me your best call” with a scenario pack, a test account, and a request to inspect the final system state. You will learn more from one imperfect call that recovers correctly than from a flawless recording that never had to recover at all.

    References

  • Global Invoicing Nightmares: Hard-Won Product Lessons on EU Tax, Compliance, and Customer Value

    Global Invoicing Nightmares: Hard-Won Product Lessons on EU Tax, Compliance, and Customer Value

    I hit play on Global Invoicing – All Things Product Podcast with Teresa Torres & Petra Wille and felt an immediate jolt of recognition. We’ve all launched a feature that looked solid—until a small, overlooked detail broke everything. Their stories about global invoicing and taxes echoed challenges I’ve faced leading product for international customers: if you don’t design for the last mile of compliance, you can accidentally block the very "moment of value creation" your product promises.

    Listen to this episode on: Spotify | Apple Podcasts

    The conversation starts as a candid rant about EU tax compliance and quickly becomes a precise product management lesson: when we fail to map the entire path to customer value—down to the tiniest regulatory requirement—we can ship something “done” that still doesn’t work in the real world. That gap between intention and outcome is where good product teams live or die.

    In my experience, the nightmare of global invoicing for small online businesses is very real. Even big platforms (like Squarespace and Teachable) miss the mark on EU tax compliance, and when they do, customers feel it immediately. It’s the kind of edge case that doesn’t show up in a demo but absolutely shows up in revenue. Or as Teresa put it, “It’s not a little detail when your client won’t pay the invoice.” — Teresa Torres

    I appreciated how the episode digs into the difference between passing a regulatory checklist and actually meeting customer needs. Put plainly: the product isn’t “done” when the ticket moves to Done; it’s done when the customer completes the job—receives an acceptable invoice, pays successfully, and can reconcile it without friction. That’s why I lean hard on story mapping for regulatory work; it exposes the invisible steps where value creation can silently fail.

    Here’s how the episode resonates with my own playbook: the nightmare of global invoicing for small online businesses is a systems problem; why even big platforms (like Squarespace and Teachable) miss the mark on EU tax compliance is a prioritization and discovery problem; how Petra and Teresa navigated invoicing across borders with Ableify and LearnWorlds highlights pragmatic tool choices and trade-offs; the key difference between meeting regulations and meeting customer needs is an outcomes-over-output mindset; what product teams can learn from regulatory edge cases is how to find the seams where markets, laws, and workflows collide; how missing a single detail can block the "moment of value creation" is a reminder that value is defined by customers; and why story mapping is critical for finding gaps between "we shipped it" and "customers got value" is the method that connects all of the above.

    Practically, that means I treat regulatory features like any other high-stakes product surface: do real product discovery with affected users; co-design the happy path and the ugly edge cases; write acceptance criteria that include jurisdictional and document-level specifics (e.g., VAT numbers, invoice formats, timing rules); align with finance and legal early; and instrument the journey from invoice issued to invoice paid so we can see where real customers get stuck. This is outcomes vs output OKRs in action, and it’s one of the fastest ways to earn trust with stakeholders.

    Key takeaways worth bookmarking: Customers define value, not your compliance checklist. Regulatory work still requires discovery—you can’t skip understanding user needs. The path to value doesn’t end when your feature works; it ends when your customer succeeds. “Sweating the details” isn’t micromanagement—it’s good product management.

    Memorable quotes to bring back to your team: “If you don’t sweat the details, people choose other platforms.” — Petra Wille. “It’s not a little detail when your client won’t pay the invoice.” — Teresa Torres.

    Follow Teresa Torres: https://ProductTalk.org | Follow Petra Wille: https://Petra-Wille.com

    Mentioned in the episode: Squarespace | Stripe | Product at Heart | Teachable | LearnWorlds | Ablefy | Become a Better Product Leader: A 52-Week Transformation Journey | Product Talk Academy

    Have thoughts on this episode? Leave a comment below.

    Full transcripts are only available for paid subscribers.


    Inspired by this post on Product Talk.


    Book a consult png image
  • From Sketch to Clickable Demo: My AI Prototyping Playbook to Build Apps in Hours

    From Sketch to Clickable Demo: My AI Prototyping Playbook to Build Apps in Hours

    I’ve spent much of my career compressing the distance between a napkin sketch and something real customers can touch. At HighLevel, my product teams use generative AI to validate ideas faster, reduce risk earlier, and win stakeholder trust with evidence instead of slides. The goal isn’t to be flashy—it’s to be precise, testable, and repeatable.

    Today, you can build it before you pitch it. AI prototyping can turn ideas into clickable demos in hours. Here are some tools to try and steps to follow.

    I start every AI prototyping sprint by sharpening the problem statement and the outcome we care about. That means being explicit about the target user, jobs-to-be-done, and the riskiest assumptions. I define a minimum detectable effect (MDE) and tie it to outcomes vs output OKRs so everyone aligns on what “good” looks like before we touch a tool.

    From there, I move from sketch to interface. I capture a rough flow (whiteboard, tablet, or even paper) and generate UI variations with my AI product toolbox—tools that translate structure into components and screens. I’ll iterate on information hierarchy and copy until the narrative supports the core job, borrowing techniques from UX writing. For product managers leaning into LLMs for product managers, this phase is about speed to feedback, not perfection.

    Next, I wire data and logic. I connect a lightweight backend or spreadsheet, stitch in a CRM integration if needed, and add LLM calls through a ChatGPT connector or Claude Code. If the concept benefits from multi-step autonomy, I introduce agentic AI to orchestrate tasks across APIs. CustomGPT workflows help me encapsulate business rules so the demo behaves consistently in user paths we care about.

    Governance is not optional at this stage. I apply privacy-by-design defaults, document data governance decisions, and run a quick AI risk management pass: input validation, prompt safety, rate limits, and fallback responses. This keeps the prototype credible and prevents false positives from polluting stakeholder perception.

    With a click-through in hand, I instrument the experience so learning compounds. I drop in Amplitude analytics to track activation, task completion, and drop-off, and set up simple A/B testing when there’s a meaningful design or copy choice. This makes the prototype a learning vehicle, not just a demo.

    Then I get it in front of users—fast. Five targeted conversations will beat fifty internal opinions. I run structured product discovery interviews, observe time-to-value, and capture objections. This is where empowered product teams shine: we make changes in real time, re-run the flow, and document what moves the needle for product-led growth.

    When speed matters, I use a four-hour cadence: Hour 1 for problem framing and MDE; Hour 2 for sketch-to-UI generation; Hour 3 for data wiring and AI logic; Hour 4 for instrumentation and user walkthroughs. By the end, we have a clickable demo, preliminary analytics, and a clear decision on whether to advance, pivot, or park.

    Finally, I translate insights into a concise artifact: the hypothesis we tested, the signal we observed, the trade-offs we made, and the next sprint plan for product roadmapping and sprint planning. The point is not to be right on the first try; it’s to learn precisely, cheaply, and quickly enough to invest with conviction.

    If you adopt this approach, you’ll find that stakeholder management becomes easier, team energy rises, and your roadmap earns credibility. Build it before you pitch it, and let real interactions—not wishful thinking—do the heavy lifting.


    Inspired by this post on Product School.


    Book a consult png image
  • Cut Time to Value, Boost Retention: My Proven Playbook for Activation, Growth, and Loyalty

    Cut Time to Value, Boost Retention: My Proven Playbook for Activation, Growth, and Loyalty

    Time to value is the most reliable early indicator of long-term user retention I know. When customers experience meaningful product impact fast, they stick around, expand, advocate, and cost less to support. Over the years leading product teams, I’ve learned that speed-to-impact isn’t a nice-to-have—it’s the engine behind sustainable product-led growth and efficient go-to-market.

    Accelerate retention by reducing time to value. Learn how faster product impact drives growth, reduces costs, and keeps users engaged in the long term.

    Practically, I define time to value as the duration from first touch (or first login) to the moment a user achieves their “aha” outcome—something tangibly useful aligned to their job-to-be-done. The shorter that journey, the higher the likelihood of user activation, trial conversion, and durable engagement. This is why I obsess over onboarding, in-app guides, product tours, and the clarity of our value proposition.

    My first move is to map the Minimum Path to Value (MPV): the smallest set of actions needed to deliver a real result for a new user. I strip away everything non-essential in that path—fields, clicks, choices, and jargon. Opinionated defaults, smart templates, sample data, and single-player workflows let customers succeed in minutes, not days. The goal is to reduce cognitive load while making the next best action unmistakably clear.

    Instrumentation turns TTV from a hunch into a system. I track activation events, cohort retention, and conversion using platforms like Amplitude analytics and Pendo, with timely nudges through Intercom when users stall. I look at the distribution of TTV (not just the average), correlate it with retention analysis, and set explicit targets such as “new users reach first value within 10 minutes.” Those targets become team-level outcomes—not outputs—and we review them weekly.

    Experimentation is how we iterate toward the fastest path to value. I rely on A/B testing to compare onboarding flows, progressive profiling to delay non-critical inputs, and opinionated setup wizards to remove guesswork. Auto-generated example projects, pre-configured integrations, and guided checklists accelerate user activation without sacrificing flexibility for advanced users.

    Content and guidance matter as much as UX. Tooltips, contextual in-app guides, and short product tours should be timely, skippable, and laser-focused on the outcome, not the feature. I pair these with a concise knowledge base and short explainer videos that reinforce the same value narrative a user sees inside the product.

    Cross-functional alignment is essential. Product, marketing, sales, and customer success must rally around the same activation metric and TTV target. That alignment ensures our trial messaging, onboarding emails, and CS playbooks don’t compete—they compound. When everyone points to the same first-value moment, friction drops and adoption rises.

    Pricing and packaging can also accelerate time to value. Free trials should be long enough for users to credibly reach first value; usage-based gates should never block the MPV. I prefer to unlock everything needed to hit the “aha” moment, then meter after the value is viscerally felt—this respects the user’s time and reinforces trust.

    There’s a cost story, too. Faster time to value reduces tickets, shortens onboarding cycles, and lowers cost-to-serve. It also clarifies product discovery: when we see where users stall, we don’t guess at roadmap priorities—we let the data guide our next bet.

    In my experience at HighLevel, I’ve repeatedly seen activation rates jump when we cut time to value from days to minutes. The specific tactics vary by product, but the pattern holds: when the first outcome is undeniable and fast, retention follows—and so does efficient growth.

    If you’re looking for a starting point, try this: define one activation event that clearly signals value, instrument it end-to-end, design a Minimum Path to Value that gets new users there in under 10 minutes, and run weekly experiments until you consistently hit the target. Do that, and you won’t just improve onboarding—you’ll build a product that earns loyalty from the very first session.


    Inspired by this post on Amplitude – Best Practices.


    Book a consult png image
  • Win AI Search: Proven Playbook to Get Your Startup Recommended by ChatGPT & Perplexity

    Win AI Search: Proven Playbook to Get Your Startup Recommended by ChatGPT & Perplexity

    AI search is quickly becoming the new homepage for startups. When a buyer asks a model for the best tools, they often take the short list at face value. I treat this moment as a product surface I can influence with strategy, content, structure, and distribution—much like any other go-to-market channel.

    Early on, I set a simple objective for my team and me: "Learn how LLMs like ChatGPT and Perplexity decide which startups to recommend and what signals help a brand get discovered in AI search." That sentence became our north star for experiments, instrumentation, and content architecture.

    Here is the mental model that consistently holds up in practice. Large language models synthesize answers from a knowledge graph built from crawled content, citations, and high-signal sources. They weight consensus, clarity, recency, authority, and machine-readability. I don’t pretend to know the internals, but across hundreds of tests, the same patterns correlate with being surfaced and cited.

    First, I make our entity unambiguous. I standardize the company name, product names, and leadership bios across the site and external profiles. I implement Organization and Product markup with schema.org and link out with sameAs to authoritative profiles like LinkedIn, Crunchbase, GitHub, and key directory listings. The goal is to collapse ambiguity so AI search knows exactly who we are and which claims are attributable to us.

    Next, I publish definitive, answer-first pages. For every core query—what we do, who it’s for, outcomes, differentiators, pricing, comparisons, and integrations—I ship a page that leads with a crisp summary, then supports it with evidence, examples, and plain language. I include Q&A sections, realistic use cases, and named case studies so models can quote and ground responses in verifiable facts.

    I then make the site maximally machine-readable. I add schema.org for SoftwareApplication, Product, FAQPage, and HowTo where relevant. I keep titles, H1/H2 structure, internal links, and metadata descriptive and consistent. I expose last-modified dates, maintain an XML sitemap, and keep a visible changelog and release notes. Freshness matters—Perplexity, in particular, tends to privilege recent, well-cited material when answering time-sensitive questions.

    Citations are non-negotiable. I earn credible mentions on third-party properties, analyst lists, comparison pages, and customer reviews. I prioritize authoritative placements over volume, then make sure our site references those sources to reinforce the signal. When Perplexity cites our page alongside a respected third-party review, our inclusion rate in answers rises noticeably.

    I also design for developers, buyers, and machines at once. That means clean docs, integration pages, and transparent security and trust content. Clear API references, integration guides, and reliability notes give models concrete artifacts to summarize. Pricing, privacy, and support policies reduce uncertainty and increase the likelihood that an answer will include us.

    Measurement turns this from a hunch into a system. I run controlled content experiments, track minimum detectable effect on discovery and mentions, and instrument referral patterns from AI assistants when citations appear. I monitor which prompts surface our brand, which sources are cited, and which pages are repeatedly used as references. When we move a KPI, we codify the pattern into our playbook and scale it.

    Trust is the compounding advantage. I maintain a transparent trust center, privacy-by-design posture, and clear data governance practices. I remove vague claims, back up benefits with evidence, and keep all performance or security statements auditable. Models tend to lift brands that feel low-risk, well-documented, and widely corroborated.

    If you want a fast start, here’s the checklist I rely on. Standardize your entity and ship schema.org. Publish answer-first pages for core jobs-to-be-done, comparisons, and integrations. Earn authoritative third-party citations and reference them. Keep release notes, changelogs, and dates current. Instrument AI discovery and iterate based on what gets cited. Do this consistently, and your startup earns a fair shot at being recommended when buyers ask AI for the best options.


    Inspired by this post on Amplitude – Best Practices.


    Book a consult png image
  • Prototypes vs Products: How I De-risk Ideas Fast and Ship Reliable Value at Scale

    Prototypes vs Products: How I De-risk Ideas Fast and Ship Reliable Value at Scale

    Note: This is part of the product creator series of articles, based on the overview article, The Era of the Product Creator. This series is for anyone who wants to create a successful product—whether or not you’ve had formal training or experience in product management, product design, or engineering. Over the years, I’ve watched smart teams stumble because they treated a prototype like a product. The distinction is simple but vital: prototypes exist to learn; products exist to earn trust by delivering value reliably at scale. When we blur that line, we ship avoidable risk to customers and slow ourselves down later with rework. When I build a prototype, I’m testing assumptions as quickly and cheaply as possible. It might be a clickable Figma mock, a Wizard‑of‑Oz demo, or a quick script stitching together a ChatGPT connector with a CustomGPT workflow. It’s intentionally disposable. I expect missing edge cases, fake data, hand‑waving on latency, and limited attention to security or privacy. The only goal is to answer the riskiest questions fast. A product is a promise. It’s hardened for reliability, performance, security, and privacy‑by‑design. It’s observable with real analytics, supports CI/CD and rollback, meets accessibility guidelines, and can be maintained by empowered product teams. It has clear SLAs, incident management runbooks, and instrumentation that lets me track outcomes vs output OKRs and DORA metrics. Keeping prototypes and products separate makes us faster and safer. Prototypes accelerate discovery; products operationalize value. If I catch myself “polishing” a prototype, I pause and either discard it or define the path to production with the right engineering rigor, data governance, and stakeholder management. Here’s how I decide. In prototype mode, I timebox learning to days, not weeks, and focus on a single risky assumption—value, usability, or feasibility. I validate through qualitative research and usability tests, not vanity metrics. To graduate to product work, I require a crisp problem statement, evidence of problem‑solution fit, a technical plan for scale and observability, a privacy and threat modeling review, and a measurement plan (including minimum detectable effect) for upcoming A/B testing. AI adds new wrinkles. For gen AI and agentic AI, I evaluate model behavior offline before exposing anything to customers. That includes prompt design, context window management, guardrails to minimize hallucinations, and clear fallback strategies. I define red‑team scenarios, logging for auditability, and policies for data retention and encryption as part of AI risk management. A recent example: we prototyped an agent workflow in a day that felt magical in demos. We resisted the urge to ship. Instead, we added authentication, rate limiting, PII redaction, human‑in‑the‑loop review, observability, and in‑app guides and product tours for onboarding. Only then did we move to a limited release with a well‑defined go‑to‑market strategy and support readiness. One more trap to avoid: calling a prototype an MVP. An MVP is still a product—minimal in scope but complete enough to deliver value, gather trustworthy data, and support customers. If you wouldn’t put your name on it or support it in production, it’s a prototype, not an MVP. If you’re a product creator, align your product trios around this discipline. Use prototypes to learn quickly in discovery, and use products to deliver outcomes in delivery. That mindset protects customer trust, speeds iteration, and moves you toward product‑market fit with far less waste.

    Inspired by this post on SVPG.


    Book a consult png image
  • AI Context Engineering: A System for Product Decisions

    AI Context Engineering: A System for Product Decisions

    You give an LLM your discovery notes, a dashboard export, and a roadmap question. It returns polished recommendations in seconds. The recommendations sound plausible, yet your product trio still cannot tell which option deserves a commitment.

    The missing ingredient is usually not a better prompt. It is a decision-ready context system: a controlled way to give AI the evidence, boundaries, and outcome definition required to reason about the same product decision your team is actually making. Done well, this gives you more than a convincing answer. It gives you a traceable choice, explicit uncertainty, and a validation plan.

    Define the decision before you collect the context

    For product work, context engineering is the deliberate design of everything an AI system can use at the moment it reasons: customer evidence, metrics, goals, constraints, definitions, instructions, and prior decisions. The useful unit is not a prompt or a document. It is the decision.

    This distinction matters because an LLM can answer an underspecified request without exposing that the request was underspecified. Ask it to improve onboarding, and it can produce a credible list of patterns. That output still does not tell you which user segment matters, what improvement means, which current friction is supported by evidence, or what downside the team must avoid.

    Before pulling any context, write a decision frame that answers these questions:

    • What decision must be made? Name the commitment, not the general topic. Choose whether to change a specific onboarding step is a decision; explore onboarding is not.
    • Who is the decision for? Identify the customer segment, use case, or part of the journey. Evidence from one segment should not silently become a claim about every user.
    • What outcome should change? State the behavior or business result you want, then identify the guardrail signals that should not deteriorate.
    • What can constrain the answer? Include privacy, risk, brand, commercial, technical, and operational boundaries before ideation begins.
    • What evidence could change the choice? If no possible evidence would change the decision, you are asking AI to justify a conclusion rather than help make one.
    • What must the output enable? Specify whether you need options, a recommendation, a decision memo, an experiment plan, or a list of unresolved questions.

    Anchor this frame in outcomes rather than deliverables. Improve activation for a defined segment while protecting support load establishes a decision boundary. Build a new onboarding checklist merely names output. The first lets AI compare interventions; the second encourages it to decorate a predetermined solution.

    A practical test is to remove the proposed feature from the frame. If the decision still makes sense, you have probably described an outcome. If the frame collapses, the team may already be committed to an output.

    Build a context packet that preserves evidence quality

    A context packet is the smallest governed collection of information that allows the model and the product team to reason about the decision. It can combine customer quotes, behavioral trends, funnel friction, support conversations, and commercial constraints. The important work is to assemble, structure, compress, and challenge that evidence before asking for recommendations.

    Do not treat every input as the same kind of truth. A customer quote gives you detail about an experience, not its prevalence. Usage analytics show behavior, not necessarily motivation. Support conversations overrepresent people who contacted support. CRM data can expose commercial constraints without proving that a feature creates customer value. Labeling these boundaries prevents the model from blending different signals into false certainty.

    Use this structure for the packet:

    • Decision header: the choice, decision owner, affected segment, and action that follows the decision.
    • Outcome frame: the desired outcome, current signal, primary measurement, guardrails, and any metric definitions needed to interpret the data correctly.
    • Evidence ledger: each relevant observation with its origin, segment, time period, and scope. Keep direct observations separate from interpretations.
    • Constraints: technical dependencies, commercial commitments, privacy rules, brand boundaries, operational capacity, and known risks.
    • Contradiction register: evidence that points in different directions, including differences between customer statements and observed behavior.
    • Unknowns: missing evidence, ambiguous definitions, unrepresented segments, and assumptions the team has not validated.
    • Output contract: the form of response you need, the criteria options must address, and the unsupported claims the model must label rather than fill in.

    Compression is where many context packets either become useful or become misleading. The goal is not merely to shorten the material. It is to increase the proportion of decision-relevant signal without erasing qualifications.

    1. Normalize repeated evidence. Deduplicate copied notes and repeated tickets so repetition in the packet does not impersonate independent confirmation. Preserve any real frequency data separately.
    2. Retain the qualifiers. Do not compress away the segment, time range, denominator, metric definition, or product state that determines what an observation means.
    3. Label epistemic status. Mark material as observation, interpretation, assumption, or generated hypothesis. A concise packet should make these distinctions clearer, not blur them.
    4. Keep contradictions visible. If interviews describe one problem while behavioral data points elsewhere, preserve both signals and ask what evidence would resolve the conflict.
    5. Remove inert context. My rule is simple: if an item cannot change an option, a risk assessment, or the validation plan, it does not belong in the active packet. Keep it available outside the model context if the team may need to inspect it later.

    Apply privacy-by-design while assembling the packet, not after the model has processed it. Customer transcripts, CRM records, and support conversations can contain personal or confidential data. Use approved systems, follow applicable access controls and data terms, redact identifiers, and aggregate where the decision does not require record-level detail. If you cannot establish that the data is permitted in the AI workflow, leave it out and provide a safe summary. The downside is not a weaker prompt; it is potential exposure of customer or company information.

    Separate synthesis, strategy, and skepticism

    Asking for a summary, a recommendation, and a critique in the same instruction makes it difficult to see where evidence ends and invention begins. A stronger agentic workflow separates those jobs into distinct passes: Summarizer, Strategist, and Skeptic.

    The Summarizer creates an evidence map

    The Summarizer should organize the packet without deciding what to build. Ask it to group evidence around the decision, preserve relevant qualifiers, expose conflicts, and identify missing information. Explicitly prohibit recommendations during this pass.

    A useful Summarizer output contains the supported observations, the segments represented, the outcome signals involved, the contradictions, and the unknowns. Review this output against the packet before continuing. If the model has turned an assumption into a fact, fix the evidence map rather than hoping a later pass corrects it.

    The Strategist develops decision options

    Give the Strategist the approved evidence map, the original decision frame, and the constraints. Ask for a small, meaningfully different set of options, including the option to leave the product unchanged when that is legitimate.

    Require the same fields for every option:

    • the customer problem or opportunity it addresses;
    • the packet evidence that supports it;
    • the assumptions required for it to work;
    • the expected outcome and guardrail signals;
    • the dependencies and material trade-offs;
    • the simplest valid way to reduce its largest uncertainty.

    This format prevents one option from winning because it received a more persuasive narrative. It also makes unsupported leaps visible. If the model cannot connect an option to evidence, that option can remain an idea, but it must be labeled as a hypothesis rather than presented as a conclusion.

    The Skeptic tries to disconfirm the options

    The Skeptic should not produce generic risks. Ask it to find the strongest contrary evidence, the segment that might be harmed, the constraint most likely to invalidate the option, the metric that could be gamed, and the observation that would show the underlying hypothesis is wrong.

    Require it to distinguish counterevidence already present in the packet from new conjecture. This matters because a skeptical tone can sound rigorous even when it is unsupported.

    The same LLM can perform all three roles, but role prompts do not create independent evidence or independent reviewers. Freeze the context packet used for the loop, label every generated artifact, and keep generated claims out of the evidence ledger until a human verifies them. Role separation is a workflow control, not a guarantee of correctness.

    Stop adding passes when the workflow is only rearranging language. The loop has done its job when the team can see the supported facts, viable options, disputed assumptions, material risks, and next evidence needed to decide.

    Make the product trio the decision gate

    AI can accelerate the reasoning, but it should not become the decision owner. Bring the packet and the three-pass output into a product trio of product, design, and engineering. The purpose of that forum is not to approve the AI recommendation. It is to make the trade-offs explicit and decide what the team is prepared to learn.

    1. Verify the evidence boundary. Check whether the represented segments, product states, and metrics match the decision. Ask which customer or operational perspective is absent.
    2. Classify the important claims. Mark each claim as supported observation, team interpretation, assumption, or generated hypothesis. If nobody can trace a recommendation back to the packet, treat it as a hypothesis or remove it.
    3. Compare trade-offs on equal terms. Evaluate every option against the desired outcome, guardrails, constraints, dependencies, and learning value. Do not let the most detailed option appear strongest merely because the model wrote more about it.
    4. Choose the next commitment. The valid outcomes are to proceed, run a discovery or validation step, defer the decision, or reject the options. Assign a human owner and make clear what action the decision authorizes.
    5. Record the rationale. Convert the discussion into a concise decision memo rather than forwarding raw model output to stakeholders.

    The decision memo should include:

    • the decision and why it is being made now;
    • the target segment, desired outcome, and guardrails;
    • the evidence that carried the most weight;
    • the chosen option and the alternatives rejected;
    • the trade-offs accepted by the decision owner;
    • the assumptions and unresolved questions;
    • the validation method and disconfirming signal;
    • the owner and trigger for revisiting the decision.

    This gives stakeholders something stronger than AI-generated confidence. They can inspect what the choice rests on, where judgment entered, what could prove the team wrong, and when the decision should be reconsidered.

    Close the loop with validation and decision memory

    Even a well-grounded model output is not product validation. It is a structured hypothesis. Match the validation method to the claim and to the consequence of being wrong.

    • For a causal behavior claim: use a controlled A/B test when traffic, instrumentation, and the product experience make that appropriate. Define the primary metric, minimum detectable effect, guardrails, analysis approach, and stopping rules before reading the result.
    • For a usability or comprehension claim: use targeted customer interviews or usability evaluation with the relevant segment. AI can help organize notes, but preserve outliers and do not turn a small qualitative sample into a prevalence claim.
    • For an operational claim: use a limited release with observability, support monitoring, and an explicit rollback condition. Watch the workflow around the feature, not only the feature interaction itself.
    • For privacy, brand, regulatory, or other high-consequence constraints: complete the appropriate human review before launch. A persuasive model assessment is not a substitute for the accountable specialist or decision owner.

    For an onboarding decision, for example, the packet may contain segment definitions, observed friction, support themes, and conversion signals. The workflow can propose alternative interventions and measurement plans. The trio still chooses which hypothesis deserves a controlled test, whether the minimum detectable effect is practical, and which activation or retention signals will determine the next move.

    After validation, return the result to the context system. Record what shipped, the observed outcome, affected segments, unexpected behavior, and which assumptions held or failed. Update the decision memo and evidence ledger. Otherwise, the next AI session begins from the same stale assumptions, and the organization pays again to relearn what it already discovered.

    That accumulated decision memory is one of the most valuable outputs of context engineering. It turns AI collaboration from isolated prompting into a feedback loop connecting discovery, strategy, execution, and measurable results.

    Key takeaways

    • Frame the product decision, target segment, outcome, and constraints before asking AI for options.
    • Give the model a compressed evidence packet, not an unstructured pile of documents.
    • Keep observations, interpretations, assumptions, and generated hypotheses visibly separate.
    • Use distinct Summarizer, Strategist, and Skeptic passes to expose where reasoning changes.
    • Let a human product trio own the trade-offs, commitment, and stakeholder rationale.
    • Treat every recommendation as a hypothesis until validation produces new evidence, then feed that evidence back into the decision record.

    Choose the next real product decision that is important enough to validate and bounded enough to act on. Write its decision frame, assemble the smallest safe context packet, run the three reasoning passes, and take a decision memo into your product trio. When the result flows back into the packet, context engineering stops being a prompting technique and becomes part of how you run product.

    References