Tag: continuous discovery

  • How to Build a Continuous Discovery Habit That Survives Delivery

    How to Build a Continuous Discovery Habit That Survives Delivery

    Your team probably doesn’t lack discovery techniques. It loses discovery when delivery becomes urgent. Customer contact clusters around planning, the team commits to a solution, the calendar fills, and assumptions quietly harden into backlog items. Everyone stays busy, but no one can point to the customer evidence behind the next decision.

    Durable continuous discovery is an operating rhythm, not a research phase. The goal isn’t to conduct more interviews. It is to shorten the distance between customer reality and the decisions shaping your product. A weekly rhythm owned by a product trio can reduce rework, sharpen strategy, and keep discovery alive while delivery continues.

    See your real discovery system before changing it

    Adding a recurring customer interview to the calendar won’t fix a decision process built around handoffs. If ideas arrive from executives, become requirements in product, move to design, and reach engineering as implementation work, the interview is an extra activity attached to the side of the system. It isn’t part of how the team decides.

    Start by making the existing system visible. Map what actually happened, including the awkward shortcuts and informal approvals. Do not draw the process described in a playbook.

    1. Spend 60 minutes drawing how your team decides what to build. Show where ideas enter, who shapes them, who approves them, where customers appear, and how the team decides whether the result worked.
    2. Compare your drawing with the drawings made by product, design, and engineering. Differences are evidence that the team does not share the same decision model.
    3. Audit every product decision from last week in a 30-minute session. Include small decisions, not just roadmap commitments.
    4. For each decision, record who made it, what information informed it, and whether the team had direct customer input or received a secondhand interpretation.
    5. Mark the places where discovery and delivery reconnect. A production problem, adoption signal, support request, or implementation constraint can create a new discovery question; it should not disappear into a separate queue.

    This process map and decision audit gives you a baseline without turning discovery into a maturity score. Look for the mechanism behind the misses. Perhaps customer input arrives after commitment. Perhaps the product manager is the only person who interprets it. Perhaps the team can describe the solution but not the opportunity it addresses.

    Track a compact baseline: how recently the team had direct customer contact, which current decisions include direct input, where cross-functional decisions become handoffs, and which active solutions lack a named customer opportunity. Do not set targets yet. First identify where evidence stops influencing action.

    If the process looks reasonable but the habit still collapses, inspect the six prerequisite mindsets: outcome-oriented, customer-centric, collaborative, visual, experimental, and continuous. Turn them into diagnostic questions:

    • Outcome-oriented: Can the team name the customer or business change it is trying to create, or only the feature it plans to ship?
    • Customer-centric: Does the team hear directly from customers, or mainly through sales, support, analytics, and stakeholder summaries?
    • Collaborative: Do product, design, and engineering make decisions together, or meet mainly to exchange work?
    • Visual: Is there one shared representation of the outcome, opportunities, solutions, and assumptions?
    • Experimental: Can the team name what could make the current idea fail?
    • Continuous: Does each learning activity lead to the next question, or does discovery end with a presentation?

    Choose the weakest link as your first intervention. A team with output-based goals does not need a better interview script first; it needs an outcome that gives the interview a purpose. A team dominated by handoffs needs shared sensemaking, not another repository.

    Install a weekly loop small enough to protect

    A habit survives because its trigger, action, and output are clear. Put a recurring discovery block at a stable point in the team’s operating rhythm. Tie it to a current outcome and a live decision, not to a general ambition to understand users better.

    • Trigger: A protected calendar block recurs every week, including during active delivery.
    • Focus: The trio brings one current outcome, the decision in front of it, and the uncertainty preventing a confident choice.
    • Customer contact: The team has a direct customer touchpoint every week. That might be a customer interview, observation of a workflow, or a usability session connected to the current question.
    • Sensemaking: The trio separates what it observed from what it inferred.
    • Update: New evidence changes the opportunity solution tree or confirms why no change is warranted.
    • Commitment: The team names the next uncertainty and starts arranging the next customer contact.

    A customer touchpoint is not any meeting attended by a customer. A sales demo, account review, or advisory session dominated by presentation may be valuable, but it does not automatically answer a discovery question. The useful test is whether the customer can reveal a real behavior, need, constraint, or reaction and whether the team can ask follow-up questions.

    Prepare each touchpoint by completing this sentence: After this contact, the trio might decide whether… If you cannot finish it, the question is probably too broad. Starting with a decision also reduces the temptation to collect interesting comments that never affect the product.

    During the interaction, capture concrete observations before interpretations. Afterward, answer five questions while the context is fresh:

    1. What did the customer do, describe, or struggle to explain?
    2. What interpretation is the team placing on that observation?
    3. Which opportunity or assumption does it affect?
    4. What decision changes, if any?
    5. What remains uncertain enough to examine next?

    The distinction between observation and interpretation matters. A customer abandoning a task is an observation. Assuming that price caused the abandonment is an interpretation. If the team records only the interpretation, an early guess can become institutional memory.

    Recruiting is part of the habit, not administrative work that begins after an interview is requested. Give coordination to a named owner, maintain a rolling pool of relevant customers, and create simple paths for customer success and support to nominate people who recently experienced the problem. Start the next invitation before the current discovery cycle feels complete. Otherwise, every customer cancellation becomes a reason to skip the week.

    When delivery pressure rises, protect the trigger and narrow the activity. Ask a smaller question, review a focused prototype, or examine one step in a workflow. Do not silently replace direct contact with an internal meeting and call the habit complete. If a customer cancels, use the protected time to recruit, refine the decision question, and reschedule. Preserve the rhythm without pretending the missing evidence exists.

    Give the product trio ownership of decisions, not ceremonies

    A product trio is not three people attending the same interview. It is product, design, and engineering sharing responsibility for understanding the opportunity and choosing how to address it. Attendance can rotate. Interpretation and decision-making cannot be delegated to one function and handed back as a deck.

    Make the trio’s decision rights explicit at the start of an outcome. Record the outcome it owns, the decisions it can make autonomously, the constraints it must respect, what requires escalation, and where its evidence will remain visible. Without that contract, discovery may reveal a better direction while the roadmap continues unchanged because nobody knows who can act.

    The responsibilities below are a practical starting point, not rigid job boundaries:

    • Product keeps the outcome, strategic context, customer segment, and pending decision visible.
    • Design helps the trio expose customer behavior, frame opportunities, and choose an appropriate way to learn.
    • Engineering surfaces feasibility, system behavior, data, and implementation assumptions before the solution becomes expensive to change.
    • The trio decides what the evidence means, which option remains viable, and what uncertainty deserves attention next.

    Use a short shared debrief after customer contact. The format can remain simple:

    • Observation: What happened without interpretation?
    • Meaning: What plausible explanations fit the observation?
    • Decision: What will the trio change or preserve?
    • Unknown: What still blocks commitment?

    This prevents the loudest interpretation from becoming the team’s conclusion. It also gives engineering a role before implementation and gives design a role beyond producing artifacts.

    Leadership should ask for evidence of changed decisions, not proof that ceremonies occurred. Instead of asking how many interviews the team completed, ask which opportunity became clearer, which assumption weakened, what decision changed, and how the change connects to the outcome. Interview volume is easy to report and easy to game. Decision quality is harder to display, but it is the reason the habit exists.

    Connect discovery evidence to strategy and delivery

    A weekly customer conversation can still become theater if its evidence floats separately from strategy, roadmaps, and sprint planning. The opportunity solution tree provides a shared spine: the desired outcome sits at the top, customer opportunities sit beneath it, and candidate solutions connect to the opportunities they could address. That outcome-opportunity-solution structure keeps the team connected to why it is considering a particular feature.

    Use the tree as a decision interface, not a workshop artifact:

    • Product strategy: Put the intended outcome at the top so the team can test whether its discovery work supports the strategic direction.
    • Roadmapping: Attach candidate solutions to named opportunities. Keep alternatives visible until evidence or a real constraint justifies commitment.
    • Sprint planning: Require each significant item to trace back to an opportunity and outcome. If it cannot, surface the mismatch before implementation.
    • Customer contact: Update the affected opportunity, solution, or assumption during the debrief. Do not wait for a separate documentation session.
    • Stakeholder communication: Show what changed in the tree, why it changed, and which decision follows. This is more useful than presenting a collection of customer quotations.

    Keep a record of rejected options and the evidence or constraint behind each rejection. Otherwise, an old idea can return with a new label and consume another round of debate. The record should remain revisable: new customer behavior, technical capability, or strategic constraints can justify reopening a branch.

    Measure whether evidence enters decisions

    The safest discovery metric is not an isolated activity count. Measure the health of the loop:

    • Cadence: Did direct customer contact happen during the weekly rhythm?
    • Decision integration: Which current decision did that contact inform?
    • Shared ownership: Did the trio participate in sensemaking, even if every member did not attend the session?
    • Strategic traceability: Can a delivery item be traced to an opportunity and outcome?
    • Learning movement: Which belief, option, or assumption changed?

    A team can conduct many interviews and learn very little if every conversation validates a solution already selected. Conversely, one focused interaction can be valuable when it exposes a faulty workflow assumption and changes a pending decision. Track cadence to protect the habit, but judge value by movement in the decision model.

    Separate customer, model, and operational uncertainty in AI products

    AI product teams face a specific discovery trap: an impressive model demonstration can make technical possibility look like customer demand. Keep different uncertainties separate so one kind of evidence does not answer a different question.

    • Customer uncertainty: What job is the person trying to complete? Where does the current workflow break? Under what conditions will the person trust, verify, correct, or reject an AI-assisted result?
    • Model uncertainty: Does the system produce acceptable behavior for the intended context? Which failures matter to the user, and how will the team evaluate them?
    • Operational uncertainty: Can the product obtain the required data and permissions? Where is human review needed? How will failures be detected, explained, and supported?

    Customer contact can reveal workflow, language, trust conditions, and failure consequences. It cannot prove that the model behaves reliably. Model evaluations can reveal performance and failure patterns. They cannot prove that the workflow is valuable. Operational checks can establish feasibility and controls. They cannot prove adoption. Keep all of these linked to the same outcome while using the right evidence for each uncertainty.

    On the opportunity solution tree, write opportunities in customer terms. “Use generative AI” is a solution direction, not an opportunity. “Reduce the effort required to turn a customer conversation into an accurate follow-up” describes a customer problem that could have AI and non-AI solutions. That distinction helps the trio discover value without becoming attached to a technology.

    Fix the mechanism when the habit breaks

    What you noticeLikely mechanismWhat to change
    The team talks to customers, but the roadmap never changes.Sessions are disconnected from a live decision.Write the decision before recruiting and record what changed immediately after the interaction.
    Engineering joins only after discovery is complete.The trio label is masking a handoff.Include engineering in opportunity framing, assumption identification, and shared sensemaking. Session attendance can rotate.
    Customer sessions repeatedly fall through.Recruitment starts only after a question becomes urgent.Maintain a rolling pool of relevant customers and assign coordination to a named owner.
    The opportunity solution tree is stale.The tree is treated as presentation material.Update it during the debrief and remove or annotate branches that no longer have support.
    Discovery pauses whenever delivery accelerates.Discovery is scoped as a project rather than a continuous rhythm.Protect the weekly trigger and narrow the question or method when capacity is tight.
    Leadership keeps asking the team for certainty.The team reports activities without showing their decision impact.Show the outcome, changed opportunity or assumption, resulting decision, and remaining uncertainty.

    Do not respond to a broken habit by adding more process everywhere. Match the intervention to the failure. A recruiting problem needs a pipeline. A decision-rights problem needs leadership alignment. A stale artifact needs an update trigger. A handoff problem needs shared sensemaking.

    Key takeaways

    • Map the current decision system and audit last week’s decisions before adding a new discovery ceremony.
    • Anchor a direct customer touchpoint every week to a current outcome, decision, and uncertainty.
    • Let attendance vary when necessary, but keep interpretation and decisions jointly owned by the product trio.
    • Use the opportunity solution tree as the live connection between strategy, customer evidence, roadmap choices, and sprint work.
    • When delivery pressure rises, protect the trigger and shrink the activity instead of suspending the cadence.
    • For AI products, do not use customer enthusiasm as proof of model reliability or an evaluation result as proof of customer value.

    Put the recurring customer touchpoint on the calendar, choose the outcome and decision it must inform, and name the product trio responsible for acting on what it learns. At the end of the next weekly cycle, do not ask whether the team “did discovery.” Ask what changed in the decision and what the team needs to learn next.

    References

  • How Product Leaders Turn AI Strategy Into an Operating System

    How Product Leaders Turn AI Strategy Into an Operating System

    Your AI roadmap probably isn’t short of ideas. The hard decision is which ideas deserve production responsibility: a user promise, a quality bar, a failure path, an owner, and a reason to keep funding them after launch.

    You operationalize AI by turning those decisions into a repeatable management system. The broader shift from experiments to execution makes that system more important than any individual model choice. It lets your teams discover useful applications, ship them responsibly, teach customers how to use them, and decide from evidence whether to scale, change, or stop.

    Turn AI ambition into a portfolio of bounded bets

    An AI strategy is not a list of places where a model could be added. It is a set of choices about which customer or business problems deserve investment, how much authority AI should receive, and what evidence will justify the next commitment.

    Start every candidate with a one-page opportunity contract. If the team can describe the model but cannot complete the contract, the idea is not ready for prioritization.

    • User and moment: Name the person, the task they are trying to complete, and the point in the workflow where the difficulty occurs.
    • Current behavior: Record how the task works without the proposed feature. Use an observable baseline such as completion, elapsed time, handoffs, abandonment, rework, or cost per completed task.
    • AI contribution: State whether AI will classify, retrieve, recommend, generate, summarize, or take an action. Avoid vague phrases such as “AI-powered experience.”
    • Expected change: Identify the user behavior that should change first and the customer or business outcome that should follow.
    • Boundaries: List what the system must not decide, which data it must not use, and which users or scenarios are outside the initial release.
    • Consequence and reversibility: Describe what happens when the system is wrong and whether the user can inspect, correct, undo, or escalate the result.
    • Next evidence: Define the smallest test that could reduce the most important uncertainty. That might be a workflow prototype, customer discovery, a retrieval test, or an evaluation against representative cases.

    This contract forces an important distinction between assistance and authority. Drafting a reply for a person to review is not the same product as sending that reply automatically. Recommending an account action is not the same as applying it. The second version has a larger blast radius, a different trust requirement, and a stricter need for auditability and recovery.

    Begin with the minimum authority required to create value. Increase autonomy only when the evidence supports it. This is not timidity. It is a sequencing decision that lets you learn about quality and user behavior before accepting a larger operational risk.

    Prioritize the resulting bets across six lenses: customer value, workflow frequency, data readiness, evaluability, blast radius, and operating cost. Do not collapse them into a decorative score that hides disagreement. Use them to expose the trade-off. A frequent, valuable task may still be a poor first bet if critical failures cannot be detected. A low-risk task may be easy to ship but too marginal to earn repeat use.

    Write a stop condition at the same time as the investment case. For example: stop if the team cannot construct a credible evaluation set, if the workflow requires data the product cannot responsibly access, or if users do not reach the intended outcome after the experience and onboarding have both been tested. A portfolio becomes manageable when stopping is a designed decision rather than an admission of defeat.

    Define production readiness before the team starts building

    A prototype proves that a system can produce a compelling result once. A product must produce an acceptable result across the situations that matter, make its limitations understandable, and recover when the result is not acceptable.

    Give each AI bet a production contract before it enters committed delivery. The contract should contain:

    • The user promise: Describe what the product will help the user accomplish. Do not promise intelligence in the abstract.
    • The context boundary: Specify which product data, retrieved knowledge, instructions, tools, and prior interactions the system may use.
    • The quality dimensions: Choose criteria that fit the task, such as correctness, completeness, groundedness, policy compliance, tool execution, tone, or structured-output validity.
    • Scenario-specific thresholds: Set release criteria for meaningful segments and failure types instead of relying on one average score. The acceptable standard for brainstorming copy is not the acceptable standard for changing an account or communicating a binding decision.
    • The fallback: Define what the user sees and can do when confidence is inadequate, a tool fails, retrieval returns weak context, or the output violates a rule.
    • The operating envelope: Set the latency, reliability, and cost constraints needed for the workflow to remain viable.
    • The data rules: Record what may be retained, what must be removed, who can inspect traces, and how sensitive information is handled.
    • The instrumentation plan: Name the events, evaluation results, feedback, escalations, and outcome measures required to make the next decision.

    There is no universal quality threshold for an AI feature. The right threshold depends on the consequence of an error, the user’s ability to detect it, and the availability of a safe recovery path. Set the bar by scenario and harm, then make the release decision against that bar. An aggregate average can conceal a severe failure in a smaller but important segment.

    Build the evaluation set before tuning the experience

    Create a versioned evaluation set from the workflow you intend to support. Include ordinary cases, meaningful variations, known edge cases, and inputs that should trigger a refusal, clarification, or handoff. Label the expected outcome and the unacceptable failure. Do not require exact wording unless exact wording is part of the product requirement.

    Run that set against the initial baseline and after changes to prompts, models, retrieval, tools, policies, or orchestration. Preserve results by scenario so the team can see both improvements and regressions. A single overall score is useful for orientation; it is not enough for a launch decision.

    Automated checks work well for properties that can be specified clearly, such as output structure, required fields, tool completion, forbidden content, or citation presence. Use structured human review where quality depends on judgement. Keep the rubric stable enough to compare versions, and change it deliberately when the product promise changes.

    Design the failure experience as part of the feature

    Users do not experience your evaluation score. They experience a suggestion they cannot verify, a slow response, an action they did not intend, or a dead end after the system fails. Design those moments before launch.

    • Show the context or inputs that materially shaped the result when doing so helps the user judge it.
    • Make generated content editable before it becomes externally visible.
    • Require explicit confirmation before consequential or difficult-to-reverse actions.
    • Preserve the original state and provide rollback where the underlying workflow permits it.
    • Offer a clear manual path when the system cannot complete the task.
    • Capture corrections and escalations as learning signals without treating every user edit as proof that the system was wrong.

    Do not place sensitive production data into an unapproved model, connector, or testing tool. The downside can include unauthorized disclosure, retention outside your controls, and regulatory or contractual exposure. Use an approved environment and appropriately protected or de-identified test material while privacy and security owners validate the production path.

    Run one decision loop from discovery through scale

    AI initiatives become expensive when discovery, delivery, launch, and governance operate as separate queues. The useful unit of management is one decision loop with shared artifacts, named owners, and explicit gates.

    1. Discover the workflow: Observe the current task, its failure points, the information available at the decision moment, and the user’s existing workarounds. Validate that the problem matters before testing how impressive a model can appear.
    2. Shape a complete slice: Select the smallest workflow that can deliver an outcome, including its context, interface, recovery path, and instrumentation. A prompt without those elements is a component, not a product increment.
    3. Pass the build gate: Approve committed delivery only when the opportunity contract, production contract, evaluation set, data path, and accountable owners are credible.
    4. Deliver through normal product planning: Put evaluation cases, telemetry, fallback behavior, privacy work, and operational readiness into the roadmap and sprint scope. Do not leave them in a separate “hardening” phase after the visible feature is complete.
    5. Launch a new behavior: Use onboarding, in-app guidance, examples, and product tours to show when the capability is useful, what input it needs, and how the user should review the result. The activation event should represent completed value, not a button click.
    6. Review and decide: Compare outcomes with the baseline, inspect evaluation performance by scenario, locate adoption drop-offs, and review cost, reliability, incidents, and new risks. End with a decision to scale, revise, constrain, or stop.

    A practical ownership split keeps this loop moving. Product owns the customer outcome, scope, adoption, and portfolio decision. Engineering owns the production system, reliability, observability, and cost controls. Design owns comprehension, user control, and recovery in the experience. The evaluation owner maintains cases, rubrics, baselines, and regression visibility. Privacy, security, legal, or compliance owners define required controls according to the risk. The business or operational owner defines any human review policy and accepts changes to the real-world process.

    One directly responsible leader should assemble the evidence and drive the launch recommendation, but that role does not erase specialist approval where it is required. Record the decision, conditions, and unresolved risks. Otherwise the same debate returns at every review and nobody can tell why the system was allowed to progress.

    Use risk-tiered oversight. A reversible drafting aid with no sensitive data does not need the same review path as an agent that changes customer records, sends external communications, or initiates a financial action. Increase review, auditability, confirmation, and monitoring as authority and consequence increase. This keeps governance proportional and makes the path to approval understandable before work begins.

    At each portfolio review, use the same compact decision packet: baseline and current outcome, scenario-level evaluation movement, activation funnel, operating performance, incidents or policy exceptions, learning completed, and the next requested commitment. A polished demonstration can support the discussion, but it cannot substitute for this evidence.

    Measure value, quality, adoption, and risk separately

    AI dashboards become misleading when usage, answer quality, customer value, and system health are blended into one success number. They answer different questions and lead to different decisions. Keep the layers separate, then connect them with a driver tree.

    LayerQuestionUseful measuresDecision it informs
    Customer or business outcomeDid the workflow become meaningfully better?Task completion, resolution, conversion, elapsed time, rework, or cost per successful outcomeWhether the use case deserves continued investment
    User behaviorAre eligible users reaching and repeating the value?Eligibility, exposure, first attempt, successful completion, repeat use, abandonment, fallback, and escalationWhether to change positioning, onboarding, interaction design, or workflow placement
    System qualityIs the result fit for the intended task?Scenario pass rate, human rubric results, groundedness where required, tool success, structured-output validity, and critical-failure countWhether to change context, retrieval, prompts, models, tools, or scope
    OperationsCan the product deliver the experience sustainably?Latency, reliability, retries, failure rate, incidents, and cost per successful taskWhether architecture and unit economics support scale
    Risk and controlAre safeguards working at the level of authority granted?Policy exceptions, unauthorized actions, sensitive-data events, confirmations, rollbacks, and human escalationsWhether to add controls, reduce authority, constrain availability, or pause

    Build the adoption funnel around the real workflow: eligible user, meaningful exposure, first attempt, successful outcome, and repeat use when the need occurs again. Define the repeat window from the natural frequency of the task. A daily workflow and a quarterly workflow cannot share a useful retention window.

    Do not mistake interaction volume for value. More messages can mean the user is retrying after poor results. A low cost per response can hide an expensive task that requires several responses and a manual correction. Favor successful outcomes per eligible user and cost per successful outcome, then use interaction-level metrics to diagnose what happened inside the journey.

    The metric layers also tell you where to intervene:

    • If evaluation quality is acceptable but activation is weak, inspect discoverability, positioning, onboarding, and whether the feature appears at the right workflow moment.
    • If first use is strong but successful completion is weak, inspect inputs, context retrieval, interaction design, tool execution, and recovery.
    • If completion is strong but repeat use is weak, verify that the use case is naturally repeatable and that the experience created enough value to displace the old behavior.
    • If adoption is strong but critical failures or operating costs are outside the contract, constrain the release while you fix the production system. Popularity does not neutralize risk or poor economics.
    • If the outcome improves, scenario evaluations remain acceptable, users return when the need recurs, and operating constraints hold, you have evidence to expand availability or authority.

    This is how measurement becomes a funding mechanism rather than a reporting ritual. Each signal points to a different action, and each review produces a clear next commitment.

    Key takeaways for your next AI portfolio review

    • Treat every AI idea as a bounded product bet with a named user, baseline workflow, expected outcome, authority level, and stop condition.
    • Require a production contract covering quality, evaluation, fallback, data, economics, instrumentation, and failure recovery before committed delivery begins.
    • Build privacy, evaluation, telemetry, onboarding, and operational readiness into the roadmap and sprint scope instead of postponing them until launch.
    • Grant the minimum authority needed to create value, then expand autonomy only when quality, adoption, control, and operational evidence support it.
    • Measure customer outcomes, user behavior, system quality, operations, and risk as connected but distinct layers.
    • End every review with an explicit decision to scale, revise, constrain, or stop, plus the evidence required for the next decision.

    At your next portfolio review, choose one leading AI candidate and refuse to discuss the model first. Write the opportunity contract, define its production bar, assign the owners, and identify the first complete workflow you can measure. If those decisions are clear, the technology has a path to become a product. If they are not, another prototype will only postpone the real work.

    References

  • Master the Five Stages of Software Experience Maturity and Prioritize What to Fix First

    Master the Five Stages of Software Experience Maturity and Prioritize What to Fix First

    Experience quality compounds just like code quality. To align teams and accelerate outcomes, I rely on a clear, five-stage software experience maturity model to assess where we are, why we’re there, and how to advance. It turns fuzzy debates into concrete product strategy and reinforces a product-led growth mindset.

    Find out where you stand—and what to fix first—with this maturity framework.

    Why a five-stage model? It gives product, design, engineering, and go-to-market a shared language for trade-offs, helps us move from opinions to evidence, and ties day-to-day improvements to outcomes vs output OKRs. Instead of spreading effort thin, we sequence the right bets at the right time and build momentum with measurable wins.

    Here’s how I apply it in practice. I start with a brief, honest self-assessment across the customer journey: onboarding clarity, user activation moments, in-app guides and product tours, UX writing, support loops, reliability, and analytics coverage. Then I layer in learnings from continuous discovery and product discovery—interviews, usage patterns, and support transcripts—so we see the experience as customers do, not just as we intended.

    When it comes to what to fix first, I prioritize prerequisites over polish. If the value proposition isn’t clear, onboarding is confusing, or activation is inconsistent, we address those before adding new features. I instrument the funnel end-to-end, establish a minimum detectable effect (MDE) for A/B testing, and ensure we can answer basic questions about who activates, who retains, and why.

    Measurement is non-negotiable. I pair retention analysis and activation metrics with qualitative signals to avoid local maxima. Amplitude analytics helps reveal behavioral patterns, while Pendo and in-app guides close gaps in comprehension and guidance. Intercom and CRM integration with HubSpot connect product signals to account health, so we can see how experience maturity drives revenue and retention.

    Operationally, I anchor the roadmap to a small set of experience outcomes, link them to product strategy, and review progress in cadence with leadership. This approach builds product management leadership muscle: sharper stakeholder management, clearer trade-offs, and faster feedback loops. Most importantly, the team sees how each improvement ladders up to a better, more durable user experience.

    If you’re mapping your own path across the five stages, start by sizing the gaps that block activation and retention, commit to a few high-leverage fixes, and measure relentlessly. With a shared maturity model, your team gains focus, your customers feel the difference, and your product compounds value with every release.


    Inspired by this post on Pendo – Best Practices.


    Book a consult png image
  • Why I’m All-In on INDUSTRY 2025: 5 Powerful Reasons For Product Leaders at The Product Conference

    Why I’m All-In on INDUSTRY 2025: 5 Powerful Reasons For Product Leaders at The Product Conference

    INDUSTRY 2025: The Product Conference is circled on my calendar for good reason. In my role leading product management at HighLevel, I look for events that sharpen strategy, accelerate learning, and connect me with operators who ship. This one consistently delivers on all three, and 2025 promises to raise the bar for product management leadership.

    Join Pendo at INDUSTRY in Cleveland, Ohio.

    First, I expect deeply actionable product strategy insights—beyond platitudes. I’m prioritizing conversations on outcomes vs output OKRs, product roadmapping and sprint planning, and how great teams articulate a crisp value proposition while maintaining points of parity that matter. I’m going in with specific questions on product-market fit lessons and how to systematize strategic bets without stifling discovery.

    Second, the surge of AI in product work is too important to observe from the sidelines. I’m comparing approaches across AI Strategy, LLMs for product managers, prompt engineering, and eval-driven development—especially in retrieval-first pipeline patterns. My focus: where AI genuinely improves product discovery, in-app guides, and customer support ai strategy, and where it risks adding complexity without outcomes.

    Third, the community is unmatched for conference networking and pragmatic learning. I’m intentional about meeting product trios who run continuous discovery at scale, as well as leaders who’ve cracked stakeholder management under pressure. These are the moments where competitive differentiation is born—through candid stories of what didn’t work and why.

    Fourth, I’m eager to stress-test data practices that power product-led growth. I’ll be exchanging notes on retention analysis, unified analytics platform decisions, user activation, and how teams integrate qualitative feedback with event data to inform roadmaps. I’m also interested in how practitioners leverage platforms like Pendo, Amplitude analytics, Intercom, and HubSpot to reduce time-to-insight and craft effective product tours and in-app guides.

    Fifth, I treat INDUSTRY as a checkpoint for leadership growth. I’m looking for fresh takes on empowering product teams, first principles decision making, organizational development, and the IC to manager transition. The best sessions don’t just inspire; they give me two moves I can apply with my team on Monday.

    To make the most of the week, I’m applying a continuous discovery mindset: arrive with clear learning goals, capture portable frameworks, and translate at least two insights into experiments before wheels-up. If you’re focused on product strategy, product discovery, and product-led growth, we’ll have plenty to compare and build on together.

    I’ll be in Cleveland ready to learn, share, and connect with peers who care about craft and outcomes. If you’re attending, let’s compare notes on what’s working, what’s stalled, and how we can raise the bar for product management leadership in 2025 and beyond.


    Inspired by this post on Pendo – Perspectives.


    Book a consult png image
  • My Proven Experimentation Playbook for AI PMs: Faster Learning, Safer Launches, Bigger Wins

    My Proven Experimentation Playbook for AI PMs: Faster Learning, Safer Launches, Bigger Wins

    I build AI products with a simple conviction: disciplined experimentation beats intuition. Over the years, I’ve refined a practical playbook that helps my teams learn faster, reduce risk, and turn every release into a smarter next step.

    Product experimentation isn’t luck; it’s a method. Learn how top AI product managers test, measure, and grow smarter with every release.

    I begin every effort with a crisp hypothesis, an expected user or business outcome, and unambiguous success criteria tied to outcomes vs output OKRs. Before writing a line of code, I define primary metrics and guardrails so we know what “good” looks like—and what to stop.

    When the change affects UX, pricing, or activation flows, I favor A/B testing with the statistical rigor to back decisions. We calculate the minimum detectable effect (MDE), choose appropriate randomization units, and pre-register the analysis plan to avoid p-hacking. This gives the team the confidence to scale wins and sunset underperformers quickly.

    AI features demand a tailored approach, so I run eval-driven development before any user sees a variant. We curate golden datasets, score candidate prompts and models, and stress-test failure modes. This is where LLMs for product managers matters: prompt templates, context window management, and a retrieval-first pipeline are all evaluated for quality, latency, and cost-to-serve. I treat “hallucination rate,” safety violations, and bias as first-class metrics under AI risk management.

    To de-risk launches, we ship behind feature flags with CI/CD, monitor DORA metrics, and roll out in stages. Product trios own problem framing to solution delivery, which shortens feedback loops and preserves accountability. If early signals drift from our hypotheses, we pause, adjust, and re-run—no sunk-cost thinking.

    Measurement is non-negotiable. I instrument user journeys end-to-end with Amplitude analytics, track activation and retention analysis, and map behavior to learning objectives. We consolidate logs and events into a unified analytics platform so qualitative insights from customer research pair cleanly with quantitative trends.

    Continuous discovery keeps the engine running. Weekly customer conversations, in-product feedback, and lightweight prototypes ensure we validate needs, not just solutions. The output flows into product discovery, product roadmapping and sprint planning, and a reusable AI product toolbox that scales across teams.

    Finally, I protect the culture that makes experimentation work: we celebrate invalidated hypotheses, document decisions, and optimize for outcomes over output. That’s how empowered product teams sustain product-led growth—even as complexity grows.

    If you’re building AI features today, adopt this playbook to maximize learning velocity, minimize risk, and compound advantage. The method is straightforward: form strong hypotheses, test with rigor, measure what matters, and let evidence—not HiPPOs—guide the roadmap.


    Inspired by this post on Product School.


    Book a consult png image
  • Quantitative Metrics vs. Qualitative Insight: How I Balance Data and Discovery to Grow Products

    Quantitative Metrics vs. Qualitative Insight: How I Balance Data and Discovery to Grow Products

    Quantitative metrics tell the story in numbers; qualitative ones whisper why it matters. Both shape how products grow. Here’s what you need to know.

    In my day-to-day, I rely on quantitative metrics to surface what’s changing in the business and where we need to focus. Activation rate, conversion through the onboarding funnel, feature adoption, retention analysis, and LTV/CAC give me a precise read on performance. I also keep an eye on DORA metrics to understand delivery health and deployment frequency, but I never mistake those for customer outcomes. Numbers spotlight signal—but they rarely explain causality on their own.

    That’s where qualitative analysis earns its keep. Customer interviews, usability studies, win/loss debriefs, support transcripts, and community feedback give me the context behind the charts. Tools like Pendo help me layer in in-app guides and micro-surveys to capture intent and friction in the flow. This combination turns raw data into decisions that actually move the product strategy forward.

    My operating cadence is simple: weekly dashboards to monitor quantitative metrics, ongoing continuous discovery to collect qualitative insight, and a monthly synthesis to reconcile both with our outcomes vs output OKRs. The aim is to move from opinions to evidence, and from anecdotes to patterns. When quant and qual agree, we execute confidently; when they diverge, we design the smallest experiment to learn fast.

    I use a three-question decision tree to choose the method. First, are we exploring or validating? Exploration leans qualitative; validation leans quantitative. Second, do we have enough volume for statistical power? If yes, I’ll run A/B testing with a clear minimum detectable effect (MDE) to avoid false positives. If not, I’ll rely on targeted qualitative discovery until we can instrument a meaningful test. Third, will this decision meaningfully impact our product-led growth or user activation goals? If it will, we invest in both measurement and discovery to reduce decision risk.

    Here’s a concrete example. We once saw a sudden drop in user activation. The quantitative dashboard flagged a step-function change at onboarding step three, but it couldn’t explain why. A quick round of qualitative interviews revealed that our tooltip design buried a critical permission request. We shipped a Pendo-powered in-app guide variant and ran an A/B test to validate the fix. Activation rebounded within a week, and 30-day retention followed suit.

    There are common pitfalls I actively avoid. Chasing vanity metrics that don’t ladder up to outcomes. Conflating shipping speed with customer value by over-indexing on DORA metrics. Overfitting with A/B testing when the MDE is unrealistic for our traffic. And on the qualitative side, mistaking a compelling anecdote for a representative sample without triangulating evidence.

    If you’re looking to tighten your practice, start with a lightweight playbook: instrument core events in Amplitude analytics; define a small set of outcomes vs output OKRs; schedule recurring customer conversations as part of continuous discovery; tag qualitative insights so patterns surface over time; and pair every material UX change with either a well-powered experiment or a clear qualitative learning goal. This creates a unified analytics and discovery loop that compounds.

    Ultimately, quantitative metrics help me prioritize with clarity, while qualitative analysis helps me decide with confidence. When you weave them together, you not only ship faster—you ship the right thing, for the right reason, at the right time.


    Inspired by this post on Product School.


    Book a consult png image
  • Healthcare Product Benchmarks That Matter: Actionable Metrics and Playbooks From Our Report

    Healthcare Product Benchmarks That Matter: Actionable Metrics and Playbooks From Our Report

    I rely on product benchmarks to align teams, sharpen strategy, and accelerate outcomes—especially in healthcare, where stakes are high and complexity is real. Over the years, I’ve learned that the right metrics create clarity across product, engineering, compliance, and go-to-market, enabling faster, safer decisions that translate into measurable impact.

    Discover exclusive data and strategies from our Product Benchmark Report. Compare the healthcare technology industry’s performance across key product metrics.

    When I evaluate a healthcare product’s health, I focus on a few essentials: activation rate and time-to-value for new users, weekly active usage and feature adoption for clinicians and admins, and cohort-based retention analysis to understand whether value compounds over time. I also look at funnel friction (onboarding drop-off, failed setup steps), support load per account, and reliability signals that influence trust—because in healthcare, trust fuels growth.

    Benchmarks turn those metrics into context. They help me answer, “Are we good, or just lucky?” By comparing our numbers to industry peers, I can prioritize the few bets that matter, set outcomes vs output OKRs, and guide empowered product teams to focus on the highest-leverage improvements.

    Operationally, I instrument products with a unified analytics platform and tools like Amplitude analytics and Pendo to track user activation, feature adoption, and in-product journeys. Pairing that with continuous discovery keeps insights fresh, while A/B testing and clear minimum detectable effect (MDE) thresholds ensure we ship with statistical confidence.

    In practice, my playbook for healthcare product-led growth is straightforward: simplify onboarding with targeted product tours and in-app guides, tighten the first-win loop to reduce time-to-value, and eliminate blockers surfaced by behavioral analytics. Then, reinforce the loop with lifecycle messaging, role-specific education, and clear value propositions for clinicians, operations teams, and executives.

    Of course, none of this works without strong governance. Data governance and regulatory compliance aren’t just guardrails; they’re growth enablers. Clear audit trails, privacy-by-design, and reliable incident management build the trust that keeps adoption high and churn low.

    If you’re ready to benchmark your roadmap against the market, this report gives you the clarity to spot gaps, the language to align stakeholders, and the metrics to execute with precision. Use it to calibrate your product strategy, guide your next set of experiments, and confidently scale what works across the healthcare technology ecosystem.


    Inspired by this post on Amplitude – Perspectives.


    Book a consult png image
  • How to Connect Voice of Customer to Behavioral Analytics

    How to Connect Voice of Customer to Behavioral Analytics

    You have interview notes, support tickets, sales objections, app reviews, and in-product feedback. Yet the roadmap discussion still comes down to which customer complained most recently or which stakeholder tells the most persuasive story.

    The way out is not another survey. Connect each voice-of-customer theme to the behavior of the people who expressed it. You can then see whether the problem changes activation, task completion, adoption, retention, or conversion; identify where the friction occurs; and decide whether the opportunity deserves roadmap space.

    Start with the decision, not the feedback backlog

    VOC becomes useful when it can change a decision. Before analyzing a theme, ask what you would do differently if the concern proved material. Would you redesign an onboarding step, improve reporting performance, simplify permissions, clarify pricing, or leave the current experience alone?

    If the answer is unclear, the theme is not ready for prioritization. It may still be worth tracking, but it should not become a roadmap item merely because it appears frequently.

    Write the theme as a behavioral hypothesis:

    Customers who encounter or mention [theme] while attempting [job] are more or less likely to [observable behavior] within [relevant window] than comparable customers who do not.

    VOC-to-behavior hypothesis template

    A useful hypothesis contains six parts:

    • Population: The users or accounts eligible to encounter the problem.
    • Job: What they were trying to accomplish, not merely the page they visited.
    • VOC theme: The friction expressed in neutral language, such as onboarding confusion or performance slowness.
    • Behavioral signal: The action or pattern you expect to observe, such as abandonment, backtracking, repeat clicks, or slow task completion.
    • Outcome: The activation, adoption, conversion, or retention metric that could move.
    • Window: The period in which that behavior and outcome are meaningful for your product.

    For example, a complaint that a flow is too complex can become a testable expectation: affected users will take longer on a step, move backward more often, depend more heavily on tooltips, or abandon the funnel at a particular screen. Those observations will not explain the customer’s motivation on their own, but they will reveal whether the stated friction has a visible behavioral footprint.

    This distinction matters. Feedback explains how customers interpret an experience. Analytics records what happened. Neither is sufficient alone. Treat the comment as a hypothesis and observable product behavior as the evidence that tests it.

    Build a shared spine between what customers say and do

    You cannot reliably connect VOC to behavior when the two systems describe customers, product areas, and outcomes differently. The work begins with a shared measurement spine: consistent identities, timestamps, product concepts, and definitions.

    Instrument the moments that represent value

    Do not begin by tracking every click. Begin with the moments that determine whether a customer reaches value:

    • The start and end points used to calculate time-to-first-value.
    • The steps and completion event in the onboarding funnel.
    • The first meaningful use of a core feature.
    • The repeated behaviors that indicate adoption rather than experimentation.
    • The conversion event that represents a real commitment.
    • The activity and return criteria used in retention analysis.

    Each event needs an explicit trigger, a user or account identity, a timestamp, and the contextual properties required for segmentation. In a business product, retain both user-level and account-level identity where your data rules permit it. A frustrated user may submit the ticket, while account retention and revenue are measured elsewhere.

    Definitions deserve the same discipline as instrumentation. If onboarding completion means reaching one screen to Product and completing a different workflow to Customer Success, the resulting cohort comparison will settle nothing. Record the definition, owner, applicable population, and known exclusions for every decision metric.

    Amplitude analytics, Pendo, or another unified analytics platform can support funnels, cohorts, and retention curves. The platform does not remove the need for a clean event taxonomy. Better charts built on inconsistent events only make the wrong conclusion look more convincing.

    Normalize VOC without stripping away its meaning

    Customer feedback arrives in incompatible forms: a support ticket describes a blocked task, a sales note records an objection, an app review compresses several problems into one comment, and an in-product response refers to the screen the customer is currently viewing. A shared theme taxonomy makes those inputs comparable.

    For each feedback record, capture the minimum fields needed to analyze it:

    • The original wording or a reference to it, so the nuance remains recoverable.
    • A neutral theme and, where necessary, a more specific subtheme.
    • The product area and job the customer was attempting.
    • The date, touchpoint, and customer or account identifier available under your privacy and data-governance rules.
    • The customer’s lifecycle stage, plan, role, or other context needed to define an eligible comparison group.
    • Whether the customer described a symptom, proposed a solution, or did both.

    That final distinction prevents a common roadmap error. A request for another button is a proposed solution. The underlying problem may be that the current action is hard to discover, too slow, or unavailable to the customer’s role. Preserve the request, but tag the friction separately. Otherwise, you will count preferred implementations rather than customer problems.

    Keep the taxonomy small enough that different people apply it consistently. Split a theme only when the distinction would produce a different cohort, root-cause investigation, or product decision. A label that never changes analysis is administrative detail, not useful structure.

    Turn each VOC theme into a fair cohort comparison

    Once the datasets share identities and definitions, build a cohort containing the users or accounts associated with a theme. Then compare that group with customers who were genuinely capable of encountering the same experience.

    Use this sequence:

    1. Define the expressed cohort. Include customers associated with the theme during a stated period. Preserve the feedback date so you can distinguish behavior before and after the comment.
    2. Define eligibility. Exclude customers who could not access the feature, workflow, plan, permission level, or product version involved.
    3. Create the comparison cohort. Use customers with a similar lifecycle stage and opportunity to perform the job, but without the same recorded theme.
    4. Align the observation window. Give both cohorts the same opportunity to complete the funnel, activate, adopt the feature, or return.
    5. Locate the behavioral difference. Compare funnel steps, task time, navigation patterns, feature adoption, conversion, and retention where each is relevant.
    6. Segment the result. Check whether the effect is concentrated by role, plan, account type, entry path, or another product-relevant dimension.
    7. Return to the qualitative evidence. Review the wording and relevant sessions around the point where behavior diverges. This is where the probable cause becomes specific enough to design against.

    The comparison group matters as much as the expressed cohort. Users who contact support are not a random sample. They may be more engaged, more experienced, more valuable, or simply more willing to report problems. A behavioral difference therefore shows an association worth investigating; it does not prove that the theme caused the outcome.

    Timing creates another trap. A customer may open a ticket because a task already failed. If you combine activity from before and after the ticket, the analysis can confuse the cause, the failure, and the attempt to recover. Anchor the timeline to the relevant exposure or task attempt, and use the feedback timestamp as context rather than automatically treating it as the beginning of the problem.

    Interpret repeated actions carefully as well. Repeat clicks can indicate an unresponsive control, uncertainty about whether a request registered, or deliberate power use. Backtracking may reflect confusion or a legitimate comparison workflow. Pair the pattern with funnel position, timing, interface state, and customer language before naming the root cause.

    Your output should be an evidence statement, not a dashboard tour. A strong statement identifies the eligible segment, the observed difference, where it appears, the outcome associated with it, and the remaining uncertainty. That is enough for a product trio to decide whether to investigate, intervene, or stop.

    Prioritize the behavioral gap and validate the fix

    Raw feedback volume is a weak prioritization rule because it has no denominator. A theme can generate many tickets because the workflow is widely used, because the problem is severe, or because the affected customers are unusually vocal. Reach, behavioral impact, and proximity to a meaningful outcome separate those possibilities.

    Build a compact opportunity case for each material theme:

    • The eligible population and the portion associated with the theme.
    • The behavior gap between the expressed and comparison cohorts.
    • The funnel, activation, adoption, conversion, or retention outcome connected to that gap.
    • The segment in which the effect is concentrated.
    • The probable root cause and the evidence supporting it.
    • The smallest intervention capable of testing that cause.
    • The primary metric, guardrails, and uncertainty that remain.

    A practical sizing model is: eligible population multiplied by the observed behavior gap multiplied by the value of recovering the affected outcome. Use a range when the inputs are uncertain. The purpose is not to manufacture a precise forecast. It is to expose whether your business case depends on broad reach, a large outcome gap, a valuable segment, or an assumption that still needs evidence.

    Do not rank opportunities by the size of the gap alone. A large drop in a low-value side path may matter less than a smaller gap immediately before activation. Conversely, a retention difference may be associated with the theme without being caused by it. Confidence intervals and explicit assumptions help keep opportunity sizing proportional to the evidence.

    When you ship, test the causal claim you actually care about. State the eligible population, intervention, primary metric, guardrails, and minimum detectable effect before looking at results. Use an A/B test when random assignment is practical. If you must rely on a staged rollout or observational comparison, label the result accordingly and keep plausible alternative explanations visible.

    Success is not a warmer survey response by itself. The behavior implicated by the original theme should move: fewer relevant drop-offs, less unnecessary backtracking, faster task completion, stronger activation, or better retention. Sentiment can confirm that the experience feels better, but the original behavioral hypothesis should still be tested.

    What a complete feedback-to-outcome loop looks like

    One reporting experience illustrates the sequence. Customers described reporting as slow. The behavioral trail contained long load times and repeated clicks on filters, which narrowed the problem beyond the broad complaint. The response combined simpler defaults, prefetching important queries, and clearer loading states. In that case, the changes reduced perceived wait time by 42% and improved day-7 retention for the affected cohorts.

    That result is a case-specific outcome, not a benchmark to paste into another business case. The transferable lesson is the chain of evidence: customer language identified the experience, behavioral data located the friction, the intervention addressed the probable mechanism, and the affected cohort supplied the right place to measure retention.

    Make this chain part of the operating cadence. Use a weekly listening review with the product trio to classify emerging themes and flag missing instrumentation. Use a monthly synthesis to join mature themes with usage data, refresh opportunity cases, and retire claims that behavior does not support. When a change ships, return to the original expressed cohort and the relevant outcome window rather than declaring success from aggregate usage.

    Key takeaways

    • Start with the roadmap decision a VOC theme could change, then express the theme as a behavioral hypothesis.
    • Give feedback and product events a shared spine: consistent identities, timestamps, product areas, jobs, and outcome definitions.
    • Compare customers who expressed a theme with customers who had the same opportunity to encounter the experience.
    • Align observation windows and lifecycle stages before interpreting funnel, activation, adoption, or retention differences.
    • Treat cohort differences as evidence of association, not automatic proof of causation.
    • Prioritize the affected population, behavior gap, outcome value, and strength of evidence rather than ticket volume alone.
    • Validate the proposed mechanism with an experiment and a predetermined minimum detectable effect whenever random assignment is practical.

    At your next listening review, choose the VOC theme consuming the most roadmap attention. Write one behavioral hypothesis, identify the eligible cohort, and compare one outcome that would make the problem worth solving. If you cannot complete that chain, the next priority is not another feature request. It is the missing identity, event, definition, or feedback tag preventing you from making the decision responsibly.

    References

  • Inside Google’s Product Model: Hard-Won Lessons to Build Empowered, Outcome-Driven Teams

    Inside Google’s Product Model: Hard-Won Lessons to Build Empowered, Outcome-Driven Teams

    I’ve been systematically exploring how the product model shows up inside iconic companies. After studying “The Product Model at Spotify” and “The Product Model at Amazon,” I’m turning my lens to Google—specifically, how the product operating model, product culture, and product strategy manifest in practice and what we can pragmatically take back to our own organizations.

    When I talk about the product model, I’m looking at the machinery that connects strategy to outcomes: empowered product teams, clear decision rights, tight product trios, continuous discovery, data-informed bets, and an operating cadence that enables learning at speed. My goal here is to unpack how those elements come together at Google and translate them into repeatable patterns you can adopt.

    At a high level, I focus on how teams are empowered to solve problems rather than ship outputs, how outcomes vs output OKRs clarify what matters, and how experimentation (from rapid prototyping to A/B testing) de-risks decisions before they scale. I also examine how engineering and product partner to balance platform scalability with customer value, and how stakeholder management reinforces alignment without slowing teams down.

    Why does this matter? Because the product model is a lever for resilience and speed. When product strategy is explicit and the operating model is built for learning, organizations multiply the impact of talented people. That’s how small, focused teams repeatedly deliver outsized results—even in complex, regulated, or high-scale environments like Google.

    In the sections that follow, I’ll synthesize what I see as the core patterns behind Google’s approach and distill them into actionable guidance: how to structure product trios, how to run continuous discovery alongside delivery, how to set and calibrate OKRs for outcomes, and how to evolve your product culture so empowered product teams can do their best work. My aim is not to idolize a model, but to extract what’s portable and help you adapt it to your context.


    Inspired by this post on SVPG.


    Book a consult png image
  • How Product Leaders Should Plan Their 2026 Conference Calendar

    How Product Leaders Should Plan Their 2026 Conference Calendar

    Your 2026 conference budget should not begin with a list of famous events. It should begin with a decision your product organization is struggling to make: where AI belongs in the roadmap, why discovery is not changing priorities, how product operations should reduce friction, or which growth problem deserves executive attention.

    That shift turns conference planning from a travel exercise into a portfolio of strategic bets. You can choose events for the decisions, relationships, and operating changes they can support – and decline the ones that offer interesting content without a credible path to action.

    Decide what each conference must change

    There is no universally best product conference. The right choice depends on the uncertainty you need to reduce and what you are prepared to do with the answer.

    The real cost is larger than registration and travel. It includes the attendee’s attention, the decisions someone else must cover, the interruption to active work, and the follow-through required after the event. An inexpensive ticket can therefore be a poor investment, while a more demanding trip can be defensible when it helps unblock a consequential product decision.

    Before approving a booking, require a short conference thesis with these fields:

    • Decision: What active product, AI, growth, hiring, or operating-model decision should improve?
    • Uncertainty: What does the team not know well enough to decide confidently?
    • External value: Which practitioners, perspectives, or examples are difficult to access inside the company?
    • Action: What will the attendee recommend if the current hypothesis becomes stronger, weaker, or more conditional?
    • Owner: Who has the authority to act on what is learned?

    For an AI product leader, a useful thesis might be: I need operator evidence that helps us decide whether the next investment belongs in a customer-facing capability, an internal workflow, or the evaluation and governance layer beneath both. That is specific enough to shape session choices and conversations. Learn more about AI is not.

    The same standard applies to networking. Meet product leaders is too vague to guide behavior. Compare how product executives assign decision rights between a central AI platform group and embedded product teams gives the attendee a real question, a relevant peer profile, and a reason to follow up.

    Key takeaways

    • Choose a conference for an active decision, not for its reputation alone.
    • Build a portfolio across strategic priorities instead of sending everyone to similar general-interest events.
    • Use date collisions to force explicit tradeoffs about content, peer access, and geography.
    • Send the person closest to the uncertainty, provided that person has enough authority to apply what they learn.
    • Measure changed decisions and operating behavior, not notes collected or badges scanned.

    Build a portfolio around strategic priorities

    Conference names are useful screening signals, but they are not proof of agenda quality. Start by matching event themes to company priorities. Then inspect the organizer’s current agenda, speakers, attendee profile, format, and location before committing money or executive time. Verify the date and venue directly before purchasing because event details can change.

    The following map helps narrow the field. It is a planning tool, not a ranking.

    Strategic priority2026 events to investigateQuestion that should govern selection
    AI products and intelligent interfacesAI Product Summit on April 15 in San Jose, ProductCon AI on August 5 as a virtual event, and ACM IUI from March 23-26 in PaphosDo you need product operating practices, broad virtual access, or deeper thinking about intelligent user interfaces?
    Executive product leadershipGartner Product Leadership Conference on March 9-10 and Chief Product Officer Summits in New York, Palo Alto, Amsterdam, and San FranciscoWhich leadership decision needs calibration with peers who have comparable scope and accountability?
    Product operationsProduct Operations Summits in New York on March 26-27, Amsterdam on May 12-13, and San Francisco on September 22-23Are you trying to improve decision flow, planning, discovery infrastructure, tooling, or cross-functional accountability?
    Product-led growthProduct-Led Summits across Washington, Austin, New York, Denver, Amsterdam, Seattle, London, San Francisco, Berlin, Sydney, Boston, and TorontoWhich growth problem – activation, adoption, expansion, pricing, or organizational ownership – is important enough to justify attendance?
    Discovery, UX, and research capabilityACM CHI, UX360, UXLx, UXDX, UXPA International, uxcon, and Leading DesignDoes the team need a new method, stronger leadership practice, or better integration of research into product decisions?

    A balanced portfolio can contain different kinds of bets. An anchor event can provide a broad external scan and senior relationships. A specialist event can address a narrow operating problem. A virtual event can provide targeted content without travel. A regional event can deepen a market-specific network. You do not need every type; use only the ones connected to current priorities.

    This matters because superficially similar conferences can serve different jobs. A product-led growth event may be useful when the team is redesigning activation or commercial ownership. It is probably redundant when the same attendee recently explored the same questions and no resulting work has reached implementation. Repetition is valuable only when the audience, market, or decision has changed.

    AI-themed events deserve an especially strict filter. A prominent AI label does not tell you whether the agenda will help with real product work. Look for evidence that sessions address evaluation, data readiness, workflow adoption, reliability, governance, product economics, and organizational change – not only model demonstrations. If AI hiring is the objective, confirm that the likely participants include the builders or leaders you need to understand, rather than assuming a large technology audience will produce relevant candidate relationships.

    Keep some conference capacity uncommitted. Later announcements, agenda changes, and newly urgent company problems can make an event that looked optional more relevant than one selected during annual budgeting. A complete calendar created too early can leave no room for better information.

    Use calendar collisions to make sharper choices

    The 2026 schedule contains several useful collisions. They are not merely logistical problems. They reveal whether your selection criteria are strong enough to distinguish one event from another.

    On March 26 in New York, the Chief Product Officer Summit, Product Operations Summit, and Product-Led Summit occur on the same day, with the latter two continuing through March 27. One attendee cannot meaningfully cover all three. A product executive working through organizational design may favor the leadership audience. A product operations owner rebuilding planning and decision infrastructure has a different reason to be there. A growth leader should not choose either merely because colleagues are attending.

    Amsterdam creates a similar choice on May 28-29, when the Chief Product Officer Summit and Product-Led Summit run on the same dates. If both themes matter, split coverage only when each attendee has a distinct brief and a real route to implementation. Dividing the team without separate objectives simply doubles the expense and produces overlapping summaries.

    San Francisco offers a cluster rather than a direct collision: the Chief Product Officer Summit takes place on September 17, followed by Product-Led and Product Operations Summits on September 22-23. Combining them may reduce duplicated travel, but it also creates time away between events. The itinerary is justified only if the later event answers a separate, important question and the attendee can use the intervening time productively.

    Europe’s June schedule shows why geographic proximity can be misleading. Mind the Product runs in London on June 15-16, Growth Minded Superheroes is in Frankfurt on June 16, UX360 EU is in Berlin on June 23-24, Product-Led Summit is back in London on June 24-25, and Product at Heart follows in Hamburg on June 26. These cities can look like one efficient conference run on a map. In practice, conflicting dates, transfers, context switching, and prolonged absence can overwhelm the incremental learning.

    Use these rules when events overlap or cluster:

    1. For a same-city overlap, choose by problem. Split coverage only when the briefs are materially different and both attendees can act afterward.
    2. For a nearby-city cluster, calculate absence as well as travel. Add transfers, working days between events, and recovery time to the decision.
    3. For a cross-region collision, protect the portfolio. Select the event serving the more important company priority, not the one creating more fear of missing out.
    4. For a virtual alternative, distinguish content from access. Choose virtual participation when learning is the main objective; protect in-person travel for relationships or interactions that genuinely require presence.

    Also block the attendee’s return capacity before approving the trip. If the calendar is packed with internal meetings immediately after the event, synthesis and follow-up will be displaced by routine work. That turns an expensive learning opportunity into a folder of notes no one uses.

    Send the right attendee with an operating brief

    The most senior available person is not automatically the right attendee. Neither is the person most eager to travel. Match the attendee to the decision, the conversations required, and the authority needed afterward.

    Conference objectiveBest-positioned attendeeAuthority or support required afterward
    Reframe AI product strategyThe product or AI leader accountable for the portfolio decisionAbility to change priorities, commission validation work, or define a new investment thesis
    Improve product operationsThe leader or operator who owns planning, decision flow, tooling, or product ritualsA sponsor willing to change responsibilities, forums, or operating mechanisms
    Strengthen continuous discoveryA product, design, or engineering representative close to active customer and delivery workA real product area in which to apply the method and partners prepared to participate
    Calibrate executive leadershipThe product executive who owns the relevant organizational or stakeholder decisionAccess to the executive forum where the operating change will be decided
    Inform AI hiringThe person accountable for role design, assessment quality, or the hiring decisionPermission to update the role scorecard, sourcing thesis, or interview process

    A practitioner sent to an executive event may hear useful ideas but lack access to the people or forums needed to apply them. An executive sent to a method-heavy event may return with broad principles while missing the implementation detail. When the objective spans levels, assign a primary attendee and an internal sponsor instead of assuming one person can represent every perspective.

    The operating brief should travel with the attendee. Include:

    • Outcome statement: The decision or operating change this event should inform.
    • Question set: The questions that sessions and conversations must help answer.
    • Current hypothesis: What the team presently believes, including the conditions that might make it wrong.
    • Conversation map: Relevant peer operators, complementary functions, speakers, customers, partners, or prospective hires to look for.
    • Agenda rules: Sessions that directly answer the brief, useful alternatives, and content that can be skipped without regret.
    • Capture format: Claim, evidence, context, transferability, open question, and follow-up.
    • Coverage plan: Decisions delegated during the attendee’s absence and the escalation path for anything that cannot wait.
    • Return forum: The operating meeting where recommendations will be considered, not merely presented.

    Networking becomes easier when the attendee has a real problem to discuss. A useful introduction contains four things: your role, the situation you are working through, the precise question, and a modest request to compare approaches. That gives the other person something concrete to respond to. A generic request to connect transfers all the work to them.

    Do not optimize for the number of conversations. Optimize for relevance and continuity. A smaller set of exchanges tied to active work is more useful than a long contact list with no reason for another interaction. Before leaving the event, record why each follow-up matters and what question should move forward.

    Measure return through decisions and behavior

    Conference return is often reduced to attendance, notes, leads, or an internal presentation because those outputs are easy to count. None proves that the company made a better decision or changed how it operates.

    Use a decision ledger instead. Before the event, record the active decision, current hypothesis, unresolved evidence, owner, and forum where action can be authorized. After the event, update the same record with the strongest signal, its origin, the context in which it appeared to work, the resulting confidence change, and the next action.

    Evidence quality matters. A vendor claim, a conference-stage success story, a private conversation with an operator, and a pattern repeated across unrelated practitioners should not receive equal weight. The attendee should label which kind of signal they captured and identify what still needs internal validation. Conference learning can shape a test or decision; it should not bypass product judgment.

    Return channelEvidence that countsWeak proxy
    Decision qualityA documented product, AI, growth, or operating decision was changed, accelerated, narrowed, or deliberately held pending better evidenceA polished event summary
    ExecutionAn experiment, discovery activity, hiring change, or operating practice has an owner and a checkpointGeneral enthusiasm about trying new ideas
    NetworkA relevant relationship continues around an active problem, with a clear reason for follow-upContacts or badge scans collected
    TalentThe team improves a role thesis, evaluation approach, market map, or ongoing candidate relationshipA stack of resumes without fit assessment
    Knowledge transferColleagues can apply the insight to current work and understand the conditions under which it may failSlides placed in a shared folder

    Not every conference needs direct revenue attribution. Executive calibration, hiring insight, and discovery capability can create value indirectly. Forcing a speculative revenue number onto those outcomes creates false precision. It is still reasonable to demand a visible chain from attendance to evidence, from evidence to a decision, and from the decision to owned work.

    Cancel or reassign attendance when that chain is missing. A booked event should be reconsidered if the final agenda no longer matches the priority, the attendee cannot articulate a decision thesis, the same questions were recently explored elsewhere, or no one has capacity to act on the result. Sunk planning effort is not a reason to spend more time and money.

    Open your planning calendar with one real company priority, then review the published 2026 product conference dates. Shortlist the event whose audience and agenda best match the unresolved decision. Assign the decision owner, write the operating brief, protect the return forum, and only then approve the booking. If an event cannot survive that sequence, decline it and preserve the capacity for one that can.

    References

  • Year-End Reflection for Product Leaders: Values, Themes, and the 100‑Wishes Reset

    Year-End Reflection for Product Leaders: Values, Themes, and the 100‑Wishes Reset

    I’ve been closing the year with a deliberate reflection ritual for more than a decade, and this season I found fresh energy for it after listening to an insightful conversation with Teresa Torres and Petra Wille on All Things Product. Their approaches mirror the evolution many product leaders experience: moving from rigid annual goal-setting to values-led themes, longer time horizons, and a healthier respect for spaciousness. In my own practice, that shift has created better focus, less pressure, and far more meaningful outcomes.

    Prefer to listen? You can find this episode here: Spotify | Apple Podcasts. I took notes with my team in mind and translated the discussion into a simple, values-driven framework that any product organization can adopt.

    Why does annual reflection matter for product people? Because our work lives at the intersection of ambiguity, trade-offs, and time. If we only measure ourselves by shipped output or quarterly OKRs, we overlook the compounding value of learning, relationships, and judgement. I treat this ritual as a strategic reset: a chance to surface patterns, adjust expectations, and recommit to outcomes over output.

    My own reflection habit started scrappy—paper notebooks, messy timelines, and even artful visualizations inspired by Dear Data by Giorgia Lupi & Stefanie Posavec. Like Petra, I’ve found that tactile, analog artifacts unlock insights I miss in a spreadsheet. Over time, I’ve kept the spirit and simplified the mechanics: a “what went well” review, a short list of hard lessons, and a handful of decisions that paid off—or didn’t.

    The biggest evolution for me has been moving from rigid annual goals to values and themes. I still run OKRs, but I use them to track progress, not identity. The lens of process vs. outcome goals—reinforced by ideas from Atomic Habits—helped me set fewer, better commitments. For example, instead of “launch X by Y,” I’ll emphasize the cadence of customer discovery, the health of the product trio, and the quality of decisions made along the way.

    One exercise that changed my practice is the “100 wishes” list. It’s powerful—and surprisingly difficult. Pushing past 30 or 40 wishes forces me to name latent interests and long-range intentions I rarely say out loud. Combined with decade-level themes, the list helps me balance ambition with patience. I don’t try to do it all next year; I use it to spotlight direction, not deadlines.

    I also review patterns across years: Where did over-scheduling create hidden costs? When did I protect focus time and what did that unlock? Paul Graham’s Maker’s Schedule, Manager’s Schedule remains a useful calibration tool here. And when I feel the pull toward constant throughput, I revisit Stefan Sagmeister’s The Power of Time Off (TED Talk) to remind myself why strategically creating space often yields the most valuable ideas.

    Of course, not every year follows plan—and that’s normal. Reflection helps me spot unrealistic expectations early and let them go. When setbacks hit, I’ll rewatch Dealing with Setbacks and re-ground in continuous discovery. The question isn’t “Did we do everything?” but “Did we learn fast, protect customer value, and make trade-offs aligned with our values?” That’s how empowered product teams compound impact.

    My sharing philosophy has become more nuanced over time. Some reflections are public to invite dialogue and accountability; others stay private so I can process honestly. I’ve found it helpful to publish what I’m saying no to, capture a theme for the year ahead, and keep the rest for myself and my team. This balance preserves motivation while still contributing to the broader product management leadership community.

    If you’re designing your own ritual, consider this lightweight flow: review wins and tough calls, write your “100 wishes,” extract a few values-based themes, then translate those into process goals for Q1. Revisit monthly, not just annually. If you like structured prompts, Chris Guillebeau’s How to Conduct Your Own Annual Review from The Art of Nonconformity offers a practical template you can adapt to your context.

    For deeper dives and complementary ideas, I bookmarked these as part of my year-end reset: What I’m Saying No to This Year—And Why, Ask Teresa: My Leaders Still Want Roadmaps with Timelines—What Should I Do?, Scaling Impact: A Look at the Year Ahead (2022), Let’s Connect in 2025: A Look at the Year Ahead, The Interview Coach, and Petra’s own year-ahead reflections (here and her 2026 version). I also recommend revisiting the prior conversation on leadership and change: Role of Leadership in Transformations.

    I’d love to hear how you approach your end-of-year reflection. What questions bring you the most clarity? Which practices help you set an intentional, values-driven path for the next year? Share your process—I’m always looking to learn from other product creators and leaders.


    Inspired by this post on Product Talk.


    Book a consult png image
  • From Concierge to AI Marketing Engine: Inside Mowie’s Document Hierarchy Playbook

    From Concierge to AI Marketing Engine: Inside Mowie’s Document Hierarchy Playbook

    I’m constantly asked by SMB owners: What if your small business could have a full marketing team—automated content calendars, customer segmentation, and channel-specific posts—without the headcount? That question is no longer hypothetical; it’s precisely the promise behind Mowie, and the way they got there is a masterclass in practical AI product development.

    I recently listened to Chris O'Connor (CEO) and Jessica Valenzuela (Co-Founder) of Mowie, an AI marketing platform built for small and medium-sized businesses in restaurants, retail, and e-commerce. Their story starts with a concierge marketing service—doing the work by hand for overwhelmed owners—and evolves into a fully automated AI product.

    They walk through their "document hierarchy" approach: how Mowie crawls the web to build a "dossier" about each business, infers customer segments and marketing pillars, and generates quarterly content calendars with channel-specific posts. As a product leader, this is the kind of retrieval-first pipeline that consistently outperforms naive prompt chaining because it builds durable context before generation.

    They also unpack the technical challenges of structuring unstructured data and the evolution from rigid schemas to loosely structured markdown. In my experience with LLMs for product managers, markdown becomes a flexible intermediate representation that’s easy to diff, trace, and feed back into models without brittle parsing.

    Equally important, they use customer feedback—from calendar approvals to regeneration requests—as their primary evaluation signal. That’s eval-driven development in practice: close the loop with lightweight evals that reflect genuine user intent, not proxy metrics.

    The planning model is elegant: the three mini-calendars—public events, business-specific events, and recommended campaigns—roll up into a coherent plan that eliminates the blank-page problem and enables steady, predictable execution.

    Crucially, they’re building traceability so customers can see which context documents influenced their content. This kind of transparency increases trust, accelerates edits, and supports governance in regulated categories where auditability matters.

    Onboarding and data collection stay pragmatic: let the system crawl first, ask humans only for deltas, and progressively profile over time. It’s a pattern I advocate in continuous discovery and AI workflows—keep humans in the loop without overwhelming them, and make the right action the easy action.

    Early on, they used Simon Sinek's Golden Circle framework to validate demand and sharpen messaging. Framing the "why" before the "what" helps teams maintain a crisp value proposition and tighten their go-to-market strategy.

    Performance measurement goes beyond vanity metrics by connecting marketing performance back to point-of-sale data for attribution. The ability to tie campaigns to revenue events is the bridge from clever content to accountable outcomes.

    What’s next is equally compelling: deeper attribution, omnichannel expansion, and digital out-of-home displays. For SMBs, that points to a unified analytics platform spanning email, social, and in-store touchpoints—exactly where modern marketing is headed.

    My takeaways for builders: invest in a retrieval-first pipeline with a resilient document hierarchy; prefer loosely structured markdown over rigid JSON when dealing with messy inputs; design human-in-the-loop controls that double as evals; and always connect activity to business outcomes. That’s how you turn an idea into a repeatable system that scales.

    If you want to explore further, start here: Mowie AI — AI marketing platform for SMBs. For early validation and storytelling, revisit Simon Sinek's Golden Circle.


    Inspired by this post on Product Talk.


    Book a consult png image