Category: Product Management

  • How to Evaluate AI Voice Support in Real-World Conditions

    How to Evaluate AI Voice Support in Real-World Conditions

    You have a shortlist of AI voice support products, a polished recording, and a decision that could affect thousands of customer conversations. The hard question is not whether an agent can sound convincing during one ideal call. It is whether the system stays useful when a caller interrupts, corrects themselves, asks an ambiguous question, waits on a backend system, or needs a human.

    You can answer that question before a broad rollout. The method is to test complete support outcomes, introduce controlled complications, score failures separately from conversational polish, and use the result to define a limited production pilot.

    Evaluate the support outcome, not the performance

    A natural voice can create an impression of competence before the agent has done anything useful. Pleasant pacing, expressive speech, and a quick opening matter, but they cannot compensate for retrieving the wrong account, misunderstanding the request, or claiming that an action succeeded when it did not.

    Treat the unit of evaluation as a completed support job. Depending on the intent, that job may require the agent to identify the caller, understand the request, retrieve the right information, explain the answer, perform an authorized action, confirm the resulting state, and send a follow-up or transfer the conversation. If you score only the spoken answer, you leave most of the product untested.

    One live Fin Voice call illustrated this end-to-end standard in about 90 seconds: the agent verified identity, retrieved account information, managed an interruption, presented options, completed a workflow, and sent a follow-up email. That sequence is a useful model for constructing a test. It is not, by itself, proof of reliability across other calls.

    Before anyone places a test call, write an outcome contract for each scenario:

    • Caller goal: What is the person trying to accomplish?
    • Starting state: What customer, account, order, subscription, or case data exists before the call?
    • Available evidence: Which knowledge, policies, and records may the agent use?
    • Permitted actions: What may the agent change, create, send, cancel, or escalate?
    • Required clarification: Which missing or conflicting facts must be resolved before an answer or action?
    • Completion evidence: What observable state proves that the request was resolved?
    • Unacceptable outcome: What error would make the call a failure even if the conversation sounded good?

    This contract prevents a common scoring mistake: confusing non-transfer with resolution. A call can remain inside the AI channel and still leave the customer with a wrong answer, an incomplete action, or no idea what happens next. Conversely, an intentional transfer can be the correct resolution when the agent reaches a policy, permission, or confidence boundary.

    Build scenarios around the ways real calls become difficult

    Start with support intents your operation actually receives. Prioritize intents that are frequent, expensive to handle, important to customer trust, or dependent on multiple systems. Do not begin with trivia questions that merely demonstrate broad language-model knowledge. You are evaluating support execution.

    For every core intent, create a straightforward case and several controlled variants. Keep the customer objective constant while changing one condition at a time. That makes a failure diagnosable instead of merely disappointing.

    A practical scenario matrix

    • Clean path: The caller gives the relevant facts in a clear order. This establishes whether the basic workflow works at all.
    • Missing information: Omit a detail the agent needs. Check whether it asks a focused question instead of guessing or restarting the intake.
    • Ambiguous intent: Use wording that could map to two support issues. The agent should disambiguate before retrieving data or taking action.
    • Mid-call correction: Let the caller change an account detail, date, product, or preferred option. Check whether the corrected fact replaces the old one throughout the workflow.
    • Interruption: Speak while the agent is answering. Observe whether it stops cleanly, understands the new input, and continues from the right point.
    • Backend delay: Introduce a slow retrieval or action. Evaluate how the agent manages the wait and whether it distinguishes a pending operation from a completed one.
    • Backend failure: Make a required system unavailable or return an error. The agent should not fabricate a result or promise completion it cannot verify.
    • Policy boundary: Ask for something the agent is not allowed to do. Test the explanation, alternatives, and escalation path.
    • Human request: Ask directly for a person. Verify that the agent follows the configured policy without turning the handoff into an argument.
    • Listening conditions: If your deployment must support different languages, accents, devices, or noisy environments, test each condition explicitly rather than treating one clear studio call as representative.

    Give testers the goal, account state, and one complication. Do not script every sentence. A fully written dialogue tests whether the agent can follow the dialogue you anticipated; a goal-based scenario tests whether it can manage the conversation the caller actually creates.

    Keep a few variants undisclosed until the live session. This is not a trick. It prevents the evaluation from becoming a memorized path while still keeping every test fair and reproducible. Record the exact variant afterward so another evaluator can run it again.

    Run the call through the systems you expect to deploy

    An unedited live call is more informative than a produced recording, but live alone is not enough. A live test can still use ideal data, a simplified integration, a practiced caller, and a workflow that avoids the hard parts of your environment.

    Ask to run the scenario through a path that resembles the intended deployment:

    1. Place a normal phone call through the proposed telephony route. If production will use call forwarding, test the forwarding path rather than a direct internal endpoint.
    2. Use a safe test account containing representative records, permissions, and history.
    3. Require the agent to retrieve data from the backend system that will be authoritative in production.
    4. Introduce the chosen interruption, correction, ambiguity, delay, or error during the live conversation.
    5. Require a real test action where it is safe to do so, not a verbal description of what the agent would have done.
    6. Inspect the backend state after the call. Confirm that the correct record changed once, with the expected values.
    7. Verify every promised follow-up, case creation, notification, or handoff outside the voice channel.
    8. Retain the recording, transcript, timestamps, tool activity, and final system state for scoring.

    This is especially important when an agent can take consequential actions. A fluent confirmation is not evidence that the action happened. The system of record is the evidence.

    Repeat important scenarios with different wording and a different caller. One successful run demonstrates that the capability can work. Repeated variants reveal whether the capability depends on a narrow phrase, a rehearsed cadence, or an unusually forgiving path.

    Key takeaways

    • Score complete resolution, including backend state and follow-up, rather than voice quality alone.
    • Change one condition at a time so you can identify why a call failed.
    • Test interruptions, corrections, ambiguity, system delays, system errors, and escalation.
    • Measure different kinds of waiting separately; a lookup pause and a turn-detection problem are not the same defect.
    • Treat a successful demo as evidence for a pilot, not permission for an unrestricted rollout.

    Score conversation, reasoning, and operational closure separately

    A single overall rating hides the information you need to make a product decision. The call may sound awkward but reach the correct outcome, or sound excellent while making a dangerous mistake. Separate the evaluation into three layers.

    LayerWhat to inspectEvidence of a passTypical failure
    Conversation mechanicsTurn detection, interruption handling, pacing, response length, and intelligibilityThe caller can speak naturally, correct the agent, and follow the response without fighting for the floorThe agent talks over the caller, leaves confusing silence, or delivers answers too long to retain by ear
    Decision qualityIntent recognition, clarification, use of account context, policy application, and answer accuracyThe agent asks only for missing information, uses the correct evidence, and avoids unsupported conclusionsThe agent guesses, asks redundant questions, ignores a correction, or applies the wrong policy
    Operational closureIdentity checks, tool calls, state changes, confirmation, follow-up, and escalationThe verified backend state matches the caller’s request and the agent’s final explanationThe agent claims success without a completed action, changes the wrong record, duplicates work, or drops context during handoff

    Use a simple 0-2 score for each criterion: 0 for failed or unsupported, 1 for completed with material caller effort or recovery, and 2 for correct and usable. The scale is deliberately small. Evaluators can usually distinguish failure, friction, and success more consistently than they can defend the difference between seven and eight on a ten-point scale.

    Do not average away critical errors. A wrong account action, failed identity control, fabricated completion, or forbidden disclosure should remain visible as a release blocker even if many low-risk calls receive high scores. Record both the criterion scores and the count of critical failures.

    Break latency into moments the caller can feel

    Latency is not one number. Capture at least three moments: the time the agent takes to recognize that the caller has finished, the time it spends reasoning or waiting for a system, and the time needed to begin and complete the spoken response.

    • End-of-turn delay: A long delay after every caller turn makes the exchange feel unresponsive and can encourage both sides to start speaking at once.
    • Reasoning or retrieval delay: A pause can be appropriate when the agent is checking account data or invoking a backend workflow. Brief pauses were audible during live subscription and backend checks, which is more informative than editing those waits out.
    • Response delivery: A fast start does not help if the answer becomes a long monologue. Voice responses need structure and pacing that work for listening, not merely text that sounds acceptable when read.

    Ask what is happening during a pause. If the system is doing useful work, the next statement should reflect that work and the action log should verify it. If the pause is long enough to make a caller wonder whether the call has dropped, the experience needs an appropriate progress cue. If the agent answers instantly but guesses, speed is concealing a quality problem.

    Review individual timings as well as an average. A generally responsive agent with occasional severe stalls creates a different operational problem from one that is consistently a little slow. Your test recordings and timestamps should make both patterns visible without inventing a universal pass threshold that ignores the complexity of the workflow.

    Make recovery and escalation part of the product test

    The strongest voice experiences are not the ones that never encounter confusion. They are the ones that recover without making the caller restart. Recovery is therefore a capability to test, not an embarrassing exception to hide.

    Interrupt the agent in the middle of an answer. Correct a fact it has already used. Add a second request after the first appears resolved. Say that an explanation was unclear. Ask for a human. These moves reveal whether the agent maintains conversational state or merely produces plausible turns one at a time.

    During recovery, look for specific behavior:

    • It stops speaking promptly when the caller takes the turn.
    • It identifies what changed instead of repeating the whole interaction.
    • It replaces corrected information rather than carrying both versions forward.
    • It asks a narrow clarification when the next action is uncertain.
    • It does not claim to understand when the transcript or subsequent action shows otherwise.
    • It preserves verified context and the reason for contact when a human takes over.
    • It tells the caller what will happen next instead of ending on an internal routing label.

    Tone belongs in this test, but not as a beauty contest between synthetic voices. Evaluate whether pacing, brevity, acknowledgement, and word choice suit the moment. A caller correcting a billing detail needs a clear acknowledgement and an accurate update, not theatrical empathy. A caller who sounds uncertain may need a shorter explanation and a confirming question. Tone is the behavior of the conversation, not just the timbre selected in a settings menu.

    Escalation should also count as a valid outcome when it is timely and informed. Define which conditions require a handoff, which allow one, and what context must travel with it. Then test the handoff from the caller’s side. If the customer reaches a person but has to repeat identity, intent, and every attempted step, the routing technically worked while the support experience failed.

    Turn the evaluation into a controlled pilot decision

    A strong live evaluation earns the right to run a pilot. It does not justify sending every eligible call to the agent. Production introduces variation in callers, data quality, traffic, integrations, and issue combinations that a demonstration cannot reproduce fully.

    I would require five gates before approving even a limited external pilot:

    1. Capability gate: Every must-have intent has completed its end-to-end workflow, including at least one controlled complication.
    2. Critical-risk gate: No unresolved failure can expose the wrong account, bypass a required check, perform an unauthorized action, or report a false completion.
    3. Conversation gate: The agent can handle interruptions, corrections, clarification, and explicit human requests without trapping the caller in a loop.
    4. Operations gate: Your team can configure terminology, guidance, escalation behavior, greetings, voice, and deployment controls for the intended support environment.
    5. Learning gate: Owners can inspect recordings, transcripts, tool activity, outcomes, and failures, then change the knowledge, workflow, policy, or conversation design responsible.

    Start the pilot with a reversible slice of traffic and a clear human fallback. Select intents whose correct outcome can be verified in your systems. Define who reviews failed and escalated calls, who can pause the rollout, and who owns each class of fix. An answer-quality issue, a telephony issue, and a backend integration issue require different owners even when the caller experiences all three as one bad call.

    Expand only when observed calls meet the outcome contracts you wrote before the demo. If the definition of success keeps changing after failures appear, the evaluation is no longer protecting the decision.

    For your next vendor session, replace “show me your best call” with a scenario pack, a test account, and a request to inspect the final system state. You will learn more from one imperfect call that recovers correctly than from a flawless recording that never had to recover at all.

    References

  • Global Invoicing Nightmares: Hard-Won Product Lessons on EU Tax, Compliance, and Customer Value

    Global Invoicing Nightmares: Hard-Won Product Lessons on EU Tax, Compliance, and Customer Value

    I hit play on Global Invoicing – All Things Product Podcast with Teresa Torres & Petra Wille and felt an immediate jolt of recognition. We’ve all launched a feature that looked solid—until a small, overlooked detail broke everything. Their stories about global invoicing and taxes echoed challenges I’ve faced leading product for international customers: if you don’t design for the last mile of compliance, you can accidentally block the very "moment of value creation" your product promises.

    Listen to this episode on: Spotify | Apple Podcasts

    The conversation starts as a candid rant about EU tax compliance and quickly becomes a precise product management lesson: when we fail to map the entire path to customer value—down to the tiniest regulatory requirement—we can ship something “done” that still doesn’t work in the real world. That gap between intention and outcome is where good product teams live or die.

    In my experience, the nightmare of global invoicing for small online businesses is very real. Even big platforms (like Squarespace and Teachable) miss the mark on EU tax compliance, and when they do, customers feel it immediately. It’s the kind of edge case that doesn’t show up in a demo but absolutely shows up in revenue. Or as Teresa put it, “It’s not a little detail when your client won’t pay the invoice.” — Teresa Torres

    I appreciated how the episode digs into the difference between passing a regulatory checklist and actually meeting customer needs. Put plainly: the product isn’t “done” when the ticket moves to Done; it’s done when the customer completes the job—receives an acceptable invoice, pays successfully, and can reconcile it without friction. That’s why I lean hard on story mapping for regulatory work; it exposes the invisible steps where value creation can silently fail.

    Here’s how the episode resonates with my own playbook: the nightmare of global invoicing for small online businesses is a systems problem; why even big platforms (like Squarespace and Teachable) miss the mark on EU tax compliance is a prioritization and discovery problem; how Petra and Teresa navigated invoicing across borders with Ableify and LearnWorlds highlights pragmatic tool choices and trade-offs; the key difference between meeting regulations and meeting customer needs is an outcomes-over-output mindset; what product teams can learn from regulatory edge cases is how to find the seams where markets, laws, and workflows collide; how missing a single detail can block the "moment of value creation" is a reminder that value is defined by customers; and why story mapping is critical for finding gaps between "we shipped it" and "customers got value" is the method that connects all of the above.

    Practically, that means I treat regulatory features like any other high-stakes product surface: do real product discovery with affected users; co-design the happy path and the ugly edge cases; write acceptance criteria that include jurisdictional and document-level specifics (e.g., VAT numbers, invoice formats, timing rules); align with finance and legal early; and instrument the journey from invoice issued to invoice paid so we can see where real customers get stuck. This is outcomes vs output OKRs in action, and it’s one of the fastest ways to earn trust with stakeholders.

    Key takeaways worth bookmarking: Customers define value, not your compliance checklist. Regulatory work still requires discovery—you can’t skip understanding user needs. The path to value doesn’t end when your feature works; it ends when your customer succeeds. “Sweating the details” isn’t micromanagement—it’s good product management.

    Memorable quotes to bring back to your team: “If you don’t sweat the details, people choose other platforms.” — Petra Wille. “It’s not a little detail when your client won’t pay the invoice.” — Teresa Torres.

    Follow Teresa Torres: https://ProductTalk.org | Follow Petra Wille: https://Petra-Wille.com

    Mentioned in the episode: Squarespace | Stripe | Product at Heart | Teachable | LearnWorlds | Ablefy | Become a Better Product Leader: A 52-Week Transformation Journey | Product Talk Academy

    Have thoughts on this episode? Leave a comment below.

    Full transcripts are only available for paid subscribers.


    Inspired by this post on Product Talk.


    Book a consult png image
  • From Sketch to Clickable Demo: My AI Prototyping Playbook to Build Apps in Hours

    From Sketch to Clickable Demo: My AI Prototyping Playbook to Build Apps in Hours

    I’ve spent much of my career compressing the distance between a napkin sketch and something real customers can touch. At HighLevel, my product teams use generative AI to validate ideas faster, reduce risk earlier, and win stakeholder trust with evidence instead of slides. The goal isn’t to be flashy—it’s to be precise, testable, and repeatable.

    Today, you can build it before you pitch it. AI prototyping can turn ideas into clickable demos in hours. Here are some tools to try and steps to follow.

    I start every AI prototyping sprint by sharpening the problem statement and the outcome we care about. That means being explicit about the target user, jobs-to-be-done, and the riskiest assumptions. I define a minimum detectable effect (MDE) and tie it to outcomes vs output OKRs so everyone aligns on what “good” looks like before we touch a tool.

    From there, I move from sketch to interface. I capture a rough flow (whiteboard, tablet, or even paper) and generate UI variations with my AI product toolbox—tools that translate structure into components and screens. I’ll iterate on information hierarchy and copy until the narrative supports the core job, borrowing techniques from UX writing. For product managers leaning into LLMs for product managers, this phase is about speed to feedback, not perfection.

    Next, I wire data and logic. I connect a lightweight backend or spreadsheet, stitch in a CRM integration if needed, and add LLM calls through a ChatGPT connector or Claude Code. If the concept benefits from multi-step autonomy, I introduce agentic AI to orchestrate tasks across APIs. CustomGPT workflows help me encapsulate business rules so the demo behaves consistently in user paths we care about.

    Governance is not optional at this stage. I apply privacy-by-design defaults, document data governance decisions, and run a quick AI risk management pass: input validation, prompt safety, rate limits, and fallback responses. This keeps the prototype credible and prevents false positives from polluting stakeholder perception.

    With a click-through in hand, I instrument the experience so learning compounds. I drop in Amplitude analytics to track activation, task completion, and drop-off, and set up simple A/B testing when there’s a meaningful design or copy choice. This makes the prototype a learning vehicle, not just a demo.

    Then I get it in front of users—fast. Five targeted conversations will beat fifty internal opinions. I run structured product discovery interviews, observe time-to-value, and capture objections. This is where empowered product teams shine: we make changes in real time, re-run the flow, and document what moves the needle for product-led growth.

    When speed matters, I use a four-hour cadence: Hour 1 for problem framing and MDE; Hour 2 for sketch-to-UI generation; Hour 3 for data wiring and AI logic; Hour 4 for instrumentation and user walkthroughs. By the end, we have a clickable demo, preliminary analytics, and a clear decision on whether to advance, pivot, or park.

    Finally, I translate insights into a concise artifact: the hypothesis we tested, the signal we observed, the trade-offs we made, and the next sprint plan for product roadmapping and sprint planning. The point is not to be right on the first try; it’s to learn precisely, cheaply, and quickly enough to invest with conviction.

    If you adopt this approach, you’ll find that stakeholder management becomes easier, team energy rises, and your roadmap earns credibility. Build it before you pitch it, and let real interactions—not wishful thinking—do the heavy lifting.


    Inspired by this post on Product School.


    Book a consult png image
  • Cut Time to Value, Boost Retention: My Proven Playbook for Activation, Growth, and Loyalty

    Cut Time to Value, Boost Retention: My Proven Playbook for Activation, Growth, and Loyalty

    Time to value is the most reliable early indicator of long-term user retention I know. When customers experience meaningful product impact fast, they stick around, expand, advocate, and cost less to support. Over the years leading product teams, I’ve learned that speed-to-impact isn’t a nice-to-have—it’s the engine behind sustainable product-led growth and efficient go-to-market.

    Accelerate retention by reducing time to value. Learn how faster product impact drives growth, reduces costs, and keeps users engaged in the long term.

    Practically, I define time to value as the duration from first touch (or first login) to the moment a user achieves their “aha” outcome—something tangibly useful aligned to their job-to-be-done. The shorter that journey, the higher the likelihood of user activation, trial conversion, and durable engagement. This is why I obsess over onboarding, in-app guides, product tours, and the clarity of our value proposition.

    My first move is to map the Minimum Path to Value (MPV): the smallest set of actions needed to deliver a real result for a new user. I strip away everything non-essential in that path—fields, clicks, choices, and jargon. Opinionated defaults, smart templates, sample data, and single-player workflows let customers succeed in minutes, not days. The goal is to reduce cognitive load while making the next best action unmistakably clear.

    Instrumentation turns TTV from a hunch into a system. I track activation events, cohort retention, and conversion using platforms like Amplitude analytics and Pendo, with timely nudges through Intercom when users stall. I look at the distribution of TTV (not just the average), correlate it with retention analysis, and set explicit targets such as “new users reach first value within 10 minutes.” Those targets become team-level outcomes—not outputs—and we review them weekly.

    Experimentation is how we iterate toward the fastest path to value. I rely on A/B testing to compare onboarding flows, progressive profiling to delay non-critical inputs, and opinionated setup wizards to remove guesswork. Auto-generated example projects, pre-configured integrations, and guided checklists accelerate user activation without sacrificing flexibility for advanced users.

    Content and guidance matter as much as UX. Tooltips, contextual in-app guides, and short product tours should be timely, skippable, and laser-focused on the outcome, not the feature. I pair these with a concise knowledge base and short explainer videos that reinforce the same value narrative a user sees inside the product.

    Cross-functional alignment is essential. Product, marketing, sales, and customer success must rally around the same activation metric and TTV target. That alignment ensures our trial messaging, onboarding emails, and CS playbooks don’t compete—they compound. When everyone points to the same first-value moment, friction drops and adoption rises.

    Pricing and packaging can also accelerate time to value. Free trials should be long enough for users to credibly reach first value; usage-based gates should never block the MPV. I prefer to unlock everything needed to hit the “aha” moment, then meter after the value is viscerally felt—this respects the user’s time and reinforces trust.

    There’s a cost story, too. Faster time to value reduces tickets, shortens onboarding cycles, and lowers cost-to-serve. It also clarifies product discovery: when we see where users stall, we don’t guess at roadmap priorities—we let the data guide our next bet.

    In my experience at HighLevel, I’ve repeatedly seen activation rates jump when we cut time to value from days to minutes. The specific tactics vary by product, but the pattern holds: when the first outcome is undeniable and fast, retention follows—and so does efficient growth.

    If you’re looking for a starting point, try this: define one activation event that clearly signals value, instrument it end-to-end, design a Minimum Path to Value that gets new users there in under 10 minutes, and run weekly experiments until you consistently hit the target. Do that, and you won’t just improve onboarding—you’ll build a product that earns loyalty from the very first session.


    Inspired by this post on Amplitude – Best Practices.


    Book a consult png image
  • Win AI Search: Proven Playbook to Get Your Startup Recommended by ChatGPT & Perplexity

    Win AI Search: Proven Playbook to Get Your Startup Recommended by ChatGPT & Perplexity

    AI search is quickly becoming the new homepage for startups. When a buyer asks a model for the best tools, they often take the short list at face value. I treat this moment as a product surface I can influence with strategy, content, structure, and distribution—much like any other go-to-market channel.

    Early on, I set a simple objective for my team and me: "Learn how LLMs like ChatGPT and Perplexity decide which startups to recommend and what signals help a brand get discovered in AI search." That sentence became our north star for experiments, instrumentation, and content architecture.

    Here is the mental model that consistently holds up in practice. Large language models synthesize answers from a knowledge graph built from crawled content, citations, and high-signal sources. They weight consensus, clarity, recency, authority, and machine-readability. I don’t pretend to know the internals, but across hundreds of tests, the same patterns correlate with being surfaced and cited.

    First, I make our entity unambiguous. I standardize the company name, product names, and leadership bios across the site and external profiles. I implement Organization and Product markup with schema.org and link out with sameAs to authoritative profiles like LinkedIn, Crunchbase, GitHub, and key directory listings. The goal is to collapse ambiguity so AI search knows exactly who we are and which claims are attributable to us.

    Next, I publish definitive, answer-first pages. For every core query—what we do, who it’s for, outcomes, differentiators, pricing, comparisons, and integrations—I ship a page that leads with a crisp summary, then supports it with evidence, examples, and plain language. I include Q&A sections, realistic use cases, and named case studies so models can quote and ground responses in verifiable facts.

    I then make the site maximally machine-readable. I add schema.org for SoftwareApplication, Product, FAQPage, and HowTo where relevant. I keep titles, H1/H2 structure, internal links, and metadata descriptive and consistent. I expose last-modified dates, maintain an XML sitemap, and keep a visible changelog and release notes. Freshness matters—Perplexity, in particular, tends to privilege recent, well-cited material when answering time-sensitive questions.

    Citations are non-negotiable. I earn credible mentions on third-party properties, analyst lists, comparison pages, and customer reviews. I prioritize authoritative placements over volume, then make sure our site references those sources to reinforce the signal. When Perplexity cites our page alongside a respected third-party review, our inclusion rate in answers rises noticeably.

    I also design for developers, buyers, and machines at once. That means clean docs, integration pages, and transparent security and trust content. Clear API references, integration guides, and reliability notes give models concrete artifacts to summarize. Pricing, privacy, and support policies reduce uncertainty and increase the likelihood that an answer will include us.

    Measurement turns this from a hunch into a system. I run controlled content experiments, track minimum detectable effect on discovery and mentions, and instrument referral patterns from AI assistants when citations appear. I monitor which prompts surface our brand, which sources are cited, and which pages are repeatedly used as references. When we move a KPI, we codify the pattern into our playbook and scale it.

    Trust is the compounding advantage. I maintain a transparent trust center, privacy-by-design posture, and clear data governance practices. I remove vague claims, back up benefits with evidence, and keep all performance or security statements auditable. Models tend to lift brands that feel low-risk, well-documented, and widely corroborated.

    If you want a fast start, here’s the checklist I rely on. Standardize your entity and ship schema.org. Publish answer-first pages for core jobs-to-be-done, comparisons, and integrations. Earn authoritative third-party citations and reference them. Keep release notes, changelogs, and dates current. Instrument AI discovery and iterate based on what gets cited. Do this consistently, and your startup earns a fair shot at being recommended when buyers ask AI for the best options.


    Inspired by this post on Amplitude – Best Practices.


    Book a consult png image
  • Prototypes vs Products: How I De-risk Ideas Fast and Ship Reliable Value at Scale

    Prototypes vs Products: How I De-risk Ideas Fast and Ship Reliable Value at Scale

    Note: This is part of the product creator series of articles, based on the overview article, The Era of the Product Creator. This series is for anyone who wants to create a successful product—whether or not you’ve had formal training or experience in product management, product design, or engineering. Over the years, I’ve watched smart teams stumble because they treated a prototype like a product. The distinction is simple but vital: prototypes exist to learn; products exist to earn trust by delivering value reliably at scale. When we blur that line, we ship avoidable risk to customers and slow ourselves down later with rework. When I build a prototype, I’m testing assumptions as quickly and cheaply as possible. It might be a clickable Figma mock, a Wizard‑of‑Oz demo, or a quick script stitching together a ChatGPT connector with a CustomGPT workflow. It’s intentionally disposable. I expect missing edge cases, fake data, hand‑waving on latency, and limited attention to security or privacy. The only goal is to answer the riskiest questions fast. A product is a promise. It’s hardened for reliability, performance, security, and privacy‑by‑design. It’s observable with real analytics, supports CI/CD and rollback, meets accessibility guidelines, and can be maintained by empowered product teams. It has clear SLAs, incident management runbooks, and instrumentation that lets me track outcomes vs output OKRs and DORA metrics. Keeping prototypes and products separate makes us faster and safer. Prototypes accelerate discovery; products operationalize value. If I catch myself “polishing” a prototype, I pause and either discard it or define the path to production with the right engineering rigor, data governance, and stakeholder management. Here’s how I decide. In prototype mode, I timebox learning to days, not weeks, and focus on a single risky assumption—value, usability, or feasibility. I validate through qualitative research and usability tests, not vanity metrics. To graduate to product work, I require a crisp problem statement, evidence of problem‑solution fit, a technical plan for scale and observability, a privacy and threat modeling review, and a measurement plan (including minimum detectable effect) for upcoming A/B testing. AI adds new wrinkles. For gen AI and agentic AI, I evaluate model behavior offline before exposing anything to customers. That includes prompt design, context window management, guardrails to minimize hallucinations, and clear fallback strategies. I define red‑team scenarios, logging for auditability, and policies for data retention and encryption as part of AI risk management. A recent example: we prototyped an agent workflow in a day that felt magical in demos. We resisted the urge to ship. Instead, we added authentication, rate limiting, PII redaction, human‑in‑the‑loop review, observability, and in‑app guides and product tours for onboarding. Only then did we move to a limited release with a well‑defined go‑to‑market strategy and support readiness. One more trap to avoid: calling a prototype an MVP. An MVP is still a product—minimal in scope but complete enough to deliver value, gather trustworthy data, and support customers. If you wouldn’t put your name on it or support it in production, it’s a prototype, not an MVP. If you’re a product creator, align your product trios around this discipline. Use prototypes to learn quickly in discovery, and use products to deliver outcomes in delivery. That mindset protects customer trust, speeds iteration, and moves you toward product‑market fit with far less waste.

    Inspired by this post on SVPG.


    Book a consult png image
  • AI Context Engineering: A System for Product Decisions

    AI Context Engineering: A System for Product Decisions

    You give an LLM your discovery notes, a dashboard export, and a roadmap question. It returns polished recommendations in seconds. The recommendations sound plausible, yet your product trio still cannot tell which option deserves a commitment.

    The missing ingredient is usually not a better prompt. It is a decision-ready context system: a controlled way to give AI the evidence, boundaries, and outcome definition required to reason about the same product decision your team is actually making. Done well, this gives you more than a convincing answer. It gives you a traceable choice, explicit uncertainty, and a validation plan.

    Define the decision before you collect the context

    For product work, context engineering is the deliberate design of everything an AI system can use at the moment it reasons: customer evidence, metrics, goals, constraints, definitions, instructions, and prior decisions. The useful unit is not a prompt or a document. It is the decision.

    This distinction matters because an LLM can answer an underspecified request without exposing that the request was underspecified. Ask it to improve onboarding, and it can produce a credible list of patterns. That output still does not tell you which user segment matters, what improvement means, which current friction is supported by evidence, or what downside the team must avoid.

    Before pulling any context, write a decision frame that answers these questions:

    • What decision must be made? Name the commitment, not the general topic. Choose whether to change a specific onboarding step is a decision; explore onboarding is not.
    • Who is the decision for? Identify the customer segment, use case, or part of the journey. Evidence from one segment should not silently become a claim about every user.
    • What outcome should change? State the behavior or business result you want, then identify the guardrail signals that should not deteriorate.
    • What can constrain the answer? Include privacy, risk, brand, commercial, technical, and operational boundaries before ideation begins.
    • What evidence could change the choice? If no possible evidence would change the decision, you are asking AI to justify a conclusion rather than help make one.
    • What must the output enable? Specify whether you need options, a recommendation, a decision memo, an experiment plan, or a list of unresolved questions.

    Anchor this frame in outcomes rather than deliverables. Improve activation for a defined segment while protecting support load establishes a decision boundary. Build a new onboarding checklist merely names output. The first lets AI compare interventions; the second encourages it to decorate a predetermined solution.

    A practical test is to remove the proposed feature from the frame. If the decision still makes sense, you have probably described an outcome. If the frame collapses, the team may already be committed to an output.

    Build a context packet that preserves evidence quality

    A context packet is the smallest governed collection of information that allows the model and the product team to reason about the decision. It can combine customer quotes, behavioral trends, funnel friction, support conversations, and commercial constraints. The important work is to assemble, structure, compress, and challenge that evidence before asking for recommendations.

    Do not treat every input as the same kind of truth. A customer quote gives you detail about an experience, not its prevalence. Usage analytics show behavior, not necessarily motivation. Support conversations overrepresent people who contacted support. CRM data can expose commercial constraints without proving that a feature creates customer value. Labeling these boundaries prevents the model from blending different signals into false certainty.

    Use this structure for the packet:

    • Decision header: the choice, decision owner, affected segment, and action that follows the decision.
    • Outcome frame: the desired outcome, current signal, primary measurement, guardrails, and any metric definitions needed to interpret the data correctly.
    • Evidence ledger: each relevant observation with its origin, segment, time period, and scope. Keep direct observations separate from interpretations.
    • Constraints: technical dependencies, commercial commitments, privacy rules, brand boundaries, operational capacity, and known risks.
    • Contradiction register: evidence that points in different directions, including differences between customer statements and observed behavior.
    • Unknowns: missing evidence, ambiguous definitions, unrepresented segments, and assumptions the team has not validated.
    • Output contract: the form of response you need, the criteria options must address, and the unsupported claims the model must label rather than fill in.

    Compression is where many context packets either become useful or become misleading. The goal is not merely to shorten the material. It is to increase the proportion of decision-relevant signal without erasing qualifications.

    1. Normalize repeated evidence. Deduplicate copied notes and repeated tickets so repetition in the packet does not impersonate independent confirmation. Preserve any real frequency data separately.
    2. Retain the qualifiers. Do not compress away the segment, time range, denominator, metric definition, or product state that determines what an observation means.
    3. Label epistemic status. Mark material as observation, interpretation, assumption, or generated hypothesis. A concise packet should make these distinctions clearer, not blur them.
    4. Keep contradictions visible. If interviews describe one problem while behavioral data points elsewhere, preserve both signals and ask what evidence would resolve the conflict.
    5. Remove inert context. My rule is simple: if an item cannot change an option, a risk assessment, or the validation plan, it does not belong in the active packet. Keep it available outside the model context if the team may need to inspect it later.

    Apply privacy-by-design while assembling the packet, not after the model has processed it. Customer transcripts, CRM records, and support conversations can contain personal or confidential data. Use approved systems, follow applicable access controls and data terms, redact identifiers, and aggregate where the decision does not require record-level detail. If you cannot establish that the data is permitted in the AI workflow, leave it out and provide a safe summary. The downside is not a weaker prompt; it is potential exposure of customer or company information.

    Separate synthesis, strategy, and skepticism

    Asking for a summary, a recommendation, and a critique in the same instruction makes it difficult to see where evidence ends and invention begins. A stronger agentic workflow separates those jobs into distinct passes: Summarizer, Strategist, and Skeptic.

    The Summarizer creates an evidence map

    The Summarizer should organize the packet without deciding what to build. Ask it to group evidence around the decision, preserve relevant qualifiers, expose conflicts, and identify missing information. Explicitly prohibit recommendations during this pass.

    A useful Summarizer output contains the supported observations, the segments represented, the outcome signals involved, the contradictions, and the unknowns. Review this output against the packet before continuing. If the model has turned an assumption into a fact, fix the evidence map rather than hoping a later pass corrects it.

    The Strategist develops decision options

    Give the Strategist the approved evidence map, the original decision frame, and the constraints. Ask for a small, meaningfully different set of options, including the option to leave the product unchanged when that is legitimate.

    Require the same fields for every option:

    • the customer problem or opportunity it addresses;
    • the packet evidence that supports it;
    • the assumptions required for it to work;
    • the expected outcome and guardrail signals;
    • the dependencies and material trade-offs;
    • the simplest valid way to reduce its largest uncertainty.

    This format prevents one option from winning because it received a more persuasive narrative. It also makes unsupported leaps visible. If the model cannot connect an option to evidence, that option can remain an idea, but it must be labeled as a hypothesis rather than presented as a conclusion.

    The Skeptic tries to disconfirm the options

    The Skeptic should not produce generic risks. Ask it to find the strongest contrary evidence, the segment that might be harmed, the constraint most likely to invalidate the option, the metric that could be gamed, and the observation that would show the underlying hypothesis is wrong.

    Require it to distinguish counterevidence already present in the packet from new conjecture. This matters because a skeptical tone can sound rigorous even when it is unsupported.

    The same LLM can perform all three roles, but role prompts do not create independent evidence or independent reviewers. Freeze the context packet used for the loop, label every generated artifact, and keep generated claims out of the evidence ledger until a human verifies them. Role separation is a workflow control, not a guarantee of correctness.

    Stop adding passes when the workflow is only rearranging language. The loop has done its job when the team can see the supported facts, viable options, disputed assumptions, material risks, and next evidence needed to decide.

    Make the product trio the decision gate

    AI can accelerate the reasoning, but it should not become the decision owner. Bring the packet and the three-pass output into a product trio of product, design, and engineering. The purpose of that forum is not to approve the AI recommendation. It is to make the trade-offs explicit and decide what the team is prepared to learn.

    1. Verify the evidence boundary. Check whether the represented segments, product states, and metrics match the decision. Ask which customer or operational perspective is absent.
    2. Classify the important claims. Mark each claim as supported observation, team interpretation, assumption, or generated hypothesis. If nobody can trace a recommendation back to the packet, treat it as a hypothesis or remove it.
    3. Compare trade-offs on equal terms. Evaluate every option against the desired outcome, guardrails, constraints, dependencies, and learning value. Do not let the most detailed option appear strongest merely because the model wrote more about it.
    4. Choose the next commitment. The valid outcomes are to proceed, run a discovery or validation step, defer the decision, or reject the options. Assign a human owner and make clear what action the decision authorizes.
    5. Record the rationale. Convert the discussion into a concise decision memo rather than forwarding raw model output to stakeholders.

    The decision memo should include:

    • the decision and why it is being made now;
    • the target segment, desired outcome, and guardrails;
    • the evidence that carried the most weight;
    • the chosen option and the alternatives rejected;
    • the trade-offs accepted by the decision owner;
    • the assumptions and unresolved questions;
    • the validation method and disconfirming signal;
    • the owner and trigger for revisiting the decision.

    This gives stakeholders something stronger than AI-generated confidence. They can inspect what the choice rests on, where judgment entered, what could prove the team wrong, and when the decision should be reconsidered.

    Close the loop with validation and decision memory

    Even a well-grounded model output is not product validation. It is a structured hypothesis. Match the validation method to the claim and to the consequence of being wrong.

    • For a causal behavior claim: use a controlled A/B test when traffic, instrumentation, and the product experience make that appropriate. Define the primary metric, minimum detectable effect, guardrails, analysis approach, and stopping rules before reading the result.
    • For a usability or comprehension claim: use targeted customer interviews or usability evaluation with the relevant segment. AI can help organize notes, but preserve outliers and do not turn a small qualitative sample into a prevalence claim.
    • For an operational claim: use a limited release with observability, support monitoring, and an explicit rollback condition. Watch the workflow around the feature, not only the feature interaction itself.
    • For privacy, brand, regulatory, or other high-consequence constraints: complete the appropriate human review before launch. A persuasive model assessment is not a substitute for the accountable specialist or decision owner.

    For an onboarding decision, for example, the packet may contain segment definitions, observed friction, support themes, and conversion signals. The workflow can propose alternative interventions and measurement plans. The trio still chooses which hypothesis deserves a controlled test, whether the minimum detectable effect is practical, and which activation or retention signals will determine the next move.

    After validation, return the result to the context system. Record what shipped, the observed outcome, affected segments, unexpected behavior, and which assumptions held or failed. Update the decision memo and evidence ledger. Otherwise, the next AI session begins from the same stale assumptions, and the organization pays again to relearn what it already discovered.

    That accumulated decision memory is one of the most valuable outputs of context engineering. It turns AI collaboration from isolated prompting into a feedback loop connecting discovery, strategy, execution, and measurable results.

    Key takeaways

    • Frame the product decision, target segment, outcome, and constraints before asking AI for options.
    • Give the model a compressed evidence packet, not an unstructured pile of documents.
    • Keep observations, interpretations, assumptions, and generated hypotheses visibly separate.
    • Use distinct Summarizer, Strategist, and Skeptic passes to expose where reasoning changes.
    • Let a human product trio own the trade-offs, commitment, and stakeholder rationale.
    • Treat every recommendation as a hypothesis until validation produces new evidence, then feed that evidence back into the decision record.

    Choose the next real product decision that is important enough to validate and bounded enough to act on. Write its decision frame, assemble the smallest safe context packet, run the three reasoning passes, and take a decision memo into your product trio. When the result flows back into the packet, context engineering stops being a prompting technique and becomes part of how you run product.

    References

  • Build a Company You’ll Run Forever: Bootstrapping vs VC, PMF, and the Art of ‘Eating Glass’

    Build a Company You’ll Run Forever: Bootstrapping vs VC, PMF, and the Art of ‘Eating Glass’

    I’ve spent my career building products and teams that I intend to steward for the long haul, and I’m drawn to founders who treat company-building as a craft you can practice forever. In this analysis, I break down a journey that crystallizes what it takes: going from a teenage wholesale hustle to an API-first healthcare clearinghouse, and in the process, learning why execution isn’t a moat, why venture capital is “going pro,” and how “eating glass” can become a durable advantage.

    Here’s the arc that anchored my thinking: a founder who, at 16, turned $2,500 into a wholesale empire; later bootstrapped a wildly profitable auto-parts business; then sold it to tackle “the most complicated problem” he’d ever encountered: business-to-business transaction exchange. He spent years building EDI infrastructure, threw away the entire codebase eight times, and found extraordinary traction in healthcare. The company recently raised a $70M Series B co-led by Stripe and Addition. The throughline is a consistent, high-agency approach to product management and go-to-market strategy, guided by first principles decision making.

    The first customer is often the trickiest—not because demand doesn’t exist, but because the product’s value proposition, points of parity, and competitive differentiation are still coalescing. I push teams to do founder-led GTM early, speak in the user’s language, and orchestrate high-signal conversations that expose real switching costs. That’s how we avoid mistaking polite interest for product-market fit.

    Bootstrapping forces rigor, but it also means being “constrained by capital.” There’s a ceiling to the speed at which you can iterate, validate, and scale. Venture capital, in the right context, is like “going pro”: you trade a bit of optionality for time, talent density, and a faster feedback loop. I often see confusion between ownership vs. control; structurally, you can design for alignment while still moving with the urgency a competitive market demands.

    One theme I return to with my own teams: execution is never actually a moat. Processes can be copied. Culture can be mimicked superficially. What can’t be easily replicated is the willingness to do the unglamorous, compounding work—what the founder here called “eating glass.” It’s the daily discipline of simplifying the system, instrumenting the edge cases, and standing up operational excellence that compounds into true competitive differentiation.

    When product-market fit hits in enterprise infrastructure, it can feel like “the snake swallowing a deer.” Capacity, process, and architecture are stretched to their limits all at once. I’ve experienced the same pattern: everything slows down so the organization can re-architect for scale. The trick is to make those constraints visible—measure service levels, queuing, and error budgets like you would in a production system—so you’re not flying blind.

    Some of the strongest product-management instincts I’ve seen borrow from discount retail and Toyota. From discount retail, we learn to obsess over unit economics, operational throughput, and ruthless simplification. From the Toyota production system, we adopt Kanban / TPS (Toyota), continuous improvement, and respect for constraints. In software terms, this becomes fast deployment frequency, small batch sizes, and defect prevention at the source—because “All software is a cascade of miracles.”

    Scaling decision-making is where most teams stall. I favor clear ownership, lightweight written narratives, and a bias for first principles decision making over committee compromise. That structure lets high-agency individuals move quickly while keeping cross-functional stakeholders aligned on outcomes vs output OKRs. It’s how you build empowered product teams without sacrificing focus.

    Hiring is where philosophy becomes practice. I resonate with the onboarding mantra “everything’s your fault now”—not as blame, but as an invitation to own outcomes end to end. I look for high-agency people who demonstrate systems thinking and the capacity to simplify. Manager hiring should lag role clarity; bring in managers when coordination overhead is the limiting factor, not when it merely feels uncomfortable.

    Longevity comes from founder-approach fit as much as product-market fit. Build a company you don’t want to leave by aligning operating cadence, decision rights, and cultural norms with how you actually work best. Maintain conviction in unconventional practice when the evidence supports it, while remembering that “Reality has a surprising amount of detail.” The more I zoom in on the real work—interfaces, edge cases, workflows—the more the right design emerges.

    In healthcare EDI, that realism matters. HIPAA overview (HHS) sets the compliance baseline. Payer integrations with Aetna, Blue Cross Blue Shield, and Cigna demand reliability and deep domain fidelity. Cloud and back-office ecosystems—from AWS and NetSuite to Slack, Microsoft Teams, Zapier, and Clay—shape the surrounding workflow. Lessons from Amazon, Target, Walmart, and Costco inform operational rigor; supply chain analogies from Ford Motor Company and GM clarify interface contracts. Porter’s five forces helps frame market structure; perspectives from Jeff Bezos and Peter Thiel sharpen strategic posture.

    If you’re building for the long run, here’s the blueprint I use with product leaders: validate painfully specific jobs-to-be-done before you scale; prefer founder-led GTM until messaging closes the intent-to-adoption gap; instrument throughput and quality like a production system; invest in people who treat ambiguity as a chance to lead; and don’t confuse speed with hurry. When the “snake swallowing a deer” moment arrives, re-architect deliberately, protect your margins, and let operational excellence carry you from product discovery to durable product-led growth.

    References and resources: Aetna: https://www.aetna.com/, Amazon: https://www.amazon.com/, AWS: https://aws.amazon.com/, Blue Cross Blue Shield: https://www.bcbs.com/, Change Healthcare: https://www.changehealthcare.com/, Cigna: https://www.cigna.com/, Clay: https://www.clay.com/, Costco: https://www.costco.com/, Ford Motor Company: https://www.ford.com/, GM: https://www.gm.com/, HIPAA overview (HHS): https://www.hhs.gov/hipaa/index.html, Jeff Bezos: https://x.com/JeffBezos, Kanban / TPS (Toyota): https://global.toyota/en/company/vision-and-philosophy/production-system, Microsoft Teams: https://www.microsoft.com/microsoft-teams, NetSuite: https://www.netsuite.com/, O’Reilly Auto Parts: https://www.oreillyauto.com/, Peter Thiel: https://x.com/peterthiel, Porter’s five forces: https://www.isc.hbs.edu/strategy/pages/the-five-forces.aspx, “Reality has a surprising amount of detail”: https://johnsalvatier.org/blog/2017/reality-has-a-surprising-amount-of-detail, Slack: https://slack.com/, Stedi: https://www.stedi.com/, Summit Racing: https://www.summitracing.com/, Target: https://www.target.com/, Walmart: https://www.walmart.com/, Zapier: https://zapier.com/


    Book a consult png image
  • Turn Claude Code Into a Trusted Teammate: My 3-Layer Memory System You Can Copy

    Turn Claude Code Into a Trusted Teammate: My 3-Layer Memory System You Can Copy

    "Can you critique the landing page for my new Story-Based Customer Interviews course?" That simple ask used to kick off hours of back-and-forth where I fed an AI the same context over and over—only to get generic feedback that wouldn’t land with my audience or fit my products. As a product leader, that inefficiency was unacceptable; as a writer, it was just plain frustrating.

    Not anymore. Today, Claude not only critiques my work, it helps me produce it. It generates marketing copy—in my voice. It helps me write blog posts. It knows what search terms are relevant to my business and helps me optimize my articles for SEO and now AEO. It helps me with competitive research, academic research, and discovery research. And it does all of this with little prompting from me.

    I don’t upload files to a web-based project. I don’t manage elaborate prompt libraries. I don’t repeat myself. I ask for help and Claude knows exactly what to do. The shift happened when I learned how to give Claude Code a memory. Claude now knows who my target customer is, the key value propositions I focus on, the specific opportunities each product addresses, my revenue model, my marketing channels, and so much more.

    Dark-mode slide with monospaced white text outlining an SEO plan: add CLAUDE.md to an AI glossary as the entry point, with bullets on article focus, audience, and search architecture for Give Claude Code a Memory.
    A dark-themed strategy slide for the post Stop Repeating Yourself: Give Claude Code a Memory, showing how to lead with a CLAUDE.md glossary page, write clearly for nontechnical readers, and link glossary and article to boost discovery and engagement.

    With that memory, I consistently get high-quality output tailored to my audience and aligned to my products and services. I don’t retype the same context; Claude just remembers. In this article, I’ll show you exactly how I set up that memory. It relies on Claude Code (which requires a Pro subscription), and it’s worth it. If you’re new to Claude Code, start with "Claude Code: What It Is, How It’s Different, and Why Non-Technical People Should Use It."

    Here’s the underlying problem: with large language models, every conversation starts from scratch. Yes, ChatGPT can remember some things and Claude can search past conversations, but practically speaking each new thread wipes the slate clean. If I were working on a new landing page, I’d normally need to upload target customer context, product details, primary and secondary value propositions, FAQ questions and answers, plus testimonials and logos for social proof—every single time.

    Dark-theme screenshot of the Claude interface with a large prompt field, model selector set to Sonnet 4.5, and quick-action buttons for Write, Learn, Code, Life stuff, and Claude’s choice on the home screen.
    Start fast with Claude’s home screen: Sonnet 4.5 is ready, and quick actions for writing, learning, and coding sit beneath a clean prompt box—ideal for showing how memory cuts repetition and streamlines daily development.

    Projects in web-based tools help a bit, but they introduce a new dilemma. When I move to the next landing page targeting the same customer but a different product and value proposition, do I start a new Project (tedious) or keep expanding the old one (which muddies the context window and degrades output quality)? The good news: Claude Code solves this by giving the model a precise, durable memory without overloading any single conversation.

    Claude Code can read files on my local machine, which is an understated superpower. I use those files to create a persistent, reusable memory that works across all chats and Projects. Files can be mixed and matched, so I give Claude exactly what it needs for the task at hand—and nothing more. For a first landing page, I reference the target customer and the relevant product; for the second, I reuse the same target customer file and point to the new product file.

    Screenshot of a macOS Notes window in dark mode showing an AI-assisted review of producttalk.org, listing Fetch and Read steps and a "Homepage Evaluation" for a first-time B2C visitor.
    Dark-mode Notes screenshot captures Claude Code in action: it fetches producttalk.org, reads context files, and delivers a concise homepage evaluation—showing how memory streamlines repeated analysis tasks.

    When you give an LLM the exact right context, output quality jumps. More context only helps if it’s the right context. For a landing page, Claude needs to know about the current product and perhaps related products for differentiation—but it doesn’t need to know about unrelated offerings. Structure your memory so Claude gets precisely what’s required.

    Once I did this, Claude shifted from “intern who needs handholding” to trusted advisor and capable teammate. It doesn’t guess at my value propositions—I’ve already told it. It writes in my voice because it has my writing guide and samples. It knows who owns which course and which use cases map to which features. The setup takes a bit of upfront work, but it compounds: update a file when something changes and you’re done. Most of this information already lives in your system; the trick is making it easy for Claude to use.

    Diagram of the Claude Code interface with a terminal-style dashboard. Arrows show Global Preferences (~/.claude/CLAUDE.md), Project Preferences (Project/CLAUDE.md), and Custom Files feeding memory into the coding chat.
    See how Claude Code stops repetition: global and project CLAUDE.md files, plus custom reference docs, flow into the editor so the assistant remembers your preferences and context while you code and run commands.

    Because the files live on my machine, I own the system. No vendor or device lock-in. I decide when and who to share with. I can work with Claude on one project and ChatGPT on another—both can rely on the same file-based memory strategy. It’s an AI strategy that scales with product discovery, accelerates go-to-market content, sharpens competitive differentiation, and supports product-led growth.

    Here’s how I design the memory: I use three layers. Claude Code already encourages global preferences and Project-specific instructions, but the third layer—reference context—is where the real power lives.

    Dark-mode screenshot of a macOS editor showing a 'Claude Code Preferences' markdown file with sections on writing conventions, planning protocol, and feedback for collaborating with Claude.
    Peek inside a markdown playbook for Claude Code: concise rules for writing, multi-level planning, and clear feedback that turn repeated reminders into reusable memory and smoother, faster coding sessions.

    Layer 1: Global Preferences (Always on). The first time I launched Claude Code, I created a CLAUDE.md file at ~/.claude/CLAUDE.md. This is where I keep the cross-project rules of engagement—how I like to work with Claude. Mine includes: Always create a plan for me to review before you start any work; Give me direct feedback (no hedging, no gentle suggestions); Use bullet points for summaries; Ask clarifying questions one at a time so I can give complete answers; No emojis unless I explicitly ask for them. Claude Code automatically loads this file at the start of every session, so I never restate my preferences.

    Layer 2: Project-Specific Instructions. Different projects have different rules. In my writing workspace, the Project CLAUDE.md sets the roles (I’m the primary writer; Claude is my thought partner and editor), defines a multi-round review flow (content → structure → accuracy → typos), prioritizes human readability over SEO, and points to my writing style guide. In my task management system, I include how my Trello integration works, file naming conventions for tasks, and how to process research papers into summaries. In my code projects, I specify the technology stack (Node.js vs. Python), testing framework (Jest for Node.js, pytest for Python), code style and conventions, project architecture and directory structure, and which dependencies and libraries to use. Each project directory has its own CLAUDE.md, and Claude automatically loads the relevant file when I’m working there.

    Dark-themed text editor screenshot of a markdown file titled 'Claude Instructions,' featuring sections for session setup, working relationship, editor responsibilities, and research and development guidelines.
    Peek inside a markdown playbook for collaborating with Claude—covering session setup, roles, editorial standards, and research steps—to show how saved instructions create consistent results without repeating yourself.

    Layer 3: Reference Context (Pull as Needed)—the real power. LLMs have a context window—a limit to how much they can process at once. Even within that limit, loading too much degrades performance due to “context rot.” The remedy is ruthless context management: small, targeted files that load only when needed. Keep CLAUDE.md files concise and focused on rules and workflows. For detailed knowledge, create separate reference files and list them in your CLAUDE.md so Claude knows they exist and when to fetch them. When I ask for help creating a landing page, Claude knows to use my business profile, the product file, and my target customers context.

    Here’s what most people miss: you don’t cram everything into global or Project files. You maintain small, reusable reference files that Claude only loads on demand. In my walkthrough, I share exactly which context files I created and why; how I got Claude Code to help me create them; how I break them into small, reusable components so Claude gets precisely what it needs; how I keep everything up to date; and step-by-step instructions so you can set up a similar memory system.

    Diagram of three markdown files (business-profile.md, story-based-customer-interviews.md, target-customers.md) feeding into a Claude Code IDE panel, showing context files powering an AI assistant.
    Three project notes funnel into Claude Code, turning reusable context into working output. This visual shows how saving key docs as memory lets the AI pick up where you left off and skip repetitive prompting across tasks.

    Let’s dive in.


    Inspired by this post on Product Talk.


    Book a consult png image
  • How to Build a Product Positioning and Messaging Strategy

    How to Build a Product Positioning and Messaging Strategy

    Your homepage promises an all-in-one platform. The sales deck leads with automation. The product demo focuses on analytics. Each claim may be true, but together they force the buyer to work out what you are, whom you serve, and why you matter. That is a positioning failure, not a copy problem.

    The way out is to separate the strategic choice from its expression. First decide which customer and buying situation you intend to win. Then build a messaging system that carries that decision from the first impression through sales, onboarding, and product use. The method below gives you the artifacts, tests, and operating rules to do both.

    Separate the strategic decision from the words

    Positioning, messaging, and copy are related, but they solve different problems:

    • Positioning decides the target segment, urgent customer job, category, primary alternative, promised outcome, meaningful difference, and proof.
    • Messaging decides which parts of that position to emphasize, in what order, for each audience and stage of the buying journey.
    • Copy turns the message into a headline, sales talk track, pricing-page explanation, onboarding prompt, or product-tour step.

    This distinction tells you where to intervene. If leaders disagree about the customer or alternative, a headline workshop will only conceal the disagreement. If the position is clear but buyers do not understand it, the messaging hierarchy needs work. If the hierarchy is sound but one page underperforms, you may have a copy or execution problem.

    I use a simple diagnostic: ask the product, marketing, sales, and customer-success owners to complete the following prompts independently. Do not let them discuss wording first.

    • The customer I most want to win is…
    • They look for a solution when…
    • The progress they need is…
    • They would otherwise use, assemble, or tolerate…
    • They should choose this product because…
    • The evidence that makes that claim credible is…

    Compare the nouns and decisions in the answers, not their polish. If one person names agencies, another names sales teams, and another says any growing business, you do not have a shared target. If the alternatives range from a direct competitor to spreadsheets and doing nothing, the team is framing different buying decisions. Resolve those differences before approving new copy.

    The output of this diagnosis should be a short list of strategic questions, each with an owner and an evidence gap. That is far more useful than a document full of compromise language.

    Build the position from evidence, not ambition

    Choose a segment that behaves like a good customer

    A broad market description is not a target segment. Modern teams, small businesses, and enterprises are labels, not choices. A usable segment combines a buyer or user, an operating context, a trigger, and a need that is unusually important in that context.

    Start with behavioral evidence from activation, retention, and expansion. Look for cohorts that reach meaningful value, continue using the product, and deepen their commitment. Then investigate why. A large cohort that requires heavy persuasion and struggles to retain may be a less attractive positioning target than a smaller cohort that recognizes the problem immediately.

    Write a segment brief with four fields:

    • Who: the buying role, user, or accountable leader.
    • Context: the company type, workflow, maturity, or constraint that changes the value of the product.
    • Trigger: the event that turns a background inconvenience into a priority.
    • Exclusion: a plausible customer for whom the product is not the best fit.

    The exclusion is important. If you cannot say who should not buy, the segment is probably still too broad. Specificity does not make the total market disappear. It gives your message a place to land.

    Name the progress, not the product output

    Customers do not wake up wanting a dashboard, an AI assistant, or another system of record. They want to make a decision sooner, remove a risky handoff, create predictable pipeline, reduce manual work, or gain control over an outcome they already own.

    Complete this sentence using the customer’s language: After adopting this product, the customer can do what they could not do reliably before? The answer should describe progress in the customer’s world. A capability belongs in the explanation of how the result happens, not in the result itself.

    Tie the promise to a business result the customer already tracks, but do not add a number merely to make the claim sound concrete. A quantified promise requires evidence that supports the same segment, use case, and conditions. Until you have that evidence, state the direction of value plainly and use verified proof lower in the message.

    Define the category, alternative, difference, and proof

    The buyer needs a familiar frame before your differentiation can matter. A category tells them what kind of decision they are making. Points of parity tell them you meet the minimum conditions for consideration. Differentiation tells them why you should win after you qualify.

    DecisionQuestion to answerCommon failure
    CategoryWhat familiar kind of solution is this?Inventing a label the buyer must decode before understanding the product.
    Points of parityWhat must be true for the product to make the shortlist?Leading with table stakes as if they were differentiation.
    AlternativeWhat would the customer use, assemble, or tolerate without this product?Assuming the only alternative is a named competitor.
    DifferentiationWhich valuable outcome or mechanism is meaningfully better?Using adjectives that any competitor could copy.
    ProofWhat evidence supports the exact claim?Offering confidence, popularity, or technical detail that does not prove the promise.

    The primary alternative may be a competitor, a generic platform, a manual workflow, a collection of tools, or the decision to do nothing. Name the one that appears in the buying situation you are targeting. Your differentiation is meaningful only in relation to that alternative.

    Proof can take several forms: measured customer outcomes, time-to-value evidence, product behavior, implementation evidence, data-governance controls, privacy-by-design, or cybersecurity commitments. Match the proof to the anxiety created by the claim. If you promise speed, prove speed. If you promise control, prove governance. A list of impressive but unrelated facts will not close the credibility gap.

    Use this compact positioning structure once those choices are clear:

    For [specific customer in a defined context] who needs [urgent progress], [product] is a [familiar category] that delivers [customer outcome]. Compared with [primary alternative], it [meaningful difference], supported by [relevant proof].

    Positioning statement template

    Treat every bracket as a decision, not a place for the most flattering phrase. Mark each clause as evidence, assumption, or aspiration. Evidence can enter the approved statement. An assumption becomes a test. An aspiration belongs in product strategy until the product and proof can support it.

    Before moving on, apply six checks:

    • Does the segment exclude anyone you could plausibly sell to?
    • Would the target customer recognize the triggering problem?
    • Does the category reduce the explanation burden?
    • Is the alternative one customers actually consider?
    • Would the difference still matter if a competitor copied the wording?
    • Does the proof establish the claim rather than merely decorate it?

    If a competitor can paste your statement onto its homepage without changing the meaning, you have described the market, not your position.

    Turn one position into a messaging system

    A positioning statement is an internal decision tool. It is rarely the exact sentence that should appear on every customer-facing surface. Buyers need the same strategic story expressed at different levels of depth.

    Build the message in this order:

    1. Category cue: help the buyer place the product on a familiar mental shelf.
    2. Core outcome: state the progress that makes the product worth considering.
    3. Mechanism: explain how the product creates that outcome differently from the alternative.
    4. Proof: supply evidence for the claim and mechanism.
    5. Objection response: address the trade-off, risk, or missing parity point most likely to stop the decision.
    6. Next step: ask for an action that fits the buyer’s current level of intent.

    This order prevents two common errors. Leading with features makes the buyer infer the value. Leading with a grand outcome and no mechanism makes the claim sound ungrounded. The combination of outcome, mechanism, and proof gives the message both relevance and credibility.

    For an intent-data product, a message unit could work like this:

    • Claim: Act on buying intent while it is still useful.
    • Mechanism: Translate live product-usage signals into prioritized opportunities and the appropriate next action.
    • Proof: Insert only verified evidence, such as observed time to value, measured conversion results, documented governance, or customer validation.

    The example does not need faster, smarter, seamless, or revolutionary. Those words add no information unless a mechanism and evidence give them a precise meaning.

    Next, create a message map for each audience that participates in the decision. Use the same position, but change emphasis:

    • Economic buyer: business consequence, strategic fit, financial logic, and adoption risk.
    • Operational user: workflow improvement, usability, time to value, and what changes in the working day.
    • Technical or trust evaluator: integration, data handling, governance, privacy, security, and operational control.

    For each audience, record the trigger, desired outcome, current alternative, core claim, supporting mechanism, accepted proof, likely objection, and appropriate call to action. That becomes the brief for a landing page, demo, campaign, or onboarding flow.

    Do not create a new position for every persona. If an executive hears an efficiency story, an operator hears a feature story, and a technical evaluator hears an infrastructure story with no common outcome, the account receives three products. Keep the strategic claim stable and translate the consequence, mechanism, and proof for the listener.

    Consistency does not mean identical copy. It means every message helps the customer reach the same conclusion about whom the product is for, what it changes, and why it is the better choice.

    Test for customer movement, not internal applause

    A message that wins a leadership vote has passed a preference test. It has not passed a market test. Validation should show whether the intended customer understands the position, believes it, and takes a more valuable next step.

    Write the hypothesis before changing the asset:

    For [target segment] at [journey stage], emphasizing [message decision] instead of [current framing] will improve [customer behavior] because [expected change in understanding or motivation].

    Messaging experiment hypothesis

    Then run the test with the following controls:

    1. Capture the current baseline and the audience definition.
    2. Change one meaningful message decision, not the message, design, offer, and traffic source at the same time.
    3. Choose a primary metric that reflects progress at that stage of the journey.
    4. Add guardrails for downstream quality, retention, or unwanted customer mix.
    5. Set the decision rule before reviewing the result.
    6. Record what changed, what happened, for whom it happened, and what the result does not establish.

    The metric must match the surface:

    • Acquisition page: qualified conversion is more useful than raw visits or attention.
    • Sales conversation: look for clearer problem recognition, fewer category misunderstandings, relevant objections, and progression to the agreed next step.
    • Onboarding: measure activation and completion of the behavior tied to the promised value.
    • In-product message: measure the meaningful action after the prompt, not merely a tooltip click.
    • Expansion motion: look for adoption and commercial movement in the segment the message was intended to reach.

    A higher click-through rate with weaker qualified conversion is not a positioning win. It may mean the new wording creates curiosity but attracts the wrong expectation. Follow the behavior far enough to see whether the message improves customer fit rather than only top-of-funnel volume.

    You can pressure-test messaging across landing pages, onboarding flows, in-app guidance, sales talk tracks, and nurture sequences. Amplitude can help inspect behavioral cohorts; Pendo and Intercom can support in-product delivery and measurement; HubSpot can connect lifecycle messages with funnel behavior. The tool is secondary to a clean hypothesis and a metric that reflects the decision you are trying to improve.

    If traffic is too limited for a reliable A/B test, use customer interviews, comprehension checks, sales-call analysis, and structured message reviews to learn why language works or fails. Treat that evidence as directional. Interview feedback can reveal confusion, relevance, and objection mechanisms, but it should not be relabeled as causal conversion lift.

    Keep a decision log. For every experiment, store the segment, surface, control, variant, hypothesis, primary metric, guardrails, result, interpretation, and next decision. Without that record, teams repeatedly test synonyms while forgetting the strategic assumption underneath them.

    Read results diagnostically. A message that improves acquisition but not activation may be setting an expectation the product does not fulfill. A message that works for one retained cohort but fails for another may reveal that the target segment is too broad. A claim that repeatedly requires explanation may indicate a poor category choice. The purpose of testing is not to defend the original language; it is to improve the decision system.

    Make positioning part of the product operating system

    Positioning decays when it lives only in a launch deck. Sales adapts the story to objections, marketing optimizes individual campaigns, product ships capabilities, and onboarding inherits old promises. Each local choice can seem reasonable while the overall narrative drifts.

    Create one canonical positioning brief with:

    • An accountable owner, version, approval date, and current validation status.
    • The target segment, trigger, and explicit exclusions.
    • The urgent job and customer outcome.
    • The category and required points of parity.
    • The primary alternative and competitive difference.
    • Approved proof for each claim, including any conditions or limits.
    • The message hierarchy and audience-specific message maps.
    • Known objections, prohibited unsupported claims, and open assumptions.
    • Links to experiment results and the decisions they changed.

    The brief should govern product decisions as well as communication. When reviewing roadmap work, ask whether the item strengthens the promised outcome, closes a parity gap that blocks consideration, compounds the reason to choose the product, or creates proof for a claim customers already value. Work that does none of these may still be necessary, but it needs a different strategic justification.

    This prevents differentiation from becoming a slogan unsupported by investment. If the product claims a uniquely fast path to value while roadmap decisions add setup complexity, the market will eventually believe the experience rather than the headline.

    Roll the position through the connected customer journey. Update the homepage, pricing explanation, sales discovery, demo narrative, onboarding, product tours, in-app guidance, customer-success materials, and nurture sequences that rely on the old framing. Prioritize the surfaces where the intended segment makes or validates its decision. A new promise on the homepage paired with an old demo and unrelated onboarding creates more confusion than a controlled, coherent rollout.

    Give one owner authority to maintain the canonical brief, while making product, marketing, sales, and customer success responsible for contributing evidence. Version meaningful changes. A headline iteration does not require a new strategic version; changing the target segment, category, alternative, outcome, or differentiation does.

    Review the position when evidence changes, not merely because the calendar says it is time. Useful triggers include a major product launch, entry into a new segment, a shift in the alternative customers choose, a parity gap that changes shortlist eligibility, new proof that strengthens the promise, or a persistent mismatch between acquisition, activation, retention, and expansion.

    Do not rewrite the position after every losing copy test. A failed expression and a failed strategic premise are different diagnoses. Change the position only when the evidence shows that the customer, problem, category, alternative, promise, or reason to believe has changed.

    Key takeaways

    • Resolve disagreements about the customer, buying trigger, category, and alternative before debating headlines.
    • Choose a segment using activation, retention, and expansion behavior, then document whom the position excludes.
    • Build every major message from an outcome, a distinctive mechanism, and proof that supports the exact claim.
    • Test messaging against meaningful customer behavior and downstream quality, not internal preference or clicks alone.
    • Use the approved position to guide roadmap trade-offs, go-to-market assets, onboarding, and future experiments.

    Take your current positioning statement and label every clause as evidence, assumption, or aspiration. Pick the assumption that would most change the strategy if it proved false, and design the next customer or behavioral test around it. Validate the decision before rewriting every surface. Once it holds, carry the same position all the way into the product experience.

    References

  • The Product Playbook: Measuring Agent Performance with Pendo and Agent Analytics to Drive ROI

    The Product Playbook: Measuring Agent Performance with Pendo and Agent Analytics to Drive ROI

    I treat agent performance analytics as a strategic product lever, not a back-office metric. When I combine Pendo’s product signals with Agent Analytics from our support systems, I get a unified view of where users struggle, how agents intervene, and which in-app experiences accelerate resolution. That visibility lets my team drive product-led growth and improve customer experience while lowering support costs.

    Increase revenue, cut costs, and reduce risk with Pendo’s Software Experience Management platform. Optimize the entire software experience to drive adoption and improve engagement.

    In practice, I build a clear scorecard that blends both product and support KPIs: first response time, resolution rate, first contact resolution, CSAT, containment/deflection rate, average handle time, ticket volume per active account, onboarding completion, user activation, and time-to-value. This balanced view ensures we reward not just speed, but durable outcomes that reduce repeat contacts and improve retention.

    To make the data actionable, we connect our CRM integration, ticketing events, and Pendo product analytics in a unified analytics platform. That gives me cohort-level clarity—who needed help, what they were doing before opening a ticket, how agents responded, and whether users stayed engaged afterward. With clean instrumentation and consistent taxonomies, Agent Analytics becomes a reliable operating system for both product and support leadership.

    I then use in-app guides, tooltips, and product tours to proactively address the top friction points that drive ticket volume. Through A/B testing, we compare cohorts exposed to guided workflows versus control groups, measuring deflection, faster task completion, and downstream conversion. When a guide meaningfully reduces tickets for a given workflow, we promote it from experiment to standard onboarding, and we feed those learnings back into our roadmap.

    The real unlock comes from tying outcomes to business impact. I track how improvements in resolution quality and self-serve adoption influence expansion revenue, support cost per account, and risk signals like churn propensity. Retention analysis helps us validate whether reduced friction and better agent coaching translate into sustained engagement and healthier accounts.

    Operationally, Agent Analytics helps me coach teams with precision. I spotlight high-performing behaviors, identify knowledge gaps, and standardize winning playbooks directly in the product via in-app guidance. This approach empowers agents, shortens onboarding for new hires, and keeps our best practices current as the product evolves.

    None of this works without trust. We apply privacy-by-design principles and strong data governance, ensuring that analytics, coaching, and automation respect user consent and data minimization standards. With that foundation, we can scale confidently—experiment faster, learn from every interaction, and continuously improve the software experience.

    If you’re getting started, begin by baselining your agent and product KPIs, ship one high-impact guide to deflect a top ticket driver, and review results weekly. Within a quarter, you’ll have a repeatable loop: diagnose friction, test an in-app solution, measure deflection and satisfaction, and reinvest the gains into the next set of improvements.


    Inspired by this post on Pendo – Best Practices.


    Book a consult png image
  • From Engineer to Product Manager: A Practical Transition Plan

    From Engineer to Product Manager: A Practical Transition Plan

    You may already be doing the parts of engineering that sit closest to product management: questioning a requirement, clarifying the user problem, challenging an unnecessary feature, or helping design and product make a difficult trade-off. The uncertainty is whether those moments add up to PM readiness – and whether changing careers means discarding the technical credibility you worked hard to earn.

    They don’t prove that you’re ready, but they give you a strong starting point. The safest path is to test the role before you depend on the title. Own a bounded customer problem, work through discovery and prioritization, ship a small bet, and make the resulting evidence visible. That gives you a transition plan based on demonstrated product judgment rather than potential alone.

    Change the scoreboard from implementation to impact

    Engineering and product management overlap, but they aren’t measured the same way. An engineer is expected to make a solution reliable, maintainable, secure, and feasible. A PM is expected to determine which problem deserves attention, why it matters now, what evidence supports the decision, and how the team will know whether its bet worked.

    The first transition is therefore moving from shipping outputs to driving measurable user or business outcomes. That doesn’t make delivery unimportant. It changes the role delivery plays: a feature becomes a hypothesis about how to create value, not the finish line.

    When you encounter a request such as “build bulk editing,” don’t start by turning it into tickets. Rewrite it as a product decision:

    • User and context: Which segment encounters the problem, and during which workflow?
    • Observed problem: What are people trying to accomplish, and where does the current experience fail them?
    • Current behavior: What workaround or alternative do they use now?
    • Desired outcome: Which user or business measure should change if the problem is solved?
    • Hypothesis: Why should this particular intervention change that measure?
    • Smallest useful test: What can you ship or simulate to reduce the most important uncertainty?
    • Decision rule: What evidence would make you continue, change direction, or stop?

    This framing exposes weak roadmap items quickly. If you can’t identify the affected segment, current behavior, baseline signal, or decision rule, the team doesn’t yet have a product bet. It has a solution looking for justification.

    Technical depth remains useful. You can detect hidden dependencies, challenge unrealistic scope, and understand where platform choices restrict future options. The trap is allowing feasibility to dominate desirability and business value. A solution can be technically elegant, delivered on time, and still leave the customer problem untouched.

    Run a 90-day transition experiment in your current role

    An internal move is usually easier to de-risk because you already understand the product, architecture, delivery process, and organizational context. Instead of asking your manager to approve a permanent career change based on intent, propose a bounded 90-day product experiment with an outcomes dashboard and a weekly stakeholder update.

    Choose a problem that matters but doesn’t require control of the entire roadmap. It should have an identifiable user, an observable pain point, a plausible measure of success, and enough room for a small intervention. Avoid a project whose scope is already fixed. Coordinating predetermined delivery may demonstrate execution, but it gives you little opportunity to show discovery, prioritization, or product judgment.

    PhaseWork to ownEvidence to preserve
    First 30 daysMap the users, workflow, current alternatives, relevant metrics, stakeholders, and decision process. Define the problem boundary and establish the baseline signal.A one-page problem brief, workflow map, initial dashboard, interview plan, and written scope.
    By day 60Run focused discovery, combine interview patterns with quantitative signals, compare possible interventions, and build a hypothesis-led roadmap.Discovery notes, customer language, an opportunity tree, rejected options, trade-offs, and a prioritized experiment.
    By day 90Deliver a thin slice, observe the result, follow up with affected users, and recommend whether to continue, revise, or stop.A before-and-after dashboard, decision log, updated roadmap, outcome narrative, and lessons that change the next decision.

    Set the operating agreement before the trial begins. Write down what you own, which decisions you can make, who remains accountable for the broader roadmap, and how much engineering work you will retain. A minimal engineering contribution can reduce the immediate staffing risk, but minimal must be explicit. Otherwise, you can end up carrying a full engineering workload while attempting a second full-time role.

    Your weekly update should be short enough that leaders will read it and structured enough that they can intervene:

    • The outcome you are trying to influence.
    • What you learned from users or data.
    • Which assumption became stronger or weaker.
    • The decision made and the trade-off accepted.
    • The next uncertainty to reduce.
    • Any decision or support needed from the recipient.

    This cadence does more than report activity. It demonstrates that you can turn incomplete information into a clear decision without hiding uncertainty. It also prevents the trial from becoming invisible work that everyone appreciates but nobody recognizes as product ownership.

    Practice the three skills engineering may not have forced you to build

    Technical competence can help you enter the conversation, but it won’t compensate for weak discovery, vague positioning, or poor stakeholder management. Those are the areas to practice deliberately during the transition.

    Product discovery: investigate behavior before proposing a solution

    Engineers are trained to solve well-defined problems. Product discovery tests whether the apparent problem is real, important, and worth solving for a particular segment. The distinction matters because confident solution design can make a weak assumption look mature.

    Use interviews to reconstruct actual behavior rather than solicit approval for an idea. Useful prompts include:

    • Walk me through the last time you tried to complete this task.
    • What triggered the need?
    • Where did the workflow slow down or break?
    • What did you do next?
    • What workaround have you adopted?
    • What was the consequence of leaving the problem unresolved?

    Avoid leading with a proposed feature or asking whether someone would use it. People can be polite, imaginative, and optimistic about hypothetical behavior. Recent examples, current workarounds, and actual consequences give you firmer evidence.

    Don’t turn each interview into a roadmap vote. Look for repeated situations, motivations, obstacles, and alternatives. Then check those patterns against quantitative signals such as activation, conversion, retention behavior, or support volume. Qualitative evidence explains what may be happening; quantitative evidence helps you understand its reach and movement.

    Product positioning: make the value segment-specific

    A technically capable product can still fail to communicate why anyone should change behavior. Positioning forces you to choose whose problem matters and why your approach is preferable to the status quo.

    Draft a simple statement: For [specific segment] struggling with [observable problem], this capability helps them achieve [meaningful outcome], unlike [current alternative], because [relevant distinction].

    Each bracket requires evidence. If you describe the user as everyone, the segment is too broad. If the outcome is easier or better, it is too vague. If you can’t name the current alternative, you may not understand the real competition, which is often an established workaround rather than another product.

    Stakeholder management: communicate decisions, not activity

    A PM rarely controls every team needed to produce an outcome. You must create alignment through context, evidence, and explicit trade-offs. That is different from satisfying every stakeholder request. Stakeholder agreement can help delivery, but it does not prove customer value.

    Build updates around the decision:

    • What decision is required?
    • Which outcome does it affect?
    • What evidence is relevant?
    • Which viable options were considered?
    • What does each option trade away?
    • What do you recommend, and why?
    • Who owns the next action?

    Remove implementation jargon unless it materially changes the decision. Executives need the consequence of a dependency, not a tour of the dependency graph. Engineers need constraints and reasoning, not a priority handed down without context.

    Practice these skills inside a product trio involving product, design, and engineering. The trio gives you access to different forms of judgment while preventing product discovery from becoming a solo PM exercise. Agree on decision rights and sponsorship at the start so you don’t become an unofficial PM with responsibility but no authority.

    Turn the work into evidence that survives an interview

    A long ticket history doesn’t demonstrate product judgment. Your portfolio has to show how you reduced uncertainty, made a choice under constraints, aligned the people needed to act, and learned from the result.

    Build each case study around a decision rather than a feature:

    • Context: Who was the user, what were they trying to do, and why did the problem matter?
    • Uncertainty: What did the team not know at the beginning?
    • Evidence: Which customer and product signals changed your understanding?
    • Alternatives: What other options were credible, including doing nothing?
    • Choice: What did you prioritize, and what did you deliberately decline?
    • Delivery: How did you reduce scope while preserving a useful test?
    • Outcome: What changed in activation, conversion, support demand, or another relevant measure?
    • Learning: What did the result change about the next roadmap decision?

    Attach the supporting artifacts only after the narrative is clear. Useful evidence includes a one-page problem brief, anonymized discovery notes, customer language, an opportunity solution tree, a hypothesis-led roadmap, an outcomes dashboard, and a before-and-after roadmap snapshot. The artifacts support your judgment; they shouldn’t force the interviewer to reconstruct it.

    Be precise about causality. If several initiatives were running at once, say that your work influenced an outcome rather than claiming it caused the entire change. If the target metric didn’t move, don’t bury the result. Explain which assumption failed, what you stopped doing, and how the evidence improved the next decision. Honest learning is a stronger PM signal than a polished success story with implausibly clean attribution.

    For an internal transfer

    Package your trial as a proposal your manager and product leader can evaluate. Include the problem boundary, success measure, product trio, weekly update rhythm, retained engineering commitment, artifacts you will produce, and the decision to be made at the end of the 90 days. This turns a vague request for a chance into a controlled staffing and product experiment.

    For an external search

    Prepare two deep case studies: one centered on discovery and another on delivery. The discovery case should show how you challenged the initial framing and reduced uncertainty. The delivery case should show how you handled constraints, aligned stakeholders, protected the outcome while reducing scope, and shipped.

    Expect follow-up questions about trade-offs: What did you say no to? Which assumption worried you most? Why was the thin slice sufficient? What evidence would have reversed your decision? What did you do when stakeholders disagreed? If your answer is only that the team completed the roadmap, you are still presenting yourself as a delivery coordinator. The stronger signal is that a decision changed because you understood the customer, business, and system more clearly.

    Key takeaways

    • Your engineering background is an advantage, not proof of PM readiness. Use it to improve decisions, not to dominate the solution.
    • Replace feature completion as your scoreboard with a clearly defined user or business outcome.
    • Build experience before changing titles by owning one bounded problem through a 90-day internal trial.
    • Use a weekly update to expose evidence, assumptions, trade-offs, decisions, and requests for help.
    • Practice discovery, positioning, and stakeholder management deliberately; technical fluency won’t substitute for them.
    • Make your portfolio decision-centered, quantify the outcomes you influenced, and represent causality honestly.
    • Prepare one discovery-led case and one delivery-led case for external interviews.

    Your next move isn’t rewriting your resume. Choose one user pain in a product you already understand. Write a one-page problem brief, identify the product and design partners you need, define the outcome you will track, and ask a sponsor to support a bounded trial. Let the title follow the evidence.

    References