Tag: product discovery

  • From Zero to One in B2B Marketing: My Proven SaaS Playbook for Growth, Hiring, and Attribution

    From Zero to One in B2B Marketing: My Proven SaaS Playbook for Growth, Hiring, and Attribution

    Early-stage B2B marketing is where momentum is made or lost. In my product leadership work, I’ve seen that getting from zero to one requires uncommon focus, founder-led GTM discipline, and a tight feedback loop between product, sales, and marketing. In this narrative, I share the playbook I use—and the patterns I took from top operators—to help SaaS teams build credibility fast, compound learnings, and scale repeatable growth.

    Alex Kracov is the CEO and Co-Founder at Dock, and the former VP of Marketing at Lattice. Alex joined Lattice as the first marketer and third employee, and he helped to grow the business from seed to 1850+ customers. Prior to Lattice, Alex was a consultant at Blue State Digital — the team that elected President Obama and orchestrated projects at Google. Since leaving Lattice in 2021, Alex co-founded Dock, a B2B platform that has streamlined the customer buying experience for clients like Loom, Origin, and Instabug.

    Here’s the agenda I use to guide founders and early marketing leaders: the 2023 SaaS marketing playbook; how to start your early-stage B2B marketing; how to prioritize resources across multiple marketing bets; how to think about attribution; Lattice’s unorthodox million-dollar marketing campaign; how to hire for early marketing roles; what makes a standout marketer; and advice for building your first website.

    When I spin up early-stage B2B marketing, I start by defining the shortest path to signal. That means a crisp ICP, problem-first messaging, and one or two channels where our buyers already congregate. At this stage, I bias toward founder-led discovery calls, live product walkthroughs, and tight content that proves outcomes—not features. This creates the raw material for positioning, case studies, and a credible top-of-funnel narrative.

    Short-term versus long-term goals must be explicitly balanced. I set near-term pipeline and learning targets (e.g., qualified conversations per week, time-to-insight from experiments) alongside long-term brand assets (evergreen content, customer proof, category POV). The rule of thumb I apply: stabilize one growth motion before layering the next, so we don’t overfit to noise or dilute the message.

    Allocating resources across marketing bets is a portfolio problem. I structure it as 70/20/10: 70% on the core motion that’s already working, 20% on adjacent bets with clear hypotheses, and 10% on contrarian experiments that could unlock step-change distribution. Weekly syntheses convert experiment data into decisions—double down, redesign, or retire.

    On attribution, I’m pragmatic. Early on, precision is less valuable than directionality. I pair multi-touch analytics with qualitative inputs (self-reported attribution, sales notes, community signals). The question I ask: which narratives and channels consistently show up in won deals? That blend avoids over-crediting the last click and keeps us honest about how trust is actually formed in B2B.

    Your first website is a conversion engine and a trust anchor. The first thing people should see on your website is the problem you solve, the outcomes you deliver, and a frictionless way to see the product in action. I recommend a tight hero message, social proof above the fold, a short demo video or interactive experience, and clear CTAs for both buyers who are ready now and those who need to explore.

    Brand and positioning mature with evidence. I translate discovery insights into a simple hierarchy: category, problem, unique insight, product proof, outcomes. At Lattice, strong brand clarity met operational excellence; at Dock, product-led collaboration sells the value by making the buying experience itself the demo. In both cases, the lesson stands: great B2B brands tell a truth buyers can quickly verify.

    Bold bets can be force multipliers. Lattice’s unorthodox million-dollar marketing campaign underscores a principle I use sparingly but decisively: when the narrative, timing, and distribution are aligned, a high-conviction investment can set the agenda for your category. The bar is high. The insight must be non-obvious, the creative durable, and the measurement plan rigorous.

    Hiring for early marketing roles, I optimize for learning velocity, narrative craft, and cross-functional empathy. The ideal first marketer is a full-stack generalist who can research, write, ship, analyze, and partner with sales and product. Experience matters, but potential—ownership, curiosity, systems thinking—often outperforms. I scale the team once one motion is repeatable and there’s a clear backlog of work we can’t tackle without specialization.

    Conferences and communities are underrated if used deliberately. I set specific objectives (target accounts, partners, customer content) and treat events as field research and content engines. Every conversation informs messaging; every meeting has a next step; every session becomes a clip, post, or asset. The outcome is pipeline plus reusable proof.

    My 2023 SaaS marketing stack emphasizes speed to insight: product analytics to observe behavior; a CRM and marketing automation platform to orchestrate journeys; lightweight data pipelines for attribution; a CMS for shipping content fast; and collaboration tools that put buyers and sellers in the same workspace. What matters most is not the logo set—it’s the operating cadence that converts data into action.

    If you’re going from zero to one, keep it simple: validate your ICP, ship a compelling narrative, pick one channel to master, and measure what buyers say and do. Sequence beats scope. Credibility compounds. And the best marketing is a mirror of a product that solves a painful, urgent problem—beautifully.

    Timestamps: [00:00:00] Intro [00:02:45] The challenges and opportunities in early-stage B2B marketing [00:05:13] How to think about short-term versus long-term marketing goals [00:07:31] Allocating resources across marketing bets [00:09:13] Signs your marketing is working [00:11:20] The most underutilized marketing strategy [00:13:03] Creating your company’s first website [00:14:22] How Lattice formed its brand messaging and positioning [00:18:22] Dock’s innovative approach to marketing software [00:20:14] The first thing people should see on your website [00:23:10] Lattice’s most successful early-stage marketing tactics [00:28:05] Determining which marketing strategies are still relevant [00:30:25] Lattice’s unorthodox million-dollar marketing campaign [00:33:26] Why Alex had an outsized impact at Lattice [00:37:05] Lessons from his first marketing hires [00:39:41] When to scale your marketing team [00:40:55] Building an effective early-stage marketing team [00:42:30] A tough conversation with the CEO & Co-founder of Lattice [00:44:46] Achieving early-stage marketing alignment [00:46:20] Transitioning from employee to entrepreneur [00:49:19] Getting the most out of conferences [00:50:47] Selecting marketing channels in the early stages [00:52:44] Hiring marketers for experience versus potential [00:56:34] The 2023 SaaS marketing stack [00:58:19] Advice for Zero to One marketing [00:60:46] What successful B2B marketing looks like

    Referenced: Dock: https://www.dock.us/ Lattice: https://lattice.com/ Jack Altman: https://www.linkedin.com/in/jackealtman J Zac Stein: https://www.linkedin.com/in/jzacstein

    Where to find Alex Kracov: Twitter: https://twitter.com/kracov/ LinkedIn: https://www.linkedin.com/in/alexkracov Website: https://www.kracov.co/

    Where to find Brett Berson: Twitter: https://twitter.com/brettberson LinkedIn: https://www.linkedin.com/in/brett-berson-9986094/


    Book a consult png image
  • Supercharge Your Engineering Org: Alignment, AI, and Productivity from Adobe to Etsy

    Supercharge Your Engineering Org: Alignment, AI, and Productivity from Adobe to Etsy

    I obsess over building high-velocity engineering organizations that ship meaningful outcomes. When I evaluate what reliably moves the needle—across startups and scaled enterprises—it always comes back to alignment, disciplined management, and a modern view of engineering productivity. Recently, I revisited a set of insights that crystallize these themes and translate them into practical rituals any leader can adopt.

    Kellan Elliott-McCrea is a Head of Engineering at Adobe, overseeing Frame.io, a newly acquired video review and collaboration platform. He is known for his experience and expertise as an engineering leader. He was previously a VPE at Dropbox, and CTO at Etsy where he built and led a team of 300 people, from tech and platform reboot through to IPO. Kellan also built and scaled teams at Flickr, and has a coaching and advising practice for companies looking to supercharge their engineering teams.

    Here’s what we dig into when we talk about world-class engineering orgs: how software engineering has changed in the last 10-15 years; the future of software engineering, and the impact of AI; the importance of alignment and tactics for achieving it; how to think about and enable engineering productivity; lessons on culture from Adobe, Dropbox, and Flickr; concrete tips for being a better manager; and rituals for building business literacy throughout an org.

    Let’s start with a reality I see in my own work: engineering teams are bigger than they were a decade ago, despite dramatically better tools and platforms. The reason isn’t inefficiency—it’s scope. Today’s products carry higher bars for reliability, privacy, security, compliance, and multi-surface experience. The coordination surface area has exploded. That’s why operating models must evolve: clear interfaces between teams, standardized decision-making, and reliable cross-functional rhythms are no longer nice-to-haves—they’re throughput constraints.

    Alignment, then, is the ultimate speed multiplier. I’ve learned the hard way that slow teams are rarely under-skilled; they’re misaligned. “Slow teams are misaligned teams.” To counter this, I anchor on a few tactics: articulate a clear strategic narrative (why now, why us, why this), commit to outcomes vs output OKRs, and institutionalize decision logs so debates don’t reset every sprint. When teams know the customer problem, the business bet, and how their work ladders up, the flywheel starts turning.

    On engineering productivity, I avoid vanity metrics and favor a portfolio: flow and focus (interruptions, WIP), system signals (lead time, deployment frequency, change fail rate), and outcome alignment (how progress maps to customer value and revenue impact). Tools matter—DX investment in CI/CD, observability, and paved roads—yet the largest gains usually come from simplifying priorities and reducing cross-team coupling. Fewer, better bets will beat “more tickets shipped” every time.

    The future of software engineering is inseparable from AI. In my practice, I treat gen ai and gen ai for product prototyping as core accelerators: copilots for code and tests, scaffolding services that convert specs to boilerplate, and retrieval-augmented knowledge that collapses the gap between tribal lore and action. The key is to measure impact at the team level—cycle time, defect escape, and learning velocity—so AI augments engineering judgment rather than creating hidden complexity.

    Culture is the compounding edge. Lessons on culture from Adobe, Dropbox, and Flickr converge on a few essentials: invest in psychological safety and clarity of purpose, operationalize blameless learning, and make information radically accessible. “How Complex Systems Fail, by Richard I. Cook, MD” is a touchstone here—complexity punishes organizations that rely on heroics and rewards those that build resilient systems and shared mental models.

    For managers, I return to a short, durable list. Schedule real one-on-ones that prioritize coaching over status. Write more than you speak; clarity scales through documents. Run crisp, time-boxed decision forums with pre-reads and owners. Close the loop on feedback—especially in moments of disagreement—by documenting trade-offs and naming the decider. These concrete tips for being a better manager build trust, accelerate decisions, and enable autonomy.

    Every high-performing engineering org I’ve led invests in business literacy as a first-class ritual. I recommend monthly “Finance 101” briefings, customer support ride-alongs, and deal reviews to connect engineers to revenue realities. Pair that with tactics and rituals for enabling effective teams—weekly written updates, demo-driven reviews, and pre-mortems—and you get sharper prioritization and far better cross-functional coordination.

    Why so few companies successfully go multi-product? Most underinvest in platforms, shared services, and explicit funding models for internal APIs. The remedy: treat platforms as products with clear roadmaps, SLAs, and customer empathy; align incentives so teams don’t fork capabilities in the rush to ship; and adopt technical governance that favors standardization where it compounds and freedom where it differentiates.

    For compensation and career architecture, I pressure-test common models by asking: does this design reward the behaviors we say we want? If we value outcomes, impact, and enabling others, the ladders should reflect it. When the incentives match the mission, the org learns faster and scales cleaner.

    Referenced:

    Adobe: https://www.adobe.com

    Dropbox: https://www.dropbox.com/

    Flickr: https://www.flickr.com/

    Frame: https://www.frame.io/

    How Complex Systems Fail, by Richard I. Cook, MD: https://how.complexsystems.fail/

    How Etsy Grew their Number of Female Engineers by Almost 500% in One Year https://review.firstround.com/How-Etsy-Grew-their-Number-of-Female-Engineers-by-500-in-One-Year

    Where to find Kellan Elliott-McCrea:

    Twitter: https://www.twitter.com/kellan

    LinkedIn: https://www.linkedin.com/in/kellanem

    Website: https://kellanem.com/

    Personal blog: https://laughingmeme.org/

    My bottom line: if you want to supercharge your engineering org, anchor on alignment, measure what matters, and leverage AI to elevate—not replace—engineering judgment. Do that, and you’ll turn coordination costs into compounding advantages that show up in customer value, velocity, and morale.


    Book a consult png image
  • Building Products in a Post-LLM World: Hard-Won Lessons, Skeptic Busters, and Team Playbooks

    Building Products in a Post-LLM World: Hard-Won Lessons, Skeptic Busters, and Team Playbooks

    The ground rules for product development have changed in the post-LLM world. I’m sharing a practical, first-person playbook—lessons I’ve pressure-tested in my own product org—to help you build AI-native products with confidence, cut through hype, and deliver outcomes that compound.

    Sprig is an AI-powered user insights platform that has raised over $88m. Today’s discussion features two key individuals in Sprig’s journey so far: Ryan Glasgow, Sprig’s CEO and founder; and Kevin Mandich, Sprig’s Head of Machine Learning. Before Sprig, Ryan was an early PM at GraphScience, Vurb, and Weeby (all of which were acquired), and Kevin was an ML Engineer at Incubit, and a Post-Doctoral Researcher at UC San Diego.

    In today’s episode, we discuss: Key lessons from the Sprig founding story; Product development in the pre vs. post-LLM world; How to overcome AI skepticism; How to evaluate new models and how to know when to switch; Why you need an ML engineer; Sprig’s “AI Squad” team structure; How Sprig upskills all team members on AI.

    Founding story takeaways I keep returning to: conviction compounds when paired with continuous discovery. Early on, prioritize direct customer signal over elegant architectures. I’ve seen the fastest learning loops come from a tight PM–ML partnership that prototypes quickly, validates with real users, and refactors only after signal stabilizes. The Jobs to Be Done Framework: https://hbr.org/2016/09/know-your-customers-jobs-to-be-done remains my favorite lens to separate what the model can do from what the customer actually needs done.

    Pre vs. post-LLM product development requires a mindset shift. Pre-LLM, we wrote deterministic systems and pushed the edge with models like Google’s BERT model: https://en.wikipedia.org/wiki/BERT_(language_model). Post-LLM, we design probabilistic systems, treat prompts like code, and invest in evaluation harnesses from day one. I routinely prototype with Chat GPT: https://chat.openai.com and scaffold experiments with Langchain: https://www.langchain.com/ to compress discovery cycles. The key is shipping guardrails and UX affordances that make non-determinism feel trustworthy.

    On AI skepticism, I don’t argue—I demonstrate. I target one painful workflow, build a narrow, high-precision solution, and expose transparent failure modes with a human-in-the-loop escape hatch. This reframes AI from magic to leverage. In customer-facing settings (think customer support ai strategy), we measure deflection and satisfaction together so automation never outpaces user psychology.

    Evaluating new models—and knowing when to switch—demands a clear rubric: task quality (ground-truthed), latency at p95, unit economics, privacy/compliance, and operational reliability. I run shadow evaluations before swapping production dependencies, then phase changes behind flags with canaries and backstops. Tools like Auto-GPT: https://github.com/Significant-Gravitas/Auto-GPT are useful for ideation, but I never skip rigorous offline and online evaluation before a cutover.

    Why you need an ML engineer: the fastest teams pair a product manager who owns the problem framing with an ML engineer who owns the feasibility frontier. This duo translates ambiguous jobs into measurable tasks, instrumented datasets, and iterative model/UX improvements. In my experience, this partnership reduces time-to-learning more than any single tooling decision.

    Sprig’s “AI Squad” team structure mirrors what I’ve seen work: a cross-functional pod with a PM, ML engineer, data engineer/analyst, design, and platform partner. The squad ships thin slices end-to-end, owns their eval suite, and meets weekly to review errors, edge cases, and customer feedback. We track outcomes vs output OKRs to ensure velocity serves impact—not the other way around.

    Upskilling the entire team on AI is non-negotiable. I’ve had success with lightweight rituals: weekly demo hours, prompt libraries maintained in Jira: https://www.atlassian.com/software/jira, red-team exercises to uncover failure patterns, and internal brown bags where engineers and PMs teach each other. Small, frequent exposure beats heavyweight training.

    For deeper exploration and hands-on experimentation, I reference: Auto-GPT: https://github.com/Significant-Gravitas/Auto-GPT; Chat GPT: https://chat.openai.com; Google’s BERT model: https://en.wikipedia.org/wiki/BERT_(language_model); Jira: https://www.atlassian.com/software/jira; Jobs to Be Done Framework: https://hbr.org/2016/09/know-your-customers-jobs-to-be-done; Langchain: https://www.langchain.com/; Sprig: https://sprig.com/.

    Timestamps: (02:50) Intro (04:57) What attracted Kevin to Sprig (05:53) Kevin’s background before Sprig (07:56) How Ryan gained conviction about Kevin (09:55) Key technical challenges and how they solved them (18:46) How to overcome AI skepticism (21:47) The early difficulties of building an ML-enabled product (25:06) Evaluating new models and knowing when to switch (35:09) Using Chat GPT (37:23) Product development in the pre vs. post-LLM world (39:53) The impact of AI hype on Sprig’s product development (45:36) Balancing AI automation with user-psychology (48:47) Do recent LLMs reduce Sprig’s competitive advantage? (51:00) The importance of “selling the vision” to customers (54:40) How Sprig structures teams (57:25) How Sprig upskills all team members on AI (60:25) 3 key tips for companies trying to navigate AI (66:05) Major limitations with LLMs right now (70:27) The future of AI and the future of Sprig

    Three guiding principles I use daily: first, reduce surface area—start with one high-value job and earn trust with reliability. Second, treat evaluation as a product—version prompts, log failures, and continuously retrain on your own data distributions. Third, design for collaboration—pair AI with human judgment and transparent controls so users feel empowered, not replaced. Post-LLM success isn’t about chasing models; it’s about building resilient systems, teams, and learning loops.


    Book a consult png image
  • Inside Rewind AI’s Playbook: PMF Breakthroughs, Bold Twitter Fundraise, and the Future of AI

    Inside Rewind AI’s Playbook: PMF Breakthroughs, Bold Twitter Fundraise, and the Future of AI

    I sat down with Dan Siroker to explore the product, fundraising, and AI strategy lessons behind Rewind AI’s rapid rise — and to reflect on what I would adopt in my own product management practice today. Dan Siroker is the co-founder and CEO at Rewind AI, a personalized AI powered by everything you’ve seen, said, or heard. Dan launched Rewind to an emphatic response on Twitter, and used a public pitch video to fundraise at a $350m valuation. Prior to starting Rewind, Dan co-founded Optimizely, which reached $120m ARR before being acquired by Episerver, a content management company. Dan was also the Director of Analytics for Obama’s first presidential campaign.

    What stood out immediately was Rewind’s journey to Product Market Fit and how deliberately the team instrumented learning loops. As a product leader, I pay close attention to how founders reduce ambiguity: narrow the target segment, ship thin slices, measure engagement cohorts, and iterate fast. Rewind’s early focus on utility and trust — not novelty — created the conditions for PMF while the team resisted the temptation to over-scope.

    I was especially interested in how Rewind works and how the team managed scope while building a category-creating product. By focusing on personalized recall powered by on-device intelligence and a clear privacy narrative, they avoided the common trap of trying to solve everything for everyone. My own rule of thumb is to enforce brutal prioritization around the highest-intent jobs-to-be-done, then earn the right to expand. That same discipline shows up in Rewind’s cultural mantra for shipping and validating fast.

    Lessons from Optimizely echo throughout. Being a second-time founder sharpens pattern recognition — from building high-clarity cultural values to operationalizing product-market fit. I’ve found that codifying operating principles early helps a team move faster with fewer collisions, and Dan’s approach to open feedback and public learning raises the bar for transparency.

    On product positioning as a category creator, the team leaned into outcomes over features, which is critical when the mental model is new. Rather than compete in a features arms race, they framed a compelling before-and-after: instant, searchable memory that augments cognition. In my experience, that level of narrative clarity drives founder-led GTM and accelerates word-of-mouth.

    We also dug into where to build in AI, and what makes a “wrapper” thin versus thick. My take: thin wrappers add shallow convenience on top of foundation models; thick wrappers integrate proprietary data, workflow depth, distribution advantages, and durable UX moats. Founders should aim for thick wrappers with unique data flywheels, not commodity interfaces easily displaced by platform shifts.

    Operationalizing Product Market Fit remains a craft. I routinely use leading indicators like activation rate, day-7/day-30 retention for key actions, and sentiment via structured PMF surveys. Rahul Vohra’s framework for measuring and optimizing Product Market Fit: https://review.firstround.com/how-superhuman-built-an-engine-to-find-product-market-fit is a proven playbook. Pair that with cohort-based instrumentation and tight audience segmentation to reveal the “sharpest edge” of value.

    On AI hype, we aligned on a pragmatic view: real value accrues where latency, accuracy, and privacy meet workflow depth. Apple’s Silicon: https://www.macrumors.com/guide/apple-silicon/ and on-device acceleration will keep unlocking new consumer experiences, while ChatGPT: https://chat.openai.com/ has reset expectations for natural interfaces. The cautionary tales of Google Glass: https://en.wikipedia.org/wiki/Google_Glass and Google Wave: https://en.wikipedia.org/wiki/Google_Wave remind me that timing, social acceptability, and use-case clarity matter as much as technical novelty.

    Data privacy is now a core buying criterion, not a checkbox. I see a clear trend toward local-first approaches, explicit consent, and user agency — especially for products that touch memory, identity, and personal archives. Framing value through Maslow’s Hierarchy of Needs: https://www.simplypsychology.org/maslow.html helps prioritize trustworthy utility over gimmicks.

    Dan’s one-of-a-kind Twitter fundraising strategy was a masterclass in founder-led GTM. By sharing a public pitch and engaging directly with early users and supporters, he compressed feedback cycles and aligned community, product, and capital. For reference, see Dan’s public Twitter fundraise: https://twitter.com/dsiroker/status/1646895452317700097 and Dan’s Rewind demo tweet: https://twitter.com/dsiroker/status/1638799931891920897. The transparency extended to leadership practice as well, with Dan publicly sharing his own 360 performance reviews: https://twitter.com/dsiroker/status/1689763756459675650 — a bold move that builds trust.

    I’m watching what’s next for Rewind with interest, particularly around thicker integrations, extensibility, and collaboration patterns. In the next decade, I expect assistive AI to become ambient, multimodal, and context-aware — an ever-present copilot that feels less like a tool and more like an extension of cognition.

    Referenced: Apple’s Silicon: https://www.macrumors.com/guide/apple-silicon/

    Referenced: ChatGPT: https://chat.openai.com/

    Referenced: Dan publicly sharing his own 360 performance reviews: https://twitter.com/dsiroker/status/1689763756459675650

    Referenced: Dan’s public Twitter fundraise: https://twitter.com/dsiroker/status/1646895452317700097

    Referenced: Dan’s Rewind demo tweet: https://twitter.com/dsiroker/status/1638799931891920897

    Referenced: Google Glass: https://en.wikipedia.org/wiki/Google_Glass

    Referenced: Google Wave: https://en.wikipedia.org/wiki/Google_Wave

    Referenced: Maslow’s Hierarchy of Needs: https://www.simplypsychology.org/maslow.html

    Referenced: Optimizely: https://www.optimizely.com/

    Referenced: Paul Graham: https://twitter.com/paulg

    Referenced: Rahul Vohra’s framework for measuring and optimizing Product Market Fit: https://review.firstround.com/how-superhuman-built-an-engine-to-find-product-market-fit

    Referenced: Rewind AI: https://www.rewind.ai/

    Referenced: Scribe (which morphed into Rewind): https://www.scribe.ai/about

    Where to find Dan Siroker: Twitter: https://twitter.com/dsiroker

    Where to find Dan Siroker: LinkedIn: https://www.linkedin.com/in/dsiroker

    Where to find Dan Siroker: Personal website: https://siroker.com/

    Where to find Dan Siroker: Blog: https://medium.com/@dsiroker

    My takeaway for founders and product leaders: obsess over segmentation, instrument for learning, and tell a crisp narrative that earns trust. Thick wrappers, privacy-first design, and founder-led GTM are how you win the next wave of AI.


    Book a consult png image
  • Goal-Setting for AI Products: How I Plan, Prioritize, and Confidently Ship in a Nonlinear GenAI World

    Goal-Setting for AI Products: How I Plan, Prioritize, and Confidently Ship in a Nonlinear GenAI World

    I build and ship AI products in an environment where the frontier changes weekly, so my planning system has to be adaptive, evidence-driven, and unapologetically outcome-focused. In this piece, I share the frameworks I use to set goals for generative AI, balance research with product execution, and scale responsibly — drawing sharp lessons from one of the most influential applied AI companies operating today.

    Consider Runway, an applied AI research company shaping the next era of art, entertainment, and human creativity. Runway has raised $237m and was one of Time Magazine’s “100 most influential companies” in 2023. Runway has been a persistent viral sensation in recent years, and is behind many of the most famous AI demos online.

    The earliest stages of an AI company often begin with research breakthroughs, scrappy prototypes, and clever distribution. In practice, that means leveraging containerization (https://aws.amazon.com/what-is/containerization/) and Docker (https://www.docker.com/) to package models reproducibly, showcasing work where practitioners already gather — Hugging Face (https://huggingface.co/), Hugging Face Spaces (https://huggingface.co/spaces), and Hugging Face Model Hub (https://huggingface.co/docs/hub/models-the-hub) — and tapping infrastructure like Replicate (https://replicate.com/) to get demos into people’s hands. Early, magical use cases — like the Green screen tool by Runway (https://runwayml.com/green-screen/) — teach us which problems are both technically feasible and viscerally valuable.

    I’ve learned to be cautious about “The limitations of being “customer-driven” when building in AI”. Traditional product discovery assumes needs are legible and solutions are relatively deterministic. In generative AI, user desire often follows model capability, not the other way around. The job is to triangulate: run tight user loops to validate perceived value, instrument objective model quality, and explore novel interaction patterns that customers can’t yet articulate. I treat this as a portfolio of discovery bets — some customer-led, some capability-led, all evaluated against clear outcome thresholds.

    Balancing research development with product development requires organizational design that prevents context-switching tax while preserving velocity. I pair research pods with product pods, supported by forward deployed engineers and domain PMs who translate evaluation metrics into user-visible milestones. Safety and content moderation sit on the critical path, not as afterthoughts — think policy definition, classifier tooling, abuse red teaming, and clear escalation playbooks. This balance is how you move from a great demo to a dependable product without losing momentum.

    Goal-setting amidst constant change in AI starts with outcomes vs output OKRs. I write OKRs in terms of user impact and model performance thresholds — for example, target ranges for latency, quality scores against a golden dataset, or creator retention — then let teams choose the highest-leverage outputs (data pipelines, fine-tuning, UX improvements) to get there. Why I don’t plan very far ahead: I treat the annual view as a vision and bet map, the quarterly view as a constrained slate of outcomes, and the 6–8 week cycle as the execution heartbeat. AI roadmaps are hypotheses; evaluation harnesses and launch gates are the truth.

    Community is a force multiplier. Forming a vocal community and fostering community requires real access and real listening: early release cohorts, office hours, and transparent changelogs. How they picked users for early release matters — diversity of use cases, sophistication of workflows, and willingness to give crisp feedback. Expanding past the first 100 users of Gen-2 demands readiness: evaluation parity across modalities, scalable infra, and safety coverage. Done well, this motion compounds learning while building authentic advocacy.

    For founders, my advice echoes the core lessons above. Start with a narrow, high-intent wedge and prove durable value fast; let founder-led GTM compress the feedback loop; instrument everything from day one; and resist the urge to over-plan features before you’ve nailed outcomes. Product-market fit lessons in AI often arrive via small, fast experiments — not grand, long-range plans. Ship thin slices that demonstrate unmistakable value, then iterate toward a system, not a single feature. When in doubt, shorten the loop and improve the evaluation harness.

    People often ask: Will AI replace video editors? My view is that AI will replace zero editors who master these tools — and many who don’t. The winners blend taste, storytelling, and generative leverage. The products we build should honor this reality: design for control, iteration, and co-creation, not just automation.

    If you’re mapping the progression of tech and use-cases, a few public references are instructive: Runway Gen-1 (https://research.runwayml.com/gen1) and Runway Gen-2 (https://research.runwayml.com/gen2) show how capability unlocks new workflows and demand. Runway’s 30 AI Magic Tools (https://runwayml.com/ai-magic-tools/) illustrates portfolio thinking — a suite of composable powers rather than a monolith.

    For builders focused on gen ai for product prototyping through production: keep your demo muscle strong, your evaluation stronger, and your outcomes strongest. Invest in community, treat safety as a feature, and let your OKRs steer what ships — not the other way around.


    Book a consult png image
  • Engineering Leadership That Scales: Strategy, Velocity, and Org Design from Carta, Stripe, Uber, Calm

    Engineering Leadership That Scales: Strategy, Velocity, and Org Design from Carta, Stripe, Uber, Calm

    I’m often asked how I translate lessons from hypergrowth engineering organizations into practical playbooks for product and platform teams. In this piece, I unpack the patterns I’ve seen repeatedly work—anchored by what I admire about Will Larson’s approaches at Carta, Calm, Stripe, and Uber—and how I apply them to build resilient, high-velocity orgs. Will Larson is a case study in modern engineering leadership. As CTO at Carta—an ownership and equity management platform—he helped guide the company after it raised at a $7.4b valuation in 2021. Before that, he was CTO at Calm, founded Stripe’s Foundation Engineering org, and led Uber’s Platform Engineering people and strategy. He’s also the author of Staff Engineer and An Elegant Puzzle, both essential reads for leaders leveling up from line management to org design. When I craft an engineering strategy, I start by writing down a small set of clear principles. This isn’t performative; it’s an alignment mechanism. Principles reduce decision thrash, make trade-offs explicit, and help teams navigate ambiguity without constant escalation. I’ve found the discipline of writing them down upfront pays off 10x in execution quality later. For the strategy document itself, I structure it so anyone can understand the why, what, and how in one sitting. A useful pattern: a sharp problem definition, a few guiding policies, and a concise set of coherent actions. That scaffolding keeps the strategy legible and actionable across functions—especially as it ladders into product roadmaps, platform investments, and talent plans. Every engineering strategy has two parts. First, compounding capabilities: the platform, tooling, and architecture that unlock future velocity. Second, targeted bets: focused initiatives that advance near-term outcomes. Neglect either and you either stall out later (too many quick wins, no compounding) or fail to ship value now (all compounding, no customer impact). Turning strategy into action requires ruthless translation. I map each guiding policy to a small number of initiatives with owners, milestones, and outcome metrics—not output. This is where outcomes vs output OKRs matter: measure the user or business result, not just the deliverable. It’s also where you surface dependencies early and avoid the Hidden Variable Problem that quietly derails timelines. I’m particularly intrigued by Carta’s unique “navigator” model, which blends technical leadership with cross-functional guidance to accelerate execution while preserving autonomy. In my experience, similar patterns work when leaders are explicitly accountable for both system health and product outcomes—reducing the gap between platform decisions and customer value. Engineering velocity is explainable, measurable, and optimizable. I anchor on DORA and the research from Accelerate (book), and I complement it with the SPACE (framework) to account for satisfaction and collaboration, not just delivery. The story I tell executives is simple: pick a few canonical measures, instrument them consistently, and then drive the feedback loops—branching strategy, CI/CD hygiene, change size, and operational excellence. Choosing the right metrics for an engineering org matters as much as the metrics themselves. I use a balanced set: delivery (lead time for changes, deployment frequency), quality (change failure rate, availability), and flow (work in progress, batch size). Then I pair these with narrative context so the numbers inform decisions rather than become a game to win. On policy, nuance beats orthodoxy. Great leaders define clear, default rules while acknowledging real-world exceptions. I’ve learned to document the policy, define who can grant exceptions, and track exception volume to spot design flaws. The goal isn’t rigidity—it’s predictable operations with a safe on-ramp for edge cases. Micromanagement is a symptom, not a root cause. Telling someone “don’t micromanage” is often counterproductive. Instead, I focus on what’s missing—trust, clarity, or visibility. If leaders can see the plan, the risks, the checkpoints, and the demo cadence, they don’t need to hover. If they still do, fix incentives and accountability, not just behavior. I avoid management anti-patterns by watching for early signals: policies without principles, roadmaps without strategy, meetings without decisions, or dashboards without actions. The best engineering executives pair systems thinking with crisp communication. They’re close enough to the details to ask sharp questions, yet disciplined enough to scale through managers and staff engineers. Executive communication is an asymmetric game. I tailor the message to the decision horizon: one slide for the ask, one for the trade-offs, one for the plan and risks. The Minto Pyramid (framework) helps—lead with the answer, then support it. In meetings, the fastest way to derail progress is to lack a clear owner, a time box, or pre-reads. Fix those and you reclaim hours every week. For presentation feedback, I’ve found a cadence that works: clarify the objective, highlight the single biggest risk, and eliminate anything that doesn’t move the decision forward. A bad sign with direct reports is when updates are status-only and insight-light; I coach toward “what changed, why it changed, and what you need.” For early-career engineers, the most durable advantage is compounding learning: pick hard problems, write more than you think you should, and seek out leaders who invest in your growth. For team development, I borrow a simple model: staff your keystones, instrument your systems, and build a culture where the best ideas win, not the loudest voices. If you want to explore the foundations behind these practices, start here. Accelerate (book): https://www.amazon.com/Accelerate-Software-Performing-Technology-Organizations/dp/1942788339 Good Strategy, Bad Strategy (book): https://www.amazon.com/Good-Strategy-Bad-Difference-Matters/dp/0307886239 DORA: https://dora.dev/ SPACE (framework): https://queue.acm.org/detail.cfm Minto Pyramid (framework): https://untools.co/minto-pyramid Carta: https://www.carta.com/ Calm: https://www.calm.com/ Stripe: https://www.stripe.com/ JavaScript: https://www.javascript.com/ KAFKA: https://kafka.apache.org/ Ruby on Rails: https://rubyonrails.org/ To go deeper on Will’s writing and perspective, these are great starting points. Twitter/X: https://twitter.com/lethain LinkedIn: https://www.linkedin.com/in/will-larson-a44b543/ Personal website/blog: https://lethain.com/ An Elegant Puzzle (book): https://www.amazon.com/Elegant-Puzzle-Systems-Engineering-Management/dp/1732265186 Staff Engineer (book): https://staffeng.com/book
    Book a consult png image
  • Inside Bard’s Playbook: How to Ship AI Fast, Build Ethically, and Outlearn Competitors

    Inside Bard’s Playbook: How to Ship AI Fast, Build Ethically, and Outlearn Competitors

    I spend a lot of time helping teams reconcile two pressures that define modern product management: ship fast enough to learn and compete, but slow enough to be safe, ethical, and useful. Studying Bard offers a crisp blueprint for navigating that tension and leveling up how we build with Generative AI. Jack Krawczyk is a Senior Director of Product at Google, building Bard. Bard is Google’s collaborative, conversational, and experimental AI tool that’s bridging the gap between humans and bots, while addressing ethical considerations around AI. After joining the project in 2020, Jack helped ship Bard in less than four years. Bard sources information directly from the web, and now enables users to inquire about and summarize YouTube videos. From a product management lens, the most valuable takeaway is the sequencing: problem definition → principled constraints → rapid public learning with clear guardrails. I’ve seen this order de-risk speed. When we anchor teams on a tight product thesis and ethical framework, we unlock faster iteration without drifting into feature theater. Shipping early—especially with a Large Language Model (LLM)—can feel risky. Yet the decision to open Bard to the public quickly reflects a disciplined bias toward learning velocity. In my experience, the longer we delay real-world feedback with LLMs, the more our internal assumptions calcify. Early exposure surfaces edge cases, calibrates safety systems, and drives better prioritization than any lab-only evaluation can. Ethics in AI is not a separate workstream; it’s a product requirement. I anchor cross-functional reviews on harm modeling, transparency, and user agency. Bard’s framing makes this explicit: collaborative, conversational, experimental—language that signals co-creation and responsible exploration rather than unfettered automation. That positioning matters for trust and sets expectations for both quality and limitations. Differentiation in AI assistants increasingly hinges on live context and modality. Bard sources information directly from the web, and now enables users to inquire about and summarize YouTube videos. In practice, this moves Bard beyond static Q&A toward dynamic sensemaking. I advise teams to ask: what fresh, authoritative context can our system responsibly ingest to reduce hallucinations and increase actionability? On development speed, I look for a culture that marries ambition with measurable risk reduction. That means small, end-to-end vertical slices; evaluation harnesses aligned to user outcomes, not model vanity metrics; and weekly red-teaming that actually changes the roadmap. Outcomes vs output OKRs are critical here—optimize for quality-adjusted learning per unit time, not just feature count. Early user research should be embedded, not episodic. I’m a proponent of forward deployed engineers paired with product and research to observe failure modes in the wild and close the loop quickly. With LLM-based experiences, qualitative signals (confusion, trust breaks, cognitive load) often precede quantitative ones; instrument both and let them inform each other. Deciding when to ship comes down to clear thresholds. I pressure-test launch criteria with two prompts: what would change my mind tomorrow, and what could break if we’re right but too early? For AI features, I also require recovery paths—explanations, undo, source attribution—so that small misses don’t become trust-ending moments. As for the competitive landscape—Bard versus ChatGPT, and others—users ultimately reward utility, reliability, and workflow fit. I encourage teams to pick a sharp use case, lean into their unique distribution or data advantage, and prove value in minutes, not weeks. “Generative AI” is table stakes; reliable outcomes in a real job-to-be-done is differentiation. Zooming out, I see three fronts shaping the future of LLM, Generative AI, and AGI: model capability, grounding and retrieval quality, and product ergonomics. Most teams overinvest in capability and underinvest in grounding and UX. The fastest wins often come from better retrieval, tighter prompts, and clearer affordances—not just a larger model. For aspiring AI developers, start narrow and instrument deeply. Pick a workflow with painful status quo, ship a thin slice, measure correctness and confidence, and iterate with real users. For non-LLM companies, the mandate is different: augment your core product where AI reduces friction or unlocks frequency—don’t bolt on a chatbot because everyone else did. For product leaders, AI changes the craft in two ways. First, prototyping is faster—use this to expand the option space early. Second, evaluation requires new muscles—build an experimentation and safety stack that blends qualitative red-teaming with quantitative reliability and cost controls. The leaders who thrive will combine taste with statistical rigor. If you want to go deeper, these references are useful: Bard: https://bard.google.com/; ChatGPT: https://chat.openai.com/; Duet AI: https://cloud.google.com/duet-ai; Free courses on machine learning by Andrew Ng: https://www.andrewng.org/courses/; Google Assistant: https://assistant.google.com/; Introducing Google Assistant to Bard: https://blog.google/products/assistant/google-assistant-bard-generative-ai/; Large Language Model (LLM): https://en.wikipedia.org/wiki/Large_language_model; Meena: https://blog.research.google/2020/01/towards-conversational-agent-that-can.html. In sum, the Bard blueprint reinforces a simple truth: ship with a thesis, learn in public with care, and let principled constraints accelerate—not slow—your path to product-market fit. That’s how we create value fast, build ethically, and stay ahead in the next era of AI.
    Book a consult png image
  • What Makes or Breaks Executive Hires: My Lessons on Fit, Red Flags, and Measuring Success

    What Makes or Breaks Executive Hires: My Lessons on Fit, Red Flags, and Measuring Success

    Executive hiring is one of those rare decisions that can bend a company’s trajectory. In my role leading product management at a high-growth SaaS company, I’ve seen the difference between a leader who compounds value and one who quietly drains momentum. That’s why I was eager to examine what actually makes (or breaks) these bets, and to share a practical lens you can use to improve executive hiring outcomes.

    I sat down with Eeke de Milliano for a focused conversation on the realities of executive hiring, leadership transitions, and measuring success. We dig into the “buy or build a leader” decision, how to avoid common red flags, and what it takes to set executives up to thrive in hyper-growth environments.

    Eeke de Milliano is the Head of Global Product at Stripe, helping drive innovation and success in the company’s product line. Before this role, she was Head of Product at Retool and co-founded Constellate. Eeke previously spent 6 years as Product Lead at Stripe, working with the company during their hyper-growth era.

    In today’s episode, we discuss how to rigorously assess executive hiring fit, including the challenges companies face when hiring new executives and the most common red flags and pitfalls I see teams miss under time pressure. We also explore practical advice for measuring success, especially when outcomes vs output get muddled in the first 90–180 days.

    A recurring theme for me is that learning your own strengths is an underrated piece of the process. If you don’t understand the leadership leverage you already have on the team, you’ll over-hire for breadth or under-hire for depth. Great executive hiring clarifies the complementary edge you need—then measures it.

    On the buy vs build decision: early signals matter. If you’re “buying” an external leader, pre-align on scope, authority, and what great looks like before day one. If you’re “building” from within, design a clear on-ramp and operating cadence so the leader can scale without drowning. In both cases, my mental model is to instrument leading indicators (team health, decision velocity, stakeholder trust) well before lagging business metrics fully show up.

    Two red flags I always watch for: first, leaders who default to playbooks without interrogating context; second, leaders who cannot articulate how they measure success beyond activity and output. In hyper-growth, pattern-matching is useful—but uncalibrated pattern-matching is dangerous.

    The human dynamics matter just as much as the strategy. What creates dysfunctional exec relationships is often misaligned interfaces: unclear decision rights, overlapping charters, or incentives that reward local maxima. High-functioning executive teams are like parents—a united front in public, with candid debate in private, anchored to shared principles and measurable outcomes.

    Referenced:

    ASML: https://www.asml.com/en

    Claire Hughes Johnson: https://www.linkedin.com/in/claire-hughes-johnson-7058/

    Constellate: https://constellate.team/

    John Collison: https://www.linkedin.com/in/johnbcollison/

    Mike Maples Jr.: https://www.linkedin.com/in/maples/

    Patrick Collison: https://www.linkedin.com/in/patrickcollison/

    Retool: https://retool.com/

    Stripe: https://stripe.com/

    Will Gaybrik: https://www.linkedin.com/in/william-gaybrick-5730347/

    Where to find Eeke:

    LinkedIn: https://www.linkedin.com/in/eeke-de-milliano-3b05a629/

    Timestamps:

    (00:00) Should you ‘buy or build’ a leader

    (03:45) Why do executive hires fail so often?

    (09:35) Why the stakes are so high for leadership hires

    (12:26) The hardest document Eeke ever wrote

    (14:06) Two red flags in a new hire

    (17:27) An example of an outstanding leader

    (21:40) What creates dysfunctional exec relationships

    (22:38) The three steps towards hiring successful leaders

    (30:30) What you should know about outside hires

    (33:12) Eeke’s advice for easing leadership transitions

    (42:06) How to notice success patterns

    (47:21) Why high-functioning executive teams are like parents

    (52:02) The most surprising lesson from Eeke’s first stint at Stripe

    (55:11) The leadership data Eeke wishes we had


    Book a consult png image
  • Persuasive Leadership for Founders: My Take on Wes Kao’s Playbook to Influence and Win

    Persuasive Leadership for Founders: My Take on Wes Kao’s Playbook to Influence and Win

    Influence starts with clarity. That’s the throughline I return to when I’m coaching founders and product leaders, and it’s why I keep revisiting the frameworks that sharpen how we communicate, persuade, and lead under pressure. Recently, I synthesized several powerful ideas that map directly to the realities of startup execution and product management leadership—ideas I’ve seen transform how teams align, how roadmaps get prioritized, and how outcomes (not outputs) become the default.

    Wes Kao is an executive coach, advisor, and instructor, best known for her newsletter on high-impact communication, and for co-founding course platform Maven and the AltMBA with Seth Godin. Across her career, Wes has helped leaders communicate with clarity and conviction, whether it’s rallying a team, pitching investors, or influencing stakeholders.

    From a founder’s seat, or in a VP of Product role, the question is always the same: How do I become more persuasive, play to my strengths, and raise the bar for myself and my team? Here’s how I’ve put these principles into practice—and what I recommend.

    First, I rely on a “personality-message fit” mindset. The goal isn’t to copy someone else’s style; it’s to package your message so it amplifies your natural strengths. If you’re analytical, use structure and crisp logic. If you’re a storyteller, build vivid narrative arcs around data. In product reviews, I’ve seen the same idea land (or fall flat) entirely based on whether the delivery aligned with the speaker’s authentic style.

    Charisma is often misunderstood. It’s not about volume or showmanship—it’s about presence, intent, and calibration. Authenticity isn’t performative; it’s the consistency between your values and your behavior. In practice, that looks like stating trade-offs plainly, owning uncertainty, and being consistent in how you make decisions. Teams don’t need theatrics; they need reliability and conviction.

    Clarity in communication is the single highest ROI skill in leadership. Start with your ideal outcome: what do you want your audience to think, feel, and do? Then reverse-engineer your message. I frame every major communication around outcomes vs output, just as I would with OKRs. This shifts the discussion from activity (“we shipped”) to impact (“we moved this metric”). When the outcome is explicit, the argument becomes self-reinforcing—and far more persuasive.

    Power dynamics shape how your message is received. Different stakeholders hear the same words through very different lenses. In board updates and investor pitches, calibrate not just content but posture: what decision are you asking for, what risks are you proactively naming, and what constraints are you strategically acknowledging? Influence often hinges less on brilliance and more on aligning incentives and expectations.

    On the perennial question—should you work on weaknesses or double down on strengths?—I’ve found the most durable gains come from role-strength fit. Eliminate spiky weaknesses that are career-limiting (for example, unreliable follow-through), but invest disproportionately in the strengths that create asymmetric value. This is how leaders move from competent generalists to compelling, irreplaceable operators.

    Effective self-reflection is a force multiplier. A deceptively powerful prompt I use with teams is: What do you resent? Resentment often points to violated boundaries, unclear roles, or recurring misalignments. Surface it, re-contract responsibilities, and redesign rituals. This isn’t soft work; it’s operational hygiene that protects focus and velocity.

    When someone tells you to “be more strategic,” they’re rarely asking for more slideware. They want clearer time horizons, sharper prioritization, and better sequencing. I lean on stack ranking to make trade-offs explicit. If everything is a priority, nothing is. Show what’s first, what’s second, and what you’re explicitly saying no to—and why. Strategy is the discipline of exclusion.

    Two ideas I return to often: how formative programs start and how craft gets defined. The origin story behind community-driven learning like the AltMBA reminds me that great products are built with a point of view and a tight feedback loop. Defining your craft—naming it, practicing it, and holding a higher standard for it—creates a culture where excellence becomes normal, not exceptional.

    If you’re a founder or product leader, a practical way to apply all of this next week is simple: decide the outcome, tailor the message to your natural style, acknowledge power dynamics up front, and stack rank your asks. Then, debrief with the team: What landed? What didn’t? What will we do differently next time? Communication is a craft, and like any craft, standards rise with deliberate practice.

    AltMBA: https://altmba.com/

    Maven: https://maven.com/

    Seth Godin: https://www.sethgodin.com/

    Udemy: https://www.udemy.com/

    Where to find Wes:

    LinkedIn: https://www.linkedin.com/in/weskao


    Book a consult png image
  • When a Focused Product Wedge Is Ready to Become a Platform

    When a Focused Product Wedge Is Ready to Become a Platform

    Your wedge is working. Customers are buying, sales keeps hearing adjacent requests, and the larger platform opportunity suddenly looks close. This is where an otherwise disciplined roadmap can become a collection of modules held together by a broad narrative.

    The decision is not whether the market could use more products. It is whether your current advantage can make the next product easier to build, easier to adopt, and harder to replace. You need evidence of reuse before you need a platform roadmap.

    A focused wedge is a precise promise, not a small product

    A product wedge is the narrowest complete solution that gives a specific customer a compelling reason to change behavior. It is not a stripped-down version of a future platform. It must solve an important job from trigger to outcome, even if the underlying product is technically complex.

    That distinction matters. A shallow product offers a few features. A focused product may include integrations, compliance logic, observability, onboarding, support, and difficult infrastructure, but every part reinforces the same customer promise.

    Guideline’s wedge was not simply a smaller retirement product. Payroll integration, compliance automation, transparent pricing, and auto-enrollment worked together to make a 401(k) plan easier for small and medium-sized businesses to adopt and operate. Linear’s performance, reliability, simplicity, and workflow design similarly served one demanding audience: high-performance software teams. Both products contained substantial depth without losing coherence.

    Write your wedge as an operating contract before discussing expansion:

    • Primary user: Who experiences the problem and uses the product?
    • Economic buyer: Who approves the purchase, and what budget or priority makes the purchase possible?
    • Trigger: What event causes the customer to look for a solution now?
    • Job: What painful, repeatable work must be completed?
    • Outcome: What changes for the customer when the product works?
    • Distribution path: Where does the customer already look, buy, or work?
    • Quality floor: Which dimensions, such as accuracy, reliability, speed, security, or compliance, cannot be compromised?

    If different leaders answer these questions differently, the wedge is not yet stable enough to support expansion. The next planning cycle should tighten the core, not add a platform theme.

    A good wedge also creates concentrated learning. Reducto found traction by solving the complete problem of turning difficult documents and spreadsheets into structured data AI teams could use. Owner learned through the urgent operating reality of independent restaurants rather than beginning with a generic small-business platform. In each case, narrow scope improved the quality of customer evidence and made the next capability easier to see.

    Earn expansion through repeated variation around a stable core

    Customers will ask for features long before you are ready to become a platform. A request proves that somebody wants something. It does not prove that the capability belongs in your product, that other customers will adopt it, or that building it will create leverage.

    The strongest platform signal is repeated variation around a stable job. Customers want the same outcome, but their inputs, rules, integrations, approval paths, or review requirements differ. That pattern can justify reusable primitives. A stream of unrelated jobs from unrelated buyers usually points to a services business or several separate products, not a platform.

    Classify every expansion request before it enters the roadmap:

    • Core gap: The request is necessary to deliver the wedge’s existing promise. Treat it as core product work.
    • Adjacent workflow: The request sits immediately before or after the core job and serves the same user or buyer. Investigate it as a possible expansion.
    • Reusable variation: The request changes how the core job is configured, connected, evaluated, or governed. Look for a platform primitive.
    • Customer-specific exception: The request matters to one account but has no visible reuse path. Price and manage it as bespoke work, or decline it.
    • Separate market: The request introduces a different user, buyer, workflow, distribution motion, or risk model. Treat it as a new wedge that must earn its own evidence.

    This taxonomy prevents a common error: interpreting every enterprise requirement as platform validation. Large prospects can expose important needs, but their contract value does not make their workflow representative.

    Expansion gateEvidence that supports expansionWarning that the wedge needs more work
    Core healthTarget customers activate, receive the promised outcome, and continue using the core without extraordinary intervention.Expansion is being used to compensate for weak activation, retention, reliability, or positioning.
    Repeated demandThe same adjacent problem appears across relevant customers in their own language and workflow.Demand comes mainly from one strategic account, a sales objection, or internal enthusiasm.
    Capability reuseExisting data, integrations, trust, workflows, or technical primitives materially reduce the work required.The new capability needs a separate architecture, data model, operating process, and support motion.
    Commercial continuityThe existing buyer understands the value and can adopt through the current go-to-market path.A new buyer, budget, sales narrative, procurement process, or channel is required.
    Core protectionThe team can name guardrails for reliability, time-to-value, release cadence, and customer support.The plan assumes the core can absorb more complexity without explicit limits.

    Do not approve the expansion merely because several gates look promising. Resolve any critical warning first. A new compliance obligation, a different buyer, or a separate operating model can outweigh several superficial similarities.

    Build the platform beneath the product before you market it

    A bundle gives customers more things to buy. A platform makes additional use cases cheaper and faster to deliver because they share durable capabilities. That leverage should exist in the product and operating model before it appears in positioning.

    Useful platform primitives tend to sit below the visible feature layer. Depending on the product, they may include connectors, normalized schemas, permissions, policy rules, workflow orchestration, validations, identity controls, audit trails, observability, review queues, or billing infrastructure. The exact list matters less than whether the same capability serves distinct customer outcomes without being copied and maintained separately.

    Persona’s move from an identity verification MVP toward a horizontal platform required turning customer-specific work into reusable systems. Reducto’s expansion logic similarly centered on transferable capabilities such as connectors, schemas, validation, review, lineage, and auditability. Guideline created leverage by doing difficult infrastructure work early, particularly payroll integration and compliance automation. These capabilities are not decorative platform features. They are the machinery that makes adjacent experiences possible.

    Use a services-to-software loop when the pattern is still emerging:

    1. Deliver the new outcome end to end for a relevant customer, even if parts of the implementation are manual.
    2. Record every exception, custom rule, data transformation, integration dependency, and support intervention.
    3. Separate stable behavior from customer-specific variation.
    4. Turn stable behavior into a shared primitive with clear inputs, outputs, ownership, telemetry, and tests.
    5. Keep variable behavior configurable only where customers genuinely need different choices. Prefer strong defaults elsewhere.
    6. Use the primitive in the core experience as well as the adjacency. If the core cannot consume it cleanly, the abstraction may be premature or misplaced.
    7. Check whether the next implementation becomes simpler. If effort and exception volume keep rising, you are accumulating services work rather than platform leverage.

    Forward-deployed work is valuable when it produces reusable artifacts: an adapter, evaluation case, acceptance test, workflow primitive, implementation playbook, or observability requirement. Without that exit condition, customer proximity can quietly become permanent customization.

    You also need a principled way to decline revenue. Persona’s early decision to turn down a $5,000 deal rather than violate a product tenet captures the issue. A deal can be commercially real and strategically expensive. If it adds a parallel architecture, unique support promise, or enduring exception for one customer, calculate the continuing complexity rather than looking only at the initial contract.

    AI reuse requires more than a shared model

    AI teams are especially vulnerable to false platform signals. Reusing the same model, prompt framework, or orchestration library does not mean two use cases share a product platform. The real question is whether they can reuse the data contracts, evaluation method, quality thresholds, permissions, observability, review workflow, and failure-handling model.

    If every adjacency needs different ground truth, a different tolerance for error, new human reviewers, separate governance, and a new output schema, it may be a separate product even when the underlying model is identical. Treat evaluation and operational controls as platform primitives. Otherwise, model reuse can hide growing product fragmentation.

    Before exposing an AI capability as a platform service, make its quality legible. Define the evaluation set, observable failure states, escalation path, versioning behavior, and human-review boundary. A platform customer needs to know not only how to call the capability, but also when its output should not be trusted.

    Choose the next adjacency by leverage, then protect the core

    Score continuity before market size

    A large adjacent market is tempting because it improves the strategy narrative immediately. It does not reduce the execution risk. Start with continuity: how much of the current customer relationship and product advantage carries into the new job?

    DimensionHigh-leverage adjacencyLow-leverage expansion
    User continuityThe same person encounters the adjacent problem during the existing workflow.A different role must learn, operate, and advocate for the product.
    Buyer continuityThe existing buyer owns the outcome and can justify the additional spend.The product enters a different budget, executive priority, or procurement path.
    Workflow continuityThe new job happens immediately before, during, or after the core job.The connection exists mainly in a market map or executive narrative.
    Capability continuityThe adjacency reuses data, integrations, permissions, trust, or operational primitives.Most of the system must be designed, built, secured, and supported independently.
    Distribution continuityThe current channel, sales motion, partnership, or product loop reaches eligible customers.The team needs a new audience, category story, channel, and acquisition model.
    Risk continuityThe existing compliance, reliability, and support model covers the added workflow.The adjacency creates materially different financial, legal, privacy, or safety exposure.

    Use the map as triage, not as a mathematical forecast. A strong candidate should show continuity across the dimensions that are expensive or slow for your company to recreate. Any major break should appear explicitly in the investment case.

    The safest expansion sequence usually moves through increasing organizational distance:

    1. Deepen the wedge: Improve the completeness, reliability, or time-to-value of the original outcome.
    2. Extend the workflow: Solve a closely connected job for the same user and buyer.
    3. Expose reusable capabilities: Let internal teams, customers, or partners configure and combine proven primitives.
    4. Enter a new segment or vertical: Reuse the platform in a market that may require different positioning, distribution, or domain controls.
    5. Pursue a different buyer or job: Treat this as a new wedge with its own discovery and product-market fit burden.

    This sequence is not mandatory, but skipping levels should be a conscious strategic bet. Owner’s multi-product opportunity is strongest when each capability deepens value for the same restaurant operator. Reducto can move horizontally when document connectors, schemas, and review workflows transfer across industries. A market adjacency is attractive only when the underlying leverage survives the move.

    Measure leverage, not the size of the release

    Revenue growth alone cannot tell you whether expansion is working. New revenue can coexist with slower onboarding, heavier support, declining reliability, and a fragmented roadmap. Track three layers of evidence:

    • Core guardrails: Activation, time-to-value, retained usage, reliability, release cadence, support demand, and delivery of the original customer outcome.
    • Expansion outcomes: Adoption among eligible customers, attach rate, usage after activation, improvement in the customer’s workflow, retention behavior, and willingness to pay without forced bundling.
    • Platform leverage: Time required to launch another use case, reuse of existing primitives, implementation effort, exception volume, operational burden, and the amount of customer-specific code or process.

    Set the decision thresholds from your own baselines before launch. There is no universal attach rate or reuse target that proves platform readiness. The important discipline is to define what improvement, acceptable cost, and core degradation would mean before results are available.

    Organize the roadmap around the same distinction. Customer-facing outcomes belong in one view; reusable capability investments belong in another. Link them explicitly. Every proposed platform investment should name the customer outcome that first requires it, the next credible consumer, the primitive being reused, and the core guardrail it must protect.

    Run the transition as a falsifiable product bet

    Do not begin with a platform launch date. Begin with a decision brief that makes the expansion easy to disprove. This changes the conversation from executive conviction to product evidence.

    1. Restate the wedge contract. Make the current user, buyer, trigger, job, outcome, distribution path, and quality floor explicit.
    2. Build a demand log. Use customer interviews, sales calls, support conversations, implementation notes, usage behavior, and renewal feedback. Record the underlying job rather than copying feature requests.
    3. Classify the demand. Separate core gaps, adjacent workflows, reusable variations, customer-specific exceptions, and separate markets.
    4. Map current primitives. Identify which data, integrations, workflows, controls, and trust assets can genuinely be reused. Mark assumptions that still need testing.
    5. Select the thinnest complete adjacency. It must deliver an end-to-end outcome while exposing the most important reuse assumptions.
    6. Test through close customer work. Keep product, engineering, go-to-market, and support near the implementation. Capture exceptions and turn recurring work into artifacts.
    7. Review core guardrails and platform leverage. Look for faster subsequent delivery, lower exception volume, sustained use, and no unacceptable damage to the wedge.
    8. Choose the next state deliberately. Deepen the core, continue validating the adjacency, extract a shared primitive, scale the expanded product, or stop.

    Write stop conditions into the brief. Pause or narrow the expansion if core reliability deteriorates, onboarding becomes materially harder, customers adopt only through discounts or bundling, implementation exceptions keep increasing, the buyer changes, or the new workflow requires an independent go-to-market and support system. These are not temporary inconveniences to hide inside execution. They are evidence that the expansion thesis may be wrong.

    Outcome-based goals make this review cleaner. Instead of committing to launch a module or publish an API, define the customer behavior and operating leverage you expect. Then attach guardrails for the core. The release is an experiment; sustained customer value and reusable capability are the result.

    Key takeaways

    • A focused wedge solves a complete, urgent job for a specific user and buyer. It can be technically deep without becoming broad.
    • Repeated variation around the same outcome is a platform signal. Unrelated requests from different buyers are not.
    • Build reusable connectors, schemas, controls, workflows, evaluations, and observability before selling a platform narrative.
    • Prefer adjacencies that preserve the user, buyer, workflow, capabilities, distribution, and risk model.
    • Measure core health, expansion adoption, and platform leverage separately. Revenue by itself can conceal rising complexity.
    • Treat every expansion as a falsifiable bet with explicit assumptions, guardrails, and stop conditions.

    At your next roadmap review, ask for the wedge contract, demand classification, primitive map, leverage case, core guardrails, and stop conditions. If those artifacts do not exist, the next step is discovery, not a platform launch. Expansion should make your original advantage compound; if it merely makes the product larger, keep the wedge sharp.

    References

    • Shivam.Consulting Blog – How Guideline Rewired 401(k)s: First-Principles Strategy, Gusto Edge, and Product Wins
    • Shivam.Consulting Blog – Scrappy Outbound to ‘Hyperbolic’ PMF: How a COVID Pivot Fueled Owner’s Explosive Growth
    • Shivam.Consulting Blog – How a Weekend Hack Hit 7-Figure ARR: My Product Playbook from Reducto’s Rise
    • Shivam.Consulting Blog – From Skeptic to $2B: The Hard-Won Product Playbook Behind Persona’s Platform
    • Shivam.Consulting Blog – Inside Linear: How Craft, Focus, and Small Teams Build Category-Defining Products
  • Reliable AI Product Systems: A Product Leader’s Playbook

    Reliable AI Product Systems: A Product Leader’s Playbook

    Your AI feature can look excellent in a demo and still be unfit for a customer workflow. The real launch question isn’t whether the model can produce a good answer. It is whether your product can detect a bad answer, contain the consequences, and recover without making the customer absorb the failure.

    If you’re deciding whether an AI feature is ready to scale, treat reliability as a property of the whole product system. The model matters, but so do the workflow, retrieval layer, tools, validation, interface, fallback, observability, evaluation suite, and operating process. This gives you something more useful than confidence in a demo: a release decision you can defend.

    Define the reliability contract before choosing the stack

    A reliable AI product does not need to be correct in every possible situation. It needs to deliver a defined outcome within a declared operating envelope, recognize when it has left that envelope, and take a safe next step. Reliability therefore starts with a product promise, not a model benchmark.

    Write that promise as a reliability contract before debating models, retrieval-augmented generation (RAG), agents, or fine-tuning. This is an internal product artifact rather than a legal document. Its job is to make success, failure, and fallback explicit enough to evaluate.

    DecisionWhat the contract must stateWhy it affects release readiness
    User and jobWho is using the system, what they are trying to complete, and where the AI enters the workflowThe same output can be useful in one workflow and dangerous in another
    Observable outcomeThe customer or business result that should improve, such as resolution, completion, time saved, or handoff qualityOutput quality has no product meaning unless it changes the job
    Quality criteriaThe dimensions that make an output acceptable, such as accuracy, relevance, completeness, grounding, and appropriate toneReviewers and automated graders need a shared definition of good
    Hard constraintsConditions the system must never violate, including required schemas, permissions, privacy rules, and prohibited actionsAn average quality improvement cannot compensate for a critical constraint failure
    Abstention and handoffWhen the system should ask for information, decline, use a deterministic fallback, or route to a personA known limitation becomes manageable when the next step is designed
    Operating envelopeThe accepted latency, cost, supported languages, data boundaries, and workflow conditionsA system can be accurate and still be commercially or operationally unusable
    Release evidenceThe eval results, production signals, and owner approval required for a changeThe team can distinguish a promising experiment from a production candidate

    Use the contract to challenge the premise that AI belongs in the workflow. Generation is a good fit when ambiguity is part of the job and useful outputs cannot be reduced to straightforward rules. If a deterministic method solves the problem more consistently, cheaply, or transparently, use it. A sound product decision considers whether failures can be bounded, whether latency and cost fit the workflow, and whether a graceful fallback exists before committing to an AI implementation.

    The acceptable failure envelope depends on what happens next. A drafting assistant whose output a user reviews can tolerate different uncertainty from an agent that sends a customer message, changes a record, or triggers an external action. Raise the evidence bar as reversibility decreases and consequence increases. Do not assign one generic reliability target to every AI feature in the portfolio.

    Your scorecard should keep four layers visible:

    • User outcome: Was the task completed, resolved, or meaningfully advanced?
    • Task quality: Was the result correct, relevant, complete, grounded, and usable?
    • Hard constraints: Did the system respect policy, privacy, permissions, required formats, and action boundaries?
    • Operations: Did latency, cost, availability, retrieval, and tool execution stay within the agreed envelope?

    Avoid compressing these layers into one attractive score. A high average can hide a critical policy violation, a weak customer segment, or a tool action that silently failed. Product leaders need the outcome view and the failure distribution, not just a leaderboard number.

    Design a bounded workflow around the probabilistic core

    A large language model (LLM) is probabilistic. Your entire product does not need to be. The practical design pattern is a constrained AI capability inside a more deterministic workflow: explicit inputs, limited actions, structured outputs, validation, and a defined recovery path.

    Map the workflow before optimizing the prompt. For every step, identify the input, the component making the decision, the data or tool it may use, the expected output, the validator, and the failure route. This exposes vague handoffs that a single conversational prompt can conceal.

    • Constrain inputs where the workflow already knows the relevant choices or context.
    • Break a broad instruction into steps that can be observed and evaluated separately.
    • Require structured outputs when downstream software will consume the result.
    • Put permissions, policy checks, schema validation, and business rules outside the model.
    • Validate retrieved evidence and tool results before allowing the workflow to continue.
    • Route low-evidence or invalid states to clarification, abstention, a deterministic fallback, or human review.

    This is also a user experience decision. Open chat is useful when exploration is the job, but it transfers substantial planning and prompting work to the user. A structured flow is often better when the user follows a repeatable process under time pressure. One K-5 teacher assistant moved away from an initial chatbot concept toward a workflow aligned with how teachers select and assign lessons. The lesson is not that chat is inherently weak. It is that interface freedom should match task freedom.

    RAG needs the same product discipline. Retrieval can improve grounding, attribution, and freshness, but it introduces its own failure surface. The system can misunderstand the query, retrieve irrelevant material, miss the necessary record, use stale metadata, or generate a claim that its citations do not support. Treat retrieval as a product subsystem, not a box that makes hallucinations disappear.

    Evaluate at least four retrieval behaviors separately:

    • Query handling: Did the system represent the user’s actual intent?
    • Retrieval relevance: Did the returned set contain the material needed for the task?
    • Grounding: Does the generated claim follow from the retrieved material?
    • Absence behavior: When evidence is missing or conflicting, does the system say so and take the designed fallback?

    Attribution is not decorative. In workflows where users must verify an answer, provenance is part of the value proposition. A system can sound plausible and still lose trust if the user cannot determine where a consequential claim came from. That is why attribution and transparency became core requirements for conversational developer search.

    Agentic systems add another layer because the model chooses or sequences actions. Evaluate both the final result and the path used to reach it. A polished response can conceal an unnecessary tool call, an incorrect lookup, a failed write, or an action taken with the wrong scope. The more autonomy the agent has, the more important it becomes to evaluate it as a production workflow rather than as a text generator.

    For consequential actions, keep authorization and confirmation outside the model. Pass only the permissions needed for the current task. Validate tool arguments before execution. Make retried operations safe where possible, and require user confirmation before an irreversible or externally visible step. The model may propose an action; the product system decides whether that action is allowed.

    A useful trace connects the entire decision path:

    • Request context and relevant user or tenant configuration
    • Model, prompt, policy, schema, retrieval index, and tool versions
    • Retrieved records and their metadata
    • Intermediate decisions, tool calls, tool responses, retries, and validation results
    • Final output, fallback, or handoff
    • User action, correction, feedback, and downstream outcome

    Do not interpret instrumentation as permission to retain every raw input. Traces can contain personal information, confidential business data, or sensitive retrieved content. Decide what must be captured, redact where appropriate, limit access, and align retention with the product’s data policy. Real-world traces should enter evaluation workflows only with the necessary consent and redaction controls.

    Turn observed failures into a living eval suite

    An eval suite is not a large spreadsheet of impressive examples. It is an executable definition of the reliability contract. Its most valuable cases are usually the situations that reveal how the product fails: missing context, ambiguous requests, weak retrieval, conflicting instructions, malformed tool responses, policy pressure, domain edge cases, and plausible but unsupported output.

    Start with error analysis rather than dataset volume:

    1. Collect representative tasks from product discovery, support conversations, domain experts, and appropriately handled production traces.
    2. Review complete traces, not only final responses, and label the component where each failure entered the workflow.
    3. Group recurring errors into a taxonomy such as input, retrieval, generation, tool use, policy, presentation, and handoff.
    4. For each important failure mode, add a case with the input, relevant context, desired behavior, unacceptable behavior, rubric, and scorer.
    5. Run the case repeatedly against the current baseline and candidate system so variance and regressions are visible.
    6. Keep the case after the defect is fixed. A production failure should become a permanent regression test unless retaining it would create a data or privacy problem.

    A balanced dataset uses three kinds of evidence. Golden cases capture canonical tasks with carefully reviewed expectations. Targeted synthetic cases expand coverage for rare, risky, multilingual, adversarial, or not-yet-observed conditions. Real-world traces reflect how customers actually use and misuse the product. Combining these inputs keeps the suite grounded while giving it enough long-tail coverage.

    Synthetic data is useful for stress testing, but it should not be mistaken for evidence that the workflow succeeds with customers. Use it to probe a named hypothesis: a missing field, a language variation, a prompt injection attempt, a contradictory record, or an unavailable tool. Then check whether the generated case is realistic and whether its expected behavior is unambiguous.

    Choose the scorer based on the criterion rather than using an LLM judge for everything:

    • Code-based assertions are the default for schemas, required fields, valid identifiers, permissions, numerical bounds, forbidden content patterns, citations, and tool execution status.
    • Human or subject-matter-expert review is appropriate when correctness depends on domain context, consequences are high, or the rubric is still being discovered.
    • LLM-as-judge is useful for semantic criteria such as relevance, clarity, tone, and completeness when the rubric is explicit and the judge is calibrated against human-reviewed examples.

    An LLM judge is a measurement instrument, not ground truth. Give each criterion a concrete rubric and anchor examples. Compare the judge with human ratings, inspect disagreements, and avoid asking one prompt for an unexplained overall quality score. Separate judgments such as correctness, completeness, tone, and groundedness so a failure is diagnosable.

    Protect the evaluation process from leakage. If development examples, near-duplicates, or expected answers reach the system being evaluated, a strong score can be meaningless. Track data provenance, deduplicate related cases, keep release holdouts sealed from prompt tuning, and periodically introduce a blind set that the implementation team has not optimized against. Sudden unexplained metric gains should trigger a leakage check before celebration.

    Your continuous integration and continuous delivery pipeline should include failures as well as ideal examples. Known broken cases are especially valuable because they prove whether a proposed change repairs the actual weakness and whether a later change reintroduces it. A durable debugging loop turns concrete error modes into repeatable tests and keeps those tests in CI/CD.

    Do not set acceptance criteria from an arbitrary industry number. Derive them from the reliability contract. Hard constraints need blocking checks. Nuanced quality criteria need an agreed minimum and a comparison with the current baseline. Critical cohorts need their own view. Customer outcomes need production validation because an offline answer score cannot prove that the workflow saves time, resolves the issue, or improves a handoff.

    I use a strict decision rule: an improvement in average quality cannot cancel a hard-constraint regression. It also cannot hide a material decline for a consequential use case or customer cohort. This keeps the release conversation focused on risk and user value instead of a single blended score.

    Make every release reversible, observable, and owned

    An AI release is a versioned system change. The candidate is not just a model name. It is the combination of model, prompt, orchestration, retrieval configuration, index or corpus, tool definitions, output schema, guardrails, interface, and fallback. If any part changes, the affected behavior needs evaluation.

    Use a release sequence that makes uncertainty visible:

    1. Freeze and identify the complete candidate configuration so results can be reproduced.
    2. Run deterministic checks and the relevant offline eval suite against both the candidate and the production baseline.
    3. Inspect results by failure mode, workflow, risk level, language, and important customer cohort rather than relying on the aggregate.
    4. Review changed failures manually, including cases where a score improved for the wrong reason.
    5. Use a shadow deployment when feasible to observe real inputs without letting candidate outputs affect customers.
    6. Roll out behind a feature flag or equivalent control, beginning with a bounded population and a tested fallback.
    7. Expand only when customer outcomes, quality signals, hard constraints, latency, cost, and handoff behavior remain inside the contract.

    Shadow and staged releases do different jobs. Shadowing reveals how a candidate behaves on realistic traffic without placing it in the customer path. A staged rollout reveals how users respond and whether downstream outcomes improve. Neither replaces the other, and neither replaces offline evals.

    The production dashboard should preserve the same layers used in the reliability contract. Track the outcome the feature exists to improve, quality indicators derived from sampled traces, hard-constraint events, abstentions and human handoffs, retrieval and tool failures, latency, cost, and the rate at which customers correct or abandon the result. A metric that cannot lead to a diagnosis or decision does not deserve prominent dashboard space.

    Maintain a persistent failure log. For each failure mode, record the affected workflow, severity, observed frequency, confidence in the evaluator, likely component, owner, mitigation, linked eval cases, and before-and-after evidence. Severity tells you what the failure can do. Frequency tells you how often customers encounter it. Evaluator confidence tells you whether the signal is trustworthy enough to drive a roadmap decision.

    Assign ownership to the product system

    Reliability will decay if everyone owns a fragment and nobody owns the outcome. Product management should own the user promise, outcome metrics, risk decisions, and release tradeoffs. Engineering should own reproducibility, validation, tracing, deployment controls, and recovery. Domain experts should help define correctness and adjudicate difficult cases. Legal, privacy, security, and support should shape constraints and escalation paths where their responsibilities apply.

    Operational ownership also needs a change policy. A model upgrade, prompt edit, new tool, schema change, retrieval-index refresh, policy update, or new customer segment can move behavior. Specify which evals run, who reviews the result, what blocks release, how the system is rolled back, and which stakeholders are notified. Prompts, data pipelines, rubrics, and guardrails are living product assets, and ongoing maintenance is part of the cost of the feature.

    Finally, define a stop condition. More orchestration cannot rescue every product idea. If the system cannot meet the user’s quality bar, if the fallback consumes the supposed efficiency gain, or if the differentiated value lies elsewhere, the responsible decision may be to narrow or end the feature. Stack Overflow sunset conversational search when it could not meet developer expectations and redirected attention toward a stronger data opportunity. Reliability work should improve a viable product, not make sunk cost harder to confront.

    Key takeaways

    • Define reliability as a user outcome, an operating envelope, hard constraints, and a safe failure path.
    • Keep probabilistic generation inside a bounded workflow with structured inputs, validation, permissions, and fallback.
    • Evaluate retrieval, generation, tools, and handoffs separately so the team can locate a failure instead of merely scoring it.
    • Build the eval suite from golden cases, targeted synthetic scenarios, and appropriately handled production traces.
    • Use code for deterministic requirements, calibrated judges for semantic criteria, and domain experts where context or consequence demands them.
    • Version the whole system, gate releases against the current baseline, roll out reversibly, and turn every meaningful production failure into a regression test.

    At your next roadmap review, take one live AI workflow and complete its reliability contract. Name the most consequential unresolved failure, add a trace that makes it diagnosable, convert it into an eval, and set the release rule. If you cannot describe what the product does when that case fails, the feature is not ready to scale.

    References

  • Deliberate Practice for Product Teams: How AI and On‑Demand Learning Unlock Mastery

    Deliberate Practice for Product Teams: How AI and On‑Demand Learning Unlock Mastery

    I recently tuned into a powerful conversation where Petra Wille sits down with Teresa Torres to unpack a major shift in product learning: moving from purely instructor-led cohort courses to offering on-demand options. As someone leading product management at HighLevel, I’ve wrestled with the same trade-offs—how to scale product discovery skills without compromising depth, community, or outcomes—and this discussion hit home.

    What stood out immediately is how Teresa shares why she resisted on-demand for so long, how deliberate practice has always been at the heart of her teaching, and what finally changed her mind. That framing matters. In my experience, deliberate practice is the backbone of real capability building: clear goals, targeted reps, tight feedback loops, and sustained reflection. It’s how we turn continuous discovery from a concept into a craft product teams can reliably execute.

    We also dug into the trade-offs between cohort-based vs. on-demand learning. Cohorts bring structure, accountability, and shared language—critical for team-based behavior change. On-demand learning offers flexibility, reach, and just-in-time reinforcement—key for busy product managers, designers, and engineers balancing roadmaps and research. The challenge is not choosing one over the other, but architecting a blended learning system that preserves the rigor of cohorts while using on-demand to extend practice, sustain momentum, and meet learners where they are.

    That’s where technology becomes a force multiplier. From AI-powered interview coaches to microlearning formats, we explored how AI can support behavior change and skill building without losing the human element. I’ve seen the same in my teams: when AI provides structured, rubric-based feedback on interviews, assumptions, or opportunity framing, people get expert-quality guidance at scale. Used well, this shortens the feedback cycle and increases the number of high-quality reps—without displacing peer critique or expert coaching.

    Microlearning and problem sets deserve special attention. Short, focused practice—think “Duolingo” for product discovery—helps teams internalize patterns like crafting unbiased interview prompts, distinguishing signals from stories, or iterating on interview flow. Combined with spaced repetition, these formats build muscle memory for critical skills, so discovery doesn’t stall the moment the cohort ends. In other words, on-demand isn’t a downgrade; with the right scaffolding, it can be a durability upgrade.

    Equally important, why AI should augment—not replace—human connection in discovery. No model can substitute for the trust you build with customers, the judgment you develop through messy real-world conversations, or the creative tension of team debate. My takeaway: use AI to accelerate preparation, evaluation, and deliberate practice; rely on humans for empathy, ethics, sense-making, and decision quality.

    If you’ve ever wondered how to balance flexibility, structure, and deliberate practice in product learning—or you’re just curious how AI might reshape how we build skills—this conversation is for you.

    Listen to this episode on: Spotify | Apple Podcasts

    Explore the resources and links mentioned: Follow Teresa Torres: https://ProductTalk.org; Follow Petra Wille: https://Petra-Wille.com; Product Talk Academy; Continuous Interviewing course by Teresa Torres; Story-Based Customer Interviews On Demand course by Teresa; Customer Recruiting for Continuous Discovery On Demand course by Teresa; Duolingo; Teresa’s Interview Coach; AI as a Strategic Thought Partner with UX Implications podcast episode; Teresa’s socials: X, LinkedIn, Youtube, Product Talk Blog.

    I’d love to hear your perspective. How are you blending cohort-based learning, on-demand practice, and AI coaching on your product teams? Drop your thoughts in the comments—let’s compare notes on what’s working.


    Inspired by this post on Product Talk.


    Book a consult png image