Tag: forward deployed engineers

  • Goal-Setting for AI Products: How I Plan, Prioritize, and Confidently Ship in a Nonlinear GenAI World

    Goal-Setting for AI Products: How I Plan, Prioritize, and Confidently Ship in a Nonlinear GenAI World

    I build and ship AI products in an environment where the frontier changes weekly, so my planning system has to be adaptive, evidence-driven, and unapologetically outcome-focused. In this piece, I share the frameworks I use to set goals for generative AI, balance research with product execution, and scale responsibly — drawing sharp lessons from one of the most influential applied AI companies operating today.

    Consider Runway, an applied AI research company shaping the next era of art, entertainment, and human creativity. Runway has raised $237m and was one of Time Magazine’s “100 most influential companies” in 2023. Runway has been a persistent viral sensation in recent years, and is behind many of the most famous AI demos online.

    The earliest stages of an AI company often begin with research breakthroughs, scrappy prototypes, and clever distribution. In practice, that means leveraging containerization (https://aws.amazon.com/what-is/containerization/) and Docker (https://www.docker.com/) to package models reproducibly, showcasing work where practitioners already gather — Hugging Face (https://huggingface.co/), Hugging Face Spaces (https://huggingface.co/spaces), and Hugging Face Model Hub (https://huggingface.co/docs/hub/models-the-hub) — and tapping infrastructure like Replicate (https://replicate.com/) to get demos into people’s hands. Early, magical use cases — like the Green screen tool by Runway (https://runwayml.com/green-screen/) — teach us which problems are both technically feasible and viscerally valuable.

    I’ve learned to be cautious about “The limitations of being “customer-driven” when building in AI”. Traditional product discovery assumes needs are legible and solutions are relatively deterministic. In generative AI, user desire often follows model capability, not the other way around. The job is to triangulate: run tight user loops to validate perceived value, instrument objective model quality, and explore novel interaction patterns that customers can’t yet articulate. I treat this as a portfolio of discovery bets — some customer-led, some capability-led, all evaluated against clear outcome thresholds.

    Balancing research development with product development requires organizational design that prevents context-switching tax while preserving velocity. I pair research pods with product pods, supported by forward deployed engineers and domain PMs who translate evaluation metrics into user-visible milestones. Safety and content moderation sit on the critical path, not as afterthoughts — think policy definition, classifier tooling, abuse red teaming, and clear escalation playbooks. This balance is how you move from a great demo to a dependable product without losing momentum.

    Goal-setting amidst constant change in AI starts with outcomes vs output OKRs. I write OKRs in terms of user impact and model performance thresholds — for example, target ranges for latency, quality scores against a golden dataset, or creator retention — then let teams choose the highest-leverage outputs (data pipelines, fine-tuning, UX improvements) to get there. Why I don’t plan very far ahead: I treat the annual view as a vision and bet map, the quarterly view as a constrained slate of outcomes, and the 6–8 week cycle as the execution heartbeat. AI roadmaps are hypotheses; evaluation harnesses and launch gates are the truth.

    Community is a force multiplier. Forming a vocal community and fostering community requires real access and real listening: early release cohorts, office hours, and transparent changelogs. How they picked users for early release matters — diversity of use cases, sophistication of workflows, and willingness to give crisp feedback. Expanding past the first 100 users of Gen-2 demands readiness: evaluation parity across modalities, scalable infra, and safety coverage. Done well, this motion compounds learning while building authentic advocacy.

    For founders, my advice echoes the core lessons above. Start with a narrow, high-intent wedge and prove durable value fast; let founder-led GTM compress the feedback loop; instrument everything from day one; and resist the urge to over-plan features before you’ve nailed outcomes. Product-market fit lessons in AI often arrive via small, fast experiments — not grand, long-range plans. Ship thin slices that demonstrate unmistakable value, then iterate toward a system, not a single feature. When in doubt, shorten the loop and improve the evaluation harness.

    People often ask: Will AI replace video editors? My view is that AI will replace zero editors who master these tools — and many who don’t. The winners blend taste, storytelling, and generative leverage. The products we build should honor this reality: design for control, iteration, and co-creation, not just automation.

    If you’re mapping the progression of tech and use-cases, a few public references are instructive: Runway Gen-1 (https://research.runwayml.com/gen1) and Runway Gen-2 (https://research.runwayml.com/gen2) show how capability unlocks new workflows and demand. Runway’s 30 AI Magic Tools (https://runwayml.com/ai-magic-tools/) illustrates portfolio thinking — a suite of composable powers rather than a monolith.

    For builders focused on gen ai for product prototyping through production: keep your demo muscle strong, your evaluation stronger, and your outcomes strongest. Invest in community, treat safety as a feature, and let your OKRs steer what ships — not the other way around.


    Book a consult png image
  • The Human Side of Engineering Leadership: Practical Plays to Build Creative, High-Performing Teams

    The Human Side of Engineering Leadership: Practical Plays to Build Creative, High-Performing Teams

    Engineering leadership is a human sport first and a technical sport second. The longer I lead product and engineering teams, the more convinced I am that world-class execution comes from clarity, trust, and an environment designed for deep work — not just clever architectures or more process. In my role leading product, I see the biggest unlocks happen when we combine operational excellence with empathy and purpose.

    My “utopia” — where engineers have time to create and invent — starts with sacred focus time. I protect no-meeting blocks, design sprints that include exploration, and carve out recurring capacity for prototypes and technical debt. When builders know they have sanctioned space to think, we get more product discovery, better ideas, and fewer last-minute heroics.

    Shipping software at scale is difficult, and it’s harder to ship today than ever before. Complexity from microservices, compliance, security, platform fragmentation, and AI-driven surface area expands every quarter. The counter is operational hygiene: clear ownership, ruthless scope, a predictable release cadence, excellent tooling, and a culture that values outcomes over activity.

    What makes a startup operationally sound is surprisingly simple to describe and hard to do consistently. Define decision rights, keep teams small and mission-aligned, instrument everything, and ship on a reliable train. Feature flags, dark launches, automated testing, and crisp rollbacks turn risk into routine. Most importantly, we write things down — intents, constraints, and success metrics — so execution scales beyond any single leader.

    Product managers can dramatically improve engineering culture. The fastest way is through precision: sharp problem statements, explicit success metrics, clear acceptance criteria, and honest trade-offs. I hold our team to “outcomes vs output OKRs,” framing goals by customer and business value rather than task volume. PMs should also shield makers from thrash, resolve ambiguity quickly, and bring real users into the room early and often.

    From an engineer’s perspective, good product management sounds like respect for the craft. We acknowledge performance budgets, technical constraints, and the hidden cost of complexity. We ask for estimates responsibly, show our work in decision docs, and make room for forward deployed engineers to close the loop with customers. When PMs consistently do these things, trust grows — so does velocity.

    The role of product compared to design and engineering is easy to state and easy to forget: product owns the why and what, design owns how it feels, engineering owns how it works, and all three own the outcome. I treat the PM job as system optimization across functions — removing friction, sequencing bets, and maximizing learning per unit of time. When incentives are aligned to shared outcomes, handoffs turn into collaboration.

    Micromanagement kills creativity. Declarative versus prescriptive leadership is the antidote: set the intent, define the constraints, agree on the measures of success — then get out of the way. I replace step-by-step tickets with a one-page brief and a weekly demo cadence. Guardrails create safety; autonomy creates ownership; together they create better software.

    I foster a debate culture by making disagreement safe and productive. We write RFCs, invite dissent, time-box decisions, and “disagree and commit” when the window closes. Good debates chew on assumptions, not people. The payoff is compounding judgment and a team that can argue well without leaving scars.

    Three leadership ideas I practice every week: first, default to clarity — ambiguity is the silent killer of execution. Second, manage energy, not just time — sustain the team’s battery with realistic pacing and visible wins. Third, train judgment — distinguish one-way doors from two-way doors and match decision speed to reversibility.

    Understanding employee motivation is a superpower. People move for different reasons — mastery, autonomy, purpose, progression, compensation, recognition. I map motivations explicitly in 1:1s and shape work accordingly. When someone’s day-to-day aligns with what they value most, performance and retention both rise.

    My advice on discovering what motivates people is straightforward: ask better questions and observe the work. I love asking, “What feels like play to you but looks like work to others?” I rotate responsibilities to run small experiments, then codify what sticks in growth plans. Motivation is dynamic; treat it like a product you’re constantly rediscovering.

    On org design, I review team topology every six months. Strategy changes and customer needs evolve; our organization should, too. I favor lightweight reorgs — adjusting missions and interfaces rather than wholesale reshuffles — and I use rotations to refresh learning without destabilizing roadmaps. The aim is responsiveness without chaos.

    One habit I see in successful leaders is relentless focus on outcomes. We inspect impact, not activity, using transparent scorecards, weekly business reviews, and OKR hygiene that spotlights real progress. When outcomes are the north star, prioritization becomes sane and teams feel meaning in their work.

    Sound judgment is crucial for decision-making. I separate reversible from irreversible choices, run pre-mortems, and keep a decision log so we can learn at the portfolio level. Informed speed beats perfect slow, and fast follow-ups are a feature, not a bug.

    Crystallized lessons from large-scale software environments keep proving true: cadence beats heroics, observability pays for itself, and investment in developer experience is the highest-ROI platform bet you can make. Also, write it down — clear artifacts are how complex systems learn together.

    I stay wary of becoming irrelevant. The antidote is shipping, curiosity, and deliberate learning — talking to customers weekly, pairing with engineers, and letting junior talent teach me new tools and patterns. Relevance is earned every quarter.

    If I had to pick one leadership lesson, it’s this: people remember how you made them feel. Trust, candor, and consistency create the conditions for excellence. The best strategy in the world won’t move if people don’t feel seen, safe, and challenged.

    I’ve changed my mind on a few big things: more documentation beats more meetings, hybrid can outperform in-office with the right rituals, and smaller, more frequent reorganizations are better than large, rare ones. I used to think speed and quality were a trade-off; now I think clarity gives you both.

    My growth has been shaped by thoughtful operators, designers, and engineers who taught me to balance ambition with stewardship. Their fingerprints are on these practices, and my teams’ results are better for it. The human side of engineering leadership isn’t soft — it’s the hard edge that makes everything else work.


    Book a consult png image
  • Scaling Enterprise AI That Sells: Battle-Tested Playbooks for PMF, Champions, and Agentic AI

    Scaling Enterprise AI That Sells: Battle-Tested Playbooks for PMF, Champions, and Agentic AI

    Enterprise AI is exhilarating and unforgiving. I’ve seen gorgeous demos fall apart under real-world constraints and seemingly modest pilots unlock outsized value at scale. In this reflection, I share the playbooks I rely on to build, scale, and sell generative AI in the enterprise—what actually moves deals, secures product-market fit, and sustains trust with the C-suite and the front line.

    Why is it so difficult to scale AI products for enterprise? Because the bar is higher on every dimension: data governance, security, extensibility, integration depth, reliability, and measurable ROI. An enterprise-grade, full-stack generative AI platform isn’t just a model; it’s the surrounding system—observability, evals, policy, workflow, and human-in-the-loop—that makes outcomes predictable, auditable, and safe. The fastest path to adoption is simple: deliver on-brand, on-policy content and decisions using a customer’s first-party data, and prove that quality holds up under load.

    My north star is dependability over demo magic. The number one challenge is making model output dependable across messy, high-variance enterprise inputs. I build an evaluation harness early, with gold datasets, task-specific metrics, and human adjudication. Every change ships behind guardrails and is measured against cost, latency, and quality SLOs. When governance, change management, and procurement show up (they always do), I treat them like first-class product requirements, not hurdles.

    Champions are the secret to winning complex accounts. I map the org, find operators who feel the pain daily, and quantify that pain in dollars and hours. Then I define success criteria upfront—time-to-value in under 30 days, measurable uplift (e.g., deflection, conversion, cycle time), and a plan for scale. I deploy forward deployed engineers alongside the business to co-design workflows, refine prompts and evaluators, and document before/after outcomes. Champions don’t just approve pilots; they co-author the business case and defend it.

    To win the enterprise, trust architecture matters as much as model architecture. I lead with clear answers on data residency, encryption, SSO, RBAC, DLP, and retention policies; I address whether customer data trains models, default behaviors, and opt-in controls. I offer flexible deployment (VPC or private networking when needed), transparent pricing, and SLAs with real teeth. I also integrate where work already happens—CRM, help desk, knowledge bases—so value shows up in the flow of work.

    Signs of healthy product-market fit are unmistakable: pull from lookalike buyers, multi-threaded expansions, champions who present results internally without me in the room, and usage that moves from experimentation to business-critical. I watch for weekly active usage above pilot thresholds, POCs converting to multi-year deals, and adjacent teams (Support, Marketing, Legal, RevOps) asking to onboard with minimal push. PMF feels less like persuasion and more like coordination.

    Scaling large language models for specific use cases requires ruthless focus. I constrain scope to tightly defined workflows, pair retrieval with structured knowledge, and mix model strategies (base models, fine-tunes, tools, and function calling) based on cost and latency budgets. I codify policy-as-code and deploy guardrails at the orchestration layer, not just the model layer. Continuous evaluation—both automatic and human—is the heartbeat of quality.

    My advice to AI founders in 2024 is pragmatic. Start with outcomes, not demos. Establish outcomes vs output OKRs that tie directly to revenue, cost, risk, or customer experience. Use gen AI for product prototyping to shorten discovery cycles, but graduate quickly to instrumented workflows in production. Align early with InfoSec and Legal; your speed will be gated by trust, not code. And when in doubt, ship smaller, safer increments faster.

    Healthy co-founder relationships look the same across winning companies: clear domains, fast escalation, and a shared appetite for “disagree and commit.” I keep a decision log, time-box debates, and make moments-of-truth visible to the team and board. You’ll know it’s working when you have more energy after hard conversations than before.

    The future of agentic AI is deeply enterprise: multi-agent workflows that plan, act, and verify with human oversight where it matters. The winners will combine reasoning, tool use, retrieval, and policy with audit trails that satisfy compliance while keeping velocity high. Think of it as re-engineering business processes around AI-native steps, not sprinkling AI on top of legacy workflows.

    Culture turns strategy into reality. I anchor my teams on “connect, challenge, and own.” Connect means obsess over the customer problem and internal alignment. Challenge means we red-team our ideas, run experiments, and measure impact. Own means we accept outcomes, not just output, and we iterate until the business moves. This is how a customer support ai strategy becomes a durable moat, not a slide.

    If you’re a product creator or product management leader, the above playbooks are meant to be lifted and adapted. Start where the pain is loudest, quantify the win, and let champions carry the story. The compound interest of disciplined product discovery, strong governance, and relentless evaluation is a generative AI business that sells itself—and scales.


    Book a consult png image
  • DevTools at Scale: Hard-Won Lessons on PMF, AI, and Culture from Apple, AWS, Microsoft

    DevTools at Scale: Hard-Won Lessons on PMF, AI, and Culture from Apple, AWS, Microsoft

    Building and scaling DevTools has taught me that world-class culture and relentless product focus are non-negotiable. Drawing on experiences across Amazon, Apple, and Microsoft—and hard-won lessons from startups like Unblocked and Buddybuild—I’m sharing the principles I rely on to ship great developer products at scale.

    Why building for developers is different: developers are discerning, allergic to friction, and quick to churn if the DX isn’t exceptional. That means fast setup, clear docs, ergonomic APIs, sane defaults, and deep integrations with GitHub, GitLab, Bitbucket, Confluence, AWS, and Microsoft Azure.

    I benchmark teams against gold-standard platforms like Stripe, Twilio, and Looker—tools that reward mastery, never bury the lede, and make success observable in minutes, not days.

    From the early days of Buddybuild, the signal was unmistakable: remove toil from CI/CD, shorten feedback loops, and teams will expand usage without a sales nudge. The pattern holds across DevTools: when time-to-value approaches zero, the product sells itself.

    Early signs of product market fit: organic team-to-team adoption, repeatable setup success, contribution from power users, and inbound demand you cannot keep up with. When these show up, “Why great product is everything” stops sounding like a platitude and starts reading like a P&L.

    Monetizing product market fit is straightforward if you align value and pricing units. Seat-based maps to collaboration; usage-based maps to compute, API calls, or storage; hybrid models reduce edge-case friction. Keep the packaging simple and double down on “The power of positioning.”

    AI is complicating product market fit. Gen AI accelerates gen ai for product prototyping, but it also introduces instability: model drift, hallucinations, and evaluation blind spots. I build an evaluation harness, human-in-the-loop review for risky flows, and a clear customer support ai strategy before scaling.

    Being customer-obsessed is the moat. I embed forward deployed engineers with key customers to translate real workflows into product decisions, close the empathy gap, and validate behavior in production environments.

    On decision-making, I blend product discovery with crisp documents and measurable bets: PRFAQs or design docs to clarify intent, guardrails in analytics, and outcomes vs output OKRs to keep teams aligned to impact.

    Unblocked, a developer tool that lets you talk to your codebase, points toward a future where code search, context, and refactoring converge into conversational workflows. I’m bullish on the pattern, but I stay sober about failure modes and cost-to-serve.

    Here’s my cautious take on AI: latency, privacy, and provenance matter as much as model quality. The best teams treat prompts as product, training data as liability, and evaluation as a first-class release gate.

    Hiring is where many teams stumble. Don’t over-index on competency when hiring. I optimize for learning velocity, ownership, and kindness under pressure. Competency scales output; character scales organizations.

    As a second-time founder and operator, I treat mental health like uptime. I schedule recovery, define non-negotiables, and surround myself with peers who normalize the hard days. Burnout is a systems failure, not an individual weakness.

    I don’t do demos. I prefer self-serve trials with instrumented onboarding, sample projects, and guardrails that let the product do the talking. If a prospect can’t succeed in 15 minutes, we fix the product, not the deck.

    On customer feedback, I separate noise from signal with cohorts and context. I prioritize requests that reduce time-to-value, unblock integrations, or meaningfully expand the surface area of successful use cases. That’s how to deal with customer feedback without losing strategic focus.

    To build and scale DevTools, keep the bar high and the loop tight: ship small, watch usage, learn fast. Invest in platform reliability, rock-solid SDKs and CLIs, and a developer experience that earns trust release after release.

    Resources and touchstones I revisit often:

    Apple’s acquisition of Buddybuild: https://www.cnbc.com/2018/01/02/apple-agrees-to-buy-buddybuild.html

    AWS: https://aws.amazon.com

    Bitbucket: https://bitbucket.org

    Confluence: https://www.atlassian.com/software/confluence

    GitHub: https://github.com

    GitLab: https://gitlab.com

    Looker: https://looker.com

    Microsoft Azure: https://azure.microsoft.com

    Stewart Butterfield: https://www.linkedin.com/in/butterfield/

    Stripe: https://stripe.com

    Twilio: https://twilio.com

    Unblocked: https://getunblocked.com/

    If you’re building for developers, stay ruthless about simplicity, respectful of their time, and obsessed with proof in production. That’s how durable product-market fit is earned—and monetized.


    Book a consult png image
  • From Prototype to the Pentagon: My Playbook for Winning DoD Customers and Mission Fit

    From Prototype to the Pentagon: My Playbook for Winning DoD Customers and Mission Fit

    I’ve spent years building dual-use products and partnering with teams navigating the Department of Defense. In this piece, I share how I move from prototype to program of record by aligning product strategy to real mission outcomes, building trust with end users, and translating commercial product rigor into the national security context. Commercial versus military market strategies require fundamentally different assumptions. In the private sector, we obsess over product-market fit and velocity; in defense, we obsess over “mission solution fit” and survivability in procurement. The buyer is a complex web—operators, program managers, contracting officers, and Program Executive Offices (PEOs)—and each needs a clear value story tied to mission impact, not just features or ARR. When I validate ideas for defense products, I start with deep discovery at the edge: talking with operators, understanding tactics, techniques, and procedures (TTPs), and quantifying what “better” looks like in their environment. The “Mission Model Canvas” helps me capture stakeholders, beneficiaries, and constraints that don’t exist in a typical SaaS motion. “Hacking for Defense” has been invaluable for structuring this discovery and ensuring we test assumptions against mission reality, not just market appetite. A practical guide to military sales and procurement starts by mapping decision pathways. I identify the PEOs, the Program Managers, and the acquisition timelines that govern transition. I treat each step like an enterprise sale with additional layers: requirements, testing, accreditation, and budgeting. I align demonstrations to mission milestones, ensure my roadmap accounts for integration and accreditation lead times, and keep decision-makers looped with concise, evidence-based updates. Rethinking go-to-market strategy for defense means planning for longer cycles and multi-level consensus. Instead of a simple funnel, I build a coalition: an operator champion for pilots, a program sponsor for funding continuity, and a contracting route that matches how the customer buys. The goal is to de-risk adoption across technical, operational, and procurement dimensions in parallel. Building a network in national security is a full-contact sport. I invest time in the field, put forward deployed engineers next to users, and show up at the training ranges and labs where problems are real. Trust accumulates when teams see you adapt quickly, respect constraints, and demonstrate an understanding of mission risk. That trust turns into access and, ultimately, into pull from the organization. The dual-use debate isn’t binary—it’s a portfolio decision. I’ve seen teams succeed by leading with a defense wedge when the problem is uniquely military, and others start commercially to prove traction then tailor for defense. The key is to avoid whiplash: design your architecture and compliance posture so you can serve both without fragmenting your roadmap. Behind the rise of a new generation of “defense founders” is a shift in ambition and capability. Teams are mission-driven, technically sophisticated, and comfortable operating in complex stakeholder environments. They’re building for hard problems and measuring success in operational outcomes, not just revenue milestones. “Mission solution fit” is my north star. I define it as measurable mission improvement with acceptable changes to TTPs, training, and integration. I seek evidence that units can and will use the solution under realistic constraints, that it interoperates with existing systems, and that program leadership can fund it at scale. When those signals align, transition becomes possible. Breaking new ground in military tech often means navigating institutional friction. The “The Frozen Middle” is real—layers that resist change even when leadership and operators are aligned. I plan for this by prototyping where adoption barriers are lowest, securing a senior sponsor, and demonstrating cost, schedule, and performance wins that the middle cannot ignore. The hidden challenges most startups miss tend to be non-technical. Security and accreditation aren’t documentation exercises; they’re product constraints that should shape architecture early. Interoperability isn’t a feature; it’s table stakes. And your ability to explain “why now” in the language of budget cycles can matter as much as a benchmark. Essential resources for any defense founder include “Hacking for Defense,” “The Hacking for Defense Manual,” the directory in “How to find your customer in the Dept of Defense,” the “Mission Model Canvas,” and lessons from “The lean launchpad at Stanford” and “The Secret History of Silicon Valley.” I also draw on the work of Alexander Osterwalder and Eric Ries to bridge discovery, iteration, and disciplined scale. What’s missing from Silicon Valley in this domain is patience paired with rigor. The best teams combine world-class product discovery with respect for acquisition realities. They instrument outcomes in the field, align roadmaps to funding gates, and bring forward deployed engineers to close the gap between prototype and operational capability. From prototype to the Pentagon is a repeatable path when we hold ourselves to mission outcomes, build coalitions across the acquisition chain, and design for constraints from day one. If you’re committed to national security, build with empathy for the operator, clarity for the buyer, and a roadmap that survives contact with procurement.
    Book a consult png image
  • Inside Clay’s $1.25B Playbook: Unconventional GTM, Pricing Strategy, and Enterprise Wins

    Clay’s path to a $1.25B valuation isn’t conventional—and that’s exactly why it’s instructive. Through the lens of product management and go-to-market strategy, I break down how unconventional tactics, rigorous pricing decisions, and a long game on brand combined to create real upmarket momentum. If you lead product, growth, or revenue, there’s a repeatable playbook here for blending product-led growth with enterprise sales without losing speed or signal. Varun Anand is the co-founder and Head of Operations at Clay, a GTM development environment that combines data and AI to help over 5000 companies power everything from CRM enrichment to highly targeted outreach campaigns. Clay recently announced their Series B expansion, raising $40M at a $1.25B valuation. Before Clay, Varun was the Director of Operations at Newfront and the Head of Expansion at Candid. Varun also spent four years working on Hillary Clinton’s presidential campaign. Turning traditional GTM on its head, Clay’s earliest traction didn’t come from glossy campaigns—it came from scrappy sales tactics: “WhatsApp groups, Reddit threads, and reverse demos.” I’ve seen this play repeatedly outperform paid channels early because it compounds social proof in the exact communities where power users congregate. When your ICP hangs out in niche threads, customer acquisition is a function of credibility, not CPM. On pricing, “credit-based pricing” was a pivotal decision. Equally important, the team “rejected the usage-based model.” For PLG plus enterprise, this matters: credits make value legible to buyers, reduce billing anxiety for ops and finance teams, and align with predictable, budgeted workflows. In my experience, credit models also create clearer upgrade paths when your product spans multiple use cases. Clay built a robust self-serve engine and then layered “enterprise customers on top of PLG.” This sequencing avoids the trap of hiring an enterprise team before the product is self-serve-proven. It also creates cleaner handoffs—self-serve for discovery and activation, sales for proof, procurement, and expansion. Content and brand weren’t afterthoughts. Clay made a “big bet on content” and “invested in brand from day-one.” That’s a contrarian move many teams delay, but content accelerates learning loops, reduces sales cycle time, and scales enablement far beyond headcount. In enterprise sales, a trusted brand is an asset class. Winning big accounts required creative proofs of value. “Reverse demos” flipped the script—show the customer’s data, in their workflow, with their outcomes. It’s one of the fastest routes to de-risking adoption and building trust with enterprise buyers. From there, they applied a pragmatic “land and expand model” that aligns with how large organizations actually buy. Clay highlights “3 changes that unlocked Clay’s upmarket motion.” While every company’s inflection points are unique, the meta-lesson is consistent: clarify the ICP, operationalize proof (reverse demos, ROI), and meet enterprise expectations on reliability, governance, and support—without sacrificing the PLG engine. Team construction was equally intentional. Hiring people who are “technical enough” and using a “hands-on interviewing process” raised the talent bar and reduced execution drag. I’ve found this mirrors the strength of forward-deployed mindsets: product, ops, and GTM talent who can prototype, troubleshoot, and translate customer complexity into scalable systems. Finally, Clay’s contrarian take on compensation signals a willingness to design incentives for the business they want to build, not the one the market expects. Compensation philosophies quietly shape culture, velocity, and who opts in. Referenced: Anthropic: https://www.anthropic.com/ Clay: https://www.clay.com/ Clay’s Series B expansion: https://www.clay.com/blog/series-b-expansion Eric Nowoslawski: https://www.linkedin.com/in/outboundphd/ Figma: https://www.figma.com/ Jesse Ouellette: https://www.linkedin.com/in/jesseoue/ Kareem Amin: https://www.linkedin.com/in/kareemamin/ Nick Merrill: https://www.linkedin.com/in/nick-merrill-64562310/ Notion: https://www.notion.com/ Oyster: https://www.oysterhr.com/ Pave: https://www.pave.com/ Rippling: https://www.rippling.com/ Snowflake: https://www.snowflake.com/ Verkada: https://www.verkada.com/ Webflow: https://webflow.com/ Yash Tekriwal: https://www.linkedin.com/in/yashtekriwal/ My takeaway: this is a modern GTM blueprint—prove value in the wild, price for clarity, build self-serve first, then industrialize trust for enterprise. Do that, and you can scale without losing the product signals that got you traction in the first place.
    Book a consult png image
  • Mastering AI Evals: Real-World Discovery Tactics to Ship Quality, Safe, Reliable AI

    Mastering AI Evals: Real-World Discovery Tactics to Ship Quality, Safe, Reliable AI

    I’ve been shipping GenAI features long enough to know that clever prompts and orchestration aren’t enough. What actually matters is evidence: Does the system work, for whom, and under what conditions? That’s where rigorous AI evals come in—the backbone of building reliable, safe, and continuously improving AI products.

    In a recent conversation focused entirely on evaluation, I dug into what “evals” mean in the AI/ML world, why they’re more than just quality assurance, and how to operationalize them end to end. If you want to explore the discussion, listen on Spotify: https://open.spotify.com/episode/7mSiEGSYNO4sXeGAVTJO4V or Apple Podcasts: https://podcasts.apple.com/kh/podcast/ai-evals-discovery/id1794203808?i=1000727980774. There’s also a video version on YouTube: https://www.youtube.com/watch?v=pfSIQMrWhQE.

    Here’s how I frame evals with my teams. First, define the behavior you want to see in terms real users care about. Then codify that intent as tests that run consistently. I distinguish between golden datasets, synthetic data, and real-world traces. Golden datasets capture canonical examples that represent “ground truth.” Synthetic data fills important gaps quickly and safely. Real-world traces keep you honest and reflect evolving usage.

    The most durable loop I’ve found is simple: identify error modes, turn them into evals, and automate. This is where error analysis pays off. Some checks should be purely deterministic—code-based checks that evaluate structured outputs, schemas, or policies. Others benefit from LLM-as-judge when human-like judgment matters, as long as you calibrate and continuously verify those judges with spot checks and inter-rater agreement.

    Discovery practices should inform every evaluation step. If you’re doing “Story-Based Customer Interviews,” you can derive realistic scenarios, acceptance criteria, and edge cases directly from user narratives. That context sharpens the evals and prevents you from overfitting to toy problems or proxy metrics that don’t reflect user value.

    Evals require ongoing care and feeding. Criteria drift is real—what counted as “good” six weeks ago may not satisfy users after you ship a new capability or your audience evolves. I treat the eval suite like living product infrastructure: versioned, reviewed, and owned. When we change prompts, models, or retrieval strategies, the evals run first, then we examine deltas, regressions, and surprises before anything reaches production.

    Guardrails and human oversight work hand-in-hand with evals. Guardrails enforce non-negotiables (safety, privacy, compliance), while evals measure progress against nuanced goals (relevance, helpfulness, tone). In high-stakes workflows, I combine pre-deployment evals, runtime guardrails, and spot human review. The goal isn’t to eliminate humans; it’s to focus their attention where judgment and context matter most.

    Practically, I start with a minimal eval harness that standardizes inputs and outputs—often in JSON (JavaScript Object Notation)—and writes repeatable tests. I maintain a small golden dataset, add targeted synthetic data for coverage, and stream real-world traces into the suite once we have consent and redaction in place. For subjective criteria (e.g., tone, helpfulness), I layer in LLM-as-judge with calibration. For objective checks (e.g., schema validation, policy compliance), code-based checks are my default.

    Tooling evolves quickly, but the principles hold. Whether you’re working with Anthropic or experimenting with V0 or Lovable in your prototyping stack, the eval loop stays the same: define success, test it the same way every time, and close the loop with learning. If you’re a product creator or leading forward deployed engineers, this discipline accelerates gen ai for product prototyping without sacrificing safety or quality.

    I also tie evals to outcomes vs output OKRs. Instead of “ship three prompts,” we commit to measurable outcomes like resolution rate, time-to-answer, or a target “helpfulness” score. In customer support ai strategy, we monitor real-world traces, CSAT, and handoff quality to ensure the AI augments agents rather than creating silent failure modes. That’s how evals drive product-market fit lessons instead of just dashboards.

    If you want to go deeper, explore these foundational concepts and tools: ML (Machine learning), LLM (Large language model), “AI Evals for Engineers and PMs”: https://maven.com/parlance-labs/evals, “The Product Leadership Wheel – A Framework for Defining and Growing Product Leadership at Scale”: https://www.petra-wille.com/plwheel, “How I Designed & Implemented Evals for Product Talk’s Interview Coach”: https://www.producttalk.org/2025/09/interview-coach-evals/, “Behind the Scenes: Building the Product Talk Interview Coach”: https://www.producttalk.org/2025/08/customer-interview-coach/, V0: https://vercel.com/docs/v0, JSON (JavaScript Object Notation): https://en.wikipedia.org/wiki/JSON, Anthropic: https://www.anthropic.com/, Lovable: https://lovable.dev/, and “Story-Based Customer Interviews”: https://learn.producttalk.org/course/story-based-customer-interviews.

    If this resonates, I’ll be sharing weekly lessons learned from building and evaluating AI features in the wild, plus conversations with cross-functional teams about real-world AI development. Have thoughts or a tactic that’s worked for you? Drop a comment and let’s compare notes.


    Inspired by this post on Product Talk.


    Book a consult png image
  • Reimagining Product Teams with Generative AI: A Bold, Practical Vision for the Next 24 Months

    Reimagining Product Teams with Generative AI: A Bold, Practical Vision for the Next 24 Months

    In this article, I want to talk about where I believe generative AI is going to take the roles on a product team, and the team topologies of product organizations. I’m motivated to write this both because I think a vision of where we should try to go is important, and also because I see…

    That conviction has only grown as I’ve led cross-functional teams through real deployments. The traditional boundaries between product management, design, engineering, and customer success are blurring as generative AI moves from novelty to dependable copilot. What follows is the vision I’m using to guide our roadmap, hiring, and rituals—practical, near-term, and focused on outcomes.

    First, on roles: product managers will spend less time drafting artifacts and more time validating assumptions and sequencing bets. AI will draft PRDs, summarize interviews, propose opportunity trees, and even flag risks. But we will anchor decisions on outcomes vs output OKRs, using AI to widen the option set, not to outsource accountability.

    Design will accelerate dramatically. With gen ai for product prototyping, designers can turn rough concepts into interactive flows in hours, stress-test copy for clarity, and explore accessibility states before code is written. The craft shifts toward problem framing, system thinking, and quality thresholds—where human judgment remains the differentiator.

    Engineering becomes even more product-facing. Forward deployed engineers will pair with PMs and designers at customer sites (or virtually) to co-create solutions, integrate LLMs, and harden edge cases. Model-aware engineering, evaluation harnesses, and data pipeline stewardship become core competencies, while “prompt engineering” becomes a skill embedded across functions rather than a standalone role.

    On team topology: our default unit stays the autonomous, outcome-owning squad, but we add an enablement layer. An AI platform team supplies shared services—feature stores, evaluation datasets, observability, and safety guardrails—so product teams can move fast without reinventing infrastructure. Guilds or communities of practice steward reusable prompts, patterns, and model cards across squads.

    Discovery evolves too. We’ll pair classic product discovery with AI-accelerated research: large-scale synthesis of qualitative feedback, scenario exploration with synthetic data, and rapid hypothesis testing through simulated cohorts. Human-in-the-loop remains non-negotiable; generative AI helps us see more options, but customers still tell us what’s true.

    Customer support becomes a flywheel. A thoughtful customer support ai strategy turns conversations into structured insights, feeds prioritization, and powers in-product guidance. The same signals that resolve tickets should inform discovery, experimentation, and roadmap trade-offs.

    Governance and safety must be proactive. We’ll define golden datasets, create red-team playbooks, and adopt model-level SLAs alongside product SLAs. Evaluation goes beyond accuracy to include fairness, latency, explainability, and cost, with clear escalation paths when models drift or fail.

    Measuring impact changes as well. Beyond feature delivery, we’ll track time-to-learning, reduction in cycle time, precision of targeting, and the quality of decisions AI actually improves. The goal is durable product-market fit lessons, not vanity metrics or demo-driven development.

    Here’s a pragmatic 90-day starter plan: identify two high-signal use cases where latency, cost, and safety are manageable; form a cross-functional pod with a PM, designer, forward deployed engineers, and a data partner; instrument robust evaluation gates; align on outcomes vs output OKRs; ship, learn, and codify the playbook. In parallel, stand up the minimal AI platform services your squads will reuse.

    This is a leadership challenge as much as a technical one. Product management leadership must set the bar for ethical use, invest in upskilling, and reorganize incentives around outcomes. The teams that win will treat generative AI as a force multiplier for curiosity, learning, and craftsmanship—not a shortcut around them.

    If we do this well, our product teams will be faster, more customer-obsessed, and more resilient. The tools are ready. The real question is whether we are ready to evolve how we work, measure progress, and lead.


    Inspired by this post on SVPG.


    Book a consult png image
  • From Vision to Value: How Generative AI Elevates Product Design and Product Management

    From Vision to Value: How Generative AI Elevates Product Design and Product Management

    Product, design, and AI now converge at the center of how we build value. In my role leading product teams at HighLevel, Inc., I’ve experienced firsthand how generative AI amplifies the craft of product management and product design when we keep the fundamentals tight: clear problems, measurable outcomes, and deep collaboration across disciplines.

    The mission hasn’t changed—deliver useful, usable, and trustworthy experiences—yet the means have. Generative AI expands our exploration space, speeds up iteration, and helps us reason over messy, real-world data. When we marry rigorous product discovery with thoughtful design and responsible AI strategy, we move from novelty to durable impact.

    In discovery, I use AI to frame hypotheses, generate research questions, cluster customer feedback, and synthesize interview notes—without replacing direct conversations with customers. The goal is sharper insight, faster. I define outcomes in customer language, pressure-test assumptions, and trace every proposed AI capability to a clear job to be done. These habits keep us anchored to product-market fit lessons rather than shiny demos.

    For prototyping, I pair designers with forward deployed engineers to build realistic vertical slices quickly. We practice gen ai for product prototyping by wiring prompts, system instructions, constrained outputs, and lightweight evaluators into clickable flows so we can test usefulness early. This reduces risk and helps the team learn which interaction patterns—chat, form, or guided workflows—fit the problem best, especially in product creator experiences.

    Designing AI-powered UX means embracing uncertainty without eroding trust. I favor patterns like transparent confidence cues, citations or references where possible, editable outputs, easy undo/redo, and clear pathways from draft to commit. Good empty states, contextual examples, and progressive disclosure teach users how to get high-quality results while keeping them in control.

    Quality requires a measurement backbone, not vibes. I define target tasks and build golden datasets, then run offline evaluations before online experiments. The core metrics stay consistent: task success rate, user confidence, time-to-first-value, latency budgets, and cost per resolution. We harden experiences with guardrails, hallucination checks, safe fallbacks, and escalation paths to humans when the model is uncertain.

    Responsible AI is a product requirement, not a checkbox. I design for privacy-by-default, PII minimization, and secure data handling; I track prompt and model versions; and I test for bias and accessibility from the outset. Human-in-the-loop review, auditability, and transparent change logs protect users and the business as features evolve.

    Go-to-market is part of the product. Clear onboarding, explainers, and in-product education reduce time to value. I align customer support ai strategy with telemetry so support teams can triage AI-specific issues, capture edge cases, and channel learning back into prompt libraries, data pipelines, and design improvements.

    From a leadership standpoint, I set strategic guardrails and empower autonomous teams. Product management leadership owns outcomes and decision quality; design leads shape multimodal experiences; engineering owns reliability and performance; and our AI platform team standardizes evaluation, safety, and cost controls. This clarity accelerates learning and throughput.

    Recently, we shipped an AI-assisted creation flow that reduced manual steps, improved time-to-first-value, and drove adoption among new users. The win wasn’t a clever prompt; it was disciplined product discovery, fast iteration with realistic data, and a crisp definition of success before we scaled.

    If you’re just starting, pick one high-value, low-risk use case, define success in customer terms, and build a thin vertical slice with evaluations and guardrails. Put it in front of real users, instrument everything, and iterate until the experience feels fast, predictable, and genuinely helpful.

    The intersection of product, design, and AI will keep evolving, but the bar remains the same: ship outcomes customers care about. When we combine the leverage of generative AI with sound product discovery and strong product design, we turn vision into value—reliably and repeatably.


    Inspired by this post on SVPG.


    Book a consult png image
  • Why INSPIRED Still Matters in the Generative AI Era: Access, Insights, and Practical Playbooks

    Why INSPIRED Still Matters in the Generative AI Era: Access, Insights, and Practical Playbooks

    In the Generative AI era, I keep returning to the enduring playbooks that shape great product teams. INSPIRED remains a cornerstone for how I coach on product discovery, product operating models, and product management leadership. I’ve used its principles to align cross-functional squads, empower product creators, and accelerate product-market fit lessons across both startups and scaled organizations.

    The book INSPIRED is available in hardcover, digital, and audio versions, but until now, the audio version was only available in an exclusive arrangement with Amazon, on audible.com. The audio versions of our other books have been available from all major audio book providers. The exclusive contract with Amazon has now expired, and…

    Why this matters: when knowledge moves beyond a single platform, more of our teams can absorb it in the flow of work. Distributed PMs, designers, data scientists, and forward deployed engineers can learn on their preferred apps during commutes or deep work breaks. That accessibility compounds learning velocity—especially when we’re iterating weekly on discovery insights, opportunity assessments, and bet selection.

    What’s changed in our craft is the tooling: gen ai now augments how we validate assumptions, run product discovery, and prototype. Pairing the timeless practices in INSPIRED with gen ai for product prototyping helps my teams get to evidence faster—turning ambiguous narratives into testable artifacts, instrumented experiments, and real customer signals. It also sharpens our product operating model by making continuous discovery the default behavior across the product team.

    Here’s how I operationalize this shift: I anchor a short “learning sprint” around one chapter at a time, then immediately translate insights into a concrete discovery activity (problem framing, assumption mapping, or opportunity sizing). We run a gen ai prototyping spike to visualize flows, draft UX copy, and simulate edge cases, followed by quick customer sessions to validate usefulness and usability. We capture outcomes in a working taxonomy of product-market fit lessons and update our decision logs so learning compounds sprint over sprint.

    This is also a practical boost for enablement: new hires, customer support leaders crafting a customer support ai strategy, and forward deployed engineers can now engage with the same source material on their own schedules. When the whole team shares a common vocabulary—shaped by proven practices and accelerated by gen ai—the quality of debate improves, discovery cycles compress, and execution becomes more predictable.

    If you’ve been meaning to revisit INSPIRED, this is an ideal moment. With access broadening, pick the format that fits your routine and turn insights into action the same day. Use it to pressure-test your product operating model, refine your discovery cadence, and elevate product management leadership across the organization. The combination of timeless principles and modern gen ai tools is exactly what our product teams need right now.


    Inspired by this post on SVPG.


    Book a consult png image
  • Inside Alyx: Dogfooding, Evals, and Observability That Power an Agentic AI Future

    Inside Alyx: Dogfooding, Evals, and Observability That Power an Agentic AI Future

    I’ve been deep in the work of building practical, agentic capabilities into AI products, so this story about Alyx immediately resonated with me. It’s a rare, clear-eyed look at what it actually takes to ship a useful AI agent inside an AI platform—while using that same platform to build, test, and continuously improve the agent.

    What does it really take to build an AI agent inside an AI platform—especially when you’re using that same platform to build the agent?

    Listening to SallyAnn DeLucia (Director of Product at Arize) and Jack Zhou (Staff Engineer at Arize) unpack Alyx—the AI agent that helps teams debug, optimize, and evaluate AI applications—I recognized playbooks I trust: start scrappy, dogfood relentlessly, build intuition with real users, and systematize improvement with thoughtful evals.

    Their early phase looked exactly like the messy reality many of us try to hide: Jupyter notebooks, hacked-together web apps, and weekly dogfooding sessions with their customer success team. That’s where patterns emerged, confidence was built, and the highest-leverage skills for the agent were prioritized. It’s a reminder that “vibe checks” matter at first—but you must quickly graduate to measurable, repeatable learning loops.

    In my experience, the foundation of GenAI product quality is threefold: tracing, observability, and evals. They reached the same conclusion—defining traces across tool calls and sessions, creating observability into model behavior, and layering evals to compare both micro-decisions and system-level outcomes. That discipline converts hunches into evidence and makes agent behavior improvable, not mysterious.

    What stood out was how cross-functional, boundary-spanning teams made the difference. Customer success engineers surfaced repeatable workflows. Product framed early skills. Engineering wrapped prototype tools into something coherent. Using their own platform to build Alyx accelerated intuition and de-risked launch. That’s the product loop I aim to cultivate: close to customers, close to data, and fast to learn.

    As Alyx matures, the next step is moving from “on rails” workflows to more autonomous, agentic planning loops. That evolution requires stronger tool design, richer feedback signals, and evals that reflect end-to-end user value. It’s exactly the shift I expect across GenAI: from scripted assistants to adaptive systems that reason, plan, and act with guardrails.

    Listen to this episode on: Spotify | Apple Podcasts

    Guests:

    SallyAnn DeLucia, Director of Product, Arize

    Jack Zhou, Staff Engineer, Arize

    In this episode, we cover:

    What tracing, observability, and evals really mean in GenAI applications

    How Arize used its own platform to build Alyx, its AI agent

    The role of customer success engineers in surfacing repeatable workflows

    Why early prototyping looked like messy notebooks and hacked-together local apps

    How dogfooding shaped Alyx’s evolution and built confidence for launch

    Why evals start messy, and how Arize layered evals across tool calls, sessions, and system-level decisions

    The importance of cross-functional, boundary-spanning teams in building AI products

    What’s next for Alyx: moving from “on rails” workflows to more autonomous, agentic planning loops

    My takeaways for product teams building GenAI agents are simple and hard: design tools with observability in mind; operationalize evals early even if they’re imperfect; embed customer-facing engineers in the loop to capture real workflows; and keep the first skills narrow, high-impact, and testable. If your team can move from demos to disciplined measurement quickly, you’ll accelerate product-market fit.

    Resources & Links

    Arize AI — Sign up for a free account and try Alex

    Arize Blog — Lessons learned from building AI products

    Maven AI Evals Course — The course Teresa took to learn about evals (Get 35% off with Teresa’s affiliate link)

    Cursor — The AI-powered code editor used by the Arize engineering team

    DataDog — For understanding application traces

    OpenAI GPT Models — GPT-3.5, GPT-4, and newer models used in early and current versions of Alex

    Jupyter Notebooks — A tool for combining code, data, and notes, used in Arise’s prototyping

    Axial Coding Method by Hamel Husain — A framework for analyzing data and designing evals

    Chapters

    00:00 Introduction to Sally Ann and Jack

    01:08 Overview of Arize.ai and Its Core Components

    01:44 Deep Dive into Tracing, Observability, and Evals

    03:56 Introduction to Alyx: Arize's AI Agent

    04:15 The Genesis and Evolution of Alyx

    08:51 Challenges and Solutions in Building Alyx

    24:33 Prototyping and Early Development of Alyx

    26:22 Exploring the Power of Coding Notebooks

    26:51 Early Experiments with Alyx

    27:59 Challenges with Real Data

    29:20 Internal Testing and Dogfooding

    31:55 The Importance of Evals

    35:16 Developing Custom Evals

    43:09 Future Plans for Alyx

    47:59 How to Get Started with Alyx

    Full Transcript

    Podcast transcripts are only available to paid subscribers.

    If you’re building in GenAI right now, this conversation offers a pragmatic blueprint. Start with high-signal workflows, turn qualitative insights into quantitative evals, and use tracing plus observability to make agents debuggable. That’s how scrappy prototypes become reliable systems. And if you want a tangible example, “47:59 How to Get Started with Alyx” is a helpful on-ramp.


    Inspired by this post on Product Talk.


    Book a consult png image
  • How Braintrust Nailed Product-Market Fit: Paranoia, Patience, and High-Bar Quality

    How Braintrust Nailed Product-Market Fit: Paranoia, Patience, and High-Bar Quality

    Product-market fit in the GenAI era is elusive because both the technology surface area and user expectations change weekly. That’s why Braintrust caught my eye: they set a relentless quality bar, delayed go-to-market on purpose, and used real-world evaluation pain to shape an end-to-end platform for building AI apps. In my work leading product management teams, I recognize this pattern as the difference between shipping demos and shipping durable value.

    Context matters. Ankur Goyal’s journey runs through MemSQL (now SingleStore), Impira, and Figma. Working with high-bar users at MemSQL forged a bias toward precision, performance, and reliability—traits that translate directly to AI infrastructure where flaky evals and brittle prompts can quietly erode trust. When you build for exacting users early, the feedback loop is unforgiving—and that’s a gift.

    The throughline is quality. Great software often comes from a place of “paranoia”—the productive kind that compels us to fail proofs, harden edge cases, and verify outcomes under load. In AI product development, that paranoia shows up as rigorous evals, clear data contracts, reproducibility, and measured rollouts. It’s not glamorous, but it’s how you earn compounding trust with builders and operators.

    Recruiting is strategy. The trick to recruiting well is selecting for taste, curiosity, and ownership—people who elevate the craft and sweat the engineering details. In AI-heavy products, I’ve had the most success with forward deployed engineers who live with users long enough to discover the non-obvious constraints that should drive the roadmap. Taste plus proximity beats velocity without context.

    Impulse control creates leverage. Braintrust delayed go-to-market, which is counterintuitive when the market is hot. But in a new category, premature scaling yields fake signals. The better move is to tighten the loop: instrument the “prompt playground,” pressure-test evals, validate the inner loop of building AI apps, and only then broaden access. When the core interaction is right, growth compounds; when it’s off, every feature feels like a workaround.

    Figma-era frustrations with evals became the opportunity. Anyone who has tried to standardize AI evaluations across prompts, models, and datasets knows how quickly the surface area explodes. Converting that frustration into Braintrust’s product thesis—reliable, end-to-end workflows for AI app development—speaks to a classic product discovery principle: go deep on a painful, persistent job-to-be-done before you go broad.

    How to recognize a real market opportunity: look for high-frequency workflows with measurable outcomes, teams who already duct-tape solutions, and buyers who have the budget and urgency to pull the product in. When you see repeatable pull from discerning users—and you can demonstrate quality with transparent evals—you’re approaching true PMF rather than narrative fit.

    Inside the first six months, the right posture is deliberate focus. For a platform like Braintrust, that means obsessing over the developer inner loop: data in, prompt iteration, eval rigor, versioning, approvals, and productionization. The “prompt playground” must evolve from experimentation to governance, so teams can move from clever demos to reliable deployments with confidence.

    AI continues to reshape the platform’s future. As model ecosystems shift (OpenAI and beyond) and the data plane sprawls (Databricks, Snowflake), developers want a unified surface to build, evaluate, and ship. Integrations with familiar tools like Airtable, Coda, Zapier, and Figma lower adoption friction by meeting teams where they already work, while enterprise-grade controls unlock buyers at the scale of Goldman Sachs.

    The cultural choices matter as much as the code. Make big bets with extreme clarity, or don’t make them at all. Stay mission-driven when novelty tempts distraction. Write down the customer promise and keep it tight. Hiring mistakes—especially around quality, curiosity, and ownership—compound quickly in AI product teams, so reset the bar early and protect it.

    What PMF really looks like here: customers self-discover core value, usage deepens without hand-holding, and cross-functional teams (engineering, data science, and operations) align around shared definitions of quality. Support volume becomes more about how-to than break-fix. Roadmap prioritization becomes easier because the next best feature reveals itself in the workflow data.

    My playbook takeaways for product management leadership in GenAI: prioritize eval rigor before growth, use forward deployed engineers for product discovery, specialize the prompt playground into a governed inner loop, and delay go-to-market until high-bar users pull you in. These are the same principles I apply to gen ai for product prototyping and customer support ai strategy—because durable PMF in AI still comes down to quality, focus, and earned trust.

    Referenced:

    • Airtable: https://www.airtable.com/

    • Adam Prout: https://www.linkedin.com/in/adam-prout-0b347630/

    • Braintrust: https://braintrust.dev

    • Brian Helmig: https://www.linkedin.com/in/bryanhelmig/

    • Coda: https://coda.io/

    • Databricks: https://www.databricks.com/

    • David Kossnick: https://www.linkedin.com/in/davidkossnick/

    • Figma: https://www.figma.com/

    • Goldman Sachs: https://www.goldmansachs.com/

    • Kris Rasmussen: https://www.linkedin.com/in/kristopherrasmussen/

    • Manu Goyal: https://www.linkedin.com/in/mngyl/

    • MemSQL: https://www.singlestore.com/ (now SingleStore)

    • Nikita Shamgunov: https://www.linkedin.com/in/nikitashamgunov/

    • OpenAI: https://openai.com/

    • Snowflake: https://www.snowflake.com/

    • Zapier: https://zapier.com/


    Book a consult png image