I’ve been deep in the work of building practical, agentic capabilities into AI products, so this story about Alyx immediately resonated with me. It’s a rare, clear-eyed look at what it actually takes to ship a useful AI agent inside an AI platform—while using that same platform to build, test, and continuously improve the agent.
What does it really take to build an AI agent inside an AI platform—especially when you’re using that same platform to build the agent?
Listening to SallyAnn DeLucia (Director of Product at Arize) and Jack Zhou (Staff Engineer at Arize) unpack Alyx—the AI agent that helps teams debug, optimize, and evaluate AI applications—I recognized playbooks I trust: start scrappy, dogfood relentlessly, build intuition with real users, and systematize improvement with thoughtful evals.
Their early phase looked exactly like the messy reality many of us try to hide: Jupyter notebooks, hacked-together web apps, and weekly dogfooding sessions with their customer success team. That’s where patterns emerged, confidence was built, and the highest-leverage skills for the agent were prioritized. It’s a reminder that “vibe checks” matter at first—but you must quickly graduate to measurable, repeatable learning loops.
In my experience, the foundation of GenAI product quality is threefold: tracing, observability, and evals. They reached the same conclusion—defining traces across tool calls and sessions, creating observability into model behavior, and layering evals to compare both micro-decisions and system-level outcomes. That discipline converts hunches into evidence and makes agent behavior improvable, not mysterious.
What stood out was how cross-functional, boundary-spanning teams made the difference. Customer success engineers surfaced repeatable workflows. Product framed early skills. Engineering wrapped prototype tools into something coherent. Using their own platform to build Alyx accelerated intuition and de-risked launch. That’s the product loop I aim to cultivate: close to customers, close to data, and fast to learn.
As Alyx matures, the next step is moving from “on rails” workflows to more autonomous, agentic planning loops. That evolution requires stronger tool design, richer feedback signals, and evals that reflect end-to-end user value. It’s exactly the shift I expect across GenAI: from scripted assistants to adaptive systems that reason, plan, and act with guardrails.
Listen to this episode on: Spotify | Apple Podcasts
Guests:
SallyAnn DeLucia, Director of Product, Arize
Jack Zhou, Staff Engineer, Arize
In this episode, we cover:
What tracing, observability, and evals really mean in GenAI applications
How Arize used its own platform to build Alyx, its AI agent
The role of customer success engineers in surfacing repeatable workflows
Why early prototyping looked like messy notebooks and hacked-together local apps
How dogfooding shaped Alyx’s evolution and built confidence for launch
Why evals start messy, and how Arize layered evals across tool calls, sessions, and system-level decisions
The importance of cross-functional, boundary-spanning teams in building AI products
What’s next for Alyx: moving from “on rails” workflows to more autonomous, agentic planning loops
My takeaways for product teams building GenAI agents are simple and hard: design tools with observability in mind; operationalize evals early even if they’re imperfect; embed customer-facing engineers in the loop to capture real workflows; and keep the first skills narrow, high-impact, and testable. If your team can move from demos to disciplined measurement quickly, you’ll accelerate product-market fit.
Resources & Links
Arize AI — Sign up for a free account and try Alex
Arize Blog — Lessons learned from building AI products
Maven AI Evals Course — The course Teresa took to learn about evals (Get 35% off with Teresa’s affiliate link)
Cursor — The AI-powered code editor used by the Arize engineering team
DataDog — For understanding application traces
OpenAI GPT Models — GPT-3.5, GPT-4, and newer models used in early and current versions of Alex
Jupyter Notebooks — A tool for combining code, data, and notes, used in Arise’s prototyping
Axial Coding Method by Hamel Husain — A framework for analyzing data and designing evals
Chapters
00:00 Introduction to Sally Ann and Jack
01:08 Overview of Arize.ai and Its Core Components
01:44 Deep Dive into Tracing, Observability, and Evals
03:56 Introduction to Alyx: Arize's AI Agent
04:15 The Genesis and Evolution of Alyx
08:51 Challenges and Solutions in Building Alyx
24:33 Prototyping and Early Development of Alyx
26:22 Exploring the Power of Coding Notebooks
26:51 Early Experiments with Alyx
27:59 Challenges with Real Data
29:20 Internal Testing and Dogfooding
31:55 The Importance of Evals
35:16 Developing Custom Evals
43:09 Future Plans for Alyx
47:59 How to Get Started with Alyx
Full Transcript
Podcast transcripts are only available to paid subscribers.
If you’re building in GenAI right now, this conversation offers a pragmatic blueprint. Start with high-signal workflows, turn qualitative insights into quantitative evals, and use tracing plus observability to make agents debuggable. That’s how scrappy prototypes become reliable systems. And if you want a tangible example, “47:59 How to Get Started with Alyx” is a helpful on-ramp.
I recently spent time with an episode of All Things Product that hit especially hard as we head into year-end: Petra Wille and Teresa Torres ask, “What do you want to be known for in your work?” As someone leading product management and building high-performing teams, I regularly bring this question into my Q4 conversations. It’s a powerful lens for product management leadership, career transitions, and how we show up for our customers and colleagues.
Listen to this episode on: Spotify | Apple Podcasts
In this conversation, I appreciated how clearly they unpack the nuances of impact, craft, personal brand, and values—and how those ideas shape the footprints we leave in teams, organizations, and the broader product community. Their stories and lessons learned are equal parts relatable and practical, which is exactly what we need when we’re balancing execution with reflection.
Let’s talk about “legacy.” The word can feel loaded—big, vague, and distant. I reframe it with my teams into a question we can act on now: What meaningful change did we enable for customers and our organization this quarter, and what do we want colleagues to remember about how we did it? That framing keeps us grounded in outcomes and behaviors, not just lofty aspirations.
The distinction between impact and craft is central. Impact is the difference our work makes—what changes because of our decisions. Craft is what we hone for intrinsic reward—our product discovery techniques, decision-making frameworks, and communication muscles. Early in my career, I over-indexed on impact metrics and under-invested in craft. I shipped value, but I wasn’t building the repeatable habits that elevate a product creator for the long haul. Over time, I learned that craft compounds—and it pays dividends in both product-market fit lessons and leadership credibility.
Personal brand and values also matter more than many of us admit. When the pressure is on, people remember how we decide, how we communicate trade-offs, and how consistently we anchor on customer value. I want to be known for rigorous product discovery, clarity under uncertainty, and the integrity to say “no” when it protects long-term outcomes. Those cues travel fast across an organization and quietly define our leadership legacy.
Feedback gaps can reveal blind spots—and we all have them. I proactively create multiple feedback loops: structured 1:1s, skip-levels, stakeholder debriefs after key product decisions, and customer touchpoints. I specifically ask for disconfirming evidence—what am I missing, where did my decision-making create friction, and how might I simplify? Weekly customer learning is non-negotiable for me; it keeps the team grounded and accelerates product discovery. If you need a starting point, Teresa’s work on weekly customer interviews is a solid playbook: Customer Interviews: How to Recruit, What to Ask, and How to Synthesize What You Learn.
Here are the prompts I’m using with my team for Q4 reflection. Why “legacy” can feel loaded—and better ways to frame the question. The difference between impact (what changes because of your work) and craft (what you hone for intrinsic reward). How personal brand and values influence what colleagues remember about you. Why feedback gaps can reveal blind spots—and how to proactively seek better input. Reflection prompts to carry into your Q4 (and beyond). I encourage folks to journal on these, then bring two concrete actions into our next planning cycle.
If you’re thinking about your own growth, preparing for career transitions, or simply curious how others reflect on their product practice, this episode offers both inspiration and pragmatic takeaways. I’m weaving these themes into our planning and calibrations because reflection is a force multiplier—it sharpens strategy, strengthens culture, and ultimately improves customer outcomes.
Follow Teresa Torres: https://ProductTalk.org
Follow Petra Wille: https://Petra-Wille.com
Mentioned in the episode: Petra’s Thought-Provoking Questions to Prompt Your End-of-Year Reflection
Mentioned in the episode: Xing
Mentioned in the episode: Teresa’s work on weekly customer interviews: Customer Interviews: How to Recruit, What to Ask, and How to Synthesize What You Learn
Mentioned in the episode: Petra’s guide: The Product Leader’s Guide to Giving Feedback
Join the conversation with me: What do you want to be known for in your product work this coming year? Share your thoughts below and let’s learn from one another.
Full Transcript
Full transcripts are only available for paid subscribers.
Product-market fit in the GenAI era is elusive because both the technology surface area and user expectations change weekly. That’s why Braintrust caught my eye: they set a relentless quality bar, delayed go-to-market on purpose, and used real-world evaluation pain to shape an end-to-end platform for building AI apps. In my work leading product management teams, I recognize this pattern as the difference between shipping demos and shipping durable value.
Context matters. Ankur Goyal’s journey runs through MemSQL (now SingleStore), Impira, and Figma. Working with high-bar users at MemSQL forged a bias toward precision, performance, and reliability—traits that translate directly to AI infrastructure where flaky evals and brittle prompts can quietly erode trust. When you build for exacting users early, the feedback loop is unforgiving—and that’s a gift.
The throughline is quality. Great software often comes from a place of “paranoia”—the productive kind that compels us to fail proofs, harden edge cases, and verify outcomes under load. In AI product development, that paranoia shows up as rigorous evals, clear data contracts, reproducibility, and measured rollouts. It’s not glamorous, but it’s how you earn compounding trust with builders and operators.
Recruiting is strategy. The trick to recruiting well is selecting for taste, curiosity, and ownership—people who elevate the craft and sweat the engineering details. In AI-heavy products, I’ve had the most success with forward deployed engineers who live with users long enough to discover the non-obvious constraints that should drive the roadmap. Taste plus proximity beats velocity without context.
Impulse control creates leverage. Braintrust delayed go-to-market, which is counterintuitive when the market is hot. But in a new category, premature scaling yields fake signals. The better move is to tighten the loop: instrument the “prompt playground,” pressure-test evals, validate the inner loop of building AI apps, and only then broaden access. When the core interaction is right, growth compounds; when it’s off, every feature feels like a workaround.
Figma-era frustrations with evals became the opportunity. Anyone who has tried to standardize AI evaluations across prompts, models, and datasets knows how quickly the surface area explodes. Converting that frustration into Braintrust’s product thesis—reliable, end-to-end workflows for AI app development—speaks to a classic product discovery principle: go deep on a painful, persistent job-to-be-done before you go broad.
How to recognize a real market opportunity: look for high-frequency workflows with measurable outcomes, teams who already duct-tape solutions, and buyers who have the budget and urgency to pull the product in. When you see repeatable pull from discerning users—and you can demonstrate quality with transparent evals—you’re approaching true PMF rather than narrative fit.
Inside the first six months, the right posture is deliberate focus. For a platform like Braintrust, that means obsessing over the developer inner loop: data in, prompt iteration, eval rigor, versioning, approvals, and productionization. The “prompt playground” must evolve from experimentation to governance, so teams can move from clever demos to reliable deployments with confidence.
AI continues to reshape the platform’s future. As model ecosystems shift (OpenAI and beyond) and the data plane sprawls (Databricks, Snowflake), developers want a unified surface to build, evaluate, and ship. Integrations with familiar tools like Airtable, Coda, Zapier, and Figma lower adoption friction by meeting teams where they already work, while enterprise-grade controls unlock buyers at the scale of Goldman Sachs.
The cultural choices matter as much as the code. Make big bets with extreme clarity, or don’t make them at all. Stay mission-driven when novelty tempts distraction. Write down the customer promise and keep it tight. Hiring mistakes—especially around quality, curiosity, and ownership—compound quickly in AI product teams, so reset the bar early and protect it.
What PMF really looks like here: customers self-discover core value, usage deepens without hand-holding, and cross-functional teams (engineering, data science, and operations) align around shared definitions of quality. Support volume becomes more about how-to than break-fix. Roadmap prioritization becomes easier because the next best feature reveals itself in the workflow data.
My playbook takeaways for product management leadership in GenAI: prioritize eval rigor before growth, use forward deployed engineers for product discovery, specialize the prompt playground into a governed inner loop, and delay go-to-market until high-bar users pull you in. These are the same principles I apply to gen ai for product prototyping and customer support ai strategy—because durable PMF in AI still comes down to quality, focus, and earned trust.
Referenced:
• Airtable: https://www.airtable.com/
• Adam Prout: https://www.linkedin.com/in/adam-prout-0b347630/
• Braintrust: https://braintrust.dev
• Brian Helmig: https://www.linkedin.com/in/bryanhelmig/
• Coda: https://coda.io/
• Databricks: https://www.databricks.com/
• David Kossnick: https://www.linkedin.com/in/davidkossnick/
I’ve spent the last few years watching AI reshape product roadmaps, developer workflows, and customer expectations. One idea now feels undeniable: the web must evolve to serve a new primary user—AIs. That shift changes how we think about search, reliability, governance, monetization, and ultimately, how we design products that scale with trust.
Parag Agrawal is the co-founder and CEO of Parallel, a startup building search infrastructure for the web’s second user: AIs. Before launching Parallel, Parag spent over a decade at Twitter, where he served as CTO and later CEO during a period of intense transformation, as well as public scrutiny.
I was particularly struck by how crisply this frames the next frontier for product leaders: build systems that machines can consume at massive scale without sacrificing accuracy, provenance, or trust. In particular, I was drawn to the emphasis on “deep research,” where Parallel is tackling “deep research” challenges by prioritizing accuracy over speed, and the design choices that make their APIs uniquely agent-friendly. As someone who has shipped AI features into production, that trade-off resonates—speed gets demos; accuracy earns renewals.
Here’s how I’m synthesizing the most actionable takeaways for product, engineering, and go-to-market leaders. First, design for AI as the primary customer. That means structuring content and APIs so agents can reliably reason, verify, and self-correct. Agent-friendly interfaces need deterministic schemas, explicit provenance, stable latency envelopes, and predictable failure modes. If an agent can’t trust your contract, it won’t chain your service into complex workflows, and you’ll lose the compounding effects that make AI platforms defensible.
Second, bring a systems mindset to accuracy. “Accuracy over speed” isn’t a slogan—it’s an architecture choice. In my experience, that shows up as retrieval strategies tuned for recall and precision trade-offs, multi-pass verification, and human-in-the-loop escalation paths for high-risk queries. For deep research use cases, you need to make the cost of being wrong explicit in your design and your SLAs.
Third, expect your ICP to evolve as AI matures. Early adopters may be research-heavy teams and product creators building agentic workflows. Over time, as reliability improves, your ideal customer shifts toward operational teams that demand measurable outcomes—support deflection, conversion lift, cycle-time reduction. I map these stages explicitly in the roadmap and keep pricing, packaging, and onboarding aligned to each phase.
Fourth, consider business models that keep the web open for AI while aligning incentives. If AIs are the web’s second user, publishers need fair value exchange for structured access, provenance, and usage. In practice, that could look like tiered access, usage-based pricing, attribution requirements, or revenue-sharing tied to agent-driven outcomes. The key is ensuring that openness and sustainability are not at odds.
Fifth, build engineering teams that are both pragmatic and research-aware. On my teams, I look for a balance between high-potential builders who move fast with ambiguous specs and experienced hands who can productionize novel systems. Forward deployed engineers can be a force multiplier here—embedding with customers to surface edge cases, close the verification loop, and turn qualitative insights into productized patterns.
Sixth, recognize how the software engineer’s role is evolving in an AI-assisted world. Engineers are increasingly orchestrators—composing models, retrieval layers, tools, and policies—rather than only writing business logic. That requires better observability for prompts and agents, reproducibility for experiments, and contracts that make emergent behavior inspectable and testable. This is where “uniquely agent-friendly” APIs show their value—clear contracts enable safe autonomy.
Seventh, treat launch timing as a function of trust, not just velocity. Founders often ask when to ship. My rule: launch when you can document bounds, prove repeatability on critical paths, and explain failure semantics. In AI, your narrative is your control surface—fundraising frameworks and customer conversations both benefit when you can quantify reliability, not just showcase capability.
Finally, the long-term vision matters. If agents are finally becoming useful in production, the platforms that win will combine: machine-readable content at scale, accuracy-first retrieval and verification, agent-safe API design, and sustainable economics for an open web. That’s the blueprint I’m applying to my own product strategy: build for agents, measure for trust, and align incentives so the ecosystem compounds rather than fragments.
To product leaders navigating this shift: revisit your ICP, rewrite your API contracts for agents, and make “accuracy over speed” a first-class requirement. To engineering leaders: invest in evaluation harnesses, data quality pipelines, and forward deployed engineers who can turn messy customer workflows into reusable system capabilities. The AI era rewards teams that pair ambition with discipline—and that’s where the next wave of durable advantage will be built.
Rolling out an AI Agent doesn’t just change how your team works – it changes who your team is.
I learned that in the crucible of a fast-moving launch. Before we launched Fin publicly, our Support team became its first alpha/beta tester and we had to move fast. No roadmap. No step-by-step guide. Just a powerful new technology, and a steep learning curve.
That experience is exactly what led us to create The AI Agent Blueprint – a resource we wish we’d had when we were starting out, and one we hope will give other support teams a clearer path forward.
Looking back, I won’t lie and say I was cool, calm, and confident about how to do this – I was nervous as hell. I had no idea how to implement an AI Agent and ensure it resulted in huge cost savings and stellar customer experiences.
We had older machine learning technology available to us (shout out to our first-gen chatbot, Resolution Bot), but as a complex software business, we really only used it for basic FAQs. In all honesty, we still had a way to go – both in using automation more effectively and in making the chatbot experience actually enjoyable for our customers.
So why the urgency?
When ChatGPT burst onto the scene nearly three (!!) years ago, Intercom’s Machine Learning team immediately spotted the opportunity and dived into building the world’s first (and objectively best) AI Customer Service Agent.
Suddenly, we were being asked to pilot this brand new technology with real customers and go all in ASAP. Because we were selling this powerful new functionality, we had to use it ourselves and show it off in the best possible light so customers would want to use it too. #nopressure
There was no playbook, just a lot to figure out. As a product management leader, I had to switch into rigorous product discovery while staying execution-minded.
Steady involvement, rising resolutions. From February to July, teams maintain a high 87–93 involvement range as resolution rates climb from 65 to 82—signaling how AI-driven workflows can boost support efficiency and outcomes.
How do we do a phased rollout, but scale very quickly?
How do we QA Fin’s responses and make continuous improvements?
How will we produce and manage all the content Fin needs?
What will we do about all the outdated content we already have?
What are the success metrics now? Should they be different to original Support KPIs?
Who’s responsible for the success metrics? Who manages this newcomer to our team?
It was daunting. We had to take a brand new technology, figure out how to use it, build a team around it, and move at breakneck speed to implement every new feature that rolled out. It was ambiguous, fast-moving, and a massive lift.
But we got there and the results speak for themselves: Fin is now resolving over 75% of our inbound support volume.
An isometric blueprint reveals how an AI agent powers modern support—from triage to resolution—linking chat, knowledge, and workflows so teams scale service without losing accuracy, context, or the human touch.
That outcome didn’t happen by accident. We embedded forward deployed engineers with Support, treated our AI Agent like a product creator in its own right, and used gen ai for product prototyping to tighten our iteration loops. We prioritized a customer support AI strategy that balanced containment with quality: containment rate, CSAT on AI-resolved conversations, first-response latency, and recontact rates became our core scorecard.
That success led to real change for me and my team: new roles, new responsibilities, and new career paths. I now run a whole new function that didn’t exist before: AI Support. We’ve created new and elevated roles like Conversation Designers and Knowledge Managers. Fin hasn’t just changed how we support customers – it’s transformed the structure of our team and the trajectory of our careers.
And now, we’re helping our customers do the same.
In all transparency, if I hadn’t been this close to the work, I might have waited to see how generative AI played out before committing. I might have waited for a blueprint for how to deploy and scale an AI Agent. I wish I had something like that when we got started, or even later when we had a solid foundation but needed to scale our AI strategy.
How much less scary would it be to implement an AI Agent if something like that existed?
Whether you’re just getting started or already using AI in some way, you’re not early anymore—and you shouldn’t have to figure it all out alone. Strong product management leadership, a clear change plan, and tight feedback loops are what separate experiments from outcomes.
That’s why we created The AI Agent Blueprint – a practical map for launching and scaling AI in support. It brings together everything we’ve learned from our own journey, and from working closely with our customers who are doing the same.
If you’re ready to operationalize gen ai in support, align on the right metrics, and redesign roles for the future, this blueprint will help you move from pilots to pervasive impact with confidence.
Bold, pragmatic bets separate teams that merely deliver from teams that truly accelerate. As a product leader, I’m drawn to decisions that reduce friction, empower engineers, and compound over time. Intercom’s recent investment in a new frontend direction is a standout example of this mindset—and it offers lessons any product, engineering, or design leader can apply.
Over the past two years, Intercom made one of the most significant changes an engineering organization can make: moving its core frontend from Ember to React. That choice fits a clear pattern of high-agency decision making in service of speed, quality, and developer experience.
Back in 2014, Ember was the right call for their main application. Its strong opinions and “batteries-included” approach aligned with a strategy I respect: make big decisions once, enable teams to move fast, and spend energy on customer problems instead of endless architecture debates. The result was scale few achieve—more than two million lines of code and 100,000+ pull requests merged.
I’ve been in the room when constraints outgrow the original bet. By 2023, “local builds stretched beyond 90 seconds,” and they were stuck on older framework versions that blocked adoption of modern build tools. Even with deep community engagement and contributions to the Embroider Initiative, the cost of staying put was compounding. Something had to change.
What I admire is the rigor behind their pivot. They ran workshops, health checks, and set explicit trigger conditions—then honored those triggers. When the evidence crossed the threshold, they chose a new path and framed the work with a clear, galvanizing banner: “The Future of Frontend.” That’s the kind of governance and narrative clarity that de-risks large platform shifts.
React quickly emerged as the right fit—not because of hype, but because it met practical criteria at scale. “React was already a core technology at Intercom (powering Messenger, Help Center, and our marketing site),” backed by a robust ecosystem, strong documentation, and broad familiarity internally and across the industry. Most importantly, it integrates naturally with AI-driven developer tools—a non-negotiable for the next decade of engineering productivity.
Fast forward to today, and the momentum is clear. “React is now the default for new UI development at Intercom.” That single sentence says a lot about organizational alignment and execution readiness.
The outcomes speak for themselves. “Blazing fast feedback loops: React builds in under 10 seconds locally, with sub-1s rebuilds – much faster than our Ember app’s 90+ seconds.” That kind of drop in cycle time unlocks more iteration, tighter designer–engineer collaboration, and faster learning loops.
Speed without joy is a half-win. “Higher developer velocity: Engineers consistently report being faster, happier, and more effective, particularly when paired with AI tools like Cursor, Augment, and Claude Code.” I’ve seen similar effects: once teams feel flow again, quality and ambition both rise.
Adoption at breadth matters as much as depth. “Wider adoption: Since March 2025, 10+ Product teams have shipped React features, contributing over 840 pull requests.” That level of traction signals a platform shift that’s not just technically sound but operationally viable.
The AI-enabled developer experience is the real unlock. “AI synergy: React just “clicks” with modern AI tooling. Designers and engineers are using agents to write code, generate components from Figma, and even build design playgrounds themselves.” That’s the future: product creators working in shared, generative environments where ideas move from Figma to code in minutes.
One engineer captured the productivity gain perfectly: “The work I had predicted would take me a week to achieve took me two days”. That’s not a marginal improvement—that’s a step-change.
This story isn’t just about frameworks; it’s about preparing for a decade where velocity, AI-native workflows, and developer experience determine competitive advantage. The ambition to “double our productivity over the next 12 months” requires removing friction, leaning into AI, and standardizing on tools that compound learning across teams. React is a pragmatic enabler for that journey.
I also appreciate the organizational design behind the change. A small, focused group—Team Frontend Tech—partnered tightly with Product teams to shape the new stack, build a design system, and accelerate adoption. That model creates a high-trust bridge between platform and product, which is essential for landing a migration at scale.
For leaders navigating similar crossroads, the playbook is clear: set explicit trigger conditions, articulate the future state, pick a stack that compounds with AI, and invest in a cross-functional nucleus to shepherd adoption. For engineers and designers, this is an exciting moment—one where your tools finally catch up with your ambition.
The takeaway I’m carrying forward: make the bold call when the evidence is conclusive, optimize for feedback loops and flow, and treat AI as a first-class partner in the creative process. That’s how we keep shipping fast, raise the quality bar, and focus on what really matters—solving meaningful problems for customers.
Your team has a polished GenAI prototype. The review goes well. Then the hard questions arrive: Which customer behavior should change? How often does the system fail? What context does it need? Who owns prompt and model changes? Can you release it without creating a permanent escalation queue?
The gap between an impressive demonstration and a dependable product is usually an operating-model problem. A useful GenAI discovery model connects customer evidence, system evaluations, field behavior, and production ownership. It lets you preserve the speed of AI-assisted prototyping while making each experiment answer a real product decision.
Key takeaways for GenAI product leaders
Prove value and capability separately. Customer demand does not prove that a model can perform reliably, and a strong evaluation score does not prove that anyone will change their behavior.
Begin with a workflow, an outcome, and a bounded role for AI. A generic AI use case is too loose to guide discovery or define acceptable performance.
Maintain an evidence chain. Connect each customer story to an opportunity, each opportunity to an assumption, each assumption to an evaluation scenario, and each released behavior to a product outcome.
Run customer learning and system learning in the same weekly loop. They require different methods, but they must meet at one decision: advance, narrow, change, or stop the bet.
Centralize reusable infrastructure and guardrails, not product judgment. Product teams should own the workflow and outcome while shared capabilities provide model access, evaluation tooling, observability, and policy controls.
Start with the workflow and outcome, not the model
GenAI discovery often starts in the wrong direction. A team gets access to a capable model, generates a list of possible features, and builds whichever idea looks most compelling in a demo. The team may learn that the model can produce something plausible, but it still does not know whether the capability removes meaningful friction from a real workflow.
Reverse the sequence. Start with a person trying to achieve an outcome in a specific context. Understand what that person does now, where the workflow breaks, what consequence follows, and which constraints shape a usable solution. Only then should you decide whether generation, retrieval, prediction, conversation, or agentic action is an appropriate intervention.
Write a bet brief that can be disproved
A GenAI bet should fit into a short brief before anyone commits significant engineering capacity. Use this framing:
For a specific actor in a specific workflow moment, the current behavior makes it harder to achieve a desired outcome. A bounded AI capability may improve an observable behavior, provided it stays within defined quality, control, privacy, and safety conditions.
Make every part concrete enough to test:
Actor and moment: Name who encounters the problem and where it occurs in the workflow. A role such as support agent is not enough; identify the decision or task the person is performing.
Current behavior: Describe what people actually do, including workarounds, handoffs, rework, and information they assemble manually. Ground this in past behavior rather than hypothetical interest.
Desired outcome: Define the customer or business result that should improve. Completion, adoption, rework, escalation, abandonment, and repeat use are product signals. A model-quality score is not the outcome.
AI contribution: State whether the system retrieves information, creates a draft, recommends an action, or performs an action. These are materially different product and risk propositions.
Operating boundary: Specify the data context, permissions, supported cases, human checkpoints, and fallback path. A capability without a boundary cannot be evaluated honestly.
Decision: Write what evidence would cause the team to proceed, narrow the scope, change the approach, or stop.
Suppose the opportunity is helping a support agent respond to a complex case. Producing a fluent answer is not the unit of value. The product may need to retrieve the right account context, create a grounded draft, expose its basis, let the agent correct it, and contribute to resolving the issue without avoidable rework. That framing changes what you prototype, what you evaluate, what you instrument, and what you ask users during discovery.
Choose the smallest coherent release
The right initial scope is not the smallest visible feature. It is the smallest end-to-end experience that can produce credible evidence. A narrow workflow with real context, a clear human checkpoint, outcome instrumentation, and a recovery path is more useful than a broad assistant that performs many disconnected tasks. This is how rapid prototyping becomes a method for reducing uncertainty instead of a factory for disposable demonstrations.
Scope the first release along four boundaries:
A defined user and workflow moment.
A bounded set of data and tools the system may use.
A clear level of authority over the resulting action.
A measurable product outcome and a separate quality floor.
Authority matters because the same model behavior can create different consequences. Showing relevant information is different from drafting a message. Drafting is different from recommending that it be sent. Recommending is different from sending it automatically. As authority rises and reversibility falls, require stronger evaluations, clearer permission boundaries, better auditability, and more deliberate human approval. If an action is costly, sensitive, or difficult to reverse, keep a human checkpoint and a safe fallback in the initial operating envelope.
Keep product success, system quality, and operating viability distinct. Product success asks whether behavior and outcomes improve. System quality asks whether outputs meet defined conditions. Operating viability asks whether the product can deliver that behavior with acceptable latency, cost, support burden, and failure recovery. A bet needs evidence in all three categories before a polished prototype should influence a roadmap commitment.
Build a traceable evidence chain
There are two uses of AI in this work, and leaders should not conflate them. AI can assist the discovery process by transcribing interviews, organizing material, and generating alternative interpretations. AI can also be the product capability under investigation. In the first role, its output is an analytical input. In the second, its behavior is an object of evaluation. Neither role turns model output into customer evidence.
Preserve the customer story before looking for patterns
Good discovery begins with a specific story about past behavior. Capture the person’s goal, the context, key moments in the experience, decisions, workarounds, and the needs or pain points that emerged. Synthesize that interview on its own before combining it with other interviews. This prevents a recurring operational mistake: compressing many transcripts into generic themes that are easy to present but impossible to design against.
A practical single-interview record includes participant context, a transcript-linked quote, the sequence of important moments, and opportunities expressed within that person’s situation. Cross-interview synthesis can then organize related opportunities without severing them from their origins. This separation between individual and cross-interview synthesis keeps insights contextual, actionable, and traceable.
Use AI as a notetaker or an additional analytical perspective, but establish simple provenance rules:
Every direct quote must link to the transcript location where it appears. If the wording cannot be verified, treat it as a paraphrase or discard it.
Every opportunity must identify the participant, workflow moment, and evidence from which it was derived.
Generated summaries must be checked against the original material before entering an opportunity map or decision record.
Tone, hesitation, visible confusion, and body language must come from a human observation note when they affect interpretation; a text transcript cannot preserve them reliably.
AI-generated themes may prompt a second look at the evidence, but they do not become evidence through repetition.
If the interview itself is shallow, automation will only process shallow material faster. A model cannot recover a missing motivation, workflow constraint, or decision context that nobody elicited. When synthesis produces vague insights, inspect the interview quality before rewriting the prompt.
Connect every artifact to a decision
The evidence chain should survive the entire path from discovery to production:
A customer story supports an opportunity.
The opportunity supports a product bet.
The bet contains assumptions about value, usability, feasibility, data, trust, and operations.
High-risk assumptions become experiments and evaluation scenarios.
Evaluation scenarios become regression cases when the product changes.
Field behavior connects the released capability to product and business outcomes.
Production failures feed new scenarios, product constraints, and discovery questions.
This chain prevents evidence laundering. A customer describing a difficult workflow does not prove that a proposed solution is desirable. A user liking a prototype does not prove that behavior will change. Passing an offline evaluation does not prove that the interaction fits the workflow. A successful pilot does not prove that the system is repeatable across customers. Each step answers a different question.
Decision
Minimum artifact
Evidence that advances the bet
What to do when it is missing
Is the problem worth solving?
Interview snapshots and an opportunity map
Specific past behavior, context, consequences, and constraints
Return to interviewing or workflow observation
Can GenAI contribute meaningfully?
Assumption ledger and a thin prototype
The capability performs the bounded task on representative examples
Narrow the task, change the system approach, or stop
Can users understand and control it?
Interaction prototype and failure scenarios
Users can interpret the result, correct it, and recover from failure
Change the interaction, authority level, or human checkpoint
Does it change real behavior?
Instrumented field pilot
Observed usage changes a leading product signal without unacceptable quality loss
Investigate workflow fit, adoption friction, or the original value hypothesis
Can the product be operated?
Evaluation harness, monitoring, ownership, and rollback path
Quality remains observable after release and a named owner can respond to degradation
Keep the release bounded until the operating controls exist
Use a living bet packet instead of a presentation deck
Keep the evidence chain in one lightweight packet. It should contain the current problem frame, links to customer evidence, the assumption ledger, prototype configuration, evaluation results, field observations, outcome measures, and the latest decision with its rationale. It is not a status report. It is the memory of the bet.
A new team member should be able to see why the work exists, which uncertainty is active, what failed previously, and what would change the decision. If the packet grows without changing a decision, reduce it. The purpose is traceability, not documentation volume.
Run a weekly dual-track discovery loop
GenAI discovery has two learning tracks. The customer track investigates the workflow, value, behavior, interaction, and trust. The system track investigates model behavior, context quality, tool use, failure modes, and operational constraints. Running only the first creates attractive concepts with unknown feasibility. Running only the second produces technically impressive capabilities looking for a problem.
A weekly loop is a useful default because it keeps the customer and system evidence synchronized. It also forces the team to choose a decision small enough to advance with the evidence available. The cadence is not a sequence of meetings. It is a sequence of testable changes.
Customer and workflow learning
This track uses customer interviews, workflow observation, usability tests, pilot behavior, support signals, and outcome instrumentation. Each activity should begin with the decision it is meant to inform. Do not schedule an interview merely to gather feedback. Decide whether you need to understand the current workflow, test the meaning of an opportunity, observe a recovery interaction, or determine why pilot users are abandoning the capability.
Invite engineering into customer learning when technical context changes the solution space. An engineer hearing how permissions, data fragmentation, or exceptional cases shape the workflow can identify constraints earlier than a written handoff would. The product manager should still own the connection between those constraints and the desired outcome.
System and evaluation learning
This track needs a scenario set, not a folder of memorable outputs. Build it from ordinary customer cases, meaningful edge cases, prohibited behavior, and failures already observed. Synthetic examples can widen coverage, but they should extend a foundation of real workflow evidence rather than replace it.
Each scenario should record:
The user task and its provenance in customer or field evidence.
The input, relevant context, permissions, and available tools.
Acceptable behavior, unacceptable behavior, and the scoring method.
The model, prompt, retrieval, tool, and data configuration used.
The actual result, the failure category when applicable, and any human correction.
The product consequence of the failure, not merely its technical label.
Define acceptance criteria before reviewing a preferred prototype. Otherwise, a persuasive output can cause the team to move the standard after the fact. Automate checks when correctness is directly observable, and retain structured human review where meaning, usefulness, or trust depends on context.
Treat the scenario set as part of the product, not disposable test material. Prompt versions, retrieval configuration, model changes, data quality, and tool behavior all belong to the release surface. Evaluation harnesses, prompt versioning, red-teaming, and production monitoring therefore need to begin in discovery and continue through CI/CD. A scenario that exposes a real failure during a pilot should become a regression case before the next release.
A decision-oriented weekly sequence
Name the decision. Choose the highest-risk assumption blocking progress rather than the most convenient activity.
Inspect the evidence. Review the relevant customer stories, field behavior, scenario results, and known failures.
Design the cheapest credible test. This may be another interview, a workflow prototype, a prompt or retrieval change, a human-assisted simulation, or a bounded field experiment.
Run offline scenarios first. Find obvious quality, safety, grounding, and permission failures before asking a customer to spend time on the prototype.
Observe the capability in the workflow. Watch what users accept, edit, ignore, misunderstand, override, or abandon.
Update both records. Add customer learning to the opportunity and assumption records; add system failures to the scenario set and failure taxonomy.
Make the decision explicit. Advance, narrow, change, pause, or stop. Record the evidence that caused the decision and identify the next uncertainty.
Parallel experimentation is useful only when each experiment resolves a distinct uncertainty. Generating many prototypes against the same vague question increases activity without increasing decision quality. The advantage of GenAI is a lower cost of learning; spend that advantage on better coverage of assumptions, not a larger pile of concepts.
Promote a prototype only when the evidence changes
A prototype should not graduate because it performs well in a stakeholder demonstration. Move it toward a broader pilot only when the team can show:
A valuable workflow and outcome grounded in customer behavior.
A bounded capability with explicit supported and unsupported cases.
An interaction that helps users understand uncertainty, correct results, and recover.
Instrumentation for product outcomes, system quality, overrides, and failures.
Ownership for monitoring, escalation, and rollback after release.
Keep the value gate and the quality gate separate. A system can meet its quality threshold and still fail because it adds a step, appears at the wrong moment, or solves a low-value problem. It can also attract strong usage while producing unacceptable errors. The first result sends you back to product discovery; the second requires narrowing authority, strengthening the system, or stopping the release. Blending the gates makes both diagnoses harder.
Design the organization around decision rights and learning
A fast learning loop will stall if every product choice waits for a central AI committee. It will also become fragile if each squad invents its own model access, evaluation approach, privacy controls, and incident response. The operating model must distinguish decisions that require local customer context from capabilities that should be reusable across the company.
Keep the product trio accountable for the outcome
The titles can vary, but the responsibilities cannot remain implicit:
Product leadership owns the desired outcome, opportunity framing, assumption sequence, value evidence, and decision record.
Design and research own workflow understanding, interaction behavior, user control, comprehension, correction, and recovery.
Engineering and data own system architecture, context and data quality, repeatable evaluations, observability, reliability, and rollback.
Domain, privacy, security, and risk partners define non-negotiable acceptance conditions and escalation paths where the use case requires them.
Go-to-market and customer-facing partners help identify adoption constraints and connect pilots to real workflows without substituting enthusiasm for evidence.
No single function can answer every release question. Product cannot declare model quality by itself. Engineering cannot infer customer value from technical performance. Governance cannot decide workflow desirability. Clarify who owns each decision, which evidence is required, and who can block a release when a non-negotiable condition fails.
Use forward deployed engineers where context is the constraint
Complex enterprise workflows often hide critical information in customer configurations, permissions, data conventions, and exception handling. A forward deployed engineer can work with product and design in the customer’s environment, shorten the path from observation to prototype, and turn real failures into reusable product knowledge. That field-facing role is especially useful when edge cases cannot be reproduced from a conference-room description.
The role needs a productization boundary. Otherwise, the team may prove that a skilled engineer can manually make each account successful rather than proving that the product is repeatable. Require every customer-specific exception to become one of four things: a reusable product requirement, an explicit configuration option, an evaluation scenario, or a documented unsupported condition. Field learning, end-to-end instrumentation, and responsible guardrails should strengthen the product system rather than disappear into account-specific work.
Centralize the paved road, not product judgment
A shared AI or platform capability should make the safe path easier. Depending on the organization, that can include approved model access, data controls, common logging, an evaluation framework, prompt and model versioning, reusable interaction patterns, cost and latency visibility, incident procedures, and privacy or safety templates.
The embedded product team should still own the opportunity, supported workflow, acceptance criteria, scenario relevance, product experience, pilot design, and outcome decision. A central group cannot judge those elements without the local customer context. It should provide leverage and enforce genuine non-negotiables, not become an approval queue for every prompt experiment.
Governance works best as a quality system with explicit rules:
Which use cases require review before customer exposure.
Which data classes, actions, and model behaviors are prohibited.
Which evidence must accompany a request to expand authority or reach.
Who approves exceptions and who owns the resulting risk.
What monitoring, audit, escalation, and rollback controls must exist after release.
This structure gives teams room to experiment inside known boundaries. It also prevents a common late-stage failure in which privacy, safety, or operational requirements appear only after the team has committed to an architecture and promised a launch.
Replace demo reviews with evidence reviews
A portfolio review should inspect what the team has learned, not how polished the prototype looks. Ask:
Which decision changed since the previous review, and what evidence changed it?
Which customer story and workflow moment support this opportunity?
What is the highest-consequence failure the current prototype still exhibits?
What product outcome must improve, and which quality condition must not regress?
What is the narrowest safe and coherent release that can produce field evidence?
Who owns the capability when its model, prompt, data, or tool behavior changes?
The answers reveal the right next move. No rich customer evidence means the team needs better discovery, not a better prompt. No scenario set means model comparisons are premature. Strong offline performance with weak field adoption points to workflow or value friction. A pilot that depends on hidden manual intervention is not yet a repeatable product. A capability without monitoring and rollback is not ready for production authority.
Take one active GenAI bet and replace its feature pitch with a falsifiable bet brief, a traceable evidence packet, and an explicit quality-and-value gate. Run the next portfolio conversation against those artifacts. Anything important that is missing is not paperwork to add later; it is the discovery work that should happen before the organization makes a larger commitment.
Your PM presents a bold strategy, but every difficult decision still comes back to you. Or the team ships reliably, yet the work rarely changes an important customer or business outcome.
These are different leadership problems. The first is an agency gap. The second is an ambition gap. Treating both as a generic performance issue leads to vague coaching, more oversight, and little improvement. You need to identify which capability is missing, change the conditions around it, and ask for observable evidence of progress.
Separate ambition from agency before you coach
Ambition is the drive to pursue greater impact, wider scope, or meaningful growth. Agency is the willingness and ability to own a problem, make decisions, and create momentum without repeatedly waiting for permission. Strong product managers need both capabilities, but one does not guarantee the other.
A confident presenter may have ambition without agency. A dependable delivery manager may have agency without ambition. If you praise the first person for vision and the second for output, you can reinforce the exact limitation you need each person to overcome.
Pattern
What you are likely to notice
Your leadership response
High ambition, high agency
The PM pursues consequential outcomes, reduces uncertainty, makes sound decisions, and creates momentum.
Protect autonomy, widen the problem space, and keep the outcome bar high.
High ambition, low agency
The PM describes a compelling future but stalls when evidence is incomplete, trade-offs appear, or stakeholders disagree.
Clarify decision rights, narrow the next reversible decision, and require a recommendation rather than another escalation.
High agency, low ambition
The PM delivers steadily but optimizes small requests or predetermined scope without questioning the size of the opportunity.
Reconnect the work to customer and business impact, then ask for a more consequential hypothesis.
Low ambition, low agency
The PM waits for tasks, avoids ownership, and cannot explain the outcome the work should produce.
Check the environment and expectations first. If clarity, access, and coaching do not change the pattern, examine role fit.
Do not assign someone to a quadrant from reputation or personality. Inspect recent work. Ask four questions:
You can approve an AI strategy, fund several prototypes, and still get almost no durable product change. The warning sign is familiar: demos multiply, customer impact remains hard to prove, and every release waits on roadmap, budget, handoff, and governance machinery built for more predictable software.
If that is your situation, the missing layer is an AI-era product operating model: the decisions, team boundaries, evidence, and guardrails that turn an uncertain capability into repeatable customer and business value. You do not need a parallel AI organization. You need a product system that learns quickly without giving up production quality or trust.
Redesign the unit of work around learning, not AI features
An AI assistant, agent, or workflow is not a useful unit of strategy. Those labels describe possible solutions. They do not identify whose behavior should change, which business result should move, or how the team will know the product is safe enough to expand. That distinction matters because a platform shift changes product strategy, architecture, discovery, and go-to-market decisions; it cannot be absorbed by adding AI features to an otherwise unchanged roadmap.
Make an outcome the unit of funding and accountability. A useful outcome statement has this shape: For a specific user in a specific workflow, improve a named measure from its current baseline, without crossing defined quality, trust, or business guardrails. The AI capability is one hypothesis for producing that result, not the result itself.
Require every AI bet to enter the portfolio with a one-page charter containing:
User and workflow: Who experiences the problem, what are they trying to complete, and where does the current workflow break down?
Outcome and baseline: Which customer or business measure should change, and what is its current state? If the eventual outcome will not move during discovery, name the leading indicator and explain the expected connection.
Why AI: What can an AI approach do that a rule, search experience, workflow redesign, or conventional automation cannot do adequately?
Riskiest assumptions: What must be true about value, usability, feasibility, and viability for the bet to work?
Trust boundary: What data may be used, what failure would be unacceptable, who could be affected, and what non-AI or human path remains available?
Next evidence: What is the smallest test that could materially change a decision?
Decision rule: What evidence would justify scaling, another iteration, or stopping?
The charter separates two types of uncertainty that often get mixed together. Model uncertainty asks whether the technology can perform a task under relevant conditions. Product uncertainty asks whether people will use it in a real workflow and whether that use will improve an outcome. A fluent demonstration can reduce the first uncertainty while saying almost nothing about the second.
If a team cannot name a baseline or observe the workflow, the bet may still deserve discovery funding. It does not yet deserve a production commitment. That distinction lets leaders support exploration without allowing every promising prototype to become an implied roadmap promise.
Move each bet through evidence states
Roadmap statuses such as planned, in progress, and complete describe activity. AI portfolios also need states that describe what has been learned:
Explore: The problem is credible, but the team is still testing the workflow, value proposition, technical approach, or failure boundary. Work should be small and reversible.
Prove: A solution has produced useful signals with target users. The team is testing a constrained production experience, instrumenting behavior, and validating that quality and trust controls hold outside a demo.
Scale: Customer behavior and the chosen outcome support broader investment, while known risks remain inside agreed limits. The team can now improve reliability, reach, economics, and operational readiness.
Capacity should increase as evidence improves. An executive sponsor’s confidence is not a substitute for customer behavior, and a model’s technical sophistication is not a substitute for outcome movement. Portfolio reviews should therefore ask what uncertainty was removed and what decision changed, not merely whether delivery is on schedule.
Give each outcome a durable product trio and elastic expertise
AI work can create additional dependencies on data, infrastructure, security, privacy, legal, and domain expertise. If each dependency becomes a handoff, the organization gets slower precisely when fast learning matters most. Keep a durable product trio accountable from discovery through production, then bring specialists into the decisions where their expertise changes the work.
Problem framing, outcome, viability assumptions, and evidence synthesis
Recommend whether to continue, change, scale, or stop the bet based on the charter
Product designer
End-to-end workflow, user comprehension, usability, and trust in the interaction
Choose how concepts are exposed to users and what usability evidence is required
Engineering lead
Technical feasibility, architecture, instrumentation, production quality, and operational trade-offs
Choose the technical path and release shape inside agreed constraints
Forward deployed engineer
Time-boxed customer immersion, rapid prototypes, and translation of workflow details into testable hypotheses
Choose the fastest responsible prototype for the current learning objective
Executive sponsor
Outcome priority, resource boundaries, organizational air cover, and cross-team escalation
Set the problem and constraints; avoid prescribing the solution
Security, privacy, legal, data, and domain specialists should have explicit consultation or approval points based on the consequence of the use case. They should not inherit ownership of the customer outcome. The product team remains accountable for integrating those constraints into a coherent experience.
Run an evidence cadence, not a status cadence
Give every discovery cycle one named learning question. Examples include whether users will delegate the task, whether they understand what the system did, whether the available data can support the workflow, or whether a failure can be detected before it causes harm. A prototype without a learning question is usually a demo; an experiment without a decision attached is usually activity.
For a pilot, a two-week evidence review is concrete enough to create accountability without turning every test into an approval meeting. Review the live charter, instrumented behavior, customer signals, and decision log. Ask five questions:
What did the team believe at the start of the cycle?
What did customers do, not merely say?
Which assumption became less uncertain?
Did the primary outcome or any guardrail move?
What decision changed, and what is the next critical question?
Keep the review focused on evidence. A long slide deck can hide the fact that no decision changed. A short decision log exposes that immediately.
Measure learning velocity as the time between asking a consequential question and obtaining credible evidence that changes a decision. That does not mean rewarding the raw number of experiments. Ten low-value tests can create less progress than one well-designed customer session or constrained release. Pair learning velocity with business outcomes so teams cannot optimize for experimentation while avoiding accountability for value.
Forward deployed assignments should also be time-boxed and documented. Record the workflow discovered, assumptions tested, prototype behavior, technical shortcuts, evidence collected, and production work still required. Rotate engineers through these assignments when practical. That spreads customer context and product judgment instead of concentrating both in a permanent hero team.
Govern AI bets by consequence, not by ceremony
AI governance fails when every experiment needs the same committee approval. It also fails when teams silently decide what data, errors, and customer consequences are acceptable. The useful middle ground is proportional governance: the higher the consequence and the harder the reversal, the stronger the evidence and independent review required.
Define consequence tiers in language your product, engineering, security, privacy, legal, and trust leaders accept:
Low consequence: The work is internal or tightly contained, uses approved non-sensitive data, cannot take consequential action, and is easy to reverse. The product team can usually proceed inside established policies.
Moderate consequence: The system influences a customer workflow, but its output is reviewable, the action is reversible, and a clear fallback exists. Require named product and technical owners plus the relevant privacy, security, or domain review.
High consequence: The system can move money, change access, affect eligibility, influence safety or legal rights, expose sensitive data, or take an action that is difficult to undo. Require qualified legal, security, privacy, safety, or domain review before customer exposure, along with human control and staged rollout where appropriate.
Do not treat these examples as universal legal classifications. Your specialists need to define the boundaries for the jurisdictions, customers, data, and decisions in scope. The operating-model requirement is that every team can determine the tier before building a release plan, not after the code is complete.
Use four gates from problem to scale
Problem gate: Name the user, workflow, baseline, desired outcome, and non-AI alternative. Explain why an AI approach is warranted. This prevents technology enthusiasm from becoming the problem statement.
Evidence gate: Test the system on tasks drawn from the intended workflow. Define useful behavior, known failure modes, unacceptable failure, and the evidence needed for value, usability, feasibility, and viability.
Exposure gate: Confirm data permissions, customer communication, logging, human review or fallback, support readiness, release owner, and rollback path. A successful prototype does not automatically satisfy this gate.
Scale gate: Require both outcome evidence and acceptable guardrail performance. Assign owners to unresolved failure modes before expanding reach or autonomy.
The gates should make autonomy safer, not eliminate it. Leaders set portfolio priorities and risk appetite. Specialists set non-negotiable data, compliance, security, and safety constraints. The product trio chooses the solution, experiment sequence, technical approach, and rollout details within those boundaries. If those decision rights remain ambiguous, governance meetings will repeatedly reopen product choices or teams will bypass the process to maintain speed.
Give every production AI bet a compact metric stack:
Business outcome: A measure such as activation, retention, expansion, conversion, or cost-to-serve that connects the work to enterprise value.
User behavior: Evidence that the target workflow changed, such as task completion, adoption, repeat use, escalation, or abandonment.
Quality and trust: The failure measures relevant to the use case, including human corrections, overrides, complaints, or occurrences of the unacceptable behavior defined in the charter.
Learning: Time to answer the current critical question, assumptions closed, and the decision produced by the evidence.
This is a menu, not a requirement to track every example. Choose one primary outcome and only the supporting measures needed to interpret it. If the primary outcome will take longer than the pilot to move, predeclare a leading indicator and its rationale. Do not replace a disappointing metric after the results arrive.
Clear baselines, measurable outcomes, and explicit ethical and trust guardrails let the team move faster because the boundaries are known. Vague risk language has the opposite effect: every reviewer imagines a different failure, so each decision is renegotiated from scratch.
Prove the operating model with a bounded 90-day pilot
Do not begin by announcing a company-wide AI transformation. Choose one or two problems that are important enough for leadership to care about, bounded enough for a team to affect, and observable enough to produce evidence. A pilot should test the operating model as well as the product bet.
A strong pilot candidate has:
A visible customer workflow with a specific friction point
A baseline or an attainable plan for establishing one
Access to target users throughout discovery
A path to shipping constrained increments rather than waiting for a complete platform
A meaningful connection to activation, retention, expansion, conversion, cost-to-serve, or another agreed business outcome
Dependencies that an executive sponsor can realistically unblock
A consequence level the organization can govern responsibly during the time box
Avoid picking a harmless showcase merely because it is easy to demo. It will not test difficult decision rights, customer discovery, production instrumentation, or governance. Also avoid starting with the most consequential and dependency-heavy workflow in the company. A pilot needs enough organizational reality to be credible without becoming a referendum on every unsolved platform issue.
Run the pilot in this sequence:
Publish the charter: State the problem, baseline, outcome, assumptions, consequence tier, team, decision rights, and scale-or-stop criteria on one page.
Staff a credible cross-functional team: Assign the product trio, add a forward deployed engineer where customer-side prototyping will reduce uncertainty, name the executive sponsor, and schedule specialist involvement before it becomes a blocker.
Establish evidence access: Arrange customer contact, instrument the current workflow, and create a shared place for test results and decisions.
Discover and deliver together: Explore multiple approaches, test the riskiest assumptions, and ship small increments when the evidence and consequence tier permit.
Review evidence every two weeks: Inspect customer signals, shipped behavior, outcome movement, guardrails, and decisions. Do not convert this into a project-status meeting.
Make the precommitted decision: At the 6-12-week decision window, choose to scale, iterate, or stop. Use the remainder of a roughly 90-day time box to verify repeatability, transfer the practices, or close the bet cleanly.
Define scale, iterate, and stop before results arrive
Scale: The workflow produces credible customer value, the business or predeclared leading measure is moving in the intended direction, guardrails hold, and the production path is viable.
Iterate: The problem remains important and evidence identifies a specific failed assumption or constrained next test. Iteration is not permission to continue indefinitely without a sharper question.
Stop: The value signal is weak, the workflow does not earn adoption, the economics are untenable, a critical risk cannot be controlled, or the non-AI alternative is better. Stopping is a valid return on discovery when it prevents a larger commitment.
The politics of a pilot can undermine otherwise sound work. Publish the criteria used to select the problem and team. Time-box special assignments. Do not hoard every high performer in a permanent AI lab. Show failed assumptions and changed decisions alongside successful demos. These practices make the pilot a path other teams can follow rather than evidence that only a protected group can succeed.
Scale the mechanics, not the heroics
After the pilot, codify the parts that made learning and delivery repeatable:
The one-page bet charter and evidence-state definitions
Team topology, specialist access, and forward deployed rotation rules
Decision rights for executives, product teams, and risk owners
The two-week evidence review and decision-log format
Consequence tiers, release gates, and escalation paths
Instrumentation for outcomes, behavior, quality, trust, and learning
The scale, iterate, and stop criteria
Do not standardize every discovery technique or technical implementation. Different workflows will need different tests and controls. Standardize the minimum system that makes evidence visible, decisions timely, and responsibility clear.
The real repeatability test is whether a second team can use the same mechanisms without relying on the original pilot’s personalities or executive attention. If it cannot, the organization has produced a hero story, not an operating model.
Key takeaways
Fund AI bets against customer and business outcomes, not solution labels such as assistant, agent, or copilot.
Require a one-page charter with a baseline, riskiest assumptions, trust boundary, next evidence, and precommitted decision rule.
Keep a durable product trio accountable end to end; use forward deployed engineers as time-boxed discovery accelerators.
Review evidence and changed decisions every two weeks during a pilot, rather than reviewing activity alone.
Apply stronger review as consequences and irreversibility increase, while preserving team autonomy inside explicit guardrails.
Use a roughly 90-day pilot to test repeatability, then scale the decision rights, cadence, instrumentation, and governance that another team can adopt.
Your next move is not to rewrite the entire product process. Pick one material, bounded workflow. Publish its one-page charter, staff the trio, set its consequence tier and baseline, schedule the evidence reviews, and precommit to a scale, iterate, or stop decision. The behavior leadership protects during that pilot, not the polish of its demo, is the operating model the rest of the organization will copy.