Author: Shivam Tiwari

  • Product Analytics for Retention: A Practical Operating System

    Product Analytics for Retention: A Practical Operating System

    You have a retention chart and a familiar problem: the curve is falling, the segments disagree, and every team has a different explanation. Another dashboard will not tell you what to build.

    You need a decision loop that connects retained value to observable behavior. Define the outcome, instrument the journey, locate the behavioral gap, and test the smallest change that could close it. That turns retention analytics into a product operating system rather than a monthly reporting exercise.

    Start with a retention contract, not a dashboard

    Before opening your analytics tool, finish this sentence: “For users who first do [starting action], retention means completing [valuable action] again within [return window].” If your team cannot agree on the blanks, it is not ready to interpret a retention curve.

    The starting action should identify a meaningful cohort. Account creation is often too weak because it combines curious visitors, evaluators, invited teammates, and serious users. Prefer the moment a person begins the journey you intend to improve, such as creating a project, starting an agent, or completing an initial workflow.

    The return action must represent delivered value, not convenient activity. Opening the app, viewing a page, or receiving a notification may be easy to count but weakly connected to the reason someone adopted the product. Choose an action that would make a customer notice if the product disappeared.

    Set the return window around the product’s natural use cycle. A daily workflow and an occasional administrative task should not share the same definition. Document the window, the qualifying action, the excluded users, and whether retention is measured at the user or account level. This is your retention contract.

    Next, build a driver tree connecting the retention outcome to measurable inputs. Put retained value at the top. Beneath it, map activation, repeated value-producing behavior, and the friction that can interrupt either one. This separates the lagging outcome you care about from leading signals a team can move sooner.

    For every leading signal, add a guardrail. If a change increases sessions but reduces task completion, it has created activity rather than value. If it improves first-session completion but does not affect return behavior, treat it as an onboarding improvement until the retention evidence catches up.

    Instrument the journey so the data can survive a decision

    Retention analysis breaks when event names mirror the interface instead of the customer’s progress. A click on “Continue” becomes meaningless after the button moves. An event such as workflow_started or task_completed remains interpretable across interface changes.

    For each critical event, record enough context to reconstruct what happened:

    • The user and, for collaborative products, the account.
    • The channel, surface, or entry point that started the journey.
    • The use case or object involved.
    • The event timestamp and relevant status.
    • The experiment assignment, when the experience is being tested.
    • The event version when its meaning or properties change.

    Give every retention-critical event a plain-language definition and an owner. The definition should state when the event fires, when it must not fire, which properties are required, and how duplicate or failed actions are handled. Keep cohort definitions centralized for the same reason. Product, marketing, and customer success cannot compare decisions if each team silently defines “activated” or “retained” differently.

    Validate the journey before trusting the curve. Trace real test accounts from the starting action through the value event and return action. Compare the interface state, raw events, and resulting cohort membership. Check identity transitions such as anonymous-to-signed-in usage, invitations, account switching, and merged profiles. A polished retention chart built on broken identity resolution is still broken.

    Treat the taxonomy like a product surface. Changes need review, backward compatibility, documentation, and monitoring. This work feels slower than building a dashboard, but it prevents teams from spending an entire planning cycle acting on instrumentation defects.

    Diagnose the behavioral gap before proposing a feature

    A retention curve tells you where return behavior weakens. It does not explain why. Use a fixed analysis sequence so the team does not jump from an interesting segment to a preferred solution.

    1. Inspect the curve shape. An early drop points you toward expectation-setting, onboarding, or initial value. A later decline points you toward repeat value, changing needs, or workflow friction.
    2. Segment with a hypothesis. Compare acquisition source, device, channel, use case, or customer type only when you can explain why that dimension might change the experience.
    3. Compare retained and non-retained cohorts. Look for behaviors that differ in sequence, completion, or repetition, not merely events with high volume.
    4. Build a funnel around the strongest candidate behavior. Find the step where the cohorts separate and inspect how users arrive there.
    5. Review session replay, conversation transcripts, or journey detail at that step. Look for hesitation, repeated attempts, unclear choices, missing context, and premature exits.

    This sequence moves you from outcome to segment, behavior, moment, and observable friction. Stop if the evidence cannot support that chain. A behavior that correlates with retention is a place to investigate, not proof that forcing the behavior will retain users.

    AI products make this distinction especially important. A generic greeting may produce a response without moving the user toward a task. If people hesitate, test a concise follow-up that clarifies the agent’s scope, offers two or three concrete choices, and still accepts free-form input. Measure the chain from continuation to task start, task completion, and return across the first three to five sessions. Do not optimize for extra conversation turns if users remain stuck.

    Pair behavioral evidence with continuous discovery. Analytics identifies the moment worth investigating; interviews and direct observation help explain the need, expectation, or constraint behind it. That combination produces a testable problem statement instead of a feature request decorated with data.

    Turn retention signals into controlled product bets

    Write the opportunity before discussing solutions: “When [cohort] reaches [moment], [observable friction] prevents [valuable behavior], which is associated with lower [retention outcome].” The wording forces you to name the user, the moment, the evidence, and the outcome without pretending you have already established causality.

    Then create an experiment card with:

    • A hypothesis linking the proposed change to a specific behavior.
    • The eligible cohort and trigger moment.
    • One primary retention outcome.
    • A leading indicator that can move earlier.
    • Guardrails for completion quality, errors, or unintended friction.
    • The minimum detectable effect and planned evaluation window.
    • A decision rule for stopping, iterating, rolling out, or reversing the change.

    Choose a change small enough to isolate the mechanism. If the suspected problem is uncertainty at the start of an AI interaction, test the opening sequence rather than redesigning the agent, onboarding flow, and navigation together. A smaller bet makes the result easier to interpret and cheaper to reverse.

    Review experiments on a regular product cadence. Begin with data quality, then evaluate the leading indicator, guardrails, and retention outcome in that order. Inspect the segments named in the original hypothesis rather than searching every possible cut for a favorable result. Record what the team decided, why it decided it, and what evidence would change the decision.

    Your roadmap should name the retention outcome and the behavioral driver, not promise a feature prematurely. “Increase repeat task completion for newly activated accounts” leaves room to test messaging, workflow design, defaults, or assistance. “Build a new onboarding wizard” locks the team into an answer before it has earned confidence in the problem.

    Key takeaways

    • Define retention as a cohort, a value-producing action, and a return window before interpreting any chart.
    • Use a driver tree to connect the lagging retention outcome to behaviors a product team can influence.
    • Standardize event and cohort definitions, then validate identity and journey data with real test accounts.
    • Move from curve to segment, behavior, moment, and friction before proposing a solution.
    • Use controlled, reversible experiments to distinguish a useful behavioral signal from a causal retention lever.

    Start with one journey that matters this week. Write its retention contract, trace the events, and identify the first point where retained and non-retained users behave differently. That single decision-ready path is more valuable than a broad analytics program nobody trusts.

    References

  • Old-School Selling Beats PLG in the AI Era: My GTM Playbook for 8‑Day Enterprise Deals

    Old-school, in-person selling is having a renaissance in the AI era, and I’ve seen why up close. From leading product and go-to-market teams through hypergrowth, I keep returning to one lesson: enterprise buyers still reward the teams who show up, orchestrate change management, and own outcomes end-to-end. The tech has changed; the human dynamics haven’t.

    Has the sales playbook changed in the AI era? The tools are faster and the surface area is bigger, but the core motion remains the same: “showing up” beats letting the marketplace decide. That’s why in-person enterprise rollouts still beat product-led motions, especially when the stakes include security, governance, and cross-functional adoption. You win by reducing organizational risk, not by assuming free trials will do the heavy lifting.

    Great enterprise sellers collapse silos. They sell to engineers and executives in one motion, pairing deeply technical validation with crisp business narratives. In my org, that means every high-velocity pilot has a dual thread: hands-on, eval-driven proof for the builders and a value architecture for the budget owners. When those motions run in parallel, time-to-value plummets and procurement friction fades.

    Selling to AI-native buyers who grew up on ChatGPT changes tempo, not fundamentals. The same seller, different tempo: 8 weeks vs. 8 business days. These buyers evaluate fast, expect clear ROI, and push for automation-first workflows. How AI-native buyers handle build vs. buy decisions comes down to build for differentiation and buy for acceleration. If you make procurement feel like product—frictionless, instrumented, and transparent—you’ll meet their bar.

    Process matters, but humanity wins. Building a robust sales process that still leaves room for unscripted moments is where trust is formed. I’ll never forget the story of the rep who taught a champion’s son guitar over Zoom—an unscripted moment that cemented a partnership. The lesson: raise the floor without capping the ceiling. Equip every rep with repeatable plays, then celebrate the creative instincts that make champions out of customers.

    In early GTM, why the three highest-leverage early sales hires aren’t sellers at all resonates with my experience. I prioritize a solutions engineer who can de-risk integration, a forward-deployed operator who can run the first rollout like a product manager, and a customer success lead who designs adoption paths from day zero. Together, they compress the value journey from proof to production.

    Compensation design shapes your talent market. The case for outsized commission accelerators for star sellers — and the kind of person they attract is real: magnets for competitors who close complex, multi-threaded deals and thrive with ownership. But beware: why too much process narrows the kind of seller you attract. Over-script it and you filter out the very people who can navigate ambiguity with customers.

    Under the hood, instrumenting the funnel from stage zero to close keeps the system honest. I track intent signals before pipeline, conversion by persona and use case, proof milestones, and time-to-value in production. The three pillars of GTM excellence for me are repeatable discovery, referenceable outcomes, and relentless enablement. And inside the leadership team, building peers who are 80% aligned, not 100% preserves healthy tension while keeping execution fast.

    AI is expanding the definition of enablement—whether AI is changing what good enablement looks like isn’t a theoretical question anymore. I see world-class teams arming reps with retrieval-first knowledge bases, sandbox environments, and objection libraries that evolve weekly. Meanwhile, selling against direct and implied competitors at once is the norm: your battlecard must cover “do nothing,” internal tools, adjacent categories, and new AI entrants—while you still remember why in-person enterprise rollouts still beat product-led motions for durable adoption.

    Planning horizons tighten in AI markets. How far out should a GTM leader be planning? I work a dual cadence: a rolling 6-week operating plan that’s ruthlessly tactical and a 2–3 quarter roadmap for coverage, enablement, and category storytelling. What a normal week looks like in hypergrowth blends customer time, pipeline triage, onboarding and enablement, deal engineering, and process tuning—always with one or two high-conviction bets that could bend the curve.

    References: Ahead: https://www.ahead.com; Amazon: https://www.amazon.com; Anthropic: https://www.anthropic.com; Attio: https://www.attio.com; Augment Code: https://www.augmentcode.com/; Cognition: https://cognition.ai; Cursor: https://cursor.com; Dani McCabe: https://www.linkedin.com/in/danielle-mccabe/; Datadog: https://www.datadoghq.com; GitHub Copilot: https://github.com/features/copilot; HubSpot: https://www.hubspot.com; Jeremy Powers: https://www.linkedin.com/in/jeremypowers/; JPMorgan: https://www.jpmorgan.com; Matt McClernan: https://www.linkedin.com/in/mattmcclernan/; MongoDB: https://www.mongodb.com; Nicole Rettinger: https://www.linkedin.com/in/nicole-rettinger-23b20465/; Notion: https://www.notion.com; OpenAI: https://openai.com; Parag Agrawal: https://www.linkedin.com/in/paragagr/; Parallel: https://parallel.ai; Snowflake: https://www.snowflake.com; University of Chicago: https://www.uchicago.edu; Windsurf: https://windsurf.com

    If you’re scaling an AI product today, pair a disciplined sales-led growth engine with the best of product-led growth: fast paths to proof, hands-on validation for builders, executive-level value mapping, and human moments that turn customers into advocates. That’s how you compress an eight-week cycle into five business days—and keep the expansion flywheel spinning.


    Book a consult png image
  • Supercharge Core Web Vitals with Amplitude’s Global Agent: Faster Rankings, Happier Users

    Supercharge Core Web Vitals with Amplitude’s Global Agent: Faster Rankings, Happier Users

    I measure product health by a simple equation: speed plus clarity equals trust. That’s why I prioritize Core Web Vitals and search performance together—because the fastest path to better UX and higher rankings is a closed loop between measurement, diagnosis, and action. Standardizing on Amplitude’s Global Agent with Amplitude AI Agents let my teams compress that loop from weeks to hours, and in many cases, to minutes.

    Learn how to track your web vitals and page rankings faster with Amplitude AI Agents and improve your site’s user experience and SEO rankings. That goal sounds ambitious, but with the right instrumentation and analytics workflow, it becomes a repeatable operating rhythm rather than a one-off project.

    Here’s what changed for us with Amplitude’s Global Agent: a single, consistent way to capture performance signals across pages and journeys, unified context for every session, and a lightweight footprint that doesn’t get in the way of speed. By centralizing measurement, we eliminated blind spots and gave product, growth, and engineering one shared truth for Core Web Vitals and behavioral analytics.

    My practical playbook is straightforward: 1) Establish a performance baseline for Core Web Vitals on key templates and critical user paths. 2) Segment results by device, location, acquisition channel, and content type to surface where users actually feel the friction. 3) Connect those vitals to downstream behaviors—scroll depth, engagement, and conversion—so we prioritize fixes that move business outcomes, not just lab scores. 4) Use feature flags and A/B testing to ship improvements safely and quantify uplift. 5) Close the loop with Agent Analytics to keep learnings visible and actionable.

    Operationally, we rely on anomaly detection to flag regressions early, CI/CD guardrails to prevent performance slips at deploy time, and observability plus session replay to accelerate root-cause analysis. This combination reduces mean time to resolution, protects page experience during fast iteration cycles, and helps us avoid trading UX for speed—or vice versa.

    The strategic benefit is compounding: better Core Web Vitals improve user perception and increase engagement, which strengthens SEO signals and, ultimately, page rankings. With a unified analytics platform in place, we can spotlight the few improvements that create outsized gains, then scale those patterns across the site with confidence.

    If your roadmap includes faster pages, stronger rankings, and happier users, align your teams around this simple loop: measure precisely, diagnose quickly, experiment safely, and learn continuously. Amplitude’s Global Agent and Amplitude AI Agents give you the instrumentation and insight to make that loop your competitive advantage.


    Inspired by this post on Amplitude – Best Practices.


    Book a consult png image
  • How to Operate Always-On AI Agents Without Losing Control

    How to Operate Always-On AI Agents Without Losing Control

    You want an AI agent to keep work moving after you close your laptop. The difficult part is not getting one successful overnight run. It is making the hundredth run predictable enough that you do not wake up to an embarrassing email, a corrupted task queue, or an unexplained usage bill.

    The right operating model looks less like a clever prompt and more like a small, well-managed operations team. Give each agent a narrow job, an inspectable queue, limited tools, a clear definition of done, and an explicit place to stop. That is how you gain useful autonomy without surrendering control.

    Start with a delegation contract, not a general-purpose assistant

    An always-on agent should not begin with a broad instruction such as “manage my sales work.” That leaves the model to decide what managing means, which systems it may change, and when it has enough evidence to act. The ambiguity is tolerable during an interactive session because you can correct it. It becomes operational risk when the agent runs unattended.

    Start by defining a job that produces a recognizable artifact. A sales-admin agent can prepare a briefing before a scheduled call and create proposed follow-up tasks afterward. A podcast-manager agent can assemble interview context, prepare a transcript-review document, and queue a reminder to share it. A coding-manager agent can review prior sessions and identify recurring mistakes. These are bounded responsibilities with visible outputs, not vague mandates to “help.” Three specialized agents handling podcast, sales, and coding workflows demonstrate how cleanly this pattern can separate unrelated work.

    Write the delegation contract in an identity file that the agent reads at the beginning of every run. It should answer seven questions:

    1. Who are you? Name the role, not the underlying model: sales admin, podcast manager, coding manager, or another function a person would recognize.
    2. What outcome do you own? Describe the recurring deliverable and the event that makes it useful.
    3. Where may you work? Name the exact task, output, and script folders the agent can use.
    4. What inputs may you trust? Identify the calendar, task file, transcript, session log, or other allowed input for the job.
    5. What may you change? Separate reading, drafting, creating internal files, updating tasks, and acting in external systems.
    6. What counts as complete? Specify the artifact, required fields, location, and status update expected at the end.
    7. When must you stop? Define what the agent should do when information conflicts, a tool fails, permission is missing, or the next step would affect another person.

    The last question matters most. A useful agent does not need permission to improvise its way through every obstacle. It needs a reliable way to say, “I could not complete this safely; here is the missing decision.” Treat a well-documented block as a successful operational outcome, not as agent failure.

    Keep consequential decisions outside the unattended role. The agent can prepare a customer email without sending it. It can propose changes to a deal record without changing the commercial commitment. It can summarize a coding pattern without modifying a production system. Moving from preparation to execution should be a deliberate permission decision, not an accidental side effect of adding another tool.

    Build an inspectable operating loop around four components

    The prompt is only one part of the system. Reliable agent operations need four components with distinct responsibilities: identity, scheduling, tasks, and scripts. Keeping them separate makes failures easier to locate and changes easier to review.

    Identity defines responsibility

    The identity file is the stable operating policy. It tells the agent what role it is playing, where its work lives, what it may do, and what completion looks like. Do not overload it with the details of one assignment. If the identity changes every time a task arrives, you no longer have a stable agent; you have an unreviewed prompt generator.

    The scheduler supplies a heartbeat

    The scheduler should wake the agent, point it to the correct identity and queue, and capture the result. It should not contain the business logic for podcast preparation or sales follow-up. That logic belongs in inspectable task instructions and small scripts.

    A Mac that remains online can use macOS LaunchAgents as this heartbeat. LaunchAgents run with the user’s permissions, which is operationally convenient but also defines the risk boundary: the agent may be able to reach anything the scheduled process and its tools can reach. Running scheduled agents on an always-on Mac Mini therefore makes permission design part of the architecture, not a setting to revisit later.

    Make the schedule explicit and easy to disable. Each job should have a known trigger, whether that is a recurring interval, a calendar-related event, or a periodic review. If you cannot quickly answer why an agent ran at a particular time, the scheduler is already too opaque.

    Tasks hold durable state

    Use a dedicated task folder for each agent. A Markdown file with frontmatter is enough to represent a work item while remaining readable by both a person and a tool. The frontmatter can hold machine-readable state; the body can hold the request, context, acceptance criteria, and eventual run notes.

    Choose a small lifecycle and apply it consistently. For example: queued, in progress, blocked, completed, and failed. The exact labels matter less than the transition rules:

    • A queued task is eligible to be claimed.
    • An in-progress task records which run claimed it, preventing another run from silently doing the same work.
    • A blocked task names the missing input or decision and preserves all useful partial work.
    • A completed task links to its output and records what changed.
    • A failed task records the failed operation and whether retrying it is safe.

    Give each recurring event a stable identifier. Before creating a meeting brief, transcript-review document, or follow-up task, the agent should check whether that event has already been processed. This idempotency check prevents a retry or overlapping schedule from creating duplicates.

    Do not treat chat history as the task database. Conversations are useful working context, but durable state belongs in a file or system you can inspect independently. Saving identities, task files, and scripts in a shared knowledge workspace such as Obsidian also makes the operating model portable across devices and coding assistants. Changing the model runner should not require rebuilding the job.

    Scripts expose narrow capabilities

    Scripts should perform small, deterministic operations: fetch an allowed input, create a document in a known location, normalize a transcript, or update a task field. Keep the judgement in the agent and the mechanics in scripts with explicit inputs and outputs.

    A small script is easier to inspect than a broad instruction to use the terminal however the model sees fit. It also gives you one place to add validation, duplicate checks, and error handling. When an agent repeatedly constructs the same command or edits the same file shape, promote that operation into a reviewed script rather than relying on the model to reproduce it perfectly on every run.

    Design the overnight failure path before the happy path

    Unattended automation changes the cost of a mistake. During an interactive session, a confusing output costs a correction. Overnight, the same confusion can trigger repeated work, alter several systems, or contact someone before you see it. Your design should limit the consequence of a wrong interpretation, not merely improve the probability of a correct one.

    Use a permission ladder

    Classify capabilities by consequence and grant them one level at a time:

    1. Read: inspect approved calendars, task files, transcripts, logs, or documents.
    2. Prepare: create drafts, summaries, reports, and proposed tasks inside a bounded workspace.
    3. Update: change internal records whose history can be inspected and reversed.
    4. Act externally: send messages, share files, update customer-facing systems, or invoke paid services.
    5. Perform destructive or privileged work: delete data, change access, alter infrastructure, or execute an irreversible operation.

    Most new agents should prove themselves at the read and prepare levels. Promotion should be capability-specific. An agent that reliably prepares a sales brief has not thereby earned permission to send customer communication. Reliability does not transfer automatically from one action class to another.

    For external actions, use a pending-approval state that contains the exact proposed action. You should be able to review the recipient, content, destination, and relevant context without reopening the entire run. Destructive or privileged actions should remain outside unattended execution unless you have an explicit recovery path and have deliberately accepted the consequence of failure.

    Treat external text as data, not authority

    Calendar descriptions, transcripts, web pages, emails, and documents may contain instructions that conflict with the agent’s job. The identity and task contract must outrank text found inside those inputs. An interview guest’s biography can inform a briefing; it cannot expand the podcast agent’s permissions. A meeting note can identify a follow-up; it cannot authorize the agent to send one.

    Keep credentials out of identity and task files. Give scripts access only to the credentials required for their operation, and avoid handing an agent a general browser, terminal, file system, and credential store merely because each tool is useful in isolation. The dangerous capability is often the combination.

    Make retries selective

    A retry is appropriate when the failure is plausibly temporary and repeating the operation is safe. A network timeout during a read may qualify. Ambiguous recipient identity, conflicting meeting details, missing share settings, or an unclear customer commitment do not. Retrying an ambiguity only asks the model to make the same unsupported decision again.

    Before enabling automatic retries, require the operation to pass three tests: it can detect whether it already succeeded, a duplicate would not create harm, and the number of attempts is capped. Otherwise, mark the task blocked and surface it for review.

    Put hard boundaries around usage

    Always-on does not mean continuously reasoning. It means the system is available to process eligible work on a known schedule. A run should inspect the queue, process a bounded amount of work, record its result, and exit.

    Set limits at several layers: eligible task types, work accepted per run, retries per task, tools available to the role, and provider-side spending or usage controls where available. Record usage beside the task outcome so you can distinguish an expensive valuable job from an agent that consumes resources while circling an ambiguity. Surprise charges are not only a pricing problem; they usually indicate that the operating loop lacks a stopping rule.

    Finally, maintain a kill switch you can use without asking the agent to cooperate. Disabling the schedule or revoking the narrow credential should stop future work. If stopping the system requires the same model and scripts that may be malfunctioning, it is not an independent control.

    Measure whether the agent is reducing work or relocating it

    A completed status is not proof of value. An agent can close every task while leaving you to verify facts, repair formatting, remove duplicates, and reconstruct why it made a decision. That is work relocation, not delegation.

    Evaluate the operation with measures tied to the job:

    • Usable completion rate: the share of eligible tasks that produce an output meeting the acceptance criteria without substantive rework.
    • Correction rate: how often you must change facts, recipients, permissions, status, or next steps before using the output.
    • Duplicate or false-action rate: how often the agent repeats a job or creates an action that the triggering event did not require.
    • Blocked rate by cause: which missing inputs, permissions, or unclear rules repeatedly prevent completion.
    • Time to review: the human attention required to approve, repair, or understand the result.
    • Usage per usable outcome: the model or service consumption attached to work you actually keep.

    These measures tell you what to change. A high blocked rate caused by missing context points to an input problem. Frequent factual corrections point to retrieval or acceptance-criteria problems. Duplicate work points to task identity and idempotency. High review time with otherwise correct output often means the evidence and change log are poorly presented.

    Require every run to leave a compact receipt: the task it claimed, inputs it used, scripts it invoked, files or records it changed, output location, completion status, and reason for any block. You should not need to replay hidden reasoning. You need enough evidence to verify the operation and diagnose the next failure.

    Review early runs closely and review again after changing an identity, script, tool, model, or input source. A stable task can become unstable when any one of those dependencies changes. Plain-text identities, tasks, and scripts make that change surface inspectable and versionable.

    Your agents can also improve the operating system itself. A periodic coding-manager workflow, for example, can review prior coding sessions, identify recurring dead ends, and propose changes in how future sessions are run. The important separation is that the agent proposes an improvement with evidence; the operating policy changes only after review. Self-observation is useful. Unreviewed self-modification is a different risk class.

    Expand only when the current job has earned more autonomy

    Adding agents is easy once the scheduler and folder structure exist. That convenience can tempt you to automate work whose boundaries are not ready. Scale based on operational evidence, not on the number of possible use cases you can imagine.

    A job is a strong candidate for always-on operation when it has a recurring trigger, stable inputs, an observable deliverable, clear acceptance criteria, bounded permissions, and enough repetition to justify maintaining the workflow. Preparation, follow-up capture, document setup, and periodic retrospectives fit because a person can inspect their artifacts and correct them before higher-consequence decisions are made.

    Keep work interactive when the task depends on novel judgement, unresolved organizational context, sensitive negotiation, or irreversible action. An agent may still prepare evidence and options, but the decision should remain with the person who owns the consequence.

    Before expanding an existing agent’s permissions or creating another role, check five gates:

    1. The current output is regularly usable without substantial reconstruction.
    2. Common failure modes are visible and end in safe states.
    3. Duplicate prevention and retry behavior have been exercised.
    4. Usage is attributable to tasks and bounded by stopping rules.
    5. The next capability has its own acceptance criteria and consequence review.

    Do not create one agent per application. Create one per coherent responsibility. A podcast manager may use a calendar, a document system, and a task list while retaining one outcome. Conversely, sales administration and coding retrospectives should not share an identity merely because they use the same model. Role boundaries should follow accountability, not tooling.

    Key takeaways

    • Begin with one recurring job that produces an inspectable artifact, not a general instruction to manage a function.
    • Give the agent a durable identity, a dedicated task queue, an explicit schedule, and small reviewed scripts.
    • Use task states, stable event identifiers, and completion receipts so retries and overlapping runs do not create invisible duplication.
    • Keep new agents at read-and-prepare permissions until their outputs and failure modes are consistently understandable.
    • Route ambiguity and consequential external actions to approval instead of asking the model to guess.
    • Cap eligible work, retries, tools, and usage; always-on availability should still produce finite runs.
    • Measure usable outcomes, corrections, blocks, duplicates, review effort, and usage before granting more autonomy.

    Pick one task you already repeat and write its delegation contract before choosing more tools. If you cannot define the input, output, permission boundary, completion test, and safe stopping condition on one page, the job is not ready to run while you are offline. Tighten the job first. The agent can earn broader responsibility after the operating evidence is there.

    References

  • Value-Based Pricing and Packaging: A Practical Playbook

    Value-Based Pricing and Packaging: A Practical Playbook

    If your pricing discussion keeps bouncing between competitor screenshots, delivery costs, and whatever Sales thinks the market will accept, you are not yet deciding a price. You are mixing four separate decisions: the pricing model, the pricing metric, the package, and the amount charged.

    Separate those decisions and make them in the right order. You will get a pricing system that customers can understand, Finance can model, Sales can explain, and Product can improve as real behavior replaces assumptions.

    Find the value, then choose a metric that tracks it

    Value-based pricing does not mean charging the highest number a customer will tolerate. It means connecting what the customer pays to a result the customer cares about. Your costs still determine whether the offer is sustainable, but they do not explain why the buyer should purchase it.

    Start by keeping four commonly confused decisions separate:

    DecisionQuestion it answersExample output
    Pricing modelWhat overall structure determines how the customer pays?Fixed fee, access-based, usage-based, or outcome-based
    Pricing metricWhat unit causes value and charges to scale?Account, seat, transaction, workflow, or verified outcome
    PackagingWhich capabilities, limits, and service levels belong together?Plans, allowances, add-ons, commitments, and overages
    PriceHow much will you charge for the package or metric?List price, contracted rate, and discount guardrails

    Define value in the buyer’s language

    Your first customer conversations should not begin with a proposed price. Begin with the decision the buyer is trying to make and the change the buyer expects after adopting the product. Ask for recent, concrete examples rather than opinions about a hypothetical offer.

    • What event made this problem important enough to address?
    • What happens if the buyer leaves the problem unsolved?
    • Who experiences the problem, and who controls the budget?
    • What observable change would count as success?
    • How does the buyer prove that change internally?
    • What alternatives compete for the same budget, including manual work and doing nothing?
    • What causes value to grow: more users, more activity, more completed work, better results, or lower risk?

    Turn the answers into one working statement: For [buyer], the product creates value when [observable result] improves [business or operational consequence], compared with [current alternative]. This is not positioning copy. It is a testable value hypothesis that will guide the metric and package.

    If different segments complete that sentence differently, do not average the answers into a vague promise. That is evidence that the segments may need different packages, metrics, or sales motions. A support leader buying fewer escalations and an operations leader buying more throughput may use the same product while evaluating its value in different ways.

    Turn value into a billable unit

    The pricing metric is the bridge between the value hypothesis and the invoice. For an AI support agent, for example, the model can charge only for results, while the unit is an outcome counted when the agent resolves a customer query without further help. The principle is attractive because payment moves with delivered value. The definition is difficult because every ambiguous edge case can become an invoice dispute.

    Write the metric specification before selecting the price. It should define:

    • The event that starts the measurement.
    • The event that qualifies the unit as complete.
    • Any quality threshold required before it is billable.
    • Exclusions, such as tests, spam, duplicates, abandoned work, or activity outside the contracted scope.
    • Attribution when a human, an automation, and an AI system all contribute.
    • How reversals, reopened work, refunds, and corrections affect the count.
    • What the customer can see before the count appears on an invoice.
    • Which record resolves a disagreement between product analytics and billing.

    My rule is simple: if a buyer cannot understand what will be counted and predict the direction of the next bill, the metric is not ready. Evaluate every candidate against six tests:

    • Value alignment: Does an increase in the unit normally mean the customer received more value?
    • Predictability: Can the customer forecast the unit well enough to plan a budget?
    • Auditability: Can both sides inspect the same underlying events?
    • Controllability: Can the customer influence usage or set limits without abandoning the product?
    • Operational feasibility: Can your product, data, billing, and support systems calculate the unit consistently?
    • Economic alignment: Does revenue scale sensibly relative to the cost and risk of delivering the value?

    A value-based design does not always require a literal outcome metric. A proxy can be the better choice when it is closely related to value and much easier to forecast and audit. Raw activity is a poor proxy when it can grow without improving the customer’s result. A seat is a poor proxy when adding users does not increase value. An outcome is a poor metric when success cannot be defined consistently. Choose the least complicated unit that preserves alignment.

    Before charging anyone, run the proposed rules against beta or historical events. Generate shadow invoices, inspect unusually high and low accounts, and reconcile the count from the raw event through the customer-facing bill. This exposes definitional and data problems while they are still product problems rather than financial disputes.

    Make packaging do the segmentation work

    Pricing determines how revenue scales. Packaging determines which customers select which offer. A package is therefore not a decorative feature table. It is a mechanism for matching different value patterns, operating needs, and willingness to pay without creating a custom product for every account.

    1. Segment customers by how they receive value. Company size may matter, but workflow complexity, risk, required integrations, volume, and the cost of failure can be more revealing.
    2. Identify the minimum complete experience. Every package should let its intended customer reach the core outcome; a deliberately crippled entry plan teaches the market that the product does not work.
    3. Place differentiators where their value is concentrated. Advanced governance, analytics, automation, integrations, service levels, and support may matter much more to one segment than another.
    4. Choose the relationship between access and consumption. Decide what is included, what is metered, whether unused commitments expire, how overages work, and whether customers can set caps or alerts.
    5. Test whether buyers can self-select. Show realistic scenarios, ask which package they would choose, and then ask them to explain why. Their explanation is more diagnostic than the selected tier.

    Choose modular, bundled, or hybrid architecture deliberately

    Modular pricing works best when capabilities have distinct buyers, adoption paths, and measurable outcomes. It lets a customer buy one job without funding unrelated functionality. Its weakness appears as the portfolio expands: each additional module adds another decision, metric, contract term, and sales explanation.

    Bundling works better when capabilities reinforce one workflow or when customers experience the combined result rather than the individual components. It reduces buying friction, but it can hide which capability creates value and can force smaller customers to pay for breadth they do not need.

    A hybrid can separate platform access from variable value: a base package covers the shared product, an included allowance makes the initial bill predictable, and overages or commitments let revenue grow with delivered value. Use that structure only when each component answers a different commercial question. Adding a platform fee, several meters, tier thresholds, credits, and add-ons without a clear role for each one creates a billing puzzle, not a pricing strategy.

    Look for these packaging failure signals:

    • Customers repeatedly need capabilities scattered across several tiers.
    • The entry package cannot produce the outcome used to sell it.
    • The highest tier is simply every leftover feature rather than an offer for a distinct need.
    • Two packages attract the same customer for reasons your sales team cannot explain consistently.
    • The economically best package for you is visibly wrong for the customer.
    • Customers need a spreadsheet or a salesperson to estimate a normal bill.
    • Every new capability becomes a new add-on because the portfolio has no shared packaging logic.

    Do not ask customers whether they like the package names or feature list. Give them a buying situation, expected volume, required controls, and a budget constraint. Ask them to choose, identify what feels unnecessary, and state what is missing. You are testing whether the architecture supports a decision, not whether the page looks polished.

    Measure willingness to pay only after the offer is clear

    Quantitative pricing work becomes useful only after buyers understand the model, metric, and package. Otherwise, a survey can produce a precise answer to a question the market would never ask. Use qualitative discovery to establish the buyer’s language and mental model, then carry that exact framing into willingness-to-pay testing.

    Methods such as Gabor-Granger and Van Westendorp answer different questions. Gabor-Granger-style testing helps estimate purchase willingness across proposed price points. Van Westendorp-style questions help expose perceived price boundaries, including where an offer begins to feel implausibly cheap or prohibitively expensive. Neither method discovers the value metric for you, and neither produces a universally correct price.

    A defensible survey sequence looks like this:

    1. Describe the customer problem and product outcome without promotional language.
    2. State exactly how charging works.
    3. Define the billable unit, including the success condition.
    4. Show what the package contains and what it excludes.
    5. Give the respondent a realistic usage or outcome scenario.
    6. Ask about willingness to purchase at a specific price or across a controlled sequence of prices.
    7. Capture the respondent’s role, segment, buying authority, expected volume, and current alternative so the results can be interpreted rather than merely averaged.

    A demand curve is more useful than a single average. In one outcome-priced case, stated purchase willingness moved from 69% at $0.86 per outcome to 39% at $1.42. Those figures are not benchmarks for another product. They demonstrate why the decision is strategic: moving along the curve changes expected adoption as well as revenue captured from each unit.

    A simple price multiplied by the share willing to buy can identify a survey-based revenue peak, but that point is not automatically your final recommendation. It does not, by itself, include realized discounts, differences in unit volume, cost to serve, retention, expansion, sales effort, or the value of establishing market share.

    Decide what the price is meant to accomplish before interpreting the curve:

    • If the priority is adoption, you may accept less revenue per unit to reach more qualified customers.
    • If the priority is near-term revenue, you may choose a higher point while accepting a lower attach rate.
    • If the product requires substantial support or delivery cost, margin may eliminate prices that look attractive in a demand survey.
    • If the category is unfamiliar, simplicity and predictability may be more important than extracting the theoretical maximum.
    • If the product is part of a broader platform, the effect on cross-sell, retention, and portfolio coherence may matter more than stand-alone revenue.

    Treat willingness-to-pay results as stated intent, not observed buying behavior. Segment the curve before using it. A blended result can conceal a high-value segment with strong demand and another segment that should not be targeted at all. It can also overstate confidence when respondents use the product but do not own the budget.

    Convert the demand curve into a commercial model

    The survey narrows the plausible range. The commercial model tells you whether an option can survive contact with actual customers, contracts, usage, discounts, and delivery costs. This is where a promising price becomes an operating plan.

    1. Set a candidate list price. Choose a point that reflects the demand curve and the strategic objective, not just the highest theoretical revenue index.
    2. Estimate realized price. Apply expected discounts, negotiated rates, credits, promotions, and channel effects. A list price that relies on constant exceptions is not the real price.
    3. Project units by segment. Use beta or observed usage to estimate outcomes, transactions, seats, or another billable quantity. Preserve the distribution instead of relying only on the mean.
    4. Model attach rate. Estimate what share of eligible customers will buy in conservative, base, and upside cases. Connect each case to an explicit assumption rather than a general level of optimism.
    5. Calculate customer and portfolio revenue. For a metered product, combine realized unit price with expected annual units. Then roll the result across eligible customers and segments.
    6. Include delivery economics. Subtract variable delivery costs and account for service obligations that grow with usage. For AI products, inspect how model, infrastructure, support, and exception-handling costs behave at both low and high volume.
    7. Connect the recommendation to the operating plan. Show the implications for customer count, adoption, annual recurring revenue, gross margin, expansion, and any dependencies on the rest of the portfolio.

    Stress-test the assumptions that can break the plan

    A single base case hides the shape of the risk. Change one major assumption at a time so decision-makers can see what the recommendation depends on.

    • Discount sensitivity: What happens if realized price is materially below list price?
    • Volume sensitivity: What happens when customers generate far fewer or far more units than the average?
    • Attach sensitivity: How much adoption is required before the product covers its fixed investment?
    • Cost sensitivity: Does high usage improve gross profit, or does the delivery cost scale almost as quickly as revenue?
    • Concentration risk: Does the forecast depend on a small number of unusually large customers?
    • Invoice volatility: Can normal changes in behavior create bills that customers will perceive as unpredictable?
    • Metric leakage: Are valuable events going unbilled, or are low-quality events being counted as successful outcomes?

    Inspect account-level scenarios, not just portfolio totals. A model can produce acceptable average revenue while creating obviously unreasonable bills for a small customer, a seasonal customer, or a high-volume account. Those tails often become the discount exceptions, support escalations, and renewal problems that the average concealed.

    Make the recommendation easy to challenge

    The approval memo should contain the decision and the logic required to dispute it. Include:

    • The buyer, value hypothesis, model, metric, and metric definition.
    • The proposed packages and the segment each package is designed to serve.
    • The willingness-to-pay range and how it changes by segment.
    • The recommended list price, expected realized price, and discount guardrails.
    • Conservative, base, and upside forecasts for adoption, revenue, and margin.
    • The most sensitive assumptions and the evidence supporting them.
    • Alternatives considered, why they were rejected, and what evidence would reopen them.
    • Operational dependencies across Product, Research, Data, Finance, Engineering, Sales, Customer Success, Support, and billing.

    Cross-functional review is not ceremonial. Finance can expose a margin or forecasting problem. Engineering can show that the proposed event cannot be measured reliably. Sales can identify a model buyers cannot procure. Support can anticipate disputes. Product can determine whether the metric rewards the behavior the product is supposed to create. Resolve those conflicts before the price becomes a public promise.

    Launch pricing as a controlled learning system

    Approval is the end of price design and the start of price operations. Customers experience pricing through entitlements, usage counters, contracts, invoices, renewal conversations, and support responses. A sensible strategy can fail if those surfaces disagree.

    Complete the billing path before charging

    • Write a billing specification that maps raw events to billable units and contract terms.
    • Verify entitlements, included allowances, overages, caps, credits, and exception handling.
    • Run parallel or shadow invoices and reconcile them from event log to customer-facing total.
    • Give customers a usage view that uses the same definitions and timing as billing.
    • Enable Sales with qualification rules, scenario-based pricing examples, and clear discount authority.
    • Prepare Customer Success and Support to explain the metric, diagnose discrepancies, and escalate genuine billing errors.
    • Instrument proof of value next to proof of usage so the commercial conversation is not reduced to a meter.
    • Communicate the effective date, affected products, counting rules, package changes, and available customer controls in plain language.

    Do not alter existing charges on the assumption that a product announcement overrides a contract. Review contractual commitments, renewal timing, migration rules, and customer communications before changing what an existing customer pays. An informal migration can create financial disputes and destroy trust even when the new model is better designed.

    Use behavior to diagnose the next problem

    Instrument the system from the first launch cohort. Review both commercial performance and customer experience:

    • Eligibility, attach rate, and package selection by segment.
    • List price, realized price, discount frequency, and exception rates.
    • The full distribution of billable units per customer, not just the average.
    • Revenue and gross margin by segment, package, and usage band.
    • Invoice variance and how accurately customers forecast their charges.
    • Billing questions, disputes, credits, and metric-definition escalations.
    • Activation, continued usage, achieved outcomes, expansion, contraction, renewal, and churn.
    • Sales-cycle friction caused by the model, procurement requirements, or package complexity.

    Use each signal to choose the next investigation. Low attach can point to weak qualification, unclear value, the wrong package, or the wrong price. Strong attach followed by low activity can indicate an onboarding or product-value problem. High activity with poor margin calls for an economics or discount review. Frequent disputes usually justify inspecting the metric definition, event quality, and customer visibility. These patterns are diagnostic prompts, not causal proof; pair the numbers with targeted customer and GTM conversations.

    Review the architecture, not only the number, when the product expands. Modular outcome pricing can work cleanly while each capability has a distinct result. As a platform adds capabilities, buyers may face several meters, overlapping modules, and an invoice they cannot predict. That is a signal to reconsider how access, bundles, allowances, and outcomes fit together, not merely to adjust every component independently.

    Reopen the pricing system when customers cannot forecast bills, new capabilities do not fit an existing package, discount exceptions become routine, sales explanations diverge, gross margin behaves differently from the model, or the value customers receive is no longer represented by the metric. Pricing should be treated as a living system informed by research, customer behavior, and go-to-market learning, not a launch artifact that becomes untouchable.

    Key takeaways

    • Make four decisions separately: pricing model, pricing metric, package, and price.
    • Define value using an observable customer result before asking what anyone will pay.
    • Choose a metric that aligns with value but remains predictable, auditable, operationally feasible, and economically sound.
    • Design packages around distinct value patterns and buying needs, not an arbitrary progression of feature counts.
    • Use willingness-to-pay work to build a demand curve, then combine it with usage, attach, discounts, and margin in a commercial model.
    • Validate the complete billing path before launch and use observed behavior to improve the system afterward.

    If your team is stuck debating the number, stop the meeting and complete six lines first: buyer, customer outcome, billable unit, measurement proof, package boundary, and commercial assumptions. Any line you cannot defend is the next research or modeling task. Put a price on the page only after those six lines tell one coherent story.

    References

  • Built for Your Biggest Days: How We Engineer Fair, Reliable Scale Without Downtime

    I’m getting sharper, more specific questions about scale from enterprise customers every quarter, and that’s exactly how it should be. Teams want to know how our platform behaves during their highest-volume moments — the Black Friday sales, the sporting events, the production incidents — and they want confidence their growth won’t outpace the systems they depend on. We welcome those questions. They’re the right ones to ask of any critical component of your business. Today, our systems handle serious scale. At daily peak, we see over 150,000 customer requests per second coming into the platform, with more than 70,000 asynchronous requests per second flowing through the background systems. During our busiest days of the week, we handle over five million conversations and more than 100 million comments being added across the platform. We also design for individual customer spikes, not just aggregate platform traffic. We can handle a single customer workspace spiking with hundreds of comments per second, or around 100 new conversations per second. Sustained over a full day, that would map to millions of conversations from a single customer. While those numbers matter, they age quickly. Every growing software company can publish a bigger number every year, month, week. What ultimately matters is whether the architecture has clear scaling levers, whether we understand the pressure points in the system, and whether we can add capacity before customers need it. Every system has limits. Competence is knowing where they are, measuring them, and moving them before customers reach them. Here’s how we do that in practice. We build on boring foundations because at the edges, we try hard not to be clever. We use AWS for the infrastructure primitives AWS is very good at running. We do not want our engineers spending their best energy recreating S3, load balancers, queues, or commodity infrastructure patterns. We want that energy spent on the parts of the system that are specific to our customers and our product. “That is a deliberate trade-off. It gives us fewer systems to understand, deeper expertise in the ones we do run, and more leverage when we need to scale.” This extends a principle I’ve embraced for years: run less software. The point isn’t to minimize the stack for its own sake; it’s to compound expertise. When many teams build on the same small set of technologies, our tooling, observability, and operational practice all improve together. Boring technology choices aren’t a lack of ambition — they reserve our ambition for the nuanced scaling challenges that matter. The source of truth is the hard part. You can scale stateless web traffic by adding machines, add queue consumers, and add cache. Those are real problems — just not the hardest ones. The source-of-truth database is where the most important data lives, where the hardest correctness guarantees exist, and where maintenance windows often come from. It has to be correct, fast, resilient to failover, capable of large migrations, and able to keep serving traffic while we improve it. As customers grow, it cannot require a full re-architecture every time the next ceiling appears. That is why we moved to Vitess, managed by PlanetScale. The goals were clear: improve availability, reduce operational complexity, make large table migrations safer, simplify MySQL scaling, and eliminate customer downtime from routine database maintenance and failovers. When we first laid out this direction, the largest part of the migration was still ahead of us. We completed that migration in 2025, and the benefits are now part of how we operate the platform day to day. Today, our highest-scale source-of-truth data is spread across 128 shards. The database layer handles around two million requests per second, with more than ten million cache reads per second in front of it. For the largest customers, we can isolate and scale database capacity independently, including dedicating a shard to a single customer when needed. We have not come close to needing that, which is significant. The goal of architecture like this is not to run every system at the edge of its capacity, but rather to have room to move before customers need it. Vitess gives us native sharding, query routing, online schema change capabilities, connection pooling, and resharding primitives built for this kind of workload. Instead of application code carrying all of the sharding complexity, the database layer can do more of the work. That reduces cognitive load for engineers and removes whole classes of operational risk. Ultimately, this gives us practical scaling options instead of hard architectural rewrites, and lets us do routine database improvement without planned customer-impacting maintenance windows. Search is not a hidden bottleneck for us. Search underpins core product surfaces across the platform — from vector search in our AI features to realtime reporting — and if it’s slow or unhealthy, customers feel it. Scaling isn’t just adding more machines; often the better approach is making the product do less unnecessary work. Today, our Elasticsearch clusters support a much higher-throughput product than in the past, with more than 650TB of storage, more than 1.7 trillion documents, and peaks above 40,000 requests per second. We’re serving a larger product surface more efficiently, not just running a bigger cluster. More importantly, when an index gets too large or traffic distribution turns unhealthy, we don’t want a high-risk, manual migration. We reshape Elasticsearch indexes online by partitioning by customer ID, dual-writing to old and new indexes, backfilling, validating, gradually moving customers with feature flags, and deleting the old index only when we’re confident. We’ve used this pattern for years to make large search migrations safer and more incremental — a core playbook in our platform scalability and SRE practices. Fairness is non-negotiable in a multi-tenant system. A single customer’s high-volume moment should not quietly become everyone else’s latency problem. We design for this at multiple layers. For asynchronous work, we use overflow queues and queueing strategies that prevent one high-volume workload from consuming shared capacity in a way that hurts quieter tenants. AWS SQS fair queues are one example of a primitive we use extensively. They’re designed for exactly this class of problem. When one tenant creates a backlog in a shared queue, fair queues help reduce the dwell-time impact on other tenants. We also build application-level guardrails so customer isolation doesn’t depend on every engineer remembering every rule in every code path. In a large multi-tenant Rails application, the safe path must be built into the system. The focus is primarily about correctness and customer data separation, but the broader operating principle is the same: important customer boundaries should be enforced by infrastructure and application frameworks. The same thinking applies to scale. We want customer-specific load to be visible, attributable, and controlled. When a customer spike happens, we should be able to understand it as that customer’s workload, protect the rest of the platform, and add capacity where it’s actually needed. Fin adds a new dimension to scaling. Our AI Agent Fin introduces a new set of infrastructure challenges. To provide reliable AI-powered support at scale, we need to operate across multiple model providers, route across them based on capacity and latency, and protect customer-facing workloads from lower-priority work. The details differ from traditional SaaS infrastructure, but the principle is the same: understand the bottlenecks, build clear scaling levers, and monitor the customer outcome. AI providers are not commodity storage systems, and we do not design as if they are. That is why we have invested in Fin-specific reliability systems. Fin now fully resolves over two million conversations per week. At that scale, high availability cannot depend on a single model, a single provider, a single region, or a single pool of capacity. Our LLM routing layer supports cross-vendor failover, cross-model failover, latency-based routing, capacity isolation, and load testing. We also maintain buffer capacity with major providers, with headroom to handle 2x to 3x normal Fin traffic at any point. For enterprise customers, this matters because AI support volume can spike just like human support volume — and the AI layer must absorb that spike without relying on one fragile upstream path. When customers depend on Fin to absorb a spike in support demand, the AI layer needs the same operational discipline as the rest of the platform. Performance tests help, but production traffic is reality. Real customers use products in ways no synthetic test will perfectly predict: launches, incidents, seasonal patterns, gaming events, sudden changes in end-user behavior. Those moments give us data that no lab can fully reproduce. Often, a large customer event barely moves the platform-wide graphs because our customer base is broad enough that one industry’s peak aligns with another’s quiet period. Black Friday and Cyber Monday are good examples. Many ecommerce customers are at their busiest, while many B2B SaaS customers are quieter. At the aggregate platform level, the change can be much less dramatic than people expect. “That does not mean those events are unimportant. It means we need to look at both levels: the health of the overall platform and the experience of the individual customer having the spike.” Sometimes, these events teach us something specific. In one case, a very large customer used the Messenger in a way that exercised the full Messenger lifecycle even though the visible user experience did not require it. Under normal traffic, this was fine. During a major customer-side incident, their users refreshed aggressively, generating a much larger burst of Messenger traffic than the integration actually needed. The platform stayed available, but the event exposed unnecessary work in that integration path. We built a lighter-weight integration path that served the customer’s actual use case with far less work per request, making future spikes easier to absorb. We treat large customer events this way even when there’s no broad customer impact. They’re opportunities to understand real scaling properties and make the next event safer — a habit that anchors our incident management, observability, and FinOps practices. Scale is also an operating model. The infrastructure matters, but it’s not enough. You can have the right database architecture and still hurt customers if you detect issues late, recover slowly, communicate poorly, or fail to learn from incidents. “That is why our operating model starts with customer outcomes. If the customer cannot do the job they came to do, the system is unhealthy. It does not matter how many dashboards are green.” Heartbeat metrics tell us whether customers can do the core jobs they hire us to do. They cut through infrastructure noise and answer the question that matters most during an incident: are customers able to use the product successfully? This shapes how we ship. Today, we average around 250 ships to production per workday, with an average merge-to-production time under 10 minutes. That isn’t a vanity metric — it’s part of the safety model. Smaller changes are easier to understand, easier to observe, and easier to roll back. Feature flags let us separate deployment from release. Automatic rollback and heartbeat-driven detection help us recover quickly when a change hurts customers. These are the very DORA metrics we hold ourselves to in order to balance CI/CD speed with stability. “Fast shipping is not the opposite of reliability. Done properly, it is one of the ways you stay in control of change.” The bar is high. Engineers are expected to understand the impact of their changes, watch them go live, and act quickly if something looks wrong. Resuming service is not the end of an incident. We expect teams to understand the root cause, fix the contributing systems, and prevent recurrence. That’s how scale stays safe over time. Scheduled maintenance should be extraordinary. Historically, database maintenance was a main reason for maintenance windows: upgrading a database, changing instance sizes, performing failovers, or moving large tables could require customer-impacting downtime. With the move to Vitess and PlanetScale, we changed what routine database improvement looks like. We can upgrade, scale, and improve critical database infrastructure without turning that work into planned customer-impacting downtime — and we do this in practice, not just as a goal. This matters because customers rely on our platform for live operations. If their support team, Messenger, Help Desk, or AI Agent is unavailable, the impact is immediate. Scheduled maintenance cannot be treated as a casual operational convenience. “Our posture is simple: routine infrastructure improvement should not require planned customer-impacting downtime.” Scheduled maintenance should be exceptional, non-routine, clearly communicated, and minimized in frequency, duration, and customer impact. That’s the practical benefit of the architecture work: better scaling is not only about handling more traffic, but also reducing the operational moments that might inconvenience customers. What this means for customers is simple: be skeptical of vague scale claims. The question isn’t whether a vendor says they can scale — it’s whether they can explain how, where the limits are, what they measure, how they recover, and what they’ve changed after learning from production. We understand the scaling properties of our systems, have clear levers to add capacity at the right layers, design for customer isolation and fairness, monitor customer outcomes directly, and use real production events to make the next one safer. Scale is never finished. Every large customer event, traffic spike, migration, and incident teaches us something about the real behavior of the system — and we use that data to keep improving. That’s what you should expect from a platform you depend on during your busiest moments.

    Inspired by this post on The Intercom Blog.


    Book a consult png image
  • How to Build an AI-Native Product Discovery Workflow

    How to Build an AI-Native Product Discovery Workflow

    Your discovery stack may already hold interview transcripts, support conversations, behavioral analytics, experiment results, and roadmap assumptions. Yet the decision in a product review can still depend on whoever read the most material or built the most persuasive deck.

    If adding an LLM only gives you faster summaries, the workflow is not AI-native. An AI-native discovery workflow shortens the distance from evidence to a decision while making every important claim easier to inspect. AI retrieves, structures, compares, and challenges the evidence. You remain accountable for what the evidence means and what the product team does next.

    Key takeaways

    • Begin every AI-assisted discovery run with an outcome, a metric, defined context, and a decision that someone needs to make.
    • Preserve raw evidence and give each observation a stable identifier before asking AI to synthesize it.
    • Break the workflow into bounded jobs such as retrieval, extraction, clustering, contradiction detection, and decision-brief drafting.
    • Evaluate citation accuracy, evidence fidelity, counterevidence, abstention, and access controls before the output enters a roadmap discussion.
    • Measure whether the workflow improves decision quality and product outcomes, not merely whether the model produces polished prose.

    Frame the decision before you involve the model

    Most weak discovery prompts fail before the model sees them. Analyze the interviews, summarize the feedback, and find insights are activities, not decisions. They give the model no principled way to distinguish useful evidence from interesting noise.

    Write a short decision contract first. A useful contract specifies the outcome and metric, the context and constraints, and the decision and deliverable. Those fields turn an open-ended request into a bounded discovery task.

    1. Outcome and metric: Name the user or business outcome, then define the behavior or measure that represents it. Activation, funnel conversion, and retention are not interchangeable. Include the event definition and observation window used by your analytics system.
    2. Context and constraints: State the relevant cohort, product surface, timeframe, market, known exclusions, and data-access limits. New self-serve accounts on the web can exhibit a different pattern from established accounts or customers using another surface.
    3. Decision and deliverable: Say what someone will do with the answer. Ask for a ranked opportunity brief, an interview plan, a set of competing explanations, or experiment candidates only when that format supports a real pending decision.

    Reusable decision prompt: Help me decide [decision]. The outcome is [outcome], measured as [metric definition]. Limit the analysis to [cohort, surface, timeframe, and constraints]. Retrieve evidence from [approved repositories]. Return [deliverable]. For every material claim, include the evidence identifier, any conflicting evidence, the affected segment, and what is still unknown. If the available evidence cannot support a recommendation, say so and specify what is missing.

    The last sentence matters. An AI system should be allowed to return insufficient evidence. If every run must end with a recommendation, the workflow rewards plausible completion instead of honest discovery.

    Keep the outcome separate from the proposed solution. Improve activation is an outcome. Validate an onboarding checklist is already a solution choice. When you embed the solution in the prompt, AI tends to organize the available evidence around that choice instead of testing whether another opportunity matters more.

    Use evidence-strength labels that a reviewer can verify rather than asking the model for an unsupported confidence percentage:

    • Sufficient: Direct evidence applies to the target context, and no material contradiction remains unresolved.
    • Mixed: Direct evidence and meaningful counterevidence both exist, or the pattern changes by segment.
    • Insufficient: Evidence is missing, indirect, stale for the decision, or outside the target context.

    Build a traceable evidence pipeline, not a transcript pile

    AI cannot make discovery evidence traceable if the underlying repository has already flattened observations, interpretations, and decisions into the same notes. Preserve those layers separately. My rule is simple: automate the movement and inspection of evidence before automating judgment.

    LayerWhat it containsControl that matters
    Raw evidenceInterview recordings or transcripts, support records, session evidence, and analytics query resultsKeep the original record intact, access-controlled, and addressable by a stable locator
    Evidence unitsAtomic observations with metadataSeparate exact customer language, observed behavior, and analyst interpretation
    OpportunitiesCandidate needs, frictions, or desired outcomesAttach supporting evidence, counterevidence, affected segments, and unresolved questions
    DecisionsChoices made, rejected alternatives, assumptions, and rationaleName the decision owner and preserve the evidence available at the time
    LearningExperiment results and later customer or behavioral evidenceUpdate the opportunity without erasing the earlier reasoning

    Each evidence unit should carry enough metadata to survive outside its original document:

    • A stable evidence identifier.
    • The collection date and an exact locator such as a transcript timestamp or saved analytics query.
    • The relevant user segment, product surface, and journey stage.
    • The raw observation, kept separate from the interpretation proposed by a person or model.
    • The access, retention, and sensitivity classification.
    • The opportunity, assumption, or outcome to which the evidence may relate.

    This structure prevents a common failure: a model paraphrases an interview, a later summary compresses that paraphrase, and the roadmap eventually treats the compressed interpretation as a customer fact. A reviewer should always be able to move from a claim to the evidence unit and then to the original record.

    Apply data-governance rules before ingestion. If customer conversations contain personal, confidential, or contract-restricted information, do not copy them into an AI system until its access, retention, redaction, and model-training terms match your commitments. A more convenient synthesis workflow is not worth an unauthorized disclosure.

    Retrieve the smallest useful context

    Once the evidence corpus no longer fits sensibly into a prompt, use a retrieval-first pipeline with modular prompts and observable traces. Retrieval-augmented generation should select evidence relevant to the decision contract, rather than asking a general agent to reason over everything the company knows.

    RAG is a grounding mechanism, not a truth guarantee. A fluent answer does not prove that the retriever found the decisive interview, the correct event definition, or the evidence that contradicts the dominant pattern. Configure retrieval to look for both support and contradiction, preserve evidence identifiers, respect access controls, and return no result when the available context does not meet the task.

    An opportunity solution tree can provide the shared view above this pipeline: the desired outcome connects to opportunities, solution candidates, and tests. Treat the tree as a navigable representation of current thinking. Every important node should still resolve to evidence and assumptions beneath it.

    Give AI a chain of bounded jobs

    A single agent asked to interview customers, interpret feedback, size opportunities, choose a solution, and write a roadmap has too many ways to hide a weak inference. Break the work into stages with explicit inputs and review gates:

    1. Prepare: Give AI the outcome, assumptions, and learning gaps. Let it draft non-leading interview questions. A human checks whether the guide is testing an assumption or merely inviting agreement.
    2. Convert: Extract atomic observations from approved records. Require exact locators and label customer language, observed behavior, and interpretation separately.
    3. Synthesize: Cluster candidate opportunities without erasing segment differences. Request supporting evidence, counterevidence, and unrepresented cohorts for every cluster.
    4. Connect: Use behavioral analytics to examine whether the observed pattern appears in the target cohort. Interviews can expose mechanisms and unmet needs; they should not be treated as a substitute for measuring prevalence.
    5. Challenge: Ask for rival explanations, evidence that would reverse the conclusion, and assumptions that remain untested. This stage should consume the evidence record, not just the previous summary.
    6. Draft: Produce a decision brief containing the pending decision, options, evidence, contradictions, unknowns, and proposed next test. A named human accepts, revises, or rejects it.
    7. Learn: Attach experiment and outcome evidence to the same opportunity record. Preserve what the team believed before the test so later reviewers can inspect how the decision changed.

    Pass structured artifacts between stages. If each stage receives only prose copied from the previous chat, unsupported claims can become progressively harder to distinguish from evidence.

    Buy workflow plumbing; own the decision logic

    You do not need to build every repository, connector, permission system, visualization, and observability screen. Licensing purpose-built opportunity-tree infrastructure can be the sensible choice when your differentiated work is the learning system rather than the canvas or collaboration layer.

    Keep ownership of the parts that encode how your company makes product decisions: the decision contract, evidence schema, opportunity taxonomy, prompt modules, evaluation cases, escalation rules, and approval gates. Before choosing a platform, ask:

    • Can you export the raw evidence, metadata, opportunity structure, prompts, and run traces?
    • Can access rules follow the evidence through retrieval and generation?
    • Can the system connect to your approved analytics and customer-evidence repositories without repeated manual copying?
    • Can you evaluate a prompt or retrieval change against representative past cases?
    • Can a reviewer inspect why a claim appeared and what evidence was omitted?
    • Would building this capability improve the customer outcome, or merely recreate commodity workflow infrastructure?

    Evaluate the workflow before it shapes the roadmap

    Start evals before AI-generated conclusions become routine inputs to product reviews. The evaluation set should represent the cases the workflow will actually encounter: a clear pattern, conflicting evidence, insufficient evidence, cohort-specific behavior, stale material, duplicated records, and content the requesting user is not allowed to retrieve.

    For synthesis and decision-support tasks, evaluate behavior that a reviewer can observe:

    • Citation validity: Every material claim points to a real, accessible evidence identifier.
    • Evidence fidelity: Quotations and behavioral facts remain faithful to the underlying record; interpretations are labeled as interpretations.
    • Retrieval coverage: The output includes the evidence required to assess the target opportunity, not merely the easiest matching passages.
    • Contradiction handling: Material counterevidence and segment differences are visible rather than buried.
    • Abstention: The system returns insufficient evidence when the decision cannot be supported.
    • Decision fit: The deliverable answers the stated decision instead of drifting into a generic summary or unrelated recommendation.
    • Policy compliance: Restricted evidence stays outside unauthorized retrieval, traces, and generated output.

    A strict release gate is useful here. Fail the output if it invents an evidence identifier, turns an interpretation into a quotation, ignores a material contradiction, or exposes restricted content. Those are not cosmetic defects that a polished paragraph can offset.

    Treat the prompt, retrieval configuration, model choice, taxonomy, and evaluation set as versioned artifacts. This is the practical value of eval-driven development and early observability: when behavior changes, you can identify the change that caused it and rerun representative cases before wider use.

    For each production run, retain the decision contract, evidence identifiers retrieved, prompt and retrieval versions, generated output, reviewer edits, final decision, and later outcome. That trace lets you distinguish a retrieval failure from a synthesis failure, a weak decision contract, or a reasonable decision invalidated by new evidence.

    Model-quality checks are only one layer. Also baseline and monitor the discovery workflow itself:

    • Time from a framed question to a reviewable decision brief.
    • The share of material claims with inspectable evidence.
    • Reviewer corrections to quotations, segments, event definitions, and interpretations.
    • Decisions reopened because relevant evidence was missing or misread.
    • Movement in the outcome and metric named in the original decision contract.

    Do not set improvement targets until you have a baseline for the existing process. A system can make synthesis faster while increasing correction work or encouraging premature decisions. The end-to-end measure tells you whether the saved time is real.

    Turn the workflow into a product operating system

    AI-native discovery changes the product team’s operating model only when ownership remains explicit. The product manager or product trio owns the outcome, assumptions, and decision. Research and design judgment protects interview quality and interpretive nuance. Data and engineering ownership protects event definitions, retrieval reliability, instrumentation, and access controls. AI produces candidate artifacts. The decision owner approves the action.

    Review by exception instead of rereading every generated sentence. Inspect claims marked mixed or insufficient, new opportunity clusters, segment differences, material contradictions, changed event definitions, and outputs that differ from earlier runs. This focuses human attention where judgment is most valuable without treating the model as an authority.

    Roll out the workflow through one recurring, reversible discovery decision:

    1. Choose a decision for which customer evidence and behavioral data already exist, such as prioritizing an onboarding friction or investigating a repeated support issue.
    2. Baseline the current path from question to decision, including reviewer corrections and missing-evidence failures.
    3. Create the decision contract, evidence schema, and access rules before connecting an agent.
    4. Build the evaluation set from previous clear, contradictory, insufficient, segment-specific, and restricted cases.
    5. Run the AI workflow in shadow mode beside the existing process. Compare claims, omissions, reviewer effort, and the resulting decision without allowing the generated output to act automatically.
    6. Promote bounded jobs only after they pass their gates. Evidence extraction may be ready before opportunity ranking, and opportunity ranking may be ready before solution recommendations.
    7. Expand to another workflow only when the traces are stable, reviewers understand escalation paths, and the first use case is improving the decision process rather than merely generating more material.

    At your next discovery review, do not ask what AI found. Bring one decision contract, require every consequential claim to resolve to evidence, and make the unresolved assumption visible. That is a small enough change to start immediately and a strong enough foundation for everything you automate later.

    References

  • Level Up: May 26 Claude Code Show & Tell + Final Product Discovery Fundamentals Cohort

    I’m excited to share two opportunities this season to uplevel your craft, connect with peers, and leave with practical, repeatable techniques you can apply immediately to your product work.

    We will be doing another round of Claude Code: Show and Tell on May 26th at 9am PDT. These community-driven sessions are hands-on and fast-paced—we swap proven workflows, compare prompts, and pressure-test approaches together. You’ll see how product teams are operationalizing AI workflows in real contexts and walk away with ideas you can adapt for your own roadmap and experimentation pipeline. Invites will go out to Supporting Members and CDH Members tomorrow. If you'd like to join us, keep an eye on your inbox for the invite.

    I love these Show & Tell sessions because they translate tacit knowledge into clear, reusable playbooks. Whether you’re refining evaluation loops for LLMs, streamlining discovery synthesis, or standardizing prompts for consistency, the shared rigor and camaraderie make it a high-signal hour for any product leader invested in AI workflows.

    I also want to share that I'll be teaching our June 4th – July 9th cohort of Product Discovery Fundamentals. This is the last time I'll be teaching this cohort in its current format. If you've been thinking of enrolling in this program, and want to take it with me, this is your last chance. Register here.

    Across this cohort, we’ll practice continuous discovery habits—framing opportunities, tightening assumptions, running lean experiments, and aligning product trios on evidence-backed decisions. If you want a rigorous, repeatable system for turning customer insight into confident prioritization and compelling product strategy, I’d be thrilled to have you in the room.


    Inspired by this post on Product Talk.


    Book a consult png image
  • Behavioral Customer Data for Proactive SaaS Retention

    Behavioral Customer Data for Proactive SaaS Retention

    Your cancellation dashboard can tell you who has already left. It cannot tell you which accounts are failing to reach value, why their behavior changed, or what your team should do while the relationship is still recoverable.

    That is the real purpose of behavioral customer data. You are not trying to produce a more sophisticated churn report. You are building an operating system that turns observable behavior into a reason, an owner, and a timely response.

    Start with the retention decision, not the dashboard

    A risk score has no operational value if nobody knows what to do when it changes. Before choosing events, dashboards, or models, write down the retention decisions your data must support.

    For every proposed signal, define a decision contract:

    • Trigger: What behavior changed, started, stopped, failed, or never happened?
    • Interpretation: What customer state might that behavior indicate?
    • Owner: Should product, customer success, support, solutions engineering, or billing respond?
    • Intervention: What is the smallest useful action that could remove the obstacle?
    • Success signal: Which subsequent behavior would show that the customer is back on a value path?
    • Expiration rule: When should the alert or intervention stop so the customer is not repeatedly contacted?

    This contract prevents a common failure: treating all declining activity as the same problem. A customer who cannot finish an integration needs a different response from an activated customer whose core usage suddenly drops. A payment problem is different again. Combining them into one generic churn-risk label hides the information required to help.

    The signal also needs to match the product’s natural rhythm. Daily inactivity can matter in a daily workflow, but the same rule will create false alarms for a workflow used weekly or at the end of a reporting cycle. Compare behavior with the expected use pattern for the account’s persona, plan, lifecycle stage, and use case.

    I would design backward from a small set of decisions rather than forward from every event that happens to be available. The most useful leading indicators usually describe activation, time-to-first-value, depth of feature adoption, usage momentum, friction, and expansion intent. Each tells you something different about whether value is beginning, recurring, weakening, or growing.

    Instrument the path from first value to recurring value

    Measure value at the account level

    In B2B SaaS, the person clicking is not always the entity that retains. Users perform actions, while the account usually owns the subscription. Your model therefore needs both a reliable user identity and an account identity, plus a record of which users belonged to which account when the behavior occurred.

    This distinction matters when roles differ. An administrator may configure the product once, an operator may use the core workflow repeatedly, and an executive may only view outcomes. A login-frequency rule applied equally to all three will misclassify healthy behavior as disengagement. Define the value-producing behavior for each relevant persona, then roll those behaviors into an account-level state.

    Map the customer journey around observable value states:

    • Setup: The account has supplied the prerequisites required to attempt the core workflow.
    • Activation: The account has completed a meaningful milestone that indicates initial value, not merely finished an onboarding screen.
    • Recurring value: The core workflow is being completed at a cadence consistent with the use case.
    • Adoption depth: The account is using the capabilities required to obtain more complete or durable value.
    • Friction: Attempts, errors, failed integrations, or support interactions indicate that progress is being blocked.
    • Expansion intent: Behavior indicates a new use case, broader adoption, or interest in a relevant upgrade path.

    Your activation milestone is the pivotal definition. It should represent the earliest behavior that credibly demonstrates value. Completing profile fields or dismissing a tour may be easy to measure, but neither proves that the customer accomplished the job for which the product was purchased.

    Do not force one milestone across materially different use cases. If a plan, persona, or workflow changes the way value is produced, define the appropriate milestone for that segment. You can still report a common activation outcome while preserving the underlying reason an account qualified.

    Use a minimal tracking contract

    Once the value path is clear, instrument attempts, completions, failures, and meaningful outcomes along that path. A useful event contract includes:

    • A stable event name with a documented business meaning.
    • The user and account identifiers required for identity resolution.
    • The time the behavior actually occurred, not only the time it reached the analytics system.
    • The persona, plan, lifecycle stage, and use case needed for segmentation.
    • The product object or workflow involved.
    • A normalized outcome or error category when the action can fail.
    • The event owner and the process for approving semantic changes.

    For an integration workflow, for example, separate connection attempted, connection completed, and connection failed. Attach the provider and a controlled error category. Do not attach credentials, tokens, raw request bodies, or unrestricted personal information. Those fields create security and privacy exposure without improving the retention decision.

    The foundation is a clean event taxonomy, dependable identity resolution, and privacy-by-design. Capture only what the decision requires. If support sentiment is useful, prefer a governed derived category over copying unrestricted support conversations into an analytics platform. Keep sensitive material in the controlled system that already owns it.

    Before using any event in a risk score, ask product, data, and customer success to reconstruct the same account timeline. Check for duplicate events, delayed delivery, internal or test traffic, users mapped to the wrong account, plan changes that were not propagated, and renamed events with conflicting meanings. If those teams see different stories, automation will only distribute the disagreement faster.

    It is also safer to trigger interventions from a derived account state than directly from a raw event. A raw event says that something happened. An account state says whether activation is incomplete, recurring value has weakened, an integration is blocked, or a commercial issue is unresolved. That state can carry a reason code, observation time, data-quality status, and expiration rule into the product, lifecycle messaging, or customer success workflow.

    Build a risk score people can challenge and act on

    You do not need a black-box model to begin. A transparent rule set is often more useful because product and customer success can inspect the evidence, dispute a weak assumption, and choose the correct response.

    A practical account score can combine several distinct dimensions:

    <!– wp:list {
  • Unlocking AI Agents: The Real Barrier Is Readiness—Not Capability—Here’s How to Scale

    There’s a question that runs underneath every AI Agent evaluation: what can it do?

    Two years ago, that was the right question to ask because Agents were limited and capability was a genuine constraint. The gap between what organizations needed and what the technology could deliver was wide. I felt that gap acutely in early pilots—plenty of ambition, not enough dependable execution.

    That gap has since narrowed considerably, and yet most organizations are running their Agents well below what’s technically possible. I see teams lean on answering and routing, but stop short of looking things up, taking actions, or resolving complex, multi-step problems—especially where data, process variance, or risk come into play.

    The standard explanation is that AI isn’t good enough yet—models must improve, or vendors must ship more features. But after studying organizations across industries actively expanding their AI automation, I’ve found that this explanation holds up less often than people assume. The blockers tend to be elsewhere.

    The teams I’ve observed weren’t primarily constrained by what their AI could do; they were constrained by what their organization was structured to let it do. In other words, the ceiling wasn’t the Agent’s capability—it was organizational readiness, governance, and risk tolerance.

    “Readiness” for AI breaks into five distinct types, and most organizations have some but not all of them. Below is how I assess them with product, operations, and engineering leaders.

    Content readiness is whether you can explain your product and policies clearly and consistently. Most companies can. In practice, that means up-to-date knowledge bases, unified policy language, and clear versions that Agents can cite and apply.

    Scope readiness is whether you’ve defined the edges: when should AI engage, and when should it step aside? Edge cases multiply, intent varies by customer segment, sensitive topics surface mid-conversation, but most teams can work through this with effort. Clear guardrails reduce ambiguity and shrink risk.

    Procedural readiness is where things start to get harder. This is about whether you can articulate your processes clearly enough for something other than a human with years of tacit knowledge to follow. The happy path is rarely the problem. It’s the failure paths, decision branches, variations that have never been written down because they’ve always lived in someone’s head.

    Data readiness is the first real cliff. Can you reliably identify the right user, account, or object at the moment a decision needs to be made? Is the data trustworthy in real time? Are the APIs stable, accessible, and actually connected? For most organizations, the honest answer is “partially, but we’re not always sure when it breaks.”

    Execution readiness is the highest bar. Not just technically (can the Agent make the change?) but organizationally. Who owns it when the wrong refund gets processed? Who detects it? Who recovers? Does someone with authority actually accept the risk?

    Most companies have the first two, some have the third, fewer have the fourth and fifth. When I map this with teams, we often discover that their Agent’s ceiling is really a reflection of operational maturity and data plumbing, not model quality.

    We studied companies across six industries – energy, healthcare, ecommerce, gaming, financial services, property management – all trying to expand what their Agents could do. The pattern was consistent: teams set out to automate real actions—looking up account status, processing changes, handling transactions. In most cases, the AI could technically do it, but at a certain point (somewhere between guiding a user through a process and looking something up on their behalf) they hit a wall.

    One team tried to automate application changes but couldn’t reliably identify which application to modify across their internal systems. Another explored billing automation but couldn’t access live account data due to regulatory constraints. A third needed to verify status across third-party vendor systems their Agent couldn’t reliably reach. I’ve seen similar constraints surface around CRM integration, data governance, and vendor SLAs—none of which are model issues.

    In most cases, the team redesigned around what their infrastructure could support. They moved toward guiding—walking users through processes step by step, rather than executing changes on their behalf. It worked, it resolved conversations and delivered real value, just differently than anyone planned. In customer support, this often looks like consultative flows that shorten time-to-resolution even without direct writes.

    Most Agent evaluations are built around capability. Can it handle complex queries? Does it support multiple channels? Can it integrate with our systems? These are reasonable things to evaluate for, but they produce a capability score, and that doesn’t tell you whether your organization can actually use what you’re buying.

    The teams that got to deeper automation, the ones executing actions early, didn’t have “better AI,” they had more standardized operations. Actions that were already well-defined, consistently applied, and exposed through stable systems with clear rules. Automation wasn’t inventing new behavior, it was triggering actions that were already tightly controlled elsewhere.

    Readiness enables capability, not the other way around. Which reframes the evaluation question from “can the AI do this?” to “are we actually ready for it to?”

    Something that gets lost in most conversations about AI readiness is that organizations are often further along than they assume, just not for the kind of work they were planning for. A team that set out to automate refunds but can reliably guide users through complex troubleshooting has genuine capability deployed. They’re operating at the level their readiness supports, which is a starting point, not a deficit.

    The more useful frame isn’t “are we ready?” – it’s “what are we ready for, and what specifically stands between here and the next level?” The gaps tend to be concrete: a missing API, data that lives in three systems that don’t agree, a process that’s never been documented, or an ownership question nobody has answered. These are solvable problems. They just require a different kind of investment than buying a more capable Agent.

    What nobody has worked through seriously yet is how organizations actually build readiness. Does it develop naturally through using AI at shallower levels first? Or is it mostly a function of prior decisions, like system architecture choices made years ago, operational maturity that accumulated over time, engineering investments that have nothing to do with AI? When readiness does increase, what actually changes? Does the support team develop it? Does engineering grant it? Does it require executive sponsorship and investment in infrastructure with no obvious AI label on it?

    In my experience, progress comes from a joint effort: product to define scope and guardrails, operations to codify procedures and edge cases, engineering to harden APIs and observability, and leadership to underwrite risk with clear ownership. When those pieces align, agentic AI moves from guided assistance to safe, auditable execution.

    Until there are clearer answers, the pattern is likely to continue. Companies will buy capable Agents, plan ambitious rollouts, and find that the harder work is building the organizational infrastructure. The Agents can do the work. The question is what it takes to let them.


    Inspired by this post on The Intercom Blog.


    Book a consult png image
  • How to Validate Behavioral Heatmap Accuracy Before You Act

    How to Validate Behavioral Heatmap Accuracy Before You Act

    Your heatmap puts a bright cluster on the primary call to action, and the next step seems obvious: move the button, rewrite the copy, or prioritize a mobile redesign. Pause before turning that picture into a roadmap decision. A heatmap can look coherent while representing the wrong interface state, assigning clicks to the wrong element, or combining users whose layouts are materially different.

    Behavioral heatmap accuracy is not about whether the colors look plausible. It is about whether each recorded interaction appears on the interface the user actually encountered, within the correct context, and supports the conclusion you want to draw. You need to validate that chain before you act on the pattern.

    Treat accuracy as a chain, not a single metric

    There is no single accuracy score that makes a heatmap trustworthy. Four separate conditions have to hold:

    • Capture fidelity: The background image represents the relevant product state. The release, page structure, loaded content, navigation, overlays, and experiment variant should match what generated the interactions.
    • Placement fidelity: A click is attached to the intended interface element after responsive reflow, personalization, localization, and other layout changes. A precise coordinate on the wrong screenshot is still wrong.
    • Population fidelity: The map contains the users, devices, variants, and product states relevant to your decision. An aggregate can be mathematically correct while describing an interface that no individual user experienced.
    • Inference fidelity: The visualization can support the claim being made. A click establishes an interaction, not the user’s motivation. Scroll depth establishes reach, not attention, comprehension, or persuasion.

    Reliable screenshot capture, selector-based placement, automatic device detection, and clearer scrollmaps address important failure modes in this chain. They reduce ambiguity, but they do not eliminate the need to inspect your product states, filters, selectors, and supporting evidence.

    The weakest link determines whether the map is useful. Perfect element placement cannot rescue a screenshot from an old release. Clean device segmentation cannot justify a claim about user intent. Before discussing what the hot area means, establish what was captured, where it was placed, and whose behavior was included.

    Run a validation pass before reading the colors

    Use the same validation sequence whenever a heatmap is about to influence an experiment, design change, or roadmap priority. This turns accuracy from a vague feeling into a reviewable process.

    1. Write down the decision first. Be specific: move the primary action, remove a section, change the activation path, or investigate a mobile interaction. This tells you which page states and elements require the strongest validation.
    2. Freeze the analysis scope. Record the screen or template, analysis window, release, experiment variant, device class, and user segment. If the interface changed during the selected window, split the data or identify the limitation rather than treating the period as one stable experience.
    3. Build a state matrix. List only the states that materially alter the interface: desktop and mobile layouts, relevant locales, personalized variants, authenticated and unauthenticated views, expanded and collapsed components, or overlays that cover the underlying page. You do not need every possible segment. You need every state capable of moving, replacing, hiding, or duplicating the elements involved in your decision.
    4. Compare the screenshot with each relevant state. Check the order and size of major sections, sticky navigation, banners, modals, lazy-loaded content, and conditional components. If the displayed background is stale or combines interactions from incompatible layouts, stop interpreting the map and repair the capture or filtering first.
    5. Test element placement. In a controlled recorded session, interact with the target and with nearby controls that could be confused with it. Repeat the check on the layouts that move the element. The target’s hotspot should remain attached to the target rather than to an old coordinate. Exclude the controlled session from normal analysis when your tooling allows it.
    6. Inspect critical selectors. Ask engineering to confirm that each selector identifies the intended component across the templates and states in scope. Pay particular attention to repeated cards, reused button components, translated labels, and responsive navigation. If adjacent actions collapse into one hotspot, the map is not suitable for deciding between those actions.
    7. Reconcile the picture with events and replay. Apply equivalent page, date, device, user, and variant filters before comparing evidence. Exact numerical agreement is only a reasonable expectation when the systems use the same interaction definition and filters. Otherwise, document why their coverage differs and investigate unexplained gaps.
    8. Assign a confidence grade. Mark the map as decision-grade, directional, or invalid. Decision-grade means the relevant states and placements were verified. Directional means a pattern is visible but a known limitation prevents a precise conclusion. Invalid means the visual representation is wrong for the proposed decision.

    For a critical call to action, treat any reproducible placement error as a blocker. A hotspot that sometimes lands on a neighboring control can reverse the apparent preference between the two controls. Fix the representation before discussing design implications.

    Split heatmaps when the interface or interaction model changes

    Segmentation is not merely an analytical refinement. It is part of measurement accuracy. Mobile and desktop users may see different navigation, stacking order, content length, control size, and interaction affordances. Combining them can create a vivid composite that corresponds to neither experience.

    Use a simple rule: split the map whenever a cohort can encounter different geometry, different elements, or a different way of interacting. Check these questions before aggregating:

    • Does the same element exist in every included state?
    • Does it keep the same purpose and selector?
    • Does responsive behavior move it relative to neighboring elements?
    • Does a variant, locale, or personalized state change the surrounding content?
    • Are touch and pointer interactions being interpreted in a comparable way?
    • Did a release alter the template during the selected analysis window?

    If any answer exposes a material difference, inspect separate maps first. You can compare the resulting patterns afterward, but you should not use the blended view as the primary evidence.

    Scrollmaps need the same discipline. The same depth percentage can correspond to different content when a mobile page stacks sections that sit side by side on desktop. Compare scroll behavior within consistent layouts, then map each depth region to the actual value proposition, trust element, form, or call to action shown there. Scroll reach tells you that a region became reachable within the journey; it does not prove that the person read or understood it.

    Match the decision to what the evidence can prove

    Even a technically accurate heatmap is an observation layer. It can show where interactions accumulated or how far sessions progressed. It cannot, by itself, tell you why the behavior occurred or whether a proposed design change will improve an outcome.

    Use an evidence ladder instead of promoting every hotspot directly into the backlog:

    • Heatmaps locate the pattern. They help you identify concentrated clicks, neglected controls, competing actions, and sections reached by fewer sessions.
    • Event data measures the associated behavior. Use it to determine whether the interaction registered, where it sits in the funnel, and whether it connects to the micro-conversion or product outcome you care about.
    • Session replay supplies sequence and context. Inspect what happened immediately before and after the interaction, including overlays, loading states, repeated attempts, navigation changes, and other conditions that an aggregate view hides.
    • A controlled experiment evaluates the proposed change. When the claim is that a different placement, label, or layout will improve an outcome, compare that change against a baseline rather than treating the heatmap as causal proof.

    The combination also helps you diagnose apparent contradictions. A strong hotspot with no corresponding outcome event may indicate a broken interaction, incomplete instrumentation, or an action whose result is unclear. Low interaction on content that few sessions reach is first a placement or journey question, not automatically a copy problem. High scroll reach with low interaction means the region was available to users, but it does not establish that they noticed or rejected its message. A hotspot outside the visible target is a measurement defect, not a behavioral insight.

    Translate each finding into the next appropriate action:

    • If the screenshot, selector, or segment is wrong, create an instrumentation or analytics repair.
    • If the behavior is verified but its explanation is uncertain, create a discovery question and inspect relevant replays.
    • If the behavior is verified and tied to an outcome gap, define a hypothesis and an A/B test.
    • If the evidence reveals a reproducible interaction defect, prioritize the defect without disguising it as a preference experiment.

    This language matters in product reviews. Say that you observed a pattern, verified its representation, formed a hypothesis, and selected the next test. Do not say that users prefer, understand, ignore, or want something unless your evidence can support that stronger claim.

    Key takeaways

    • A heatmap is decision-grade only when the captured state, element placement, population, and proposed inference all align.
    • Validate the critical target and its neighboring controls across every layout that can move or replace them.
    • Split device classes, variants, releases, locales, or personalized states when they produce materially different interfaces.
    • Read scroll depth as reach and click concentration as interaction. Neither measure establishes attention, intent, or causality.
    • Pair heatmaps with event data and session replay, then use a controlled experiment when your decision depends on predicted impact.

    At your next heatmap review, do not begin with the hottest color. Begin with the screenshot, segment label, release, and one critical interaction traced from capture through outcome. If that path survives validation, turn the pattern into a hypothesis or product action. If it does not, fix the measurement before it becomes roadmap evidence.

    References

  • Governed AI Analytics in Financial Services: A Playbook

    Governed AI Analytics in Financial Services: A Playbook

    You have a credible AI analytics use case, product teams want access, and risk leaders want proof that the system will not expose sensitive data or influence the wrong decision. The mistake is to settle that tension with a broad choice between “innovation” and “control.” That choice is too vague to operate.

    Start with a narrower question: what decision may this system influence, using which data, under whose authority, with what evidence afterward? Once those boundaries are explicit, you can give teams meaningful speed without asking compliance to accept an invisible risk.

    Classify the decision before you assess the AI

    Many AI reviews begin with the model: where it is hosted, how it was trained, or whether it can explain an answer. Those questions matter, but they do not establish the business risk. The same model can summarize an approved dashboard, flag an unusual transaction pattern, or help determine an outcome that affects a customer. Those are not equivalent uses.

    Classify each use case by consequence, reversibility, and action authority. Consequence asks what happens if the output is wrong. Reversibility asks whether a person can correct the result before harm occurs. Action authority asks whether the system informs a person, recommends an action, or executes one.

    Use case patternPermitted role for AIControl that matters mostBoundary to make explicit
    Descriptive analysisSummarize approved metrics or behavioral patternsData permissions and traceable metric definitionsThe output cannot create a new customer-level action
    Investigative signalSurface anomalies or suspicious patterns for reviewAnalyst validation, evidence capture, and disposition loggingA signal is not a finding or a verdict
    Product recommendationSuggest an intervention, workflow, or experimentHuman approval and outcome monitoringThe recommendation cannot bypass existing approval paths
    Customer-affecting decisionSupport a formally governed decision processDocumented oversight, explainability, and accountable human authorityThe final authority and escalation path must be unambiguous

    This classification prevents two common errors. The first is applying the heaviest possible review to every analytical assistant, which sends teams into unofficial tools and manual workarounds. The second is treating every output as “just an insight” even when a downstream workflow turns it into a customer action.

    Trace the output one step beyond the interface. If an anomaly score enters a case-management queue, changes account handling, or triggers outreach, govern that downstream effect as part of the use case. A recommendation does not become low risk merely because a person clicks the final button.

    Before development begins, write an allowed-action statement and a prohibited-action statement. For example: “The system may prioritize patterns for analyst investigation. It may not label a customer, close a case, or initiate an external action.” That pair of sentences is more operationally useful than calling the project “medium risk.”

    Risk and compliance leaders still need to map the use case to the organization’s actual legal and regulatory obligations. A product risk classification is an operating tool, not a legal conclusion. When a use case could affect access, eligibility, pricing, fraud treatment, or another consequential outcome, obtain the appropriate compliance and legal review before activation.

    Turn governance principles into an enforceable contract

    Principles such as fairness, privacy, transparency, and human oversight do not control a production workflow by themselves. Each principle needs an owner, an enforcement point, and evidence that the control operated. I treat that combination as the governance contract for the use case.

    Define the data boundary

    List the approved data domains, fields, purposes, environments, and user groups. Do not stop at “customer data” or “analytics data.” Those labels are too broad to enforce. State which attributes the system can retrieve, which identifiers it can display, whether results may be exported, and where generated outputs may be stored.

    • Purpose: the business question the data may be used to answer.
    • Permitted inputs: the approved events, attributes, aggregates, and reference data.
    • Prohibited inputs: data classes that the workflow must never retrieve or infer.
    • Permitted users: roles allowed to query, review, approve, or export results.
    • Output handling: where results may be displayed, retained, shared, or reused.
    • Failure behavior: what the system does when permission, provenance, or confidence is insufficient.

    Enforce that boundary with role-based access controls and granular permissions at retrieval time. Filtering an answer after a model has received restricted data is not equivalent to preventing access. The model, retrieval layer, analytics service, export path, and destination workflow all need to respect the same user identity and policy context.

    Assign decision rights to named roles

    A committee can set policy, but it cannot own every operational decision. Give each use case an accountable product owner, a data owner, a control owner, and a business reviewer. Clarify who can approve launch, who can change the data scope, who reviews exceptions, and who has authority to stop the workflow.

    • The product owner defines the user problem, allowed action, prohibited action, and business outcome.
    • The data owner approves the data purpose, quality expectations, permissions, and reuse limits.
    • The risk or compliance owner maps policy obligations to testable controls and reviews material exceptions.
    • The platform or security owner implements identity, access, isolation, logging, and change controls.
    • The business reviewer accepts, rejects, or escalates outputs and records why.

    Keep the decision rights close to the workflow. If a reviewer sees an unsupported conclusion, that person needs a clear way to reject it, preserve the evidence, and route the issue. If every exception disappears into a general governance inbox, the formal control will be bypassed when operational pressure rises.

    Design the audit record before launch

    An audit trail should reconstruct what happened without relying on someone’s memory. Capture the requesting identity and role, the approved purpose, the data and metric definitions used, the system configuration, the generated result, any human review, the resulting action, and later corrections or overrides.

    Logging creates its own data risk. Prompts, retrieved context, generated explanations, and reviewer notes can contain sensitive information. Protect the audit store with appropriate access, retention, and segregation rather than treating logs as harmless operational exhaust. Where policy permits, record protected references to sensitive records instead of duplicating raw payloads.

    A practical platform evaluation should test whether the system combines strong data governance, auditable AI behavior, secure scale, and a direct connection to product outcomes. A policy document that cannot be enforced in the workflow is not enough, and a platform control without an accountable operating process is not enough either.

    Put controls inside the workflows people actually use

    Governance fails when it exists as a review ceremony around the product rather than a behavior inside it. Analysts should not have to remember a separate policy every time they ask a question. The approved data scope, identity context, review step, and evidence capture should travel with the task.

    Behavioral analytics: govern the meaning as well as the data

    Behavioral analytics can reveal how customers move through onboarding, self-service, support, payments, and other product journeys. The danger is not limited to unauthorized access. An AI system can also combine valid events into a misleading interpretation of customer intent.

    Start the workflow with curated event definitions and approved business metrics. Require the output to expose the cohort definition, time context, filters, exclusions, and comparison used. The analyst should be able to inspect the path from a narrative claim back to the underlying measure before sharing it.

    Separate observation from inference in the interface. “Users in this cohort abandoned the flow after this step” is an observation tied to event data. “They abandoned because they distrusted the process” is a hypothesis. Labeling those differently prevents fluent language from turning a plausible explanation into an unsupported fact.

    Anomaly detection: route a signal into investigation, not judgment

    An anomaly means a pattern differs from an expected baseline. It does not establish fraud, customer intent, system abuse, or operational error. Treat anomaly detection as a prioritization mechanism unless a separately governed process establishes something more.

    Give the reviewer the observed deviation, relevant context, the comparison baseline, and links to permitted evidence. Capture the reviewer’s disposition: confirmed issue, expected behavior, insufficient evidence, data-quality problem, or escalation. That disposition is both an audit artifact and a feedback signal for improving the workflow.

    Watch the operational burden as closely as the detection capability. A flood of weak signals can make the nominal control less safe because reviewers rush, defer, or stop trusting the queue. Monitor false positives, unresolved escalations, overrides, and the reasons analysts reject outputs. When those indicators deteriorate, reduce scope or pause automated routing while the cause is investigated.

    Self-service analysis: give teams a governed lane

    Product managers and analysts need enough freedom to explore without sending every question through a central approval queue. Create a governed workspace containing approved metrics, documented data products, role-aware access, and restricted export paths. Let people iterate freely inside that lane while changes to data scope, decision authority, or external activation trigger a new review.

    Make the boundary visible. Users should know when an answer is based on incomplete data, when a metric is not approved for customer-level decisions, and when an output cannot be exported. A silent denial encourages workarounds; a clear denial that identifies the policy boundary gives the user a legitimate next step.

    Do not give an analytics assistant write access to operational systems merely because the integration is convenient. Insight generation and action execution are separate privileges. Connect them only when the action, reviewer, failure mode, and rollback path have been governed explicitly.

    Pilot with evidence, not a polished demonstration

    A convincing demo proves that the happy path works. A governed pilot must also prove that the system refuses the wrong request, exposes enough evidence for review, and leaves a usable record when something goes wrong.

    Choose a narrow workflow with an identifiable user, a bounded data set, a reviewable output, and a business outcome you already understand. Avoid beginning with an enterprise-wide assistant or an autonomous action layer. Broad scope makes it difficult to distinguish model behavior, data problems, permission failures, and process gaps.

    1. Write the decision contract. Record the user, purpose, permitted inputs, allowed action, prohibited action, reviewer, and stop authority.
    2. Configure the smallest useful data boundary. Include only the fields and metrics needed for the chosen workflow.
    3. Test legitimate work. Confirm that authorized users can produce an insight, inspect its basis, and complete the intended review.
    4. Test prohibited work. Attempt access with the wrong role, request excluded attributes, try an unauthorized export, and ask the system to take a prohibited action.
    5. Test ambiguity and failure. Use incomplete context, conflicting metric definitions, missing permissions, and unavailable dependencies. Confirm that the system fails visibly and safely.
    6. Reconstruct the event. Use the audit record to determine who requested the output, what information was used, what was generated, who reviewed it, and what happened next.
    7. Change the system deliberately. Update a relevant configuration or model component and confirm that approval, documentation, testing, and monitoring follow the change.

    Do not accept screenshots as evidence for controls that operate behind the interface. Ask the vendor or internal platform team to demonstrate a denied request, a permission change, a reviewer override, an exported audit record, and the behavior after a governed configuration change. The test should follow your use case and identities, not a generic demonstration tenant.

    Measure value and control health together. If the system produces faster insights but increases unreviewed actions, weakens attribution, or creates an investigation backlog, it has not delivered a durable improvement.

    DimensionQuestionUseful signals
    Business valueDoes the workflow improve a real product, growth, risk, or operational decision?Time to a validated insight, useful investigations completed, issues resolved, and attributable product outcomes
    Analytical qualityCan a reviewer verify the conclusion?Accepted and rejected outputs, unsupported claims, metric-definition errors, and missing context
    Control effectivenessDid policy operate as designed?Prohibited requests blocked, required reviews completed, permission exceptions, and audit-record completeness
    Operational healthCan people sustain the workflow?False-positive burden, unresolved escalations, overrides, rework, and reviewer backlog
    Change safetyDo updates preserve the approved boundary?Documented changes, completed regression checks, new failure patterns, and monitored post-change behavior

    Set release gates in binary language. The use case has a named accountable owner or it does not. Permissions have been tested with unauthorized identities or they have not. High-impact outputs receive the required review or they do not. Audit evidence can reconstruct an event or it cannot. Ambiguous gates become exceptions as soon as delivery pressure appears.

    When the pilot is stable, reuse the control components rather than copying the entire use case. Standard identity propagation, data classification, audit schemas, reviewer workflows, and change gates can form a shared control plane. Each new use case still needs its own purpose, decision boundary, outcome measure, and risk assessment.

    Key takeaways

    • Govern the decision the AI can influence, not just the model that produces the output.
    • Write both an allowed-action statement and a prohibited-action statement before development begins.
    • Enforce data permissions before retrieval and carry the user’s identity through analysis, export, and downstream action.
    • Treat human review as an operational workflow with evidence, dispositions, escalations, and stop authority.
    • Keep observations, hypotheses, recommendations, and customer-affecting decisions visibly distinct.
    • Test denial, ambiguity, change, and audit reconstruction alongside the happy path.
    • Track business value, analytical quality, control effectiveness, and operational burden on the same scorecard.

    Your next move is not to draft an enterprise AI policy. Pick one live analytics workflow and write its decision contract on a single page. If you cannot name the allowed action, prohibited action, data boundary, reviewer, audit evidence, and stop authority, the workflow is not ready to scale. If you can, you have the foundation for AI analytics that product teams can use and risk leaders can defend.

    References