Tag: customer support ai strategy

  • Own Your AI: 4 Essential Roles to Supercharge Support and Prevent Performance Drift by 2026

    Own Your AI: 4 Essential Roles to Supercharge Support and Prevent Performance Drift by 2026

    AI doesn’t fail because the model is bad, it fails because ownership is missing.

    When someone truly owns your AI, everything changes. Resolution and automation rates climb, the system self-improves, and the customer experience transforms in ways a dashboard alone will never show you.

    This is part three of our five-part series on customer service planning for 2026. We’ll be sharing all five editions on our blog and on LinkedIn.

    If you’d rather have them emailed to you directly as they’re published, drop your details here.

    Last week, we introduced the four roles that make AI actually work in a support organization. These roles are already showing up inside the teams who are scaling AI the fastest, and this week, we get closer to the ground.

    Here’s what these roles look like in practice — what they do, how they work, and why your AI performance will inevitably drift without them.

    AI operations lead — owns AI performance, every day. I think of this person as the air-traffic controller for our AI Agent. I treat the AI as a living system that needs ongoing supervision, evaluation, and tuning. This role is accountable for what leaders care about most: quality, reliability, and continuous improvement.

    The AI ops lead sees the whole picture: conversation quality, missing knowledge, flawed assumptions, unexpected failures, new opportunities for automation, and the subtle signals that the system is beginning to drift. In practice, that vigilance is the difference between steady gains and slow decline.

    Day-to-day, here’s what I expect from this role.

    1. Reviews AI conversations and surfaces performance patterns. The AI ops lead monitors the AI Agent’s behavior — the tone shift after a product launch, a sudden dip in resolution for a specific intent, or conversation clusters revealing new customer behavior. They scan for anomalies, trends, and early warnings, with an emphasis on what’s happening right now, not last week. Without this intentional ownership, I’ve watched a 2% dip turn into a 10% drop in days.

    2. Prioritizes fixes and improvements. Once patterns emerge, they triage fixes like a product team handles bugs. Missing or incorrect content? They route it to the knowledge manager. Behavioral issues? They adjust guidance and guardrails. Action or system issues? They partner with the automation specialist. This connective tissue turns individual fixes into compounding improvements.

    3. Defines and maintains AI guardrails. Leaders everywhere worry about AI doing things it shouldn’t. This role answers that fear by establishing clarification logic, escalation rules, “never answer” policies, and safety boundaries. The goal is predictable behavior that protects customer trust — an essential pillar of any AI Strategy and AI risk management practice.

    4. Aligns reporting with leadership. The AI ops lead reports on resolution rate, CX Score, CSAT, automation coverage, and hours saved — making the economic impact visible. That visibility is a foundational step in any credible customer support ai strategy.

    Why this role exists now. AI systems are dynamic and require constant tuning. A small dip in quality quickly becomes an operational issue, and no existing role naturally owns that. When someone does, teams feel the benefit almost immediately.

    Knowledge manager — builds and maintains the structured knowledge AI depends on. I hear the same thing from leaders again and again: AI is only as good as the content you give it. This role is rapidly evolving from classic knowledge management into knowledge strategy — part content designer, part systems thinker, part information architect. Their job is to build the knowledge scaffolding that lets AI answer accurately, consistently, and safely.

    Here’s how the knowledge manager creates leverage.

    1. Writes, maintains, and improves support knowledge — continuously. After every product change, they update articles, remove duplication, resolve contradictions, and pay down “knowledge debt” that quietly erodes accuracy. The upkeep is shaped by AI performance; when patterns expose gaps, they fix the source.

    2. Structures knowledge for AI, not for browsing. Traditional help centers are for humans skimming pages. AI needs clean intent signals, crisp formatting, and clearly structured language. The knowledge manager designs that structure as intentionally as the content itself.

    3. Works hand-in-hand with AI ops. Many performance issues stem from missing or unclear knowledge. When the AI ops lead surfaces recurring misunderstandings or low-resolution categories, the knowledge manager resolves the root cause at the source.

    4. Ensures accuracy and compliance at scale. As AI handles more sensitive situations, the knowledge manager safeguards correctness, currency, and compliance — critical for data governance and regulatory alignment.

    5. Develops a cross-functional knowledge strategy. The role creates a canonical, cross-functional source of truth that product, engineering, product marketing, go-to-market, and support (AI and human) can all rely on.

    Why this role exists now. This is one of the highest-leverage positions in an AI-first support org. Teams like Rocket Money and Anthropic are hiring knowledge managers because AI accuracy depends on the quality of knowledge feeding it. Without this role, resolution rate caps out early and never climbs.

    Conversation designer — designs how the AI speaks, clarifies, and interacts. AI isn’t just a tool customers use; it’s a representative they interact with. Tone, clarity, pacing, and conversational structure matter, especially in voice. Every word affects perceived expertise, trustworthiness, and brand. The conversation designer ensures the AI feels human-friendly without pretending to be human — the sweet spot that builds trust without misleading customers.

    In my experience, staffing conversation design early accelerates results. It changes not only how we tune AI, but how we understand the end-to-end customer experience.

    Here’s what great conversation design looks like.

    1. Shapes the AI’s tone, voice, and communication style. This role refines phrasing, tunes politeness, adjusts how confusion is handled, and shapes micro-interactions that determine whether customers feel cared for or dismissed. On voice channels, natural cadence is make-or-break.

    2. Designs flows for high-value conversations. They design how the AI clarifies intent, branches, communicates uncertainty, verifies details, escalates, hands off, and returns to the main thread without feeling mechanical — treating customer experience as a product with language as the interface.

    3. Translates procedures and complex workflows into natural language and logic. As AI runs structured procedures and actions, this role becomes a conversational system architect, translating SOPs into conditional logic with exceptions and fallbacks. For example, in Intercom, our conversation designer uses Simulations to run simulated conversations to see where the AI Agent gets confused, over-confident, or awkward, and refine flows until the interaction feels effortless end-to-end.

    4. Ensures transitions to humans feel smooth and respectful. Handoffs should provide clear context to the human agent and maintain continuity so customers never feel dropped.

    Why this role exists now. As AI becomes the primary interface, conversation design directly influences trust, brand perception, and operational outcomes. It’s a core competency for any Generative AI and LLMs for product managers program.

    Support automation specialist — builds the backend actions that allow AI to do real work. If the conversation designer shapes expression, this role shapes capability. They transform AI from an answering machine into an outcome engine by bridging AI and the systems it must safely and deterministically act on.

    Support teams increasingly expect AI to do what a human would do: refund a charge, adjust a subscription, verify an identity, update an account setting, or pull relevant data. That expectation creates a new technical role at the edge of support, ops, and engineering.

    What I rely on this specialist to deliver.

    1. Creates and maintains backend workflows the AI executes. This includes building and maintaining: Fin Tasks. Fin Procedures with embedded steps. Action flows that call internal and external APIs. Automations that span billing systems, user identity layers, CRM objects, subscription entitlements, refund tools, and more. They ensure the AI can act compliantly and predictably — the playbooks that turn intent into action.

    2. Owns the integrations required for advanced automation. Many problems require data elsewhere — billing platforms, internal databases, systems of record. The specialist ensures the AI can retrieve, validate, and use that information safely, often partnering closely on CRM integration and internal services.

    3. Partners closely with product and engineering. Some workflows require new endpoints, permission layers, safety gates, or deterministic fallbacks. This role drives those changes across the stack.

    4. Ensures reliability and safety at every step. Guardrails, validation logic, exception handling, safe execution paths — all are essential. They confirm that the AI has access to the correct data, the action matches policy, edge cases are accounted for, risky flows have deterministic constraints, and every action is auditable and reversible.

    Why this role exists now. Customers don’t want answers, they want outcomes. AI can now deliver those outcomes, but only with the right backend scaffolding. This role modernizes operational architecture and unlocks end-to-end automation.

    How these roles work together — the new operating loop. These roles aren’t silos; they’re interdependent parts of one system. The AI ops lead identifies patterns and performance gaps. The knowledge manager resolves inaccuracies or missing content. The conversation designer improves clarity, tone, and flow. The automation specialist expands the system’s ability to take action. Each improvement compounds the next, moving you from early automation to transformational resolution rates through continuous refinement.

    This loop is what separates teams that plateau early from teams that scale AI into a reliable, high-performing system — the essence of a durable AI Strategy.

    How to get started (even if you can’t hire all four roles today). Most teams phase into this model: assign partial ownership, formalize responsibilities, then specialize as AI volume grows. Here’s the progression I recommend.

    Phase 1: Assign ownership. Give each role’s core responsibilities to someone who can devote five to 10 hours weekly. Early on, support ops, enablement, senior ICs, and technically inclined teammates can anchor the work.

    Phase 2: Formalize the responsibilities. As AI resolves more queries, optimization becomes core operational work. Formalizing ownership prevents performance drift and knowledge debt.

    Phase 3: Specialize and hire. Once AI handles 50–70% of incoming volume, these responsibilities become full-time roles. Investing in specialization becomes essential infrastructure for the next scale stage.

    The bottom line. AI changes the shape of your support team. These four roles — AI operations lead, knowledge manager, conversation designer, and support automation specialist — form the backbone of the AI-first support organization. They bring order to a constantly changing environment and enable AI to deliver the outcomes leaders and customers expect heading into 2026.

    Next week, we’ll continue the 2026 planning series with a deep dive into org design models for AI-first support teams — how to structure people, workflows, and accountability in a world where AI resolves most conversations before a human ever sees them.

    To follow along with the series and have each new edition emailed to you directly, drop your details here.


    Inspired by this post on The Intercom Blog.


    Book a consult png image
  • How to Build a Conversation-Based Customer Experience Score

    How to Build a Conversation-Based Customer Experience Score

    Your dashboard says the ticket was resolved. The customer remembers repeating the problem, moving between an AI agent and a teammate, and discovering that company policy still blocked the outcome they wanted. Product, Support, and Operations can all look at the same conversation and reach different conclusions.

    If you are considering conversation-based customer experience scoring, the hard part is not asking an AI model for a rating. It is designing a measurement system that distinguishes the experience from its causes, shows people why the score exists, and sends each cause to someone who can change it.

    A useful score separates experience from ownership

    A customer experience score should answer a narrow question: how well did this interaction work for the customer? It should not silently answer a different question, such as whether the support agent performed well or whether the product team made the right policy decision.

    Those questions overlap, but they are not interchangeable. A teammate can give a clear and accurate explanation of an unpopular refund policy. The teammate’s answer quality may be strong while the overall experience remains poor. An AI agent can use a warm tone while giving an incorrect answer. The sentiment may look positive even though the handling failed. A product limitation can make resolution impossible despite excellent support work.

    This is why a credible score needs several layers:

    • Outcome: Was the customer’s request resolved, partially resolved, redirected to a workable next step, or left unresolved?
    • Answer quality: Were the responses clear, accurate, relevant, and internally consistent? Evaluate AI and human responses separately when both participated.
    • Customer effort: Did the customer repeat information, survive avoidable handoffs, chase a promised follow-up, or clarify something the company should already have understood?
    • Emotional context: Did the customer express strong frustration, anger, relief, gratitude, or delight? Treat emotion as context rather than a verdict by itself.
    • Product or service feedback: Was the customer reacting to a bug, missing capability, reliability problem, delivery failure, confusing design, or service issue?
    • Policy feedback: Was the real source of dissatisfaction a refund rule, eligibility condition, account limit, return policy, or another business decision?

    These dimensions reflect the reality that customers react to the whole interaction, including effort and product or policy constraints, not merely the final support response.

    Score the experience first. Attribute the drivers second. Assign ownership third. Reversing that order creates predictable dysfunction: teams defend their own performance, difficult conversations get excluded, and the metric becomes a political argument instead of a customer signal.

    Design the score as a diagnosis, not a black box

    Leadership may want one number for a dashboard, but the useful product is the diagnostic record underneath it. If a support leader cannot open a low-scoring conversation and see why it received that result, the number is not ready for coaching, prioritization, or executive reporting.

    The minimum record behind each score

    For every eligible conversation, preserve these fields:

    • Overall experience band: A small set of anchored labels is easier to calibrate than a decimal-heavy score that implies unsupported precision.
    • Eligibility status: Record whether the interaction was scored, excluded under a defined rule, or genuinely lacked enough information.
    • Outcome status: Resolved, partially resolved, unresolved, or unclear.
    • Answer-quality results: Separate evaluations for AI and teammate contributions where applicable.
    • Driver codes: Effort, strong emotion, product or service feedback, policy feedback, and any operational reason codes you have explicitly defined.
    • Evidence: The specific message or interaction event that supports each driver. A generated explanation without transcript evidence is an assertion, not an explanation.
    • Plain-language summary: What the customer needed, what happened, and why the experience earned its band.
    • System metadata: The scoring model, rubric, and schema versions used to produce the record.

    I would begin with anchored experience bands rather than pretending the system can distinguish tiny numerical differences. A practical rubric might distinguish a strong experience, an acceptable experience with minor friction, a weak experience with material friction or incomplete resolution, and a poor experience with an unresolved outcome, serious inaccuracy, contradiction, or excessive burden.

    The labels matter less than the anchors. Reviewers need observable criteria for each band. Phrases such as good conversation or unhappy customer leave too much room for interpretation. Criteria such as customer repeated the account history after a handoff or answer contradicted an earlier commitment can be checked against the transcript.

    Do not let emotion dominate the rubric. A customer may arrive angry because of a product outage and receive excellent assistance. Another may remain polite after receiving a materially wrong answer. Emotion can increase urgency and explain the experience, but it cannot substitute for outcome, accuracy, and effort.

    Do not average away disagreement between dimensions either. An acceptable overall score can conceal an inaccurate AI answer that a teammate later repaired. Preserve that AI-quality failure as a driver so the AI product team can add it to an evaluation set even when the customer ultimately gets a resolution.

    Make the metric reliable enough for decisions

    A score can look stable while measuring a changing subset of conversations. If short threads, low-context requests, escalations, or mixed AI-human interactions are harder to score, improvements in the average may simply reflect which conversations entered the denominator.

    Coverage therefore belongs beside the score, not in a technical footnote. Broader scoring can reveal parts of the support mix that were previously invisible, and adding previously unscored conversations can change the reported result even when operating performance has not changed.

    Define eligibility before calibration. Spam, automated notifications, internal-only threads, and interactions with no customer request may reasonably sit outside the metric. A short conversation should not be excluded merely because it is short, and a difficult conversation should not be excluded merely because the model is uncertain. Track uncertainty explicitly rather than removing inconvenient cases from view.

    Your recurring dashboard should show:

    • The share of eligible conversations that received a score.
    • The distribution across experience bands, not just an average.
    • The mix of positive and negative drivers.
    • Results split by AI-only, teammate-only, and mixed handling.
    • Relevant slices such as channel, language, issue type, conversation length, product area, and escalation path.
    • The active model, rubric, and schema versions.

    Calibration should happen against human judgment before the score becomes a target. Use a representative set containing routine resolutions, short exchanges, long investigations, escalations, emotionally charged threads, AI-only conversations, human-only conversations, and AI-to-human handoffs. Have independent reviewers apply the same rubric, examine disagreements, and rewrite any criterion that depends on intuition rather than observable evidence.

    Then test the slices separately. Aggregate agreement can hide systematic failure in one language, channel, issue class, or interaction type. The acceptable level of disagreement depends on the decision. A model used to discover recurring workflow friction can tolerate more uncertainty than one used in individual performance management.

    Keep the adjudicated examples as a regression set. Re-run them whenever you change the prompt, model, rubric, knowledge architecture, conversation parser, or driver definitions. Review newly common failure patterns as well; a frozen evaluation set eventually stops representing the work.

    Model changes require visible reporting boundaries. A more contextual scoring system may produce a one-time shift without a corresponding decline in support quality. Backfill historical conversations with the new version when that is practical. Otherwise, annotate the change on every trend view and establish a new baseline. Never splice two scoring regimes into one continuous line and ask leaders to interpret the movement as operational performance.

    Turn low scores into routed work, not dashboard theatre

    A low score is only a symptom. The driver determines who should investigate it and what kind of intervention is plausible. Sending every poor experience to the support manager guarantees that product defects, policy choices, and broken workflows will be misclassified as coaching problems.

    Primary driverWhat to inspectPrimary ownerDefault next action
    AI answer qualityInaccuracy, contradiction, irrelevant guidance, or repeated clarificationAI product and knowledge ownersCorrect the underlying knowledge or response path, then add the failure to the AI evaluation set
    Teammate answer qualityUnclear explanation, incorrect guidance, missed question, or inconsistent commitmentSupport lead or enablement ownerReview the conversation against the rubric, then improve coaching, documentation, or access to information
    Customer effortRepeated information, handoff loops, unnecessary forms, follow-up chasing, or duplicated verificationSupport operations or journey ownerMap the failing transition and remove the avoidable step, ownership gap, or workflow rule
    Product or service feedbackBug, missing capability, confusing design, reliability issue, delivery failure, or service breakdownRelevant Product, Engineering, or service ownerCluster related conversations, connect them to the product area, and decide whether the response is a fix, discovery work, or an explicit trade-off
    Policy feedbackRefund, return, eligibility, account, usage, or limit ruleBusiness or operations owner responsible for the policySeparate unclear communication from disagreement with the policy, then revise the explanation, the policy, or neither – deliberately
    Strong negative emotionThe event that triggered the emotion and whether the issue remains unresolvedTriage owner, followed by the owner of the actual causePrioritize review where appropriate, but do not treat emotion alone as proof of agent failure

    Automation should route the evidence package, not just the score. Include the conversation link, customer request, outcome, overall band, driver codes, supporting messages, scoring version, and proposed owner. That context lets the receiving team judge the issue without rereading an entire thread or trusting an opaque summary.

    Use separate operating lanes for individual cases and recurring patterns. A materially incorrect answer may need immediate review. Repeated handoff friction usually needs aggregation so Operations can see the broken transition. Product and policy feedback becomes useful when related conversations are clustered around a shared problem, while still retaining representative examples.

    Count affected conversations consistently rather than allowing a verbose customer to create many separate votes within one thread. Preserve the denominator for every filter. A driver that appears frequently in one product area may look dominant in a filtered dashboard while remaining uncommon across the full support mix.

    For recurring themes, maintain a problem record with the driver, affected journey, frequency, severity, controllable owner, proposed intervention, status, and comparable post-change cohort. This converts conversation scoring into a product and operations feedback loop. Without that record, the same issue can be rediscovered in every review without anyone becoming accountable for changing it.

    After an intervention, compare like with like: the same scoring version, eligibility rules, issue type, and relevant handling path. If the score improves but coverage falls, or the issue mix changes, you do not yet know whether the intervention worked.

    Earn the right to replace CSAT

    Conversation scoring addresses a real blind spot: survey metrics describe the customers who choose to respond, while a conversation-based system can evaluate a much broader share of eligible support volume. That makes it attractive as a replacement for CSAT, but broader coverage does not automatically make the new metric valid.

    Start in shadow mode. Continue the existing reporting while you calibrate the new score, inspect disagreements, and learn which drivers are actionable. Do not demand that the two measures match. They observe different things: one evaluates evidence in the interaction, while the other records a respondent’s self-reported reaction.

    Move the conversation score into operational reviews once teams can inspect its reasoning and route its drivers. Move it into executive reporting only after coverage, version changes, and slice-level performance are visible. Consider reducing or retiring a survey only when all of the following are true:

    • Eligibility and coverage are stable enough that changes in the denominator cannot masquerade as experience improvements.
    • The rubric has been calibrated against human review, including difficult and ambiguous conversations.
    • Explanations consistently point to transcript evidence rather than merely producing plausible prose.
    • Important channels, languages, issue classes, and AI-human handling paths have been checked separately.
    • Model and rubric changes are versioned, regression-tested, and visibly marked in reporting.
    • Driver routing produces owned work, and teams can show what they changed because of the signal.
    • Material disagreements between the conversation score and survey feedback are investigated rather than averaged away.

    Keep a higher standard for individual performance decisions. A conversation score can flag work for human QA, but it should not become an automatic employee rating merely because it covers more conversations. Product limitations, customer history, policy constraints, and model error can all affect the result. Use the driver record and human review to establish what the teammate actually controlled.

    Key takeaways

    • Measure the customer’s experience separately from the performance of the AI agent, teammate, product, policy, or workflow that shaped it.
    • Keep an overall band for scanning, but preserve outcome, answer quality, effort, emotion, feedback drivers, evidence, and version metadata underneath it.
    • Report coverage and score distribution together; an unexplained denominator change can invalidate the trend.
    • Calibrate with representative human-reviewed conversations and retest meaningful slices after every scoring change.
    • Route each driver to the owner who can change it, then measure a comparable cohort after the intervention.
    • Replace CSAT only after the conversation score has earned trust as both a measurement system and an operating loop.

    At your next customer experience review, bring one low-scoring conversation, its evidence-backed driver record, and the owner capable of changing that driver. If the meeting ends with only a debate about whether the number is fair, calibration is unfinished. If it ends with a named intervention and a valid way to examine comparable future conversations, the score is doing useful work.

    References

  • AI-Enabled Customer Support Roles: A Practical Org Design

    AI-Enabled Customer Support Roles: A Practical Org Design

    Your AI agent is resolving enough conversations that queue volume is no longer a useful blueprint for organizing the team. Yet your org chart still assumes that every customer outcome belongs to a human agent. The result is a dangerous ownership gap: everyone can recognize a poor AI interaction, but no one is clearly responsible for the content, behavior, action, or handoff that caused it.

    The decision in front of you is not simply which jobs AI will remove. It is which responsibilities become more important, who should own them now, and when that work deserves a dedicated role. You can answer those questions before adding headcount.

    The unit of work has shifted from tickets to systems

    A human-owned ticket usually has a visible assignee, a queue, and a closed state. An AI conversation can fail much earlier in the system. The policy may be stale. The right knowledge may exist but be difficult to retrieve. The response may be accurate but confusing. A backend action may fail. A handoff may reach a person without the context needed to continue.

    If you classify all of those outcomes as generic AI accuracy problems, the team will spend its time rewriting prompts while structural defects remain untouched. Diagnosis has to begin with the layer that failed.

    • Knowledge: Did the system have an accurate, current, and unambiguous basis for answering?
    • Conversation: Did it communicate clearly, follow policy, and recognize when it should stop or escalate?
    • Action: Could it complete the requested task safely and confirm the result?
    • Operations: Could the organization detect the failure, assign it, correct it, and verify that the correction worked?

    When an AI agent carries a substantial share of conversation volume, human work moves from processing individual questions toward improving the system that handles them. That does not make human support irrelevant. People still own ambiguity, exceptions, sensitive situations, and cases that require judgment. What changes is the management model around them.

    Start with a simple accountability rule: every layer needs one named owner. Several people can contribute, but shared contribution is not shared accountability. If nobody has the authority to prioritize a correction and see it through, performance will drift between support, product, engineering, and content teams.

    Assign four ownership roles before opening four requisitions

    I would not begin by hiring four specialists. I would begin by assigning four explicit responsibilities to people who already understand the customers, policies, tools, and failure modes. A dedicated job becomes necessary when the work is continuous, consequential, and repeatedly displaced by the owner’s primary role.

    AI operations lead: owns performance and improvement

    The AI operations lead is accountable for day-to-day performance. This person maintains the quality view, classifies recurring failure modes, prioritizes corrections, and coordinates changes across support, product, data, content, and engineering.

    This is not a meeting coordinator or a person who manually edits every weak response. The role needs enough operating authority to decide which problems deserve attention and enough analytical depth to separate isolated bad conversations from systemic patterns. Support operations is often a strong internal starting point because the function already understands workflows, tooling, routing, and capacity.

    The first useful deliverable is an AI performance register. For each recurring issue, record the affected customer intent, observed outcome, failure layer, accountable owner, proposed change, validation method, and current status. That register becomes the shared backlog for the AI support system.

    A good decision boundary is equally important: the operations lead prioritizes what must improve, while the relevant domain owner decides how to change knowledge, conversation behavior, or automation. Otherwise, one person becomes a bottleneck for every adjustment.

    Knowledge manager: owns what the AI is allowed to know

    The knowledge manager owns the material that grounds customer answers: help content, internal procedures, macros, snippets, policy explanations, and the relationships between them. The job is not merely publishing more documentation. It is making sure the AI has one dependable answer for each supported question.

    Conflicting content is often more damaging than missing content because the system can produce a plausible answer from the wrong instruction. Every important knowledge item should therefore have a clear owner, audience, scope, status, and review trigger. Product or policy changes should update the source of truth before teams try to compensate through prompt wording.

    The first useful deliverable is a knowledge inventory organized by customer intent. Mark where content is missing, duplicated, contradictory, overly broad, or dependent on information the AI cannot reliably access. This turns an abstract content audit into a prioritized quality backlog.

    Measure this role by whether knowledge-related failures become easier to prevent and diagnose. Page count is an output. Dependable answers are the outcome.

    Conversation designer: owns how the AI behaves

    The conversation designer defines how the AI communicates and how an interaction progresses. That includes tone, question sequencing, explanation structure, confirmation language, boundaries, escalation triggers, and the context passed to a human agent.

    This role is broader than polishing wording. A response can be factually correct and still produce a poor outcome because it is too confident, asks for information in the wrong order, buries a constraint, or continues after the situation calls for human judgment. Conversation design turns brand, policy, and customer-experience expectations into observable behavior.

    The first useful deliverable is an interaction specification for each important intent. It should define the customer’s likely goal, information required, acceptable response structure, prohibited claims, escalation conditions, action confirmation, and handoff payload. That specification gives QA something concrete to evaluate and gives the automation specialist a stable flow to implement.

    Content design, UX writing, and support enablement are natural backgrounds for this work. The essential skill is not clever phrasing. It is recognizing how small language and sequencing choices affect comprehension, trust, and completion.

    Support automation specialist: owns safe execution

    The automation specialist connects customer intent to business systems. This person builds and maintains the workflows that let an AI agent retrieve account-specific information, update a record, initiate an approved process, or complete another backend task.

    This is the role that moves AI support from answering to resolving. It also introduces a different class of risk. A weak answer can be corrected in conversation; a wrongly executed refund, cancellation, permission change, or account update can create financial loss, access problems, or corrupted state. Begin with reversible, bounded actions. Enforce identity, authorization, business rules, and transaction limits outside the language model, and preserve a human path when the system cannot establish that an action is safe.

    The first useful deliverable is an action catalog. For each action, document the eligible intent, required inputs, source system, authorization rule, success response, known failure states, recovery path, and human fallback. Do not enable an action merely because the model can call it.

    Support engineering, systems administration, solutions engineering, and tooling operations can all supply the necessary background. The role must be able to work with product and engineering without waiting for those teams to own every support-specific workflow.

    Organize around human support, AI support, and optimization

    The four roles solve ownership at the working level. You still need an organizational model that prevents AI support from becoming an isolated automation project. A practical structure uses three connected pillars: Human Support, AI Support, and Support Operations and Optimization.

    PillarPrimary responsibilityQuestions it must answer
    Human SupportResolve complex, sensitive, ambiguous, and exception-driven customer needsWhat requires judgment? What should be escalated? What are people learning that the system does not yet know?
    AI SupportOwn automated knowledge, behavior, actions, and continuous performance improvementWhere does the AI succeed or fail? What change will improve the outcome? Who can safely approve that change?
    Support Operations and OptimizationProvide tooling, analytics, enablement, QA, workflow design, and capacity planningCan performance be measured? Can failures be routed to an owner? How should human capacity change as automation coverage changes?

    The reporting lines can vary. The interfaces cannot. Before debating where each role sits, write down how work crosses the boundaries.

    • Human Support to AI Support: Frontline agents provide structured evidence about missing knowledge, failed automation, confusing language, and escalation gaps. A collection of anecdotes is not enough; feedback needs an intent and a failure category.
    • AI Support to Human Support: A handoff carries the customer’s goal, relevant context, questions already asked, actions attempted, confirmed results, and remaining uncertainty. The customer should not have to reconstruct the conversation.
    • Operations to both: Operations supplies the measurement, workflow, tools, and change process needed to turn observed failures into verified improvements.
    • Product and engineering partnership: Support owns the customer problem and operating priority. Product and engineering own changes that affect the core product, shared platform, security boundary, or technical architecture.

    Make decision rights explicit as well. Name who can publish source knowledge, change conversational behavior, enable or disable an action, alter handoff rules, accept residual risk, and declare a correction ready. Without these boundaries, teams either move recklessly or wait for broad consensus on routine changes.

    Transition the current team without hiring ahead of the work

    Most organizations do not need a separate AI department on the first day. They need visible ownership. Distributing the responsibilities across existing people lets you prove where the workload and leverage actually sit before converting responsibilities into job titles.

    1. Map recent failures to ownership layers. Classify each meaningful problem as knowledge, conversation, action, or operations. If the team cannot classify it, that ambiguity is itself an operations problem.
    2. Put a person’s name beside every layer. Avoid team names such as Support Ops or Product. A team cannot make a decision; an accountable owner can.
    3. Give each owner an artifact and decision boundary. Use a performance register, knowledge inventory, interaction specification, and action catalog so that the role produces something inspectable.
    4. Run the work on a fixed operating cadence. Review outcomes, inspect representative conversations, assign root causes, prioritize changes, and check whether previous corrections held.
    5. Formalize the role when borrowed capacity stops working. A dedicated hire is justified when the responsibility is continuous, affects important outcomes, and repeatedly loses priority to the owner’s original job.

    The existing support functions should evolve at the same time:

    • Frontline agents spend less time repeating known answers and more time resolving exceptions, preserving trust in difficult moments, and supplying structured feedback about system weaknesses.
    • Enablement teaches agents how to receive AI handoffs, identify failure layers, use AI-generated context critically, and submit feedback that another owner can act on.
    • Quality assurance expands beyond grading agent conversations. It evaluates the end-to-end customer outcome, including AI behavior, action results, escalation decisions, and continuity after handoff.
    • Workforce management plans for automation coverage and the type of work reaching people, not only gross inbound volume. Lower human volume can still demand substantial capacity when the remaining cases are more complex.
    • Support leadership becomes a player-coach responsibility. The leader must understand performance data and system behavior well enough to guide priorities while helping people move into unfamiliar work.

    Do not treat the move as a title-renaming exercise. A knowledge manager without publishing authority, an operations lead without a performance view, or an automation specialist without access to technical partners will reproduce the old model under new labels.

    This transition can also create credible internal career paths. Analytical support-operations talent can grow into AI operations. Content and enablement specialists can move toward knowledge or conversation design. Technically inclined support staff can develop into automation. Frontline experts with strong policy judgment can contribute to knowledge governance, QA, and escalation design. The best candidate is often the person who already understands where customer intent and company systems fail to meet.

    Run AI support as a product, not a side project

    An AI support system changes whenever its knowledge, instructions, workflows, integrations, policies, or underlying product changes. It therefore needs a product-like operating loop: observe an outcome, diagnose the responsible layer, change the right artifact, validate the result, and watch for regression.

    The scorecard should distinguish customer outcomes from automation activity. An impressive volume metric can hide poor resolution, unnecessary handoffs, or actions that appear successful but do not complete in the business system.

    • Resolution quality: Did the customer achieve the intended outcome, rather than merely receive a response?
    • Handoff quality: Was escalation appropriate, correctly routed, and supplied with enough context for a person to continue?
    • Action reliability: Did the requested action complete, produce the expected state, and recover safely when it failed?
    • Knowledge health: Which failures came from missing, stale, conflicting, or poorly scoped information?
    • Customer signals: Do repeat contacts, corrections, abandonment, or explicit dissatisfaction indicate that an apparently completed interaction did not work?
    • Coverage: Which customer intents are eligible for automation, and which remain deliberately human-owned?
    • Human workload: What volume, complexity, and urgency reach agents after automation and handoffs?

    Segment these measures by customer intent. A single aggregate can conceal a reliable password-reset flow beside a weak billing or cancellation flow. Intent-level views also make ownership clearer: you can connect a measurable outcome to the knowledge, conversation specification, action workflow, and escalation rule behind it.

    During an operating review, resist the urge to solve every failure by changing the prompt. First classify the root cause. Correct the source material when the knowledge is wrong. Change the interaction specification when the behavior is wrong. Repair the workflow when an action is wrong. Improve instrumentation or accountability when the organization cannot tell what happened.

    The leader’s job is to keep that loop moving. AI support needs someone who can move between customer experience, operational data, content, and technical constraints. Pure people management is insufficient, but so is pure systems administration. The effective leader coaches the people while actively shaping the system they operate.

    Key takeaways

    • Organize AI support around four accountable layers: operations, knowledge, conversation, and action.
    • Assign the responsibilities before creating dedicated positions; hire when continuous ownership can no longer fit beside an existing role.
    • Connect Human Support, AI Support, and Support Operations through explicit handoffs, feedback contracts, and decision rights.
    • Evolve enablement, QA, workforce management, and leadership around system outcomes rather than ticket throughput.
    • Measure resolution, action reliability, handoff quality, and knowledge health by customer intent, then fix the layer that actually failed.

    Your first move should be small but explicit. Pull recent AI failures, classify each one into the four ownership layers, and put a person’s name beside every layer. Then publish what each owner may change and how the team will verify that a correction worked.

    Do that before requesting a new organization chart. Once the work is visible, you will know which responsibilities can remain distributed and which have become real jobs. More importantly, your customers will no longer depend on an AI system that everybody observes but nobody owns.

    References

  • AI-First Customer Support for Sustainable Ecommerce Growth

    AI-First Customer Support for Sustainable Ecommerce Growth

    Your ecommerce support queue is growing, but cutting ticket volume is not the real decision in front of you. The harder question is which customer outcomes you can let AI own – from order questions to address changes and refunds – without creating a faster path to a wrong answer or action.

    AI-first support earns its place when it completes customer work safely, gives human agents the full context when it cannot, and produces evidence you can use to improve the buying and ownership experience. Growth does not mean forcing a sale into every conversation. It means removing avoidable friction before purchase, resolving post-purchase problems well, and turning repeated support demand into better product and operational decisions.

    Define the unit of automation as a resolved customer job

    A message is not a resolution. An answer is not always a resolution either. If a customer asks to cancel an order, sending the cancellation policy may be factually correct while leaving the actual job unfinished.

    For an AI agent to resolve that request, it must verify the customer and order, check whether cancellation is allowed, execute the permitted action, confirm the exact outcome, and recognize when an exception requires a person. This distinction matters because a deflected conversation can still represent an unresolved customer and a second contact waiting to happen.

    Start by separating support demand into four kinds of work:

    • Informational work: order status, delivery information, return-policy questions, and other requests that can be completed with a grounded answer.
    • Bounded transactional work: changing an eligible shipping address, cancelling an order, issuing an allowed refund, or performing another action with clear rules and permissions.
    • Advisory work: helping a shopper find a suitable product using current catalog data and the constraints the shopper has provided.
    • Judgment-heavy work: policy exceptions, ambiguous intent, conflicting account data, unusual financial consequences, or emotionally sensitive cases where discretion matters.

    Use a workflow map like this before choosing what to automate:

    Customer jobAI needsEvidence of completionWhen AI must stop
    Get current order informationVerified identity, correct storefront, and current order dataThe requested state is returned from the commerce systemIdentity, store, or order data is missing or inconsistent
    Change a shipping addressAn eligible order, editable fields, an authorized tool, and customer confirmationThe commerce platform accepts the new value and returns the updated orderThe order has progressed too far, the address is ambiguous, or the tool fails
    Cancel or refund an orderPolicy rules, order state, transaction permissions, and explicit confirmationThe platform confirms the exact cancellation or refund that occurredThe request is an exception, the amount is unclear, or execution is incomplete
    Choose a productCurrent catalog data and relevant shopper constraintsThe shopper receives grounded options or a clean route to human adviceRequired constraints are unknown or the catalog cannot support the recommendation

    For example, a Shopify support integration can distinguish between retrieving order information and executing actions such as address edits, cancellations, refunds, and duplicate-order workflows. That separation is the architectural principle to preserve: knowing something about an order is not the same as having permission to change it.

    Prioritize each workflow using three factors: how much customer demand it represents, how ready the required data and tools are, and how costly a wrong outcome would be. High frequency alone is a poor selection rule. A common request with unreliable data will produce common failures, while a lower-volume workflow with clear rules may be the better place to prove the operating model.

    Build shared context, bounded actions, and deliberate handoffs

    Treating AI as infrastructure and assigning clear ownership of its performance changes the design question. You are no longer adding a writing assistant to an inbox. You are creating a customer-facing system that reads business state, applies policy, calls tools, and hands work to people.

    The minimum useful context for ecommerce support usually includes verified customer identity, storefront, order and customer records, applicable policies, product or catalog information, conversation history, and the current state of any attempted workflow. Multi-store merchants need the store identifier to travel with the conversation. A valid order number in the wrong storefront is still the wrong context.

    Data architecture deserves the same attention as the model. Capabilities such as multi-store handling, synchronized custom fields, updated data mappings, and EU workspace support illustrate the practical requirements. If the AI cannot determine which record is authoritative, it should expose the conflict and stop. It should never manufacture the missing state.

    Give every action an explicit contract

    A prompt is not an adequate control for a transactional workflow. Every tool the AI can call should have an action contract that defines:

    • Preconditions: what must be true before the action is available.
    • Required inputs: which values must come from verified commerce data and which may come from the customer.
    • Permissions: which customers, agents, stores, order states, and transaction types are eligible.
    • Confirmation: the exact order, field, amount, or consequence the customer must approve.
    • Execution response: a structured success or failure state returned by the commerce platform, not a guess based on generated text.
    • Duplicate-submission protection: how the system prevents the same action from being executed twice.
    • Failure behavior: whether to retry, stop, reverse a reversible step, or hand the case to a person.
    • Audit data: what action was requested, which policy was applied, what the tool returned, and what the customer was told.

    Separate permissions by consequence. Reading authenticated order status is different from drafting a proposed change. Drafting is different from executing a reversible update. A cancellation or refund carries financial and customer-trust consequences, so it needs stricter eligibility checks, explicit confirmation, and a reliable human path for exceptions. Customer confirmation does not compensate for an ineligible order or an unreliable tool.

    The integration method does not remove these obligations. Whether a tool is exposed through a native connector, an internal API, or Model Context Protocol, the AI still needs a constrained schema, narrow permissions, deterministic validation, and an unambiguous result.

    Make escalation a designed path, not a failure bucket

    AI-first does not mean AI-only. Humans should enter when judgment adds value or when a control condition is triggered. Define those conditions before launch rather than expecting the model to improvise them.

    Escalate when identity cannot be verified, records conflict, a policy exception is requested, a consequential action falls outside permission, a tool returns an incomplete result, the customer disputes an executed action, or the customer asks for a person. A model confidence score is not enough unless you have calibrated it against the actual intents and failure costs in your environment.

    The human receiving the conversation should get a compact handoff package containing:

    • The customer’s current request and the reason for escalation.
    • The verified customer, storefront, and order identifiers.
    • A short summary of facts already established.
    • Every action attempted and the exact tool result.
    • The unresolved decision or exception.
    • Anything already promised to the customer.

    The customer should not have to reconstruct the case. When the AI has enough context to recognize that it cannot finish, passing that context forward is part of the resolution experience.

    Measure verified outcomes, system reliability, and growth impact

    Deflection is an activity measure. It tells you a human did not enter the conversation, but it does not prove the customer received the right answer, the requested action succeeded, or the issue stayed resolved. An AI-first operating model should instead emphasize resolution, impact, and system reliability.

    Define a successful automated resolution before you build a dashboard. A practical definition is: the AI correctly understood an eligible request, delivered the correct answer or completed the authorized action, communicated the outcome accurately, and did not create an avoidable repeat contact within a fixed follow-up window. Choose the window for your business and apply it consistently.

    Report coverage and success separately. A strong success rate on a very narrow set of conversations can look impressive while leaving most customer demand untouched. A broad coverage rate can hide weak execution. At minimum, track these metric layers:

    • Eligibility and coverage: the share of total conversations that match a workflow AI is allowed to handle, followed by the share it actually attempts.
    • Resolution quality: verified correctness by intent, policy adherence, repeat contact, customer dispute, and the rate of unnecessary escalation.
    • Action reliability: successful tool execution, rejected actions, duplicate attempts, incomplete results, and wrong or unauthorized changes.
    • Handoff quality: whether the right cases escalate, whether the context package is complete, and whether customers must repeat information.
    • Customer experience: time to the completed outcome and satisfaction segmented by intent and resolution path.
    • Business impact: cost per verified resolution, pre-purchase assisted conversion where attribution is credible, and downstream retention or repeat-purchase signals.

    Do not present an association as growth causation. Customers who contact support may already differ from those who do not. Use controlled experiments where they are practical, compare like-for-like intent cohorts, and treat retention as a downstream signal unless the measurement design supports a stronger claim.

    Ownership matters as much as measurement. Assign someone to own AI support as a product surface, someone to govern knowledge and policy, someone to own commerce integrations and permissions, and someone to review quality and customer harm. These are responsibilities, not mandatory job titles. A smaller organization may place several with one person, but none should be left implicit.

    During a live rollout, I would review every failed or disputed write action and sample successful actions across each active intent every operating day. Once the important failure modes are understood and performance is stable, intent-level review can move to a weekly cadence. Scope changes should still happen through an explicit release decision, not because the queue happens to be busy.

    Roll out one dependable resolution lane at a time

    The safest path to meaningful automation is not a site-wide chatbot launch. It is a sequence of narrow resolution lanes, each with grounded data, an evaluation set, clear permissions, a human fallback, and a rollback path.

    1. Establish the baseline. Group current conversations by customer intent and record volume, time to outcome, repeat contact, escalation, and the systems or policies each intent depends on.
    2. Select a narrow first lane. Favor a request with clear rules, reliable data, and low action reversibility. Authenticated order information is often a better proving ground than refunds, but your own data readiness should decide.
    3. Create an evaluation set from real, appropriately handled conversations. Include ordinary cases as well as missing orders, stale data, multi-store ambiguity, policy exceptions, tool errors, changed customer intent, and explicit requests for a person.
    4. Write expected outcomes before testing. For every case, specify whether AI should answer, act, ask for missing information, or escalate. Classify unauthorized disclosure, wrong transactional action, and missed consequential escalation as critical failures that an overall average cannot hide.
    5. Observe before granting broad action permissions. If your platform supports a draft or shadow mode, compare proposed behavior with the expected outcomes. Then launch to a limited storefront, channel, workflow, or customer cohort with active monitoring.
    6. Add one write action at a time. Confirm the action contract, permissions, confirmation language, duplicate protection, audit trail, human fallback, and rollback mechanism before expanding eligibility.
    7. Protect peak periods. Do not introduce a consequential workflow immediately before your highest-demand period unless it has already passed realistic evaluation and the operating team can disable it quickly. Keep staffing and fallback capacity based on verified workload movement, not projected deflection.

    This expansion model creates a compounding loop. Every failed or repeated conversation should produce a specific improvement task: repair missing knowledge, correct a data mapping, clarify a policy, tighten an action permission, improve the handoff, or send a recurring upstream problem to product, merchandising, fulfillment, or operations. The value is not only that AI absorbs work. It is that support demand becomes structured evidence about where ecommerce growth is leaking.

    Continue expanding only when a lane remains dependable under real conditions. Tight merchant feedback loops and peak-season planning are especially important as the agent moves from answering questions to taking actions. Pause when unresolved contacts or ambiguous cases rise. Roll back immediately when the system performs an unauthorized or incorrect consequential action.

    Key takeaways

    • Optimize for completed customer jobs, not avoided human conversations.
    • Separate information retrieval from transactional authority, and give every action a testable contract.
    • Make verified identity, storefront, order state, policy, and tool state part of the shared context.
    • Design human escalation before launch so judgment-heavy cases arrive with their context intact.
    • Report eligibility, coverage, resolution quality, action harm, and business impact separately.
    • Expand through evaluated resolution lanes with explicit release, monitoring, and rollback decisions.

    Your next move is concrete: choose one customer job, write down its required data, allowed actions, stop conditions, success evidence, and human fallback. If you cannot make those five elements explicit, the workflow is not ready for autonomous resolution. If you can, you have the first building block of an AI-first support system that can grow without asking customers to absorb the risk.

    References

  • Taming 1,000+ Vendor Emails: How Xelix’s AI Helpdesk Delivers Fast, Confident Answers

    Taming 1,000+ Vendor Emails: How Xelix’s AI Helpdesk Delivers Fast, Confident Answers

    Chaos in vendor communications is a problem I see across finance operations: sprawling accounts payable inboxes, slow response times, and missed context. That’s why this build caught my attention—not just because it’s GenAI, but because it’s a disciplined product strategy that converts email overload into measurable outcomes.

    Accounts payable inboxes can see 1,000+ vendor emails a day. Xelix’s new Helpdesk turns that chaos into structured tickets, enriched with ERP data, and pre-drafted replies—complete with confidence scores.

    I dug into the end-to-end approach with the team—Claire Smid — AI Engineer, Xelix; Emilija Gransaull — Back-End Tech Lead, Xelix; Talal A. — Product Manager, Xelix—focusing on how they scoped the problem, iterated fast, and de-risked AI in production.

    Their product thesis is refreshingly pragmatic. They prototyped with “daily slices” (Carpaccio-style) and built a retrieval-first pipeline that matches vendors, links invoices, and drafts accurate responses—before a human ever clicks “send.” That framing matters: enrichment and matching take center stage, with the model amplifying precision instead of improvising.

    We unpacked the tricky bits that make or break an AI helpdesk at scale: vendor identity matching, Outlook threading, UX pivots from “inbox clone” to ticket-first views, and the metrics that prove real impact (handling time, stickiness, auto-closed spam). The pipeline architecture and email processing choices were grounded in operational realities, not just AI aspirations.

    Several takeaways are worth pinning to any AI product roadmap. “Start narrow to win: pick high-volume, high-cost requests (invoice status & reminders).” “Enrichment > magic: accurate replies come from great retrieval/matching, not just a bigger LLM.” “Design for adoption: familiar inbox view helps onboarding, but a ticket-first UI unlocks AI features.” These are the kinds of decisions that drive adoption, trust, and ROI.

    Data enrichment challenges dominated early learning curves: stitching ERP context into tickets, handling vendor identification at scale, managing email thread continuity, and calibrating response generation for accuracy. On the generation side, the team emphasized precision over verbosity—clean responses that reflect system-of-record truth—then instrumented the experience to “Evaluate System Performance” with production-grade telemetry.

    Trust was treated as a product feature. “Measure outcomes, not vibes: track ‘messages sent from Helpdesk’, % auto-resolved.” And critically, “Confidence builds trust: show match quality and response confidence so humans know when to edit.” By surfacing match quality and confidence scores, they shortened coaching loops and made human-in-the-loop supervision feel natural, not burdensome.

    What’s next is equally compelling: “targeted generation, multiple specialized responders, and more agentic routing.” That direction aligns with agentic AI patterns I recommend for operations-heavy workflows—route first, retrieve deeply, then generate with intent. It’s a scalable path from assistive AI to autonomous resolution while maintaining governance and auditability.

    If you want a quick map of the journey, the conversation flowed from 0:00 Meet the Team: Claire, Emilija, and Talal, 00:36 Introduction to Xelix and Its Products, 01:08 Understanding Accounts Payable Teams, 01:37 Help Desk Product Overview, 03:11 Challenges Faced by Accounts Payable Teams, 04:03 AI Integration in Help Desk, 05:47 Automating Reconciliation Requests, 07:45 Development Methodology: Carpaccio, 09:11 Prototyping and Beta Testing, 12:00 Manual Tagging and Data Collection, 16:39 Focusing on High-Impact Use Cases, 18:55 User Experience and Interface Design, 24:56 Pipeline Architecture and Email Processing, 28:21 Data Enrichment Challenges, 29:04 Handling Vendor Identification, 33:33 Email Thread Management, 36:15 Generating Accurate Responses, 40:48 Evaluating System Performance, 49:20 Future Developments and Goals.

    My takeaway for product leaders: when the domain is high-volume and rules-heavy (like AP), retrieval-first beats model-first. Start with the narrowest, costliest intents; prove lift with “messages sent from Helpdesk” and “% auto-resolved”; then graduate UX from familiar to AI-native (ticket-first) once trust is earned. That’s how you turn vendor chaos into answers—reliably, scalably, and fast.


    Inspired by this post on Product Talk.


    Book a consult png image
  • How to Evaluate AI Voice Support in Real-World Conditions

    How to Evaluate AI Voice Support in Real-World Conditions

    You have a shortlist of AI voice support products, a polished recording, and a decision that could affect thousands of customer conversations. The hard question is not whether an agent can sound convincing during one ideal call. It is whether the system stays useful when a caller interrupts, corrects themselves, asks an ambiguous question, waits on a backend system, or needs a human.

    You can answer that question before a broad rollout. The method is to test complete support outcomes, introduce controlled complications, score failures separately from conversational polish, and use the result to define a limited production pilot.

    Evaluate the support outcome, not the performance

    A natural voice can create an impression of competence before the agent has done anything useful. Pleasant pacing, expressive speech, and a quick opening matter, but they cannot compensate for retrieving the wrong account, misunderstanding the request, or claiming that an action succeeded when it did not.

    Treat the unit of evaluation as a completed support job. Depending on the intent, that job may require the agent to identify the caller, understand the request, retrieve the right information, explain the answer, perform an authorized action, confirm the resulting state, and send a follow-up or transfer the conversation. If you score only the spoken answer, you leave most of the product untested.

    One live Fin Voice call illustrated this end-to-end standard in about 90 seconds: the agent verified identity, retrieved account information, managed an interruption, presented options, completed a workflow, and sent a follow-up email. That sequence is a useful model for constructing a test. It is not, by itself, proof of reliability across other calls.

    Before anyone places a test call, write an outcome contract for each scenario:

    • Caller goal: What is the person trying to accomplish?
    • Starting state: What customer, account, order, subscription, or case data exists before the call?
    • Available evidence: Which knowledge, policies, and records may the agent use?
    • Permitted actions: What may the agent change, create, send, cancel, or escalate?
    • Required clarification: Which missing or conflicting facts must be resolved before an answer or action?
    • Completion evidence: What observable state proves that the request was resolved?
    • Unacceptable outcome: What error would make the call a failure even if the conversation sounded good?

    This contract prevents a common scoring mistake: confusing non-transfer with resolution. A call can remain inside the AI channel and still leave the customer with a wrong answer, an incomplete action, or no idea what happens next. Conversely, an intentional transfer can be the correct resolution when the agent reaches a policy, permission, or confidence boundary.

    Build scenarios around the ways real calls become difficult

    Start with support intents your operation actually receives. Prioritize intents that are frequent, expensive to handle, important to customer trust, or dependent on multiple systems. Do not begin with trivia questions that merely demonstrate broad language-model knowledge. You are evaluating support execution.

    For every core intent, create a straightforward case and several controlled variants. Keep the customer objective constant while changing one condition at a time. That makes a failure diagnosable instead of merely disappointing.

    A practical scenario matrix

    • Clean path: The caller gives the relevant facts in a clear order. This establishes whether the basic workflow works at all.
    • Missing information: Omit a detail the agent needs. Check whether it asks a focused question instead of guessing or restarting the intake.
    • Ambiguous intent: Use wording that could map to two support issues. The agent should disambiguate before retrieving data or taking action.
    • Mid-call correction: Let the caller change an account detail, date, product, or preferred option. Check whether the corrected fact replaces the old one throughout the workflow.
    • Interruption: Speak while the agent is answering. Observe whether it stops cleanly, understands the new input, and continues from the right point.
    • Backend delay: Introduce a slow retrieval or action. Evaluate how the agent manages the wait and whether it distinguishes a pending operation from a completed one.
    • Backend failure: Make a required system unavailable or return an error. The agent should not fabricate a result or promise completion it cannot verify.
    • Policy boundary: Ask for something the agent is not allowed to do. Test the explanation, alternatives, and escalation path.
    • Human request: Ask directly for a person. Verify that the agent follows the configured policy without turning the handoff into an argument.
    • Listening conditions: If your deployment must support different languages, accents, devices, or noisy environments, test each condition explicitly rather than treating one clear studio call as representative.

    Give testers the goal, account state, and one complication. Do not script every sentence. A fully written dialogue tests whether the agent can follow the dialogue you anticipated; a goal-based scenario tests whether it can manage the conversation the caller actually creates.

    Keep a few variants undisclosed until the live session. This is not a trick. It prevents the evaluation from becoming a memorized path while still keeping every test fair and reproducible. Record the exact variant afterward so another evaluator can run it again.

    Run the call through the systems you expect to deploy

    An unedited live call is more informative than a produced recording, but live alone is not enough. A live test can still use ideal data, a simplified integration, a practiced caller, and a workflow that avoids the hard parts of your environment.

    Ask to run the scenario through a path that resembles the intended deployment:

    1. Place a normal phone call through the proposed telephony route. If production will use call forwarding, test the forwarding path rather than a direct internal endpoint.
    2. Use a safe test account containing representative records, permissions, and history.
    3. Require the agent to retrieve data from the backend system that will be authoritative in production.
    4. Introduce the chosen interruption, correction, ambiguity, delay, or error during the live conversation.
    5. Require a real test action where it is safe to do so, not a verbal description of what the agent would have done.
    6. Inspect the backend state after the call. Confirm that the correct record changed once, with the expected values.
    7. Verify every promised follow-up, case creation, notification, or handoff outside the voice channel.
    8. Retain the recording, transcript, timestamps, tool activity, and final system state for scoring.

    This is especially important when an agent can take consequential actions. A fluent confirmation is not evidence that the action happened. The system of record is the evidence.

    Repeat important scenarios with different wording and a different caller. One successful run demonstrates that the capability can work. Repeated variants reveal whether the capability depends on a narrow phrase, a rehearsed cadence, or an unusually forgiving path.

    Key takeaways

    • Score complete resolution, including backend state and follow-up, rather than voice quality alone.
    • Change one condition at a time so you can identify why a call failed.
    • Test interruptions, corrections, ambiguity, system delays, system errors, and escalation.
    • Measure different kinds of waiting separately; a lookup pause and a turn-detection problem are not the same defect.
    • Treat a successful demo as evidence for a pilot, not permission for an unrestricted rollout.

    Score conversation, reasoning, and operational closure separately

    A single overall rating hides the information you need to make a product decision. The call may sound awkward but reach the correct outcome, or sound excellent while making a dangerous mistake. Separate the evaluation into three layers.

    LayerWhat to inspectEvidence of a passTypical failure
    Conversation mechanicsTurn detection, interruption handling, pacing, response length, and intelligibilityThe caller can speak naturally, correct the agent, and follow the response without fighting for the floorThe agent talks over the caller, leaves confusing silence, or delivers answers too long to retain by ear
    Decision qualityIntent recognition, clarification, use of account context, policy application, and answer accuracyThe agent asks only for missing information, uses the correct evidence, and avoids unsupported conclusionsThe agent guesses, asks redundant questions, ignores a correction, or applies the wrong policy
    Operational closureIdentity checks, tool calls, state changes, confirmation, follow-up, and escalationThe verified backend state matches the caller’s request and the agent’s final explanationThe agent claims success without a completed action, changes the wrong record, duplicates work, or drops context during handoff

    Use a simple 0-2 score for each criterion: 0 for failed or unsupported, 1 for completed with material caller effort or recovery, and 2 for correct and usable. The scale is deliberately small. Evaluators can usually distinguish failure, friction, and success more consistently than they can defend the difference between seven and eight on a ten-point scale.

    Do not average away critical errors. A wrong account action, failed identity control, fabricated completion, or forbidden disclosure should remain visible as a release blocker even if many low-risk calls receive high scores. Record both the criterion scores and the count of critical failures.

    Break latency into moments the caller can feel

    Latency is not one number. Capture at least three moments: the time the agent takes to recognize that the caller has finished, the time it spends reasoning or waiting for a system, and the time needed to begin and complete the spoken response.

    • End-of-turn delay: A long delay after every caller turn makes the exchange feel unresponsive and can encourage both sides to start speaking at once.
    • Reasoning or retrieval delay: A pause can be appropriate when the agent is checking account data or invoking a backend workflow. Brief pauses were audible during live subscription and backend checks, which is more informative than editing those waits out.
    • Response delivery: A fast start does not help if the answer becomes a long monologue. Voice responses need structure and pacing that work for listening, not merely text that sounds acceptable when read.

    Ask what is happening during a pause. If the system is doing useful work, the next statement should reflect that work and the action log should verify it. If the pause is long enough to make a caller wonder whether the call has dropped, the experience needs an appropriate progress cue. If the agent answers instantly but guesses, speed is concealing a quality problem.

    Review individual timings as well as an average. A generally responsive agent with occasional severe stalls creates a different operational problem from one that is consistently a little slow. Your test recordings and timestamps should make both patterns visible without inventing a universal pass threshold that ignores the complexity of the workflow.

    Make recovery and escalation part of the product test

    The strongest voice experiences are not the ones that never encounter confusion. They are the ones that recover without making the caller restart. Recovery is therefore a capability to test, not an embarrassing exception to hide.

    Interrupt the agent in the middle of an answer. Correct a fact it has already used. Add a second request after the first appears resolved. Say that an explanation was unclear. Ask for a human. These moves reveal whether the agent maintains conversational state or merely produces plausible turns one at a time.

    During recovery, look for specific behavior:

    • It stops speaking promptly when the caller takes the turn.
    • It identifies what changed instead of repeating the whole interaction.
    • It replaces corrected information rather than carrying both versions forward.
    • It asks a narrow clarification when the next action is uncertain.
    • It does not claim to understand when the transcript or subsequent action shows otherwise.
    • It preserves verified context and the reason for contact when a human takes over.
    • It tells the caller what will happen next instead of ending on an internal routing label.

    Tone belongs in this test, but not as a beauty contest between synthetic voices. Evaluate whether pacing, brevity, acknowledgement, and word choice suit the moment. A caller correcting a billing detail needs a clear acknowledgement and an accurate update, not theatrical empathy. A caller who sounds uncertain may need a shorter explanation and a confirming question. Tone is the behavior of the conversation, not just the timbre selected in a settings menu.

    Escalation should also count as a valid outcome when it is timely and informed. Define which conditions require a handoff, which allow one, and what context must travel with it. Then test the handoff from the caller’s side. If the customer reaches a person but has to repeat identity, intent, and every attempted step, the routing technically worked while the support experience failed.

    Turn the evaluation into a controlled pilot decision

    A strong live evaluation earns the right to run a pilot. It does not justify sending every eligible call to the agent. Production introduces variation in callers, data quality, traffic, integrations, and issue combinations that a demonstration cannot reproduce fully.

    I would require five gates before approving even a limited external pilot:

    1. Capability gate: Every must-have intent has completed its end-to-end workflow, including at least one controlled complication.
    2. Critical-risk gate: No unresolved failure can expose the wrong account, bypass a required check, perform an unauthorized action, or report a false completion.
    3. Conversation gate: The agent can handle interruptions, corrections, clarification, and explicit human requests without trapping the caller in a loop.
    4. Operations gate: Your team can configure terminology, guidance, escalation behavior, greetings, voice, and deployment controls for the intended support environment.
    5. Learning gate: Owners can inspect recordings, transcripts, tool activity, outcomes, and failures, then change the knowledge, workflow, policy, or conversation design responsible.

    Start the pilot with a reversible slice of traffic and a clear human fallback. Select intents whose correct outcome can be verified in your systems. Define who reviews failed and escalated calls, who can pause the rollout, and who owns each class of fix. An answer-quality issue, a telephony issue, and a backend integration issue require different owners even when the caller experiences all three as one bad call.

    Expand only when observed calls meet the outcome contracts you wrote before the demo. If the definition of success keeps changing after failures appear, the evaluation is no longer protecting the decision.

    For your next vendor session, replace “show me your best call” with a scenario pack, a test account, and a request to inspect the final system state. You will learn more from one imperfect call that recovers correctly than from a flawless recording that never had to recover at all.

    References

  • Why We’re Building Our Next AI R&D Hub in Berlin—and Hiring 100 to Power Fin’s Growth

    Why We’re Building Our Next AI R&D Hub in Berlin—and Hiring 100 to Power Fin’s Growth

    I’m excited to share that we’re opening our next R&D hub in Berlin to support significant investment in our AI customer service platform, Intercom, and market-leading AI Agent, Fin. We intend to hire 100 people in Berlin over the year ahead across engineering, AI, data science, product, and design. This move reflects our AI Strategy, our commitment to product management leadership, and our focus on building enduring product-led growth.

    We believe that in a short number of years, the vast majority of customer service will be done by AI. Fin is already the world’s best Customer Service Agent. At Pioneer, our recent summit for AI customer service leaders in NYC, we talked about how Fin will become a true end-to-end Customer Agent, extending far beyond service. We showcased how companies like WHOOP, Anthropic, and Lightspeed are already pushing Fin in ways that help them grow their business.

    This market opportunity is massive and expanding at unprecedented pace. Our ambition is to earn our place as one of the most successful AI businesses during this wave of AI disruption, and we want more brilliant people on our team to pursue this as aggressively as possible. If you’re motivated by Generative AI, LLMs, and building real products that scale, you’ll find both challenge and impact here.

    We are already on track to be one of the fastest growing private software companies. Fin is the primary contributor to this, and is months away from passing $100m in ARR. So far, more than 7000 businesses have transformed their customer service with Fin, including German companies like electricity provider Ostrom, smart home technology provider tado°, and grocery delivery company Flink, along with global leaders like Vanta, Clay, Lovable, and Miro.

    Why Berlin? We’re drawn to the city’s rare blend of deep technical talent and rich creative culture—within a vibrant, globally connected ecosystem close to our R&D hubs in Dublin and London. It’s a place where top-tier engineers and designers thrive, and where ambitious builders from around the world want to relocate and create category-defining products.

    Orange gradient area chart with a white line and circular markers showing steady growth from about 26% to nearly 70% across monthly labels from May 2023 to Sep 2025, on a light grid with percentage ticks.
    Momentum is building: this month-by-month chart shows a consistent rise from the mid-20s to nearly 70% between May 2023 and Sep 2025—signaling strong progress as we expand engineering, AI, and automation at our new Berlin R&D hub.

    We needed a new location that would sustain the high ambition and standards held by our world-class AI teams in Dublin and London. Berlin has emerged as one of Europe’s hottest centers for AI talent, with a high density of AI-focused startups, applied research labs, and practitioners who bring exceptional literacy, optimism, and ambition. It’s the right accelerator for our AI hiring and a place to bring in brilliant minds to shape the future of our product and business.

    While Intercom’s reach is global with our headquarters in San Francisco, our R&D leadership remains anchored in Dublin, where half of the executive team sits—making Berlin both geographically and strategically an ideal next location for our growth.

    This isn’t our first time expanding our footprint; we previously bet on London and are delighted with how that’s been working. When we shared our Berlin news internally, the energy was palpable, with many teammates volunteering to help spin up the hub successfully—including colleagues who helped make London a big success, like Danny. That level of ownership and momentum is exactly what we aim to cultivate in Berlin.

    We’re looking for people who thrive in a high-intensity, high-ambition, high-standards environment and want to help build one of the world’s best AI companies. For builders like that, the opportunity for impact, growth, and career progression is extraordinary. As with London and Dublin before it, the early Berlin cohort will have a disproportionate influence on team norms, culture, and long-term outcomes. We are in the middle of a huge disruptive wave with AI, and Fin is one of the leading examples of commercially successful AI applications. Joining Intercom is an opportunity to be part of this disruptive wave, and help us build out our vision for Fin becoming the world’s best Customer Agent.

    Four panelists seated on a dark stage during an AI engineering discussion, with on-screen titles above them, at an event announcing a new R&D hub in Berlin.
    On a minimalist stage, four speakers share insights on AI research, automation, and engineering as part of a panel tied to Berlin expansion and the launch of a new European R&D hub.

    There are plenty of AI companies to join, but our technology and culture set us apart. Any AI product is only as good as the AI layer powering it. Ours is industry-leading, built by a highly talented, ambitious, and technical team of over 40 machine learning scientists, engineers, and designers in Europe who continuously optimize Fin’s performance through cutting-edge research, experimentation, and innovation. Fin’s average resolution rate increases 1% every month. That kind of steady, compounding improvement is exactly what great customer support AI strategy looks like in practice.

    We also build in public and share our progress and learnings with the AI community at large. Recently, our Chief AI Officer Fergal Reid and SVP of Engineering Jordan Neill joined leaders from Cognition, Harvey, and Perplexity in San Francisco to share real lessons, challenges, and breakthroughs from building frontier AI products. Our AI team regularly publishes their insights on the AI research blog; from optimizing inference speed and availability, to building our own proprietary models that outperform general purpose models for CX.

    Our AI group and the broader R&D org they operate within work at extraordinary scale and speed. We recognize that moving fast can’t be taken for granted—you must fight for it—and we’re doing just that, embracing the capabilities AI tooling brings us to achieve 2x the throughput. One example of this mindset in practice is us “Betting on the future of frontend at Intercom,” making a technology choice that optimizes for our teams’ ability to build high-quality product, fast.

    Our design and product teams are world-class and forward-thinking; they’re embracing AI to evolve how they work, as shared in our 3-point framework for AI-driven design and recently presented by Emmet Connolly, our SVP of Design, at this year’s Hatch conference in Berlin. As a product leader, I’m grateful to work alongside brilliant product and design thinkers—it gives me confidence that we’re solving the right problems, solving them well, and driving real impact.

    Tech conference collage with a speaker on stage beside four panels: AGI teaser on a tablet, code editor, webcam demo with hand tracking, and a simulation. Banner reads Hatch Conference 2025 Main Stage.
    From live demos to hands-on coding, this snapshot captures the momentum we're bringing to our Berlin R&D hub – AI experiments, hand-tracking prototypes, and simulation tools powering our next wave of engineering.

    We plan to open our Berlin office space in December or January. To get the office started, we’re hiring Senior Product Engineers, Machine Learning Scientists, Product Managers, Senior Product Designers, Engineering Managers, and Data Scientists immediately. If your craft sits at the intersection of LLMs for product managers, agentic AI, and empowered product teams, you’ll be right at home.

    You can learn more about our open roles, company, culture, and locations on our careers site, or feel free to reach out to me, Jordan, Fergal, or Brian directly on LinkedIn if you have any questions.

    Some of our engineering team will also be at LeadDev Berlin on November 3rd—come say hi if you’re attending.

    I’m looking forward to continuing to build Intercom as one of our generation’s best AI companies—and I’m excited for our expansion into Berlin to be a major contribution to that success.


    Inspired by this post on The Intercom Blog.


    Book a consult png image
  • Beyond Digital: How AI Transformation Builds Adaptive, Intelligent Organizations That Win

    Beyond Digital: How AI Transformation Builds Adaptive, Intelligent Organizations That Win

    Digital transformation rewired our systems; AI transformation rewires how we learn, decide, and compete. “AI transformation goes beyond automation to create adaptive, intelligent organizations. Discover why it’s the next imperative and how to measure success.” That statement captures what I experience daily: we’re moving from scripted workflows to living systems that improve with every interaction.

    When I talk about AI transformation, I’m not describing a tool rollout. I’m describing an operating model where data, models, and product strategy converge to create compounding advantage. In practice, that means agentic AI orchestrating tasks, robust data governance and privacy-by-design from day one, and empowered product teams that ship, measure, and iterate at high tempo.

    The imperative is strategic, not merely technical. Markets are compressing cycle times, and customers now expect intelligent experiences by default. Organizations that master AI Strategy and product-led growth will set the pace—using AI for competitive differentiation rather than feature parity.

    This shift changes how I build teams and backlogs. I lean on product trios, forward deployed engineers, and tight product discovery loops to reduce uncertainty early. We design for resilience and learning: human-in-the-loop feedback, clear escalation paths, and telemetry that turns every interaction into a hypothesis test.

    Governance is a first-class feature. AI risk management, data governance, and threat detection and response sit alongside performance metrics in the same dashboard. We codify guardrails—policy, provenance, and permissions—so innovation scales safely and sustainably.

    Measurement is where transformation becomes real. I anchor on outcomes vs output OKRs tied to customer value and revenue impact. At the product layer, I track activation, time-to-value, retention, and adoption by persona. For ML quality, I monitor precision/recall, coverage, hallucination rate, and model drift. In experimentation, A/B testing with a thoughtful minimum detectable effect (MDE) prevents false wins, while Amplitude analytics, Pendo, and Intercom instrumentation expose where guidance or UX writing can unlock activation.

    The fastest wins often start in service and sales. A customer support ai strategy can deflect tickets with high-resolution answers while escalating edge cases to humans with full context. CRM integration with HubSpot and a ChatGPT connector enables reps to generate next-best-actions, summarize calls, and personalize outreach—measurably lifting conversion and lowering cost-to-serve.

    On the build side, LLMs for product managers and gen ai for product prototyping accelerate discovery cycles. I use CustomGPT workflows to validate value propositions quickly, then harden successful flows with engineering. Throughout, product positioning and a crisp value proposition ensure that what we ship is understandable, differentiated, and priced to match ROI—consumption SaaS pricing when usage scales value.

    If you’re getting started, begin with a single, high-frequency journey, instrument it deeply, and publish transparent OKRs. Pair empowered product teams with clear governance, and iterate toward agentic AI experiences. The payoff isn’t a one-time launch; it’s a continuously learning system—and a culture—that compounds advantage release after release.


    Inspired by this post on Pendo – Perspectives.


    Book a consult png image
  • Intuition, White-Glove Support, and Relentless Execution: Lessons from Looker to Omni

    Intuition, White-Glove Support, and Relentless Execution: Lessons from Looker to Omni

    I’m constantly drawn to product stories where intuition, customer obsession, and raw effort compound into durable advantage. This conversation with Colin Zima crystallized that arc—from pioneering high-touch support at scale to balancing gut feel with data to ship what matters. The through-line for me: when you operationalize empathy and pair it with disciplined execution, you create momentum that’s almost impossible to copy. Colin Zima is the co-founder and CEO of Omni, a business intelligence tool that has raised over $26.9m. Prior to starting Omni, Colin was Chief Analytics Officer and VP of Product at Looker, which was acquired by Google for $2.6b. Colin was an early employee at Looker, and stood up its high-touch customer support arm, which turned into a cornerstone competitive advantage for the company. What resonated most with my own practice is how deliberate investment in white-glove customer support can become a product strategy lever—not just a service function. When you’re in a category-creating phase or displacing an entrenched incumbent, those high-touch loops are how you learn the truth fast, reduce onboarding friction, and convert early believers into reference customers. The trick isn’t whether to do it; it’s when, why, and how to sequence it so the economics still make sense as you scale. On scaling high-touch support, I look for three signals before pushing the gas: repeatability in the top 5 user pain patterns, a crisp path to tooling and self-service, and tight product feedback loops that turn today’s premium assistance into tomorrow’s default experience. That’s how white-glove support pays for itself—first as acceleration for adoption, then as inputs that harden the core product. I also emphasize role clarity and career ladders so support becomes a talent engine, not a cul-de-sac, which makes hiring for and hiring from customer support a strategic advantage. Colin’s intuition-based approach to product echoes a belief I hold closely: data is essential for validation and prioritization, but it rarely originates the leap. Intuition frames the bet; data sizes the risk; customers ground the narrative. I’ve seen the merits—speed, conviction, and differentiated UX—and the downsides when intuition goes unchecked—overfitting to edge cases or mistaking novelty for value. The balance is intellectual honesty: writing down the thesis, the counter-thesis, and the disconfirming evidence you’ll accept before you commit resources. I was especially struck by the operational rigor behind hitting goals for 24 quarters in a row. That kind of consistency doesn’t happen by accident; it comes from outcomes over output, sober forecasting, and the cultural discipline to cut or delay work that doesn’t ladder up. I coach teams to make the target visible, tie metrics to customer value, and then prune relentlessly—because the opportunity cost of “almost done” is usually invisible until the quarter slips. The founding story of Omni reminds me that category shifts rarely come from a single breakthrough. They’re the product of dozens of earned insights about where the market is going and what’s still too hard for customers today. I pay close attention to how founders maintain intellectual honesty as the narrative tightens—keeping a clear line between what we know from the field and what we’re assuming, and revisiting that line often. There’s also practical career wisdom here. When choosing which startup to join, I look for founder clarity on the core problem, the early design partners, and the distribution wedge. On founder-market fit, I care less about domain tenure and more about a pattern of shipping, learning, and adjusting fast. And Colin’s unpopular opinion on how to hire good PMs aligns with my experience: bias toward builders who can synthesize customer reality, technology constraints, and go-to-market timing—then communicate clearly and commit. If you’re building in data and analytics, these references are useful context for the ecosystem and buyer expectations: BigQuery: https://cloud.google.com/bigquery, Hotel Tonight: https://www.hoteltonight.com/, Omni: https://omni.co/, Tableau: https://www.tableau.com/. For those who want to go deeper with Colin’s thinking and product journey, you can find him here: Twitter: https://twitter.com/drinkzima?lang=en and LinkedIn: https://www.linkedin.com/in/colinzima/. My takeaway as a product leader: make white-glove customer support a strategic instrument, not a cost center; let intuition set bold direction while data governs scope; and cultivate the operational cadence that makes hitting your goals a habit, not a headline. That combination is how you compound trust with customers and ship products that stand the test of time.
    Book a consult png image
  • Scaling Enterprise AI That Sells: Battle-Tested Playbooks for PMF, Champions, and Agentic AI

    Scaling Enterprise AI That Sells: Battle-Tested Playbooks for PMF, Champions, and Agentic AI

    Enterprise AI is exhilarating and unforgiving. I’ve seen gorgeous demos fall apart under real-world constraints and seemingly modest pilots unlock outsized value at scale. In this reflection, I share the playbooks I rely on to build, scale, and sell generative AI in the enterprise—what actually moves deals, secures product-market fit, and sustains trust with the C-suite and the front line.

    Why is it so difficult to scale AI products for enterprise? Because the bar is higher on every dimension: data governance, security, extensibility, integration depth, reliability, and measurable ROI. An enterprise-grade, full-stack generative AI platform isn’t just a model; it’s the surrounding system—observability, evals, policy, workflow, and human-in-the-loop—that makes outcomes predictable, auditable, and safe. The fastest path to adoption is simple: deliver on-brand, on-policy content and decisions using a customer’s first-party data, and prove that quality holds up under load.

    My north star is dependability over demo magic. The number one challenge is making model output dependable across messy, high-variance enterprise inputs. I build an evaluation harness early, with gold datasets, task-specific metrics, and human adjudication. Every change ships behind guardrails and is measured against cost, latency, and quality SLOs. When governance, change management, and procurement show up (they always do), I treat them like first-class product requirements, not hurdles.

    Champions are the secret to winning complex accounts. I map the org, find operators who feel the pain daily, and quantify that pain in dollars and hours. Then I define success criteria upfront—time-to-value in under 30 days, measurable uplift (e.g., deflection, conversion, cycle time), and a plan for scale. I deploy forward deployed engineers alongside the business to co-design workflows, refine prompts and evaluators, and document before/after outcomes. Champions don’t just approve pilots; they co-author the business case and defend it.

    To win the enterprise, trust architecture matters as much as model architecture. I lead with clear answers on data residency, encryption, SSO, RBAC, DLP, and retention policies; I address whether customer data trains models, default behaviors, and opt-in controls. I offer flexible deployment (VPC or private networking when needed), transparent pricing, and SLAs with real teeth. I also integrate where work already happens—CRM, help desk, knowledge bases—so value shows up in the flow of work.

    Signs of healthy product-market fit are unmistakable: pull from lookalike buyers, multi-threaded expansions, champions who present results internally without me in the room, and usage that moves from experimentation to business-critical. I watch for weekly active usage above pilot thresholds, POCs converting to multi-year deals, and adjacent teams (Support, Marketing, Legal, RevOps) asking to onboard with minimal push. PMF feels less like persuasion and more like coordination.

    Scaling large language models for specific use cases requires ruthless focus. I constrain scope to tightly defined workflows, pair retrieval with structured knowledge, and mix model strategies (base models, fine-tunes, tools, and function calling) based on cost and latency budgets. I codify policy-as-code and deploy guardrails at the orchestration layer, not just the model layer. Continuous evaluation—both automatic and human—is the heartbeat of quality.

    My advice to AI founders in 2024 is pragmatic. Start with outcomes, not demos. Establish outcomes vs output OKRs that tie directly to revenue, cost, risk, or customer experience. Use gen AI for product prototyping to shorten discovery cycles, but graduate quickly to instrumented workflows in production. Align early with InfoSec and Legal; your speed will be gated by trust, not code. And when in doubt, ship smaller, safer increments faster.

    Healthy co-founder relationships look the same across winning companies: clear domains, fast escalation, and a shared appetite for “disagree and commit.” I keep a decision log, time-box debates, and make moments-of-truth visible to the team and board. You’ll know it’s working when you have more energy after hard conversations than before.

    The future of agentic AI is deeply enterprise: multi-agent workflows that plan, act, and verify with human oversight where it matters. The winners will combine reasoning, tool use, retrieval, and policy with audit trails that satisfy compliance while keeping velocity high. Think of it as re-engineering business processes around AI-native steps, not sprinkling AI on top of legacy workflows.

    Culture turns strategy into reality. I anchor my teams on “connect, challenge, and own.” Connect means obsess over the customer problem and internal alignment. Challenge means we red-team our ideas, run experiments, and measure impact. Own means we accept outcomes, not just output, and we iterate until the business moves. This is how a customer support ai strategy becomes a durable moat, not a slide.

    If you’re a product creator or product management leader, the above playbooks are meant to be lifted and adapted. Start where the pain is loudest, quantify the win, and let champions carry the story. The compound interest of disciplined product discovery, strong governance, and relentless evaluation is a generative AI business that sells itself—and scales.


    Book a consult png image
  • DevTools at Scale: Hard-Won Lessons on PMF, AI, and Culture from Apple, AWS, Microsoft

    DevTools at Scale: Hard-Won Lessons on PMF, AI, and Culture from Apple, AWS, Microsoft

    Building and scaling DevTools has taught me that world-class culture and relentless product focus are non-negotiable. Drawing on experiences across Amazon, Apple, and Microsoft—and hard-won lessons from startups like Unblocked and Buddybuild—I’m sharing the principles I rely on to ship great developer products at scale.

    Why building for developers is different: developers are discerning, allergic to friction, and quick to churn if the DX isn’t exceptional. That means fast setup, clear docs, ergonomic APIs, sane defaults, and deep integrations with GitHub, GitLab, Bitbucket, Confluence, AWS, and Microsoft Azure.

    I benchmark teams against gold-standard platforms like Stripe, Twilio, and Looker—tools that reward mastery, never bury the lede, and make success observable in minutes, not days.

    From the early days of Buddybuild, the signal was unmistakable: remove toil from CI/CD, shorten feedback loops, and teams will expand usage without a sales nudge. The pattern holds across DevTools: when time-to-value approaches zero, the product sells itself.

    Early signs of product market fit: organic team-to-team adoption, repeatable setup success, contribution from power users, and inbound demand you cannot keep up with. When these show up, “Why great product is everything” stops sounding like a platitude and starts reading like a P&L.

    Monetizing product market fit is straightforward if you align value and pricing units. Seat-based maps to collaboration; usage-based maps to compute, API calls, or storage; hybrid models reduce edge-case friction. Keep the packaging simple and double down on “The power of positioning.”

    AI is complicating product market fit. Gen AI accelerates gen ai for product prototyping, but it also introduces instability: model drift, hallucinations, and evaluation blind spots. I build an evaluation harness, human-in-the-loop review for risky flows, and a clear customer support ai strategy before scaling.

    Being customer-obsessed is the moat. I embed forward deployed engineers with key customers to translate real workflows into product decisions, close the empathy gap, and validate behavior in production environments.

    On decision-making, I blend product discovery with crisp documents and measurable bets: PRFAQs or design docs to clarify intent, guardrails in analytics, and outcomes vs output OKRs to keep teams aligned to impact.

    Unblocked, a developer tool that lets you talk to your codebase, points toward a future where code search, context, and refactoring converge into conversational workflows. I’m bullish on the pattern, but I stay sober about failure modes and cost-to-serve.

    Here’s my cautious take on AI: latency, privacy, and provenance matter as much as model quality. The best teams treat prompts as product, training data as liability, and evaluation as a first-class release gate.

    Hiring is where many teams stumble. Don’t over-index on competency when hiring. I optimize for learning velocity, ownership, and kindness under pressure. Competency scales output; character scales organizations.

    As a second-time founder and operator, I treat mental health like uptime. I schedule recovery, define non-negotiables, and surround myself with peers who normalize the hard days. Burnout is a systems failure, not an individual weakness.

    I don’t do demos. I prefer self-serve trials with instrumented onboarding, sample projects, and guardrails that let the product do the talking. If a prospect can’t succeed in 15 minutes, we fix the product, not the deck.

    On customer feedback, I separate noise from signal with cohorts and context. I prioritize requests that reduce time-to-value, unblock integrations, or meaningfully expand the surface area of successful use cases. That’s how to deal with customer feedback without losing strategic focus.

    To build and scale DevTools, keep the bar high and the loop tight: ship small, watch usage, learn fast. Invest in platform reliability, rock-solid SDKs and CLIs, and a developer experience that earns trust release after release.

    Resources and touchstones I revisit often:

    Apple’s acquisition of Buddybuild: https://www.cnbc.com/2018/01/02/apple-agrees-to-buy-buddybuild.html

    AWS: https://aws.amazon.com

    Bitbucket: https://bitbucket.org

    Confluence: https://www.atlassian.com/software/confluence

    GitHub: https://github.com

    GitLab: https://gitlab.com

    Looker: https://looker.com

    Microsoft Azure: https://azure.microsoft.com

    Stewart Butterfield: https://www.linkedin.com/in/butterfield/

    Stripe: https://stripe.com

    Twilio: https://twilio.com

    Unblocked: https://getunblocked.com/

    If you’re building for developers, stay ruthless about simplicity, respectful of their time, and obsessed with proof in production. That’s how durable product-market fit is earned—and monetized.


    Book a consult png image
  • Inside Intercom’s Bold Reboot: Lessons in AI Strategy, Ruthless Focus, and Culture

    Inside Intercom’s Bold Reboot: Lessons in AI Strategy, Ruthless Focus, and Culture

    I’ve been reflecting on a remarkable comeback story that offers sharp lessons for product leaders navigating AI disruption. Eoghan McCabe is the CEO and cofounder at Intercom, an AI customer service platform. Intercom has raised over $240M, and was last valued at $1.3B in 2018. After spending 9 years building the company, Eoghan left Intercom in 2020, but he’s since returned, reshaping Intercom and pioneering its pivot to an AI-first service. That arc—departure, return, and reinvention—captures a founder’s willingness to defy orthodoxy and act from first principles.

    What stood out to me most was the unapologetic embrace of intuition. In high-variance environments like AI and customer support, best practices lag reality. Founder intuition vs. standard practice isn’t a cliché here; it’s a capability. I’ve seen teams overfit to playbooks and underweight the signals that matter—customer truth, product discovery signals, and outcomes vs output OKRs that force clarity on what actually moves the needle.

    McCabe’s reflections since leaving Intercom highlight the value of distance. Stepping away often exposes where complexity crept in and where focus was lost. On return, the immediate moves were decisive: refocus the strategy, simplify priorities, and set a higher bar for cadence and quality. Those changes were anchored by first-principles thinking and a willingness to question everything, including sacred cows.

    The productivity step-change is telling. How Eoghan increased Intercom’s productivity by 41% wasn’t magic—it was management. In my experience, that kind of shift comes from ruthless prioritization, removing low-leverage work, and consolidating teams around fewer, outcome-aligned bets. Tactically, think tighter operating rhythms, clearer decision rights, and forward deployed engineers who sit closer to customers to collapse feedback loops—especially critical in gen ai and customer support AI strategy.

    Strategy-wise, the pivot to AI-first wasn’t about feature-chasing; it was about category leadership. AI and category disruption demand conviction. Why you can’t make small improvements in big categories is simple: customers reward step changes in outcomes, not incrementalism. In customer service, that means rethinking workflows end-to-end, not just sprinkling gen ai for product prototyping on top of legacy processes.

    Hiring was another area where the guidance was crisp. Tactical advice on hiring top talent included raising the bar on slope (rate of learning) and ownership, biasing for product creators who thrive in ambiguity, and building an executive team that can scale the operating model, not just the org chart. I’ve found this is where product management leadership shows up most clearly—pushing beyond conventional resumes to find people who can compound execution and insight.

    Culture carried equal weight. Crafting a culture of ruthless honesty and transparency isn’t about being abrasive; it’s about creating a system where truth travels fast. In practice, that looks like instrumented business reviews tied to outcomes, written decision memos that capture tradeoffs, and a shared language for escalation. It’s uncomfortable at first, then liberating—because it accelerates learning.

    Brand came in for a reality check, too. Why software branding is in crisis resonates in an era where many products sound the same, look the same, and promise the same. The antidote is clarity: a point-of-view that’s inseparable from the product experience. How Intercom thinks about brand appears to lean into differentiated behavior—speed, quality, outcomes—rather than slogans. In crowded categories, that’s what earns attention and trust.

    Under the hood, this story is a masterclass in product-market fit lessons. It reaffirms that PMF isn’t a one-time event; it’s a moving target, especially when technology paradigms shift. The companies that navigate the shift are those that re-baseline their bets, measure what matters, and ship faster with higher standards. That’s the compounding loop I try to build: focused strategy, outcome-centric execution, and continuous product discovery.

    If you’re steering an AI transformation, a few prompts I use: Are we solving for an outcome that customers will feel in minutes, not months? Where are we making bold, non-incremental bets? Which processes can we kill to regain tempo? And do our leaders model transparency in a way that accelerates truth-telling across the org?

    For further context and inspiration, here are some of the references mentioned: 37signals: https://37signals.com, Basecamp: https://basecamp.com, Brian Halligan (HubSpot): https://www.linkedin.com/in/brianhalligan, David Heinemeier Hansson (37signals, Basecamp): https://www.linkedin.com/in/david-heinemeier-hansson-374b18221, Intercom: https://www.intercom.com, Jason Fried (37signals, Basecamp): https://www.linkedin.com/in/jason-fried, Salesforce: https://www.salesforce.com, Marc Benioff (Salesforce): https://www.linkedin.com/in/marcbenioff, Zendesk: https://www.zendesk.com.

    If you want to follow Eoghan directly: LinkedIn: https://www.linkedin.com/in/eoghanmccabe/ and Twitter/X: https://x.com/eoghan. I find it valuable to track leaders who are actively rewriting the playbook in real time.


    Book a consult png image