Author: Shivam Tiwari

  • Inside Rewind AI’s Playbook: PMF Breakthroughs, Bold Twitter Fundraise, and the Future of AI

    Inside Rewind AI’s Playbook: PMF Breakthroughs, Bold Twitter Fundraise, and the Future of AI

    I sat down with Dan Siroker to explore the product, fundraising, and AI strategy lessons behind Rewind AI’s rapid rise — and to reflect on what I would adopt in my own product management practice today. Dan Siroker is the co-founder and CEO at Rewind AI, a personalized AI powered by everything you’ve seen, said, or heard. Dan launched Rewind to an emphatic response on Twitter, and used a public pitch video to fundraise at a $350m valuation. Prior to starting Rewind, Dan co-founded Optimizely, which reached $120m ARR before being acquired by Episerver, a content management company. Dan was also the Director of Analytics for Obama’s first presidential campaign.

    What stood out immediately was Rewind’s journey to Product Market Fit and how deliberately the team instrumented learning loops. As a product leader, I pay close attention to how founders reduce ambiguity: narrow the target segment, ship thin slices, measure engagement cohorts, and iterate fast. Rewind’s early focus on utility and trust — not novelty — created the conditions for PMF while the team resisted the temptation to over-scope.

    I was especially interested in how Rewind works and how the team managed scope while building a category-creating product. By focusing on personalized recall powered by on-device intelligence and a clear privacy narrative, they avoided the common trap of trying to solve everything for everyone. My own rule of thumb is to enforce brutal prioritization around the highest-intent jobs-to-be-done, then earn the right to expand. That same discipline shows up in Rewind’s cultural mantra for shipping and validating fast.

    Lessons from Optimizely echo throughout. Being a second-time founder sharpens pattern recognition — from building high-clarity cultural values to operationalizing product-market fit. I’ve found that codifying operating principles early helps a team move faster with fewer collisions, and Dan’s approach to open feedback and public learning raises the bar for transparency.

    On product positioning as a category creator, the team leaned into outcomes over features, which is critical when the mental model is new. Rather than compete in a features arms race, they framed a compelling before-and-after: instant, searchable memory that augments cognition. In my experience, that level of narrative clarity drives founder-led GTM and accelerates word-of-mouth.

    We also dug into where to build in AI, and what makes a “wrapper” thin versus thick. My take: thin wrappers add shallow convenience on top of foundation models; thick wrappers integrate proprietary data, workflow depth, distribution advantages, and durable UX moats. Founders should aim for thick wrappers with unique data flywheels, not commodity interfaces easily displaced by platform shifts.

    Operationalizing Product Market Fit remains a craft. I routinely use leading indicators like activation rate, day-7/day-30 retention for key actions, and sentiment via structured PMF surveys. Rahul Vohra’s framework for measuring and optimizing Product Market Fit: https://review.firstround.com/how-superhuman-built-an-engine-to-find-product-market-fit is a proven playbook. Pair that with cohort-based instrumentation and tight audience segmentation to reveal the “sharpest edge” of value.

    On AI hype, we aligned on a pragmatic view: real value accrues where latency, accuracy, and privacy meet workflow depth. Apple’s Silicon: https://www.macrumors.com/guide/apple-silicon/ and on-device acceleration will keep unlocking new consumer experiences, while ChatGPT: https://chat.openai.com/ has reset expectations for natural interfaces. The cautionary tales of Google Glass: https://en.wikipedia.org/wiki/Google_Glass and Google Wave: https://en.wikipedia.org/wiki/Google_Wave remind me that timing, social acceptability, and use-case clarity matter as much as technical novelty.

    Data privacy is now a core buying criterion, not a checkbox. I see a clear trend toward local-first approaches, explicit consent, and user agency — especially for products that touch memory, identity, and personal archives. Framing value through Maslow’s Hierarchy of Needs: https://www.simplypsychology.org/maslow.html helps prioritize trustworthy utility over gimmicks.

    Dan’s one-of-a-kind Twitter fundraising strategy was a masterclass in founder-led GTM. By sharing a public pitch and engaging directly with early users and supporters, he compressed feedback cycles and aligned community, product, and capital. For reference, see Dan’s public Twitter fundraise: https://twitter.com/dsiroker/status/1646895452317700097 and Dan’s Rewind demo tweet: https://twitter.com/dsiroker/status/1638799931891920897. The transparency extended to leadership practice as well, with Dan publicly sharing his own 360 performance reviews: https://twitter.com/dsiroker/status/1689763756459675650 — a bold move that builds trust.

    I’m watching what’s next for Rewind with interest, particularly around thicker integrations, extensibility, and collaboration patterns. In the next decade, I expect assistive AI to become ambient, multimodal, and context-aware — an ever-present copilot that feels less like a tool and more like an extension of cognition.

    Referenced: Apple’s Silicon: https://www.macrumors.com/guide/apple-silicon/

    Referenced: ChatGPT: https://chat.openai.com/

    Referenced: Dan publicly sharing his own 360 performance reviews: https://twitter.com/dsiroker/status/1689763756459675650

    Referenced: Dan’s public Twitter fundraise: https://twitter.com/dsiroker/status/1646895452317700097

    Referenced: Dan’s Rewind demo tweet: https://twitter.com/dsiroker/status/1638799931891920897

    Referenced: Google Glass: https://en.wikipedia.org/wiki/Google_Glass

    Referenced: Google Wave: https://en.wikipedia.org/wiki/Google_Wave

    Referenced: Maslow’s Hierarchy of Needs: https://www.simplypsychology.org/maslow.html

    Referenced: Optimizely: https://www.optimizely.com/

    Referenced: Paul Graham: https://twitter.com/paulg

    Referenced: Rahul Vohra’s framework for measuring and optimizing Product Market Fit: https://review.firstround.com/how-superhuman-built-an-engine-to-find-product-market-fit

    Referenced: Rewind AI: https://www.rewind.ai/

    Referenced: Scribe (which morphed into Rewind): https://www.scribe.ai/about

    Where to find Dan Siroker: Twitter: https://twitter.com/dsiroker

    Where to find Dan Siroker: LinkedIn: https://www.linkedin.com/in/dsiroker

    Where to find Dan Siroker: Personal website: https://siroker.com/

    Where to find Dan Siroker: Blog: https://medium.com/@dsiroker

    My takeaway for founders and product leaders: obsess over segmentation, instrument for learning, and tell a crisp narrative that earns trust. Thick wrappers, privacy-first design, and founder-led GTM are how you win the next wave of AI.


    Book a consult png image
  • From Founder-Led GTM to Repeatable Product-Market Fit

    From Founder-Led GTM to Repeatable Product-Market Fit

    You have several paying customers, a founder who can rescue almost any sales call, and a roadmap full of requests. That can feel like product-market fit. It may also be a collection of individually negotiated successes that will break the moment you add leads, sellers, or a second customer segment.

    The practical test is not whether the founder can win another deal. It is whether the same type of customer buys for the same reason, reaches value through the same core path, and stays or expands without bespoke intervention. Founder-led go-to-market should discover that pattern and turn it into a system someone else can operate.

    Founder-led GTM must reveal a repeatable unit

    Founder-led GTM has two jobs. The visible job is closing customers. The more important job is learning why a specific customer buys, what the product must do to deliver value, and which parts of the sale can be repeated.

    A founder can cross gaps that would stop a normal go-to-market motion. They can redesign the demo, promise roadmap work, adjust pricing, pull engineers into implementation, and lend personal credibility to an uncertain purchase. That flexibility is useful while the company is learning. It also distorts the signal. A deal is not evidence of repeatability if it depends on founder status, an unplanned feature, an unusual commercial exception, or invisible manual work.

    Before widening the funnel, define the unit you are trying to repeat:

    • Who: The customer segment, operating context, user, economic buyer, and disqualifying characteristics.
    • What: The acute workflow problem the customer is already trying to solve, described through the last real occurrence rather than a hypothetical future need.
    • Why now: The event, cost, risk, or operational pressure that makes the status quo unacceptable.
    • Promise: The business outcome the buyer expects, not the collection of capabilities being sold.
    • Path: The minimum sequence from setup to first proof to realized value.
    • Boundary: The conditions under which you should decline the opportunity rather than turn an outlier into roadmap policy.

    I find it useful to turn the path into a three-frame value storyboard. The first frame captures the current pain step by step. The second identifies the first moment when the customer can see that the product works. The third shows the completed workflow and the business result the buyer can verify.

    Give each frame observable evidence. The pain might be demonstrated by time spent, an error, a delayed handoff, or exposure to risk. The first-value frame needs an activation event that both the product and customer can recognize. The final frame needs an outcome in the customer’s terms. This storyboard becomes a shared contract across product, sales, implementation, and the customer. If a proposed feature does not move a customer toward one of those frames, it should not automatically enter the core roadmap.

    Early teams often mistake breadth for demand. Ten different feature requests can mean ten customers want ten different products. A narrower signal is more valuable: one capability repeatedly attracts urgency, earns willingness to pay, concentrates meaningful usage, and shortens the path to value. When those signals converge, a zoom-in decision can be stronger than expanding the feature set.

    Do not focus on a feature because customers compliment it. Look for three forms of evidence together: qualitative pull, concentrated behavior, and business impact. An adjacent request belongs in the core only when it serves the same customer, workflow, buyer, and value metric. Otherwise, treat it as a separate hypothesis.

    Run early accounts as controlled learning cohorts

    If every early customer is different, a company-wide average tells you very little. Group similar accounts into tight cohorts and assign explicit learning goals to each cohort. Keep the major assumptions stable enough to interpret the result. Changing the segment, pain, packaging, channel, and onboarding model at the same time produces activity, not knowledge.

    A disciplined founder-led loop looks like this:

    1. Qualify against the repeatable unit. Record why the account fits and every exception required to include it. An attractive logo is not a substitute for fit.
    2. Reconstruct the last instance of the problem. Ask the customer to walk through what happened, who touched the workflow, where it failed, and what the failure cost. This is more reliable than asking what features they might want.
    3. Sell the outcome to the economic buyer. The CEO is useful when the outcome and organizational change genuinely sit with the CEO. Otherwise, find the person who owns the cost, risk, or operating result. Use that conversation to test whether the value narrative survives beyond the end user.
    4. Ask for payment early. Praise and participation show interest. Payment tests whether the problem and proposed outcome justify a budget decision. Document discounts, special terms, and bundled services so revenue is not mistaken for a clean pricing signal.
    5. Deliver with high-touch support. Observe the real workflow, perform uncertain steps manually, capture edge cases, and write down each intervention. Manual delivery is productive when it creates reusable knowledge.
    6. Classify what you learned. Recurring, core, and deterministic work should move toward the product. Bounded variation can become an implementation or support playbook. One-off work that does not strengthen the core should be declined or priced and managed separately.

    This is the practical meaning of doing the job before automating it. The objective is not to build a permanent services layer around an immature product. It is to see enough of the workflow to distinguish the stable system from its edge cases.

    White-glove support can remain a strategic learning channel until three things are true: the top five pain patterns are becoming repeatable, there is a clear route to tooling or self-service, and customer feedback reaches the product team quickly enough to change the default experience. High-touch delivery is not inherently unscalable. Unclassified manual work is.

    Keep a one-page record for every account. Capture the ICP evidence, triggering event, buyer, promised outcome, commercial exceptions, activation milestones, manual interventions, realized result, and renewal or expansion signal. At the end of the cohort, compare the records side by side. The repeated pattern matters more than the most enthusiastic anecdote.

    The cohort review should end with a decision. Narrow the ICP, focus the product, revise the value narrative, change packaging, repair onboarding, or reject the hypothesis. If the review ends with a longer list of features but no changed assumption, the learning loop is incomplete.

    Measure fit with revenue, engagement, and value

    Revenue alone can reflect founder skill, heavy services, or favorable terms. Usage alone can reflect curiosity or a useful tool that is not important enough to fund. A compelling customer outcome can still fail commercially if activation, packaging, or distribution is too difficult. Product-market fit becomes more credible when revenue, engagement, and value strengthen together.

    SignalQuestion it answersUseful evidenceDecision it should inform
    RevenueWill this customer pay, remain, and expand?Pilot-to-paid conversion, logo retention, Net Revenue Retention, and expansionWhether pricing, packaging, qualification, and the commercial motion are working
    EngagementDoes the product become part of the intended workflow?Time to first value, activation milestones, and depth, frequency, and breadth of usageWhether onboarding and the core product path are becoming easier to complete
    ValueDoes usage create the result the buyer expected?Customer-specific outcomes such as cost savings, yield improvement, or risk reductionWhether the product solves a problem important enough to sustain demand

    Choose three to five REV metrics for each lifecycle stage, ensuring the set covers revenue, engagement, and value. Define thresholds by cohort and by the natural cadence of the workflow. A low-frequency process should not be judged by a daily-use standard. The relevant question is whether the intended workflow is completed when the need occurs and whether that completion produces the promised result.

    Do not blend every customer into one company average. A mature core segment can hide a weak new cohort, while a large expansion can disguise poor pilot conversion. Compare like with like and examine the movement between cohorts. You are looking for a product that becomes easier to sell, faster to adopt, and more valuable without increasing the exceptional effort around each account.

    The shape of the REV score tells you where to invest next:

    • Engagement and value are strong, but revenue is weak: Investigate pricing, packaging, qualification, and sales enablement before adding product breadth.
    • Revenue is strong, but engagement lags: Pause segment expansion and fix onboarding, the first-value moment, and the core workflow. Contract value does not compensate for a product customers fail to adopt.
    • Engagement is strong, but value is unproven: Instrument the business result and return to the economic buyer. Frequent activity is not automatically meaningful impact.
    • Value exists only after extensive manual intervention: Decide which interventions can become product defaults, repeatable services, or disqualifiers. Do not hide them inside a blended margin or implementation number.
    • All three signals improve across comparable cohorts: The motion is a candidate for transfer and controlled scaling.

    REV should function as a lifecycle diagnostic, not a badge declaring that product-market fit has been achieved forever. The balance will change as the product, segment, and buying motion mature. What matters is that the scorecard tells you which constraint to address next.

    Pass the transfer test before you scale

    A founder-led motion becomes repeatable when another capable operator can run it from documented choices rather than founder intuition. This does not mean the founder disappears from strategic accounts or stops talking to customers. It means routine progress no longer depends on the founder rescuing qualification, the demo, pricing, implementation, or value proof.

    Build the minimum operating system before adding volume:

    • An ICP with observable qualifiers, disqualifiers, trigger events, users, and economic buyers
    • An outcome narrative tied to the three-frame value storyboard
    • A discovery sequence grounded in the customer’s last real experience of the problem
    • A demo that follows the core value path rather than touring every capability
    • Pricing and packaging boundaries, including the exceptions that require approval
    • Activation and time-to-value milestones visible to product, sales, and customer success
    • An objection and proof library built from actual deals
    • Implementation and support playbooks for the recurring pain patterns
    • Escalation rules that separate a product gap, a service need, and a poor-fit customer
    • A distribution wedge that reliably reaches the defined customer

    Test the system in stages. First, let the operator observe the founder. Next, let the operator lead while the founder remains silent unless an agreed escalation condition appears. Then let the operator run a comparable opportunity without the founder. Start with lower-risk interactions, review the evidence after each stage, and update the system where it fails.

    The location of the failure points to the work. Poor qualification suggests an unclear ICP. A feature-heavy demo suggests weak positioning. Repeated implementation rescue suggests a product or onboarding gap. Inability to prove the result suggests weak value instrumentation. Hiring more sellers addresses capacity; it does not repair any of those problems.

    A scalable go-to-market system does not necessarily mean a larger sales team. For an SMB or product-led motion, distribution may compound through integrations, partner ecosystems, search, lifecycle communication, or in-product discovery. Apply the same test: can the channel repeatedly reach the intended customer, set the right expectation, activate the core workflow, and produce healthy REV signals?

    Treat every new segment as another fit search

    Repeatability in one segment does not automatically transfer to another. A move from smaller customers to larger organizations can change the buyer, urgency, security requirements, implementation path, sales process, value metric, and support model. Treat the expansion as a new product-market-fit hypothesis rather than an extra filter in the existing funnel.

    Give the adjacent segment its own ICP, storyboard, cohort, and REV thresholds. Protect the working core while the new motion is uncertain. A 70/20/10 portfolio split can be a useful starting constraint: roughly 70% of capacity hardens the core, 20% tests adjacent growth, and 10% explores longer-term bets. It is not a universal law, but it forces the cost of expansion into the open.

    Keep the roadmaps separate until the evidence shows that the same capability can serve both segments without weakening the core. Interest from a prestigious logo is not proof. Neither is a contract held together by custom implementation.

    Use the same restraint with category creation. New category language is warranted when existing labels constrain the value story, the product reliably produces a distinct outcome, and customers begin using the language without prompting. Before those signals appear, inventing a category adds an education problem to an unresolved fit problem.

    Key takeaways

    • Founder-won revenue is traction. Repeatable fit requires the same kind of customer to buy, activate, realize value, and remain without bespoke rescue.
    • Define the unit of repetition as a specific customer, painful workflow, triggering event, promised outcome, core product path, and boundary.
    • Use early accounts as controlled learning cohorts. Price early, observe the real workflow, and classify every manual intervention.
    • Measure revenue, engagement, and value together. The combination explains whether the constraint is commercial, behavioral, or tied to customer outcomes.
    • Transfer the motion in stages before adding volume. A new hire can absorb capacity only after the underlying decisions are legible.
    • Treat every segment expansion as a fresh fit search, with its own cohort and evidence, while protecting the proven core.

    Your next two weeks should produce evidence, not a larger funnel. Storyboard the core value journey, choose three to five REV measures for each relevant lifecycle stage, group current customers into comparable cohorts, and mark every commercial exception and manual intervention. Then run a cohort review and make one decision: narrow the ICP, focus the product, change packaging, repair onboarding, or transfer a repeatable step.

    If the evidence cannot support one of those decisions, the answer is not more scale. Keep the founder inside the learning loop until the motion is clear enough to teach, measure, and repeat.

    References

    • Shivam.Consulting Blog – Mastering Product-Market Fit with the REV Model: My Battle-Tested Category Playbook
    • Shivam.Consulting Blog – How a 3-Time Founding Team at Pilot Unlocked Product-Market Fit Faster – My Proven Playbook
    • Shivam.Consulting Blog – Pulling Off the Zoom-In Pivot: Luminai’s Kesava on Focus, Sales Psychology, and Product-Market Fit
    • Shivam.Consulting Blog – How I Repeatedly Find Product-Market Fit: Shippo-Inspired Playbook for Bold Product Leaders
    • Shivam.Consulting Blog – Building Zapier by First Principles: Hard-Won Growth, Distribution, and Hiring Lessons
    • Shivam.Consulting Blog – Intuition, White-Glove Support, and Relentless Execution: Lessons from Looker to Omni
  • Open-Source Commercialization: A Developer-Led Playbook

    Open-Source Commercialization: A Developer-Led Playbook

    Your repository is gaining adoption. Developers are asking for integrations, while larger companies want security reviews, support, and a managed option. The tempting response is to pick an enterprise feature, hide it behind a paywall, and call that a business model. That can just as easily weaken the adoption engine you are trying to monetize.

    Your real job is to preserve the low-friction path that developers value while charging for the new burdens that appear when usage becomes organizational: operating infrastructure, governing access, satisfying compliance requirements, guaranteeing reliability, and supporting critical workloads. The boundary between those two experiences determines whether developer adoption compounds into revenue or stalls in mistrust.

    Choose the commercial promise before choosing paid features

    Open source, a managed cloud, and an enterprise edition are not merely three packages of the same software. Each makes a different promise.

    • An open-source project gives developers autonomy. They can inspect it, run it, extend it, and decide whether it deserves a place in their stack.
    • A managed service takes operational responsibility away from the customer. The customer pays to avoid provisioning, upgrades, scaling work, multi-tenant reliability problems, and routine maintenance.
    • An enterprise offering helps an organization control risk. The buyer pays for identity, governance, compliance, support, and predictable operation across teams.

    These promises can coexist, but you should not blur them. If customers mainly want you to operate the software, a hosted product is the natural commercial surface. If they can operate it but need policy controls and contractual assurance, an enterprise package is more coherent. If value and cost both rise with workload, consumption pricing may fit better than a fixed feature tier.

    Start with a short commercialization brief. It should answer the following questions before anyone debates individual paywalls:

    1. What useful outcome must a developer be able to reach without paying?
    2. Which responsibilities become materially harder when the product moves from an individual project to a production system?
    3. Who feels that difficulty: the developer, platform team, security team, procurement function, or executive owner?
    4. Is the customer paying for software capability, transferred operations, reduced risk, or guaranteed service?
    5. Can a successful community user move to the paid product without rebuilding the implementation?

    The first answer is your community promise. Protect it. The second through fourth answers reveal the commercial job. The last answer tests whether you have a growth path or merely two products that happen to share a name.

    Write the boundary down and make ownership explicit. A visible open-core stewardship model can make decisions easier to inspect: contributors can see what belongs in the shared foundation, customers can understand what they are buying, and product teams have a durable standard for future packaging debates.

    Licensing requires separate care. Open core is a commercial architecture, not a license, and changing package boundaries does not automatically change rights granted under earlier releases. Before relicensing code, moving contributed work into a proprietary edition, or changing contributor terms, use qualified open-source legal counsel. A product decision is not a substitute for a license review.

    Put the paywall where organizational complexity begins

    A durable paywall usually appears where the beneficiary changes. The foundational workflow benefits every developer and drives distribution. Governance, compliance, managed operation, and contractual reliability primarily benefit organizations with larger systems and more downside risk.

    That gives you a practical starting map:

    Customer jobLikely commercial surfaceWhat should remain intactEvidence to seek
    Run the core workflow independentlyOpen-source projectA complete, credible path to the product’s foundational valueSuccessful setup, repeated use, extensions, and community participation
    Avoid operating the systemManaged cloud or hosted serviceThe ability to self-manage without deliberate degradationRequests for hosting, upgrades, scaling help, security operations, or migration support
    Control access and prove complianceEnterprise tierThe individual developer workflowRequirements for SSO or SAML, granular role-based access, audit logs, and policy enforcement
    Reduce production and support riskEnterprise tier or support planSelf-service documentation and a usable community experienceRequirements for advanced alerting, longer retention, premium support, or service-level commitments
    Expand a measurable workloadUsage-based or consumption pricingA low-friction entry point and transparent meteringA value metric that grows with customer outcomes and produces a bill the customer can anticipate

    This is a hypothesis map, not a universal feature list. SSO, audit logs, retention, and support can be sensible enterprise fences because they serve organizational control. They are poor fences when withholding them makes the foundational product unsafe or unusable for the very community responsible for its adoption.

    Run every proposed paywall through five tests:

    1. Beneficiary test: Does the capability mainly help an individual do the core job, or help an organization govern many people and systems?
    2. Burden test: Does delivering it create meaningful infrastructure, reliability, security, or support responsibility for your company?
    3. Value test: Can the customer explain the operational cost, risk, or delay the capability removes?
    4. Trust test: Will a reasonable maintainer see the boundary as funding a stronger ecosystem, or as weakening the open product to manufacture conversion?
    5. Migration test: Can users upgrade without changing their architecture, redoing configuration, or losing state?

    If a feature fails the beneficiary or trust test, keep it open unless you have unusually strong contrary evidence. If it passes the burden and value tests, it is a stronger hosted or enterprise candidate. If migration fails, fix that before increasing acquisition. More adoption will otherwise create more stranded users, not more qualified demand.

    Only then should you select a pricing structure. A good, better, best model works when customers progress through qualitatively different needs, such as collaboration, governance, and enterprise assurance. Usage-based pricing works when consumption is measurable, understandable, and connected to value. Outcome-based pricing requires an outcome that both sides can define and attribute; without that clarity, it turns normal product variance into a billing dispute.

    Use the customer, competition, and company lens to pressure-test the result. Customer analysis tells you which outcome deserves a budget. Competition includes the do-it-yourself alternative, not just commercial vendors. Company analysis tells you whether the price can support the infrastructure, security, support, and go-to-market obligations attached to the promise.

    Willingness-to-pay work should test decisions, not compliments. Ask prospective buyers to compare real package boundaries, identify what they could approve, and explain what would block procurement. A positive answer to a vague question about paying someday is not pricing evidence. A buyer choosing between concrete offers and naming the approval path is much closer to it.

    Turn developer adoption into a designed growth loop

    Free availability is not developer-led growth. A project grows commercially only when developers reach value, return, bring the product into a team, and encounter a paid path that solves the next problem without undoing their earlier work.

    Design that journey as a sequence of observable transitions:

    1. Discovery: A developer finds a credible example, integration, technical explanation, or community recommendation that matches a current problem.
    2. First value: The developer completes the core workflow with sensible defaults and without needing a meeting.
    3. Repeated value: The product becomes part of an actual development or production routine rather than a one-time experiment.
    4. Team adoption: Configuration, projects, dashboards, workflows, or operational responsibility begin to span more people.
    5. Organizational need: Security, governance, reliability, procurement, or managed-operation requirements emerge.
    6. Upgrade: The team moves to the commercial offer while preserving its implementation, knowledge, and momentum.

    For each transition, write the obstacle that can prevent it and the product response that removes that obstacle. Discovery may fail because the positioning is broad and the documentation does not name a concrete job. First value may fail because setup exposes infrastructure decisions before the user has seen the benefit. Team adoption may fail because permissions and shared workflows were added as afterthoughts. Upgrade may fail because the cloud product requires a new configuration model.

    Your activation definition should describe achieved value, not administrative activity. Creating an account, starring a repository, cloning code, or downloading a package proves interest. It does not prove that the product worked. Define the first meaningful result for your product and instrument that event wherever users have consented to telemetry.

    Then simplify the path to that result. Give the user a strong default. Defer optional configuration. Provide a working example that can be changed after it succeeds. Make error messages point to the next corrective action. Treat documentation, command-line output, sample projects, and migration tooling as parts of the product rather than promotional material around it.

    The proof moment depends on the product. It might be a successful deployment, a populated dashboard, a completed pipeline, or a policy enforced against a real resource. Whatever it is, make that moment fast, visible, and repeatable. Developers tolerate depth once they trust the result; complexity before proof merely consumes goodwill.

    Developer evangelism should reinforce this loop. Its job is to teach useful patterns, reveal friction, and give technical users a credible path into the community. Treating every interaction as lead capture damages that role. Product and go-to-market teams still need feedback, but they should earn it through useful documentation, transparent communication, responsive community work, and clear consent.

    The commercial transition deserves the same product discipline as onboarding. Show what changes when a team upgrades. Preserve configuration and integrations. Explain the usage metric before a bill arrives. Provide migration validation or a preview when the move carries operational risk. If a solutions engineer must manually reconstruct every deployment, you have a services dependency rather than a scalable upgrade path.

    Measure the handoff and add GTM capacity in sequence

    Repository stars, package downloads, community membership, and documentation traffic are useful reach indicators. None of them, alone, tells you whether users activated, retained, or developed a reason to buy. Keep reach separate from product value and commercial intent.

    A workable scorecard follows the user’s progression:

    • Reach: Which channels bring developers with the problem your product actually solves?
    • Activation: What share of observable new users reaches the first meaningful result?
    • Retention: Do activated users repeat the core workflow or continue operating real workloads?
    • Team adoption: Does use expand into shared projects, environments, workflows, or operational ownership?
    • Commercial intent: Are users exploring hosting, migration, security documentation, governance controls, support, or service commitments?
    • Revenue quality: Do paid customers retain usage, expand for understandable reasons, and continue receiving value from the metric you charge against?

    Self-managed open source creates an unavoidable visibility gap. Do not fill that gap by pretending public activity equals product usage or by collecting invasive telemetry. Use opt-in product signals, cloud behavior, support requests, community conversations, version adoption, and direct customer discovery as different pieces of evidence. Keep the limits of each signal visible in the dashboard.

    Sales assistance should begin when customer complexity appears, not merely when a developer downloads the product. Stronger triggers include a request to migrate a production workload, satisfy security review, coordinate several teams, obtain contractual support, implement access governance, or model a substantial managed deployment. Those signals give sales and solutions teams a real problem to solve.

    The go-to-market organization should grow in the same order as the bottlenecks:

    1. When the bottleneck is adoption, invest in product experience, documentation, onboarding, community, and developer evangelism. Adding sellers cannot compensate for a developer path that does not reach value.
    2. When the bottleneck is technical evaluation or migration, add sales-assist, solutions engineering, and forward deployed engineering. Their purpose is to resolve complex implementation risk and return patterns to the product team.
    3. When the bottleneck is repeatability and expansion, add customer success, pricing operations, and ecosystem partnerships. Their purpose is to make value delivery, billing, retention, and adjacent distribution systematic.

    Keep one feedback loop across those functions. At a fixed operating cadence, review the largest activation obstacle, the most frequent scale or governance request, failed migrations, paywall exceptions, and the reasons paid customers did not expand. Assign a single owner to each decision, then record the community promise, target buyer, evidence, value metric, migration effect, and trust risk.

    This decision log prevents the commercial boundary from becoming a collection of historical accidents. It also gives product leaders a way to revisit assumptions without reopening every philosophical argument about open source. New evidence can change a package; the underlying decision standard should remain stable.

    Key takeaways

    • Define the community promise before selecting anything to monetize. The free product must deliver a complete foundational outcome.
    • Choose a hosted offer when customers want operational responsibility transferred to you; choose enterprise packaging when they need governance, compliance, control, or assurance.
    • Gate capabilities at the point where organizational complexity begins, not at an arbitrary point in the developer’s first-value journey.
    • Use pricing tiers for qualitatively different needs and consumption pricing only when the usage metric is measurable, valuable, and predictable.
    • Measure activation, retention, team adoption, and commercial intent separately from public reach indicators.
    • Add developer education, technical sales assistance, customer success, and pricing operations as their corresponding bottlenecks emerge.

    Start with one production workflow. Mark what must remain open for a developer to succeed, what operational responsibility a hosted service could absorb, and what organizational risk an enterprise tier could reduce. Validate the paid side with the people who own those burdens before moving code or setting prices.

    If maintainers cannot explain why the boundary is fair and buyers cannot explain why the paid offer is valuable, the model is not ready. When both explanations are clear, commercialization stops being a tax on adoption and becomes the mechanism that helps adoption survive at scale.

    References

    • Shivam.Consulting Blog – Open-Source GTM Masterclass: Pricing, Packaging, and Paywalls with Grafana Labs’ COO
    • Shivam.Consulting Blog – Open Source to Revenue: How GitLab Scales Transparency, Community, and Enterprise Growth
    • Shivam.Consulting Blog – How Radical Simplification Drove Vercel’s Product-Market Fit: Lessons for PMs and Founders
    • GitLab – Stewardship and open core business model
  • Goal-Setting for AI Products: How I Plan, Prioritize, and Confidently Ship in a Nonlinear GenAI World

    Goal-Setting for AI Products: How I Plan, Prioritize, and Confidently Ship in a Nonlinear GenAI World

    I build and ship AI products in an environment where the frontier changes weekly, so my planning system has to be adaptive, evidence-driven, and unapologetically outcome-focused. In this piece, I share the frameworks I use to set goals for generative AI, balance research with product execution, and scale responsibly — drawing sharp lessons from one of the most influential applied AI companies operating today.

    Consider Runway, an applied AI research company shaping the next era of art, entertainment, and human creativity. Runway has raised $237m and was one of Time Magazine’s “100 most influential companies” in 2023. Runway has been a persistent viral sensation in recent years, and is behind many of the most famous AI demos online.

    The earliest stages of an AI company often begin with research breakthroughs, scrappy prototypes, and clever distribution. In practice, that means leveraging containerization (https://aws.amazon.com/what-is/containerization/) and Docker (https://www.docker.com/) to package models reproducibly, showcasing work where practitioners already gather — Hugging Face (https://huggingface.co/), Hugging Face Spaces (https://huggingface.co/spaces), and Hugging Face Model Hub (https://huggingface.co/docs/hub/models-the-hub) — and tapping infrastructure like Replicate (https://replicate.com/) to get demos into people’s hands. Early, magical use cases — like the Green screen tool by Runway (https://runwayml.com/green-screen/) — teach us which problems are both technically feasible and viscerally valuable.

    I’ve learned to be cautious about “The limitations of being “customer-driven” when building in AI”. Traditional product discovery assumes needs are legible and solutions are relatively deterministic. In generative AI, user desire often follows model capability, not the other way around. The job is to triangulate: run tight user loops to validate perceived value, instrument objective model quality, and explore novel interaction patterns that customers can’t yet articulate. I treat this as a portfolio of discovery bets — some customer-led, some capability-led, all evaluated against clear outcome thresholds.

    Balancing research development with product development requires organizational design that prevents context-switching tax while preserving velocity. I pair research pods with product pods, supported by forward deployed engineers and domain PMs who translate evaluation metrics into user-visible milestones. Safety and content moderation sit on the critical path, not as afterthoughts — think policy definition, classifier tooling, abuse red teaming, and clear escalation playbooks. This balance is how you move from a great demo to a dependable product without losing momentum.

    Goal-setting amidst constant change in AI starts with outcomes vs output OKRs. I write OKRs in terms of user impact and model performance thresholds — for example, target ranges for latency, quality scores against a golden dataset, or creator retention — then let teams choose the highest-leverage outputs (data pipelines, fine-tuning, UX improvements) to get there. Why I don’t plan very far ahead: I treat the annual view as a vision and bet map, the quarterly view as a constrained slate of outcomes, and the 6–8 week cycle as the execution heartbeat. AI roadmaps are hypotheses; evaluation harnesses and launch gates are the truth.

    Community is a force multiplier. Forming a vocal community and fostering community requires real access and real listening: early release cohorts, office hours, and transparent changelogs. How they picked users for early release matters — diversity of use cases, sophistication of workflows, and willingness to give crisp feedback. Expanding past the first 100 users of Gen-2 demands readiness: evaluation parity across modalities, scalable infra, and safety coverage. Done well, this motion compounds learning while building authentic advocacy.

    For founders, my advice echoes the core lessons above. Start with a narrow, high-intent wedge and prove durable value fast; let founder-led GTM compress the feedback loop; instrument everything from day one; and resist the urge to over-plan features before you’ve nailed outcomes. Product-market fit lessons in AI often arrive via small, fast experiments — not grand, long-range plans. Ship thin slices that demonstrate unmistakable value, then iterate toward a system, not a single feature. When in doubt, shorten the loop and improve the evaluation harness.

    People often ask: Will AI replace video editors? My view is that AI will replace zero editors who master these tools — and many who don’t. The winners blend taste, storytelling, and generative leverage. The products we build should honor this reality: design for control, iteration, and co-creation, not just automation.

    If you’re mapping the progression of tech and use-cases, a few public references are instructive: Runway Gen-1 (https://research.runwayml.com/gen1) and Runway Gen-2 (https://research.runwayml.com/gen2) show how capability unlocks new workflows and demand. Runway’s 30 AI Magic Tools (https://runwayml.com/ai-magic-tools/) illustrates portfolio thinking — a suite of composable powers rather than a monolith.

    For builders focused on gen ai for product prototyping through production: keep your demo muscle strong, your evaluation stronger, and your outcomes strongest. Invest in community, treat safety as a feature, and let your OKRs steer what ships — not the other way around.


    Book a consult png image
  • Engineering Leadership That Scales: Strategy, Velocity, and Org Design from Carta, Stripe, Uber, Calm

    Engineering Leadership That Scales: Strategy, Velocity, and Org Design from Carta, Stripe, Uber, Calm

    I’m often asked how I translate lessons from hypergrowth engineering organizations into practical playbooks for product and platform teams. In this piece, I unpack the patterns I’ve seen repeatedly work—anchored by what I admire about Will Larson’s approaches at Carta, Calm, Stripe, and Uber—and how I apply them to build resilient, high-velocity orgs. Will Larson is a case study in modern engineering leadership. As CTO at Carta—an ownership and equity management platform—he helped guide the company after it raised at a $7.4b valuation in 2021. Before that, he was CTO at Calm, founded Stripe’s Foundation Engineering org, and led Uber’s Platform Engineering people and strategy. He’s also the author of Staff Engineer and An Elegant Puzzle, both essential reads for leaders leveling up from line management to org design. When I craft an engineering strategy, I start by writing down a small set of clear principles. This isn’t performative; it’s an alignment mechanism. Principles reduce decision thrash, make trade-offs explicit, and help teams navigate ambiguity without constant escalation. I’ve found the discipline of writing them down upfront pays off 10x in execution quality later. For the strategy document itself, I structure it so anyone can understand the why, what, and how in one sitting. A useful pattern: a sharp problem definition, a few guiding policies, and a concise set of coherent actions. That scaffolding keeps the strategy legible and actionable across functions—especially as it ladders into product roadmaps, platform investments, and talent plans. Every engineering strategy has two parts. First, compounding capabilities: the platform, tooling, and architecture that unlock future velocity. Second, targeted bets: focused initiatives that advance near-term outcomes. Neglect either and you either stall out later (too many quick wins, no compounding) or fail to ship value now (all compounding, no customer impact). Turning strategy into action requires ruthless translation. I map each guiding policy to a small number of initiatives with owners, milestones, and outcome metrics—not output. This is where outcomes vs output OKRs matter: measure the user or business result, not just the deliverable. It’s also where you surface dependencies early and avoid the Hidden Variable Problem that quietly derails timelines. I’m particularly intrigued by Carta’s unique “navigator” model, which blends technical leadership with cross-functional guidance to accelerate execution while preserving autonomy. In my experience, similar patterns work when leaders are explicitly accountable for both system health and product outcomes—reducing the gap between platform decisions and customer value. Engineering velocity is explainable, measurable, and optimizable. I anchor on DORA and the research from Accelerate (book), and I complement it with the SPACE (framework) to account for satisfaction and collaboration, not just delivery. The story I tell executives is simple: pick a few canonical measures, instrument them consistently, and then drive the feedback loops—branching strategy, CI/CD hygiene, change size, and operational excellence. Choosing the right metrics for an engineering org matters as much as the metrics themselves. I use a balanced set: delivery (lead time for changes, deployment frequency), quality (change failure rate, availability), and flow (work in progress, batch size). Then I pair these with narrative context so the numbers inform decisions rather than become a game to win. On policy, nuance beats orthodoxy. Great leaders define clear, default rules while acknowledging real-world exceptions. I’ve learned to document the policy, define who can grant exceptions, and track exception volume to spot design flaws. The goal isn’t rigidity—it’s predictable operations with a safe on-ramp for edge cases. Micromanagement is a symptom, not a root cause. Telling someone “don’t micromanage” is often counterproductive. Instead, I focus on what’s missing—trust, clarity, or visibility. If leaders can see the plan, the risks, the checkpoints, and the demo cadence, they don’t need to hover. If they still do, fix incentives and accountability, not just behavior. I avoid management anti-patterns by watching for early signals: policies without principles, roadmaps without strategy, meetings without decisions, or dashboards without actions. The best engineering executives pair systems thinking with crisp communication. They’re close enough to the details to ask sharp questions, yet disciplined enough to scale through managers and staff engineers. Executive communication is an asymmetric game. I tailor the message to the decision horizon: one slide for the ask, one for the trade-offs, one for the plan and risks. The Minto Pyramid (framework) helps—lead with the answer, then support it. In meetings, the fastest way to derail progress is to lack a clear owner, a time box, or pre-reads. Fix those and you reclaim hours every week. For presentation feedback, I’ve found a cadence that works: clarify the objective, highlight the single biggest risk, and eliminate anything that doesn’t move the decision forward. A bad sign with direct reports is when updates are status-only and insight-light; I coach toward “what changed, why it changed, and what you need.” For early-career engineers, the most durable advantage is compounding learning: pick hard problems, write more than you think you should, and seek out leaders who invest in your growth. For team development, I borrow a simple model: staff your keystones, instrument your systems, and build a culture where the best ideas win, not the loudest voices. If you want to explore the foundations behind these practices, start here. Accelerate (book): https://www.amazon.com/Accelerate-Software-Performing-Technology-Organizations/dp/1942788339 Good Strategy, Bad Strategy (book): https://www.amazon.com/Good-Strategy-Bad-Difference-Matters/dp/0307886239 DORA: https://dora.dev/ SPACE (framework): https://queue.acm.org/detail.cfm Minto Pyramid (framework): https://untools.co/minto-pyramid Carta: https://www.carta.com/ Calm: https://www.calm.com/ Stripe: https://www.stripe.com/ JavaScript: https://www.javascript.com/ KAFKA: https://kafka.apache.org/ Ruby on Rails: https://rubyonrails.org/ To go deeper on Will’s writing and perspective, these are great starting points. Twitter/X: https://twitter.com/lethain LinkedIn: https://www.linkedin.com/in/will-larson-a44b543/ Personal website/blog: https://lethain.com/ An Elegant Puzzle (book): https://www.amazon.com/Elegant-Puzzle-Systems-Engineering-Management/dp/1732265186 Staff Engineer (book): https://staffeng.com/book
    Book a consult png image
  • Inside Bard’s Playbook: How to Ship AI Fast, Build Ethically, and Outlearn Competitors

    Inside Bard’s Playbook: How to Ship AI Fast, Build Ethically, and Outlearn Competitors

    I spend a lot of time helping teams reconcile two pressures that define modern product management: ship fast enough to learn and compete, but slow enough to be safe, ethical, and useful. Studying Bard offers a crisp blueprint for navigating that tension and leveling up how we build with Generative AI. Jack Krawczyk is a Senior Director of Product at Google, building Bard. Bard is Google’s collaborative, conversational, and experimental AI tool that’s bridging the gap between humans and bots, while addressing ethical considerations around AI. After joining the project in 2020, Jack helped ship Bard in less than four years. Bard sources information directly from the web, and now enables users to inquire about and summarize YouTube videos. From a product management lens, the most valuable takeaway is the sequencing: problem definition → principled constraints → rapid public learning with clear guardrails. I’ve seen this order de-risk speed. When we anchor teams on a tight product thesis and ethical framework, we unlock faster iteration without drifting into feature theater. Shipping early—especially with a Large Language Model (LLM)—can feel risky. Yet the decision to open Bard to the public quickly reflects a disciplined bias toward learning velocity. In my experience, the longer we delay real-world feedback with LLMs, the more our internal assumptions calcify. Early exposure surfaces edge cases, calibrates safety systems, and drives better prioritization than any lab-only evaluation can. Ethics in AI is not a separate workstream; it’s a product requirement. I anchor cross-functional reviews on harm modeling, transparency, and user agency. Bard’s framing makes this explicit: collaborative, conversational, experimental—language that signals co-creation and responsible exploration rather than unfettered automation. That positioning matters for trust and sets expectations for both quality and limitations. Differentiation in AI assistants increasingly hinges on live context and modality. Bard sources information directly from the web, and now enables users to inquire about and summarize YouTube videos. In practice, this moves Bard beyond static Q&A toward dynamic sensemaking. I advise teams to ask: what fresh, authoritative context can our system responsibly ingest to reduce hallucinations and increase actionability? On development speed, I look for a culture that marries ambition with measurable risk reduction. That means small, end-to-end vertical slices; evaluation harnesses aligned to user outcomes, not model vanity metrics; and weekly red-teaming that actually changes the roadmap. Outcomes vs output OKRs are critical here—optimize for quality-adjusted learning per unit time, not just feature count. Early user research should be embedded, not episodic. I’m a proponent of forward deployed engineers paired with product and research to observe failure modes in the wild and close the loop quickly. With LLM-based experiences, qualitative signals (confusion, trust breaks, cognitive load) often precede quantitative ones; instrument both and let them inform each other. Deciding when to ship comes down to clear thresholds. I pressure-test launch criteria with two prompts: what would change my mind tomorrow, and what could break if we’re right but too early? For AI features, I also require recovery paths—explanations, undo, source attribution—so that small misses don’t become trust-ending moments. As for the competitive landscape—Bard versus ChatGPT, and others—users ultimately reward utility, reliability, and workflow fit. I encourage teams to pick a sharp use case, lean into their unique distribution or data advantage, and prove value in minutes, not weeks. “Generative AI” is table stakes; reliable outcomes in a real job-to-be-done is differentiation. Zooming out, I see three fronts shaping the future of LLM, Generative AI, and AGI: model capability, grounding and retrieval quality, and product ergonomics. Most teams overinvest in capability and underinvest in grounding and UX. The fastest wins often come from better retrieval, tighter prompts, and clearer affordances—not just a larger model. For aspiring AI developers, start narrow and instrument deeply. Pick a workflow with painful status quo, ship a thin slice, measure correctness and confidence, and iterate with real users. For non-LLM companies, the mandate is different: augment your core product where AI reduces friction or unlocks frequency—don’t bolt on a chatbot because everyone else did. For product leaders, AI changes the craft in two ways. First, prototyping is faster—use this to expand the option space early. Second, evaluation requires new muscles—build an experimentation and safety stack that blends qualitative red-teaming with quantitative reliability and cost controls. The leaders who thrive will combine taste with statistical rigor. If you want to go deeper, these references are useful: Bard: https://bard.google.com/; ChatGPT: https://chat.openai.com/; Duet AI: https://cloud.google.com/duet-ai; Free courses on machine learning by Andrew Ng: https://www.andrewng.org/courses/; Google Assistant: https://assistant.google.com/; Introducing Google Assistant to Bard: https://blog.google/products/assistant/google-assistant-bard-generative-ai/; Large Language Model (LLM): https://en.wikipedia.org/wiki/Large_language_model; Meena: https://blog.research.google/2020/01/towards-conversational-agent-that-can.html. In sum, the Bard blueprint reinforces a simple truth: ship with a thesis, learn in public with care, and let principled constraints accelerate—not slow—your path to product-market fit. That’s how we create value fast, build ethically, and stay ahead in the next era of AI.
    Book a consult png image
  • Developer-First Growth: From Fast Activation to Revenue

    Developer-First Growth: From Fast Activation to Revenue

    You can have healthy developer sign-ups, an active community, and enthusiastic feedback while the business remains fragile. The missing link is usually not another acquisition channel. It is an explicit path from a developer’s first successful result to a team-level reason to pay.

    If you are deciding what should stay free, where to place upgrade gates, or when to add sales, make those decisions in this order: first proof, repeated use, team expansion, then monetization. That sequence keeps revenue from choking the behavior that creates demand.

    Map the complete value chain before changing your funnel

    Developer-first is a sequence of proof, not a declaration that the developer is your only customer. The hands-on user needs technical evidence. A champion needs evidence that the tool will help colleagues. A manager needs evidence of recurring team value. Security, platform, and procurement stakeholders need evidence that adoption will not introduce unmanaged risk.

    A weak growth model treats all four as one persona and asks one landing page, one trial, and one pricing plan to serve everyone. A stronger model gives each person the proof needed for the next commitment.

    DecisionQuestion to answerPossible evidence
    Value objectWhat observable output proves that the developer’s job was completed?A successful API response, a runnable project, or a diagnosed issue
    Distribution objectWhat can leave one workspace and help another person discover or understand the product?A project link, pull-request check, alert, template, or reusable configuration
    Expansion eventWhat can a teammate do that makes the product more valuable to the original user?Collaborate, take ownership, reuse a workflow, or connect another system
    Billing meterWhich unit remains understandable as customer value and delivery cost increase?Seats, API calls, compute, storage, or a hybrid of access and consumption

    Choose one primary event for each row. Do not assume they are interchangeable. A sign-up is not proof of value. An invitation is not team activation. Hitting a free limit is not proof that a customer understands or accepts the paid proposition.

    For an error-monitoring product, creating a project is setup; receiving a real issue and connecting it to an owner is much closer to value. For a coding environment, opening an editor is setup; producing a runnable artifact that another person can use is value. For an API product, generating a key is setup; completing the first valid request is proof.

    Write a one-page value chain for your product with five entries:

    1. The recurring technical problem that creates urgency.
    2. The first observable result that proves the product works.
    3. The repeated workflow that makes the product useful rather than merely interesting.
    4. The teammate action that turns individual utility into organizational value.
    5. The operational, collaborative, or risk-related need that justifies payment.

    If any entry is vague, do not compensate with more acquisition. You will only send more developers into a journey whose economic logic is still missing.

    Make the first proof fast, observable, and honest

    Developer onboarding has two clocks. The first measures time to technical proof: can the developer make the product do something real? The second measures time to a useful workflow: can the developer connect that proof to the job that brought them here?

    Treat five minutes to a clean first proof and fifteen minutes to a meaningful self-serve success as design constraints, not universal market benchmarks. Some products require deployment approvals, production data, or infrastructure changes that cannot honestly fit those windows. In that case, provide a safe sandbox for immediate proof, label it clearly, and make every remaining production step visible. Do not count synthetic sandbox activity as production activation.

    A reliable activation path has six parts:

    1. State the result before explaining the product. Tell the developer what will exist, run, or become visible at the end of the path.
    2. Ask only for prerequisites needed to produce that result. Defer profile fields, teammate invitations, and purchasing questions.
    3. Offer one recommended route. Pick a primary SDK, CLI flow, or sample project instead of presenting every option at once.
    4. Show expected output beside each command or configuration step. A developer should be able to distinguish success from silent failure without opening a support ticket.
    5. Make errors recoverable. Explain the likely cause, the corrective action, and whether retrying is safe.
    6. Point from first proof to the next real workflow: connect a repository, ingest production-like data, share the artifact, or schedule recurring execution.

    Documentation, sample projects, SDKs, CLIs, and integration setup are part of this product surface. If a quickstart breaks when a dependency changes, the failure belongs in the activation funnel just as surely as a broken button does.

    Instrument the journey with events whose names describe completed states, not interface activity. Account created and button clicked can help diagnose behavior, but first success, first workflow completed, first repeat use, artifact shared, and teammate value completed are better business events. Define the payload and eligibility rules for each event so that internal traffic, retries, imported projects, and automated tests do not inflate the result.

    Track activation rate against eligible new workspaces, then inspect median and 90th-percentile time to first proof. The median tells you how the common path behaves. The tail shows where particular languages, SDKs, integrations, environments, or account types are failing. Segment before averaging; a smooth aggregate can conceal an unusable integration.

    When activation is weak, fix the dominant failed step before adding tours, messages, or lifecycle email. More explanation cannot rescue a path that produces authentication errors, ambiguous output, or an incomplete sample.

    Turn individual success into a measurable expansion loop

    A developer-first product becomes a growth engine only when value survives the handoff to another person. That handoff can produce acquisition, account expansion, or retention, but those are different loops and should be designed separately.

    • External sharing drives acquisition when a runnable project, template, result, or public artifact exposes the product to a new developer.
    • Internal sharing drives expansion when a teammate can review, reuse, own, or improve the original developer’s work.
    • Workflow integration drives retention when the product returns through the repository, incident process, deployment flow, alerting system, or another place where work already happens.

    The sequence matters. Let the developer create value before asking for an invitation. Then make the invitation carry the relevant object and context. A message that says a teammate shared a specific issue, project, or workflow gives the recipient a job to complete. A generic invitation merely gives them another account to create.

    A practical expansion loop looks like this:

    1. A developer completes a frequent, painful task.
    2. The product creates an artifact or signal that is useful beyond that session.
    3. The developer shares it or connects it to a team workflow.
    4. A teammate performs a meaningful action on the same object.
    5. The combined workflow repeats without a sales prompt.
    6. Privacy, collaboration, capacity, administration, reliability, or support needs create a natural paid threshold.

    Measure each transition. Useful metrics include the share rate among activated workspaces, the percentage of recipients who reach first value, collaborative activation, repeat use after collaboration, and organic expansion within retained workspaces. Count a teammate only after a meaningful action; an accepted invitation without product use is not expansion.

    Review these metrics by activation cohort. If a new onboarding experience raises sign-ups but lowers repeated team use, it has created cheaper accounts rather than stronger growth. Keep individual, team, and enterprise cohorts separate because their setup requirements, usage frequency, and reasons to remain can be materially different.

    Community activity adds another useful signal. Templates, integrations, documentation improvements, and contributions from power users show where the product has become important enough for developers to invest their own effort. Treat those contributions as product discovery: repeated extensions often reveal missing platform capabilities, while repeated documentation fixes identify friction in the official path.

    Monetize the consequences of success, not the act of trying

    The free boundary should protect the behaviors that create trust and distribution. The paid boundary should appear when successful use creates more demanding requirements. Charging too early suppresses learning and sharing. Charging too late leaves the company funding collaboration, infrastructure, and enterprise obligations without capturing the value they create.

    Build packages around escalating value and risk

    A simple packaging ladder usually has distinct jobs:

    • A free or community package lets a developer learn, create, and prove the core workflow with clear limits.
    • A team package supports private work, deeper collaboration, higher capacity, shared history, and stronger workflow integration.
    • An enterprise package addresses organizational access, governance, observability, scalability, reliability, support commitments, and service-level requirements.
    • A managed service removes deployment and operational burden, with pricing that may increase as the underlying workload grows.

    These are value layers, not a requirement to publish four plans. A small product may combine them. What matters is that each upgrade tells a coherent story about the customer’s changing job rather than presenting a random collection of disabled features.

    For an open source product, the community version should complete a real developer job. The commercial offer can remove operational burden and add the controls, assurance, and service required to run the product across an organization. An intentionally crippled core may generate upgrade clicks, but it also weakens the trust and adoption that open source was meant to create. Basic product safety should not be a premium feature; organizational policy, administration, and assurance are legitimate commercial value.

    Closed-source products can use the same logic through a self-serve free tier or trial. Open versus closed is not the central question. The central question is whether a developer can establish credible value before the organization is asked to make a larger commitment.

    Match the billing unit to both value and cost

    Seat-based pricing works when collaboration and access are the main sources of incremental value. Consumption pricing works when API calls, compute, storage, or another workload unit grows with both customer value and delivery cost. A hybrid model works when customers receive persistent platform value but also create variable infrastructure expense.

    Do not expose a technically convenient meter merely because it is easy to count. Developers may understand tokens, requests, build minutes, events, or storage internally, while the buyer thinks in deployed services, completed jobs, monitored applications, or active workflows. Choose a customer-facing unit that is predictable, auditable, and close enough to the outcome that increased use feels like increased value.

    Keep four usage concepts distinct:

    • Raw usage records everything the system processes and helps with capacity planning.
    • Eligible usage removes internal work, failed attempts, duplicate retries, and activity that should not be charged.
    • Customer-visible usage is the meter shown in the product, with a definition the customer can understand.
    • Invoiced usage is the final quantity after contractual allowances, credits, and plan rules are applied.

    If those definitions drift apart, billing becomes a trust problem. Reconcile them before launching consumption pricing. Give customers a current usage view, explain what causes the meter to move, and provide estimates, alerts, or caps where unexpected consumption could create a material bill. Instrument variable cost early as well; rapid adoption is not healthy expansion if the cost to serve the workload grows faster than revenue.

    Add human assistance after product proof, not in place of it

    Developer-first does not mean sales-free. It changes when human help enters and what that help is expected to accomplish.

    • The self-serve lane should prove the basic workflow without a meeting.
    • The product-assisted lane should respond to behavioral evidence of team value, such as repeated use, teammate activity, sustained consumption, or demand for private and administrative capabilities.
    • The enterprise-assisted lane should handle migration, architecture, procurement, security review, deployment planning, and commercial terms.

    Do not route a developer to sales merely because the email domain appears valuable. A stronger product-qualified signal combines first success, repeated use, and an expansion or operational need. Human assistance should remove organizational friction after technical conviction; it should not be required to demonstrate the happy path.

    Early founder-led selling remains useful because it exposes the language customers use, the objections that block purchase, and the capabilities that repeatedly matter. I would not scale outbound until teams in the same target segment can reach value through a similar path, describe a similar urgent problem, encounter recognizable paid triggers, and complete implementation with reasonably predictable effort. That is the point at which a sales narrative can be codified rather than improvised on every call.

    Forward deployed engineers can shorten the loop for complex accounts, but each engagement needs a learning objective, a reusable output, and an exit condition. Repeated one-off code is a services dependency. Reusable integrations, defaults, diagnostics, and product improvements turn customer work into a stronger platform.

    Run one weekly operating loop across growth and revenue

    Growth, product, sales, and finance should not maintain competing versions of the journey. Use one scorecard that connects the stages:

    • Acquisition quality: eligible new workspaces by segment and entry path.
    • Activation: completion rate and median and tail time to first proof.
    • Activation quality: sandbox success versus production or production-like success.
    • Retention: repeated completion of the core workflow by activation cohort.
    • Expansion: artifact sharing, teammate value, integration depth, and organic account growth.
    • Monetization: conversion after a real paid trigger, not conversion from all registrations.
    • Unit economics: variable cost per customer-visible or billable unit, including high-cost workloads.
    • Assisted growth: which product behaviors preceded a useful sales or engineering intervention.

    Every metric needs a defined owner and a decision it can trigger. Run experiments against one bottleneck at a time, with a primary outcome and guardrails for errors, support burden, retention, and cost. Do not A/B test wording around a path whose event semantics are unclear or whose dominant problem is technical failure. Repair the product and instrumentation first.

    Key takeaways for a developer-first growth model

    • Define first proof, repeated value, team value, and paid value as separate events.
    • Use five minutes to first proof and fifteen minutes to self-serve success as design constraints where the product can honestly support them.
    • Ask for sharing or collaboration after the developer has created something worth sharing.
    • Keep learning, creation, and distribution accessible; monetize collaboration, operational burden, capacity, governance, reliability, and support.
    • Use seats for collaboration, consumption for variable workloads, and a hybrid when both create material value and cost.
    • Qualify accounts through successful behavior and expansion signals, not registration volume or email domain alone.
    • Scale sales only after the target customer, activation path, paid trigger, and implementation pattern have become repeatable.

    Open your funnel this week and trace one recent cohort from first proof to its first teammate action and first paid need. The broken connection will tell you whether to simplify onboarding, create a better sharing object, move an upgrade gate, or add human help. That is a much more useful growth agenda than buying more traffic for an unfinished journey.

    References

    • Shivam.Consulting Blog — Winning with Open Source and SaaS: My GTM Playbook, Monetization Tactics, and Founder Fit
    • Shivam.Consulting Blog — The Secret Lever Behind Replit’s Hypergrowth—and the Product Playbook You Can Reuse
    • Shivam.Consulting Blog — DevTools at Scale: Hard-Won Lessons on PMF, AI, and Culture from Apple, AWS, Microsoft
    • Shivam.Consulting Blog — How Sentry Scaled DevTools to $100M ARR: My Playbook for PMF, B2D, and Packaging
  • An Operating System for AI-Era Product and Engineering Leaders

    An Operating System for AI-Era Product and Engineering Leaders

    If your teams can produce prototypes, specifications, and code faster with AI, why does the roadmap still feel slow? The work did not disappear. It moved from creating the first draft to deciding what deserves customer and production trust.

    That shift changes your leadership job. You are no longer optimizing only for delivery capacity. You are building a system that turns uncertain AI behavior into reliable customer outcomes. That system needs sharper bets, separate exploration and industrialization modes, evidence-based operating rhythms, clear decision rights, and people who can exercise judgment without waiting for permission.

    The bottleneck has moved from production to judgment

    AI makes many artifacts cheaper to produce. A team can generate interface concepts, implementation options, test cases, documentation, and working prototypes before it has proved that the underlying problem matters. That is useful leverage, but it creates a throughput trap: more plausible work enters the system than the organization can evaluate responsibly.

    Feature count, ticket velocity, and lines of generated code become even weaker management signals in this environment. They measure activity at the stage where activity is becoming abundant. The scarce resources are customer insight, technical taste, attention, and the willingness to stop work that has not earned further investment.

    Start every meaningful AI initiative with a one-page bet brief. It should be precise enough for product, design, and engineering to disagree before code creates momentum.

    • Customer and job: Name the user, the workflow, and the moment in which the problem occurs. Avoid broad labels such as productivity assistant.
    • Outcome: State what should improve for the customer or business. A launch is not an outcome. A completed task, resolved case, retained account, or reduced source of friction can be.
    • AI responsibility: Specify what the model must classify, retrieve, decide, generate, or recommend. Also state which parts of the workflow should remain deterministic.
    • Evidence: Define the cases that will demonstrate useful behavior, including common tasks, difficult edge cases, and unacceptable failures.
    • Constraints: Make latency, cost, privacy, security, explainability, and human-review requirements visible before the team chooses an architecture.
    • Failure boundary: Describe what happens when confidence is low or the system is wrong. Name the fallback, escalation path, and person accountable for the customer experience.
    • Rollout: Identify the owner, initial exposure, feature-flag plan, rollback mechanism, and decision that the first release is meant to inform.

    This brief prevents a common category error. Product acceptance and engineering acceptance are related, but they are not identical. Product acceptance asks whether the workflow creates meaningful value. Engineering acceptance asks whether the system is reliable, observable, maintainable, secure, and economical enough for its intended use. An impressive demonstration answers neither question on its own.

    I would not approve a production AI bet whose success criteria describe only what the team will ship. The brief should make it possible to observe a customer result, inspect system behavior, and decide whether to expand, revise, or stop the investment.

    Separate exploration from industrialization

    AI work becomes expensive when leaders ask one team to discover the product and harden the platform at the same time. Exploration rewards speed, range, and cheap learning. Industrialization rewards repeatability, control, and operational discipline. Both matter, but they should not be confused.

    Explore the customer outcome

    Give a small, mission-aligned group protected time to test the riskiest assumptions. Product should bring a specific customer problem. Design should make the interaction and trust model tangible. Engineering should expose feasibility limits early. A forward deployed engineer or another technically fluent customer-facing person can shorten the loop by observing the workflow where it actually happens.

    Use prototypes to answer questions, not to create the appearance of progress:

    • Does the proposed behavior remove a real step from the user’s job, or merely relocate it to review?
    • Can the user tell when the system is uncertain, and do they know what to do next?
    • Which inputs produce useful results, and which expose brittle assumptions?
    • Does the workflow still create value after human verification time is included?
    • What did the team learn that changes the product, model, data, or distribution decision?

    Protect focus time during this phase. The team needs room to test alternatives, inspect failures, and discard work without having to defend every abandoned prototype as lost output. Use a weekly evidence demo to maintain urgency without filling the calendar with status meetings.

    Industrialize the proven behavior

    Once a workflow earns further investment, treat the AI capability as a production system rather than a model call. The system includes prompts, retrieval, data transformations, tools, permissions, deterministic checks, user controls, monitoring, and recovery paths. Reliability comes from the whole chain.

    The transition should be explicit. Before moving from exploration to industrialization, confirm that the team has:

    • a repeated customer need rather than a technology looking for a workflow;
    • an observable outcome and a credible leading signal;
    • a representative evaluation set with difficult and unacceptable cases;
    • a named owner for model quality, service reliability, and the end-to-end customer experience;
    • known latency and cost constraints for the intended level of use;
    • privacy, security, data-governance, and access-control requirements;
    • a staged release plan with feature flags, monitoring, fallback behavior, and rollback;
    • a decision rule for expanding, revising, or ending the bet.

    Automated tests should cover deterministic components. Evaluations should cover AI behavior. Observability should connect technical events to user outcomes so the team can distinguish a model-quality problem from a retrieval failure, tool error, interface problem, or poorly defined task. Version the prompts, configurations, and evaluation sets that influence behavior; otherwise, the team cannot explain why performance changed.

    Do not interpret exploration as permission to ignore safety until later. Irreversible constraints belong in the initial brief. The distinction is about the maturity of the implementation, not whether privacy, security, or customer harm matters.

    The release target should be the smallest remarkable workflow, not the largest collection of AI features. Give the user a short path to value, opinionated defaults, understandable controls, and a complete recovery experience. A narrow capability that can be trusted will teach you more than a broad copilot whose value is difficult to locate.

    Run the organization on evidence, not AI activity

    An AI team does not need a new ceremony for every new tool. It needs a tighter truth loop. The operating rhythm should move evidence from customers and production into decisions while preserving enough uninterrupted time for builders to think.

    1. Write the intent before work begins. The one-page brief records the problem, constraints, owner, and success measures. If the intent changes, update the brief instead of allowing assumptions to diverge across meetings.
    2. Protect maker time. Reserve no-meeting blocks for implementation, evaluation, and failure analysis. Keep recurring capacity for prototypes, developer experience, and technical debt so short-term AI pressure does not hollow out the platform.
    3. Hold a weekly evidence demo. Show the real workflow, not a slide about completion. Demonstrate where the system helped, where it failed, what evidence was collected, and which decision is now required.
    4. Record the decision. Capture the evidence considered, assumptions still open, trade-offs made, owner, and next review point. A decision log lets the organization improve judgment instead of repeatedly debating the same context.
    5. Inspect outcomes separately from delivery status. Review customer impact, learning, service quality, and business effect. Delivery milestones remain useful, but they should not masquerade as proof of value.

    A good evidence demo is not a performance. The team should be able to show a failed evaluation, explain what it invalidated, and receive credit for preventing a weak assumption from reaching customers. If every demo ends with a green status, the mechanism is probably rewarding confidence rather than truth.

    Scope discipline matters here. AI expands the number of ideas that appear feasible, so the backlog will grow faster than the team’s capacity to validate it. Remove low-leverage work, consolidate teams around fewer outcomes, and use customer impact as the tie-breaker. Otherwise, faster prototyping produces a larger inventory of unfinished decisions.

    Match decision speed to reversibility. A reversible interface experiment can move with guardrails and a named owner. A choice involving sensitive data, security exposure, an irreversible migration, or reputational risk deserves a pre-mortem and wider review. Treating every choice as a committee decision slows learning; treating every choice as reversible hides real risk.

    Healthy debate is part of the cadence. Invite dissent in written RFCs, challenge assumptions rather than people, time-box the decision, and commit once the window closes. Truth travels faster when high standards are delivered with respect.

    Keep decision rights clear as roles begin to overlap

    AI lets more people create artifacts outside their traditional discipline. A product manager can generate a prototype. A designer can test implementation details. An engineer can draft a product specification. That overlap can accelerate discovery, but it does not erase accountability.

    RolePrimary decision rightRequired contribution to an AI bet
    ProductWhy this problem matters and what outcome the team will pursueCustomer context, outcome metric, scope, trade-offs, evaluation acceptance, and stopping rule
    DesignHow the experience communicates value, control, confidence, and recoveryWorkflow design, feedback, error states, human handoff, and trust cues
    EngineeringHow the system works and what production standard it must meetArchitecture, data flow, evaluations, testing, observability, security, reliability, and rollback
    All threeWhether the end-to-end outcome is good enough to expandShared evidence, customer exposure, failure analysis, and an explicit recommendation

    An artifact created with AI remains subject to the decision rights of the discipline that must stand behind it. Code generated by a PM is a prototype until engineering accepts responsibility for operating it. A model-generated requirements document is not product strategy until product has resolved the customer and business choices inside it. A generated interface is not finished design merely because it looks polished.

    Lead declaratively at the team level. Set the intent, constraints, measures, and decision deadline. Do not prescribe every prompt, framework, or implementation step. Guardrails create safety; room to choose creates ownership. This is especially important when tools and techniques change faster than executive expertise.

    You should move into the details under three conditions: the bet carries an existential reliability, security, or reputation risk; it is a pivotal zero-to-one decision; or cross-functional misalignment keeps recurring despite clear ownership. Enter to diagnose the system, expose the trade-off, and model the expected standard. Then step back out. Staying in the work turns executive attention into a dependency and quietly replaces the accountable team.

    Hire for judgment before tool fluency

    AI hiring can over-index on familiarity with the latest model or framework. Tool fluency has value, but it decays quickly. In an evolving product area, prioritize adaptable builders who can reduce ambiguity, derive a solution from first principles, and learn from failed assumptions. Add deep specialists when the motion and interfaces are stable enough for specialization to compound.

    Interview for the derivation, not merely the answer. Give the candidate an ambiguous customer problem and ask them to identify the first assumption they would test, the evidence they would collect, the failure they would refuse to expose, and the point at which they would stop. Ask what would change their mind. A polished solution with no falsifiable reasoning is a warning sign.

    Develop the same judgment inside the organization. Bring product managers into sales and support workflows. Let engineers observe customers rather than receiving filtered requirements. Rotate people through adjacent responsibilities when it improves their understanding of the whole system. Ask precise what-if questions during reviews: What if the retrieval result is stale? What if the tool executes twice? What if the user cannot verify the answer? What if the cost works in a pilot but not at broad adoption?

    Do not convert faster first drafts into permanently higher commitments before the quality loop proves that the gain is real. AI can reduce effort in one stage while increasing review, integration, or operational work elsewhere. Manage the whole value stream and the team’s energy, not the speed of the most visible artifact.

    Key takeaways

    • Optimize for reliable customer outcomes and decision quality, not the volume of AI-assisted output.
    • Require a one-page bet brief that defines the customer job, AI responsibility, evidence, constraints, failure boundary, owner, and rollout.
    • Run exploration and industrialization as distinct modes with an explicit transition between them.
    • Use weekly evidence demos, protected maker time, decision logs, and outcome reviews to shorten the truth loop.
    • Keep product, design, and engineering decision rights clear even when AI allows their artifacts to overlap.
    • Hire and develop people for technical taste, first-principles reasoning, customer fluency, and rate of learning.

    At your next planning review, choose one active AI bet and force it through the one-page brief. If the team cannot name the customer outcome, representative evaluations, unacceptable failure, accountable owner, and rollback path, the bet is not ready to scale. Protect the next build block, schedule the evidence demo, and make the next investment decision from what the team learns.

    References

    • Shivam.Consulting Blog – The Human Side of Engineering Leadership: Practical Plays to Build Creative, High-Performing Teams
    • Shivam.Consulting Blog – Build Enduring Software: Minimum Remarkable Products, Customer-First Culture, and Org Design Lessons
    • Shivam.Consulting Blog – Leading Up, Down, and Across the Org: Hard-Won Lessons in Executive Effectiveness, Culture, and Speed
    • Shivam.Consulting Blog – Developing Technical Taste: My Playbook for Next-Gen Engineers, AI Strategy, and 2024 Scaling
    • Shivam.Consulting Blog – Inside Intercom’s Bold Reboot: Lessons in AI Strategy, Ruthless Focus, and Culture
    • Shivam.Consulting Blog – Mastering Altitude Shifts: Hard-Won Product Leadership Lessons from Anneka Gupta’s Journey
  • Enterprise GTM: Build One System From Pipeline to Expansion

    Enterprise GTM: Build One System From Pipeline to Expansion

    You may have a healthy pipeline and still have a broken enterprise motion. The warning signs show up after the applause: pilots do not convert, onboarding starts from scratch, the executive sponsor disappears, and renewal depends on a last-minute rescue.

    That happens when acquisition, sales, implementation, customer success, and product operate as adjacent functions instead of one value-delivery system. The fix is to manage the customer lifecycle as a chain of evidence. At every transition, you should be able to name the customer decision, the proof required, the accountable owner, and the next commitment.

    Start at renewal, then design the lifecycle backward

    An enterprise customer does not renew because the rollout completed or users logged in. The customer renews when your product has become a credible way to produce an outcome the company still cares about. That makes renewal an input to GTM design, not a post-sales event.

    Write the success thesis before you qualify the opportunity. A useful structure is: for this account and process owner, the product will change a specific workflow, produce an agreed business result, and prove that result through an observable signal during the evaluation period. Do not let placeholders such as better productivity or improved collaboration survive. If the result cannot be observed, the account will eventually debate value through anecdotes.

    Then design each lifecycle moment around the decision the customer must make. The following model works as a practical starting point:

    Lifecycle momentCustomer decisionEvidence requiredExit criterion
    SignalIs this problem important enough to investigate?Repeated usage or expressed pain, a defined workflow, and a reason to actThe use case, affected role, and process owner are named
    Qualified opportunityIs the potential company value worth time, budget, and political capital?An outcome tied to an executive priority, a credible champion, and a visible buying pathThe value hypothesis and qualification record are accepted by the account team
    EvaluationCan the product create the result under real operating constraints?A baseline, target result, evaluation method, representative workflow, and production conditionsThe customer accepts the evidence and applies the agreed decision rule
    Commercial commitmentCan the organization safely buy and deploy?Security, legal, procurement, budget, implementation, and stakeholder commitmentsThe deployment plan, commercial path, and mutual responsibilities are explicit
    ActivationCan the intended users complete the valuable workflow?Configuration, integrations, access, enablement, and a completed first-value eventThe onboarding exit criteria are met rather than merely scheduled
    Value realizationIs sustained product behavior producing the promised result?Adoption depth, outcome movement, executive validation, and owned remediation for gapsProgress is accepted by the process owner and remaining risks have owners
    Renewal and expansionIs continued or broader investment justified?Realized value, renewal intent, sponsor engagement, and evidence for an adjacent use caseThe customer makes a clear renewal, expansion, or stop decision

    Use this as a lifecycle contract across functions. For every stage, assign one directly accountable owner and name the person who receives the account at the next stage. Collaboration can be shared; accountability cannot. Customer success should not discover the value promise after the contract is signed, and product should not first learn about a deployment blocker through an escalation.

    The proof should also change by segment. A self-serve motion can lead with fast activation and transparent packaging. An enterprise motion must add trust, workflow integration, executive relevance, and a navigable buying process. Treating enterprise as a larger pricing tier leaves the hardest parts of the customer decision unowned.

    Apply the same discipline before adding sales capacity. A clear ideal customer profile, a supported value hypothesis, and a repeatable early motion should exist before you scale coverage. If the narrative, proof points, and qualification rules cannot fit into a concise operating brief, more sellers will multiply ambiguity rather than revenue.

    Qualify company value and map how the account buys

    User enthusiasm is a useful signal, but it is not an enterprise business case. A product can be loved by individuals while remaining easy for an executive to cut. Your qualification process must translate user value into company value before an opportunity receives expensive sales, solutions, and product attention.

    Build that translation as a value chain:

    • Business objective: the result an executive or process owner is accountable for.
    • Operational change: the behavior, decision, or workflow that must improve.
    • Product behavior: the specific capability and usage pattern that enables the change.
    • Measured result: the business or operational signal that will confirm progress.

    Consider an AI product used in customer support. Generated summaries, drafted responses, and active users describe product activity. Company value appears when those behaviors contribute to faster resolution, deflection, conversion, lower cycle time, or a better customer experience while meeting the required quality and risk bar. The exact outcome depends on the customer’s objective, but the distinction does not: activity is evidence only when you can connect it to a result.

    A qualified enterprise opportunity should contain more than a large logo and an interested user. Record the following before you commit significant resources:

    • The costly or strategically important problem, including how it appears in the current workflow.
    • The person who owns that process and the objective attached to it.
    • The result the customer expects and the signal that will be used to judge it.
    • The internal champion who will mobilize people, information, and decisions.
    • The executive sponsor or a concrete path to one.
    • The data, integration, security, governance, and implementation conditions.
    • The economic buyer, budget path, procurement steps, and commercial timing.
    • The reason the organization should act rather than leave the problem in place.

    Do not confuse an enthusiastic user with a champion. A real champion feels the problem, has credibility in the organization, helps you navigate resistance, and can explain the business case when you are not in the room. That last test matters. If the opportunity depends on your seller retelling the value story at every internal meeting, you have interest but not mobilization.

    Map the buying system as a champion tree rather than a flat contact list. Include the operator who lives with the problem, the manager who owns the workflow, the executive who can protect the priority, the technical owner who must trust the deployment, and the legal or procurement stakeholder who controls the path to purchase. One person may cover several roles early in a deal, but the roles usually separate as the commitment grows.

    For each stakeholder, record the desired outcome, perceived risk, evidence needed, and next commitment. This changes account planning from contact collection into decision orchestration. It also reveals dangerous gaps early: a strong operator champion with no executive access, an executive sponsor with no frontline adoption, or a business case that ignores the security owner.

    Use deal reviews to inspect missing evidence, not to hear a chronological update. Ask what company outcome is being purchased, who owns it, what has been proven, which stakeholder can still stop the decision, and what commitment should happen next. If sales, product, and customer success give different answers, repair the shared record before advancing the stage.

    Turn evaluation into an enterprise decision, not a science project

    An open-ended pilot is one of the most expensive forms of false progress. Users experiment, the vendor supplies support, and both sides collect impressions without defining the decision that the work is meant to unlock. Activity rises while commercial certainty stays flat.

    Write the evaluation brief before the customer receives access. It should include:

    • The workflow included in the evaluation and the work explicitly excluded.
    • The current baseline or a documented description of the existing state.
    • The target result, quality bar, and decision rule.
    • The participating users, process owner, executive sponsor, and technical owner.
    • The required data, integrations, permissions, and operating environment.
    • The instrumentation and review method that will produce credible evidence.
    • The security, legal, governance, and change-management conditions.
    • The decision date, decision-makers, and possible outcomes.
    • The implementation and commercial path if the evaluation passes.

    For a tightly scoped enterprise AI workflow, I prefer success criteria that make value visible in under 30 days when the data, security, and integration path make that credible. The point is not to force an artificial clock onto a complex deployment. It is to constrain the workflow enough that the customer can learn something decisive before the evaluation loses executive attention.

    AI evaluations also need technical evidence that ordinary feature demonstrations do not provide. Build an evaluation harness around representative or gold data, task-specific measures, and human adjudication. Track quality alongside cost and latency. Document failure cases, guardrails, policy enforcement, and the points where human review is required. A polished demo on curated inputs does not prove dependable operation across messy enterprise data.

    Trust requirements belong in the product and evaluation plan. Give direct answers about whether customer data trains models, what retention controls apply, where data resides, how it is encrypted, and which access controls are supported. Validate the account’s actual needs for SSO, RBAC, DLP, auditability, private networking, or a VPC rather than dropping every possible control into a generic checklist. Each requirement should have an owner and a disposition: supported, planned, handled through an approved alternative, or blocking.

    Run business validation, technical validation, and the buying process in parallel. The best-performing workflow is still unbuyable if security starts after the pilot, procurement has no contracting route, or implementation depends on an integration that nobody scoped. In regulated or public-sector environments, accreditation, interoperability, funding gates, and acquisition timing may be product constraints in their own right. Surface them before the prototype becomes politically successful but operationally stranded.

    A completed evaluation should produce evidence plus a recorded decision: proceed, proceed after named conditions are resolved, or stop. Do not accept indefinite testing as a fourth state. Continued evaluation needs a new hypothesis, a specific missing data point, an owner, and another decision point. Otherwise the pilot is absorbing resources without reducing uncertainty.

    Protect the roadmap during this process. Classify requested work as essential to the core use case, account-specific configuration, or evidence of a repeatable segment need. A prominent prospect’s willingness to ask does not make a request strategic. The product should bend when the learning strengthens the chosen enterprise wedge, not merely because a deal is visible.

    Use the same caution with discounts. A lower price can accelerate paper while concealing an unclear value case, weak qualification, or unbounded services burden. If the customer cannot explain why the outcome is worth funding, discounting changes the amount under debate without resolving the reason for buying.

    Make contract signature the midpoint of the value journey

    Closed-won is a commercial milestone, not customer success. Treating it as the finish line creates a predictable reset: the seller celebrates, the implementation team asks discovery questions again, the customer repeats context, and the time-to-value clock starts while everyone reconstructs commitments.

    A handoff is not a meeting. It is the transfer of context, promises, evidence, and accountability. For complex deployments, the implementation or customer success owner should validate the success plan before signature. The shared account record should contain:

    • The customer’s business objective, use case, baseline, and agreed result.
    • The stakeholder map, including the champion, executive sponsor, process owner, and technical owner.
    • The claims already proven and the assumptions that remain open.
    • The contractual promises, product dependencies, integrations, and security conditions.
    • The onboarding milestones and explicit exit criteria.
    • The value-review cadence and the person accountable for renewal.
    • The known risks, mitigation action, owner, and next decision.

    Define onboarding by customer capability rather than vendor activity. A kickoff call, training session, or configured account is an output. The customer exits onboarding when the intended people can perform the valuable workflow, the necessary data and integrations function, administrators can operate the deployment, and the first meaningful value event has occurred. If those conditions are not met, onboarding is still open even when the project plan says complete.

    Instrument the account in layers

    A single health score often hides more than it reveals. Track distinct evidence layers so the team can diagnose what is actually weak:

    • Product evidence: activation, breadth and depth of adoption, and use of the workflows that create value.
    • Outcome evidence: movement in the operational or business result named in the success plan.
    • Relationship evidence: champion strength, executive engagement, and access to the process owner.
    • Delivery evidence: implementation progress, unresolved dependencies, support patterns, and configuration risk.
    • Commercial evidence: renewal intent, procurement readiness, contract timing, and credible expansion demand.

    Time-to-first-value, activation, usage depth, implementation milestones, executive engagement, product-qualified account signals, and renewal intent are leading evidence. Gross retention, net revenue retention, and expansion revenue confirm what already happened. NRR can tell you that the lifecycle produced a result; it cannot tell you where the lifecycle is breaking. The preceding evidence can.

    A green usage dashboard should not overrule a missing sponsor or an unproven outcome. The reverse is also true: an executive relationship cannot compensate indefinitely for weak adoption. Durable accounts align product behavior, business value, and organizational support.

    Use repeated patterns to distinguish a product problem from an account-specific success problem. If the same friction appears across a segment, workflow, or cohort, bring product the user journey, supporting evidence, and a prioritized hypothesis. If the issue is isolated to one configuration, stakeholder group, or deployment, address enablement, implementation, and alignment first. This keeps customer success from masking systemic product debt and keeps the roadmap from absorbing every local exception.

    Use an operating rhythm that forces decisions

    Weekly risk reviews should focus on the risk hypothesis, supporting signal, intervention, owner, and next check. A list of red accounts without an intervention model is status reporting, not risk management.

    Value-realization reviews and QBRs should compare the current result with the success plan, explain which product behaviors contributed, identify barriers, and secure the next customer decision. Do not fill the meeting with feature activity that the executive cannot connect to an objective. The customer should leave knowing what changed, what remains uncertain, and what each side will do next.

    Feed the same evidence into product discovery. A disciplined Voice of Customer readout should separate recurring value drivers, repeatable friction, account-specific requests, and emerging adjacent use cases. Close the loop with customers on what changed, what will not change, and why. That improves candor while preventing the roadmap from becoming a collection of unresolved promises.

    Renewal ownership should match the business model, but it must be unambiguous. In a complex, value-expansive deployment, customer success can own or co-own renewal because it orchestrates value realization and risk. In a more transactional or quota-led motion, sales may own the commercial paper while customer success owns health and expansion signals. Either model can work. Split accountability and conflicting incentives usually cannot.

    Expansion begins only after the original wedge has earned credibility. Check that onboarding is complete, the valuable behavior has become a habit, the process owner accepts the result, the sponsor remains engaged, and the adjacent use case has its own owner and reason to act. Expansion is a new value case, not an administrative upgrade.

    Sequence expansion deliberately: deepen the critical workflow, extend it to adjacent teams or processes where the proof transfers, and broaden the product surface only after repeatability appears. Building horizontally too early dilutes the use case that created executive attention in the first place.

    Key takeaways

    • Design enterprise GTM backward from renewal. Every stage should specify the customer decision, required evidence, accountable owner, and exit criterion.
    • Translate user value into company value through a visible chain from business objective to workflow change, product behavior, and measured result.
    • Qualify the buying system as well as the use case. A champion, executive path, technical trust owner, procurement route, and reason to act are part of the opportunity.
    • Define evaluations around a decision. Baselines, success criteria, instrumentation, security, implementation, and the commercial path belong in the pilot brief.
    • Treat contract signature as the midpoint. Onboarding, value realization, renewal, and expansion should continue the same success plan rather than restart discovery.
    • Use leading evidence to manage the account before retention metrics report the outcome. Product usage alone is not proof of business value.

    Take one active enterprise account and walk it through the lifecycle table. Wherever you cannot name the evidence, exit criterion, or owner, you have found the break in your GTM system. Fix that break before adding another acquisition channel, process layer, or tranche of headcount. A coherent customer lifecycle turns enterprise growth from a series of rescues into a repeatable value-delivery discipline.

    References

    • Shivam.Consulting Blog – The New PLG Playbook: Avoid the Trap, Win Enterprise, and Break the $10B Ceiling
    • Shivam.Consulting Blog – From Zero to One: My Playbook for Building a World-Class Sales Org (Lessons from Figma)
    • Shivam.Consulting Blog – Customer Success Masterclass: How I Design, Build, and Scale a World-Class CS Org
    • Shivam.Consulting Blog – Scaling Enterprise AI That Sells: Battle-Tested Playbooks for PMF, Champions, and Agentic AI
    • Shivam.Consulting Blog – A Masterclass in Founder Conviction: Gong’s $100m ARR, PMF Breakthroughs, and AI Sales
    • Shivam.Consulting Blog – Inside Stripe, OpenAI, Retool: Hard-Won Marketing Lessons on Brand, GTM, and Scale
    • Shivam.Consulting Blog – From Prototype to the Pentagon: My Playbook for Winning DoD Customers and Mission Fit
  • From Product-Market Fit to Expansion: A Sequencing Model

    From Product-Market Fit to Expansion: A Sequencing Model

    Your core users are staying, power users are asking for adjacent workflows, and sales wants a broader story. Expansion now feels inevitable. The risk is that visible demand can come from a few enthusiastic accounts while the underlying product-market fit is still narrow, manual, or fragile.

    The decision is not simply whether to expand. You need to know what created the fit you have, which expansion model preserves that mechanism, and what evidence must appear before the new bet earns more capital. The safest next move is the shortest one that increases customer value without weakening the reason your core users chose you.

    Define the product-market fit you actually have

    Product-market fit does not belong to a company in the abstract. It exists within a specific combination of customer, job, value moment, product experience, price, and distribution motion. A product can have strong fit with one customer archetype and weak fit everywhere else. It can also retain users for one job while an apparently similar use case fails.

    Before discussing expansion, write a one-sentence fit contract:

    For [specific customer], when [trigger occurs], the product completes [important job], produces [observable outcome], and becomes part of [repeat behavior or workflow].

    That sentence forces several useful distinctions. The customer cannot be "SMBs" if the successful users are independent dental practices with a particular workflow. The job cannot be "grow revenue" if the product actually helps a sales manager build and launch an outbound campaign. The outcome cannot be "save time" unless you can identify what gets completed faster and what users do with that advantage.

    Then test the contract against behavior, not enthusiasm:

    • New users can move from setup to a first successful workflow without extraordinary intervention.
    • The same job produces repeat use in successive cohorts, rather than one burst of exploration.
    • Retention is concentrated in the customer archetype named in the contract.
    • Users tolerate some incidental friction because the core outcome is important enough to preserve.
    • Account expansion begins with usage, collaboration, or workflow depth rather than a discount engineered to inflate seat count.
    • The support burden and sales-assist requirement do not rise every time another customer adopts the core use case.

    I find it useful to label the evidence as observed, repeated, or scalable. Observed fit means a small group has found value, often with manual help. Repeated fit means several cohorts reach and repeat the same value moment. Scalable fit means that pattern survives as onboarding, selling, and support become less dependent on heroic effort. The farther an expansion moves from the original customer and job, the stronger this evidence needs to be.

    This framing also catches PMF decay. If time-to-first-value lengthens, retention weakens for the original job, or support work accumulates around the core workflow, expansion should not become a distraction from repairing the wedge. Product-market fit can change when customer behavior, infrastructure, regulation, distribution, or an underlying platform changes. Treat the fit contract as a living operating claim, not a permanent certificate.

    Make every expansion proposal pass the same gates

    An expansion idea deserves roadmap capacity only when it can answer a consistent set of questions. This prevents a large prospect, an executive preference, or an attractive total addressable market from bypassing the evidence required of every other product bet.

    1. Core gate: Which retained customer cohort and repeat job prove the current wedge? If the team cannot identify them, the immediate task is segmentation and discovery.
    2. Pull gate: What customer behavior reveals the boundary of the current product? Look for repeated workarounds, exports, manual handoffs, integration activity, invited collaborators, and adjacent tools customers already pay for.
    3. Continuity gate: Does the expansion make the existing promise faster, clearer, or more complete? If it creates a separate value proposition, acknowledge that you are considering a new product rather than pretending it is a feature.
    4. Delivery gate: Can the new cohort reach value without adding disproportionate implementation, support, compliance, or sales work? Demand that depends on bespoke service may be real, but it is not yet evidence of scalable product fit.
    5. Distribution gate: Is the user, buyer, budget, channel, and buying moment still the same? A change across several of these dimensions is a new go-to-market problem even when the software looks adjacent.
    6. Protection gate: Which core metrics must not regress, and what result will stop the bet? Name the guardrails before building so the team does not reinterpret weak evidence after launch.

    Put those answers in a one-page expansion contract. It should name the target cohort, unmet job, expected value moment, leading behavioral signal, core guardrails, owner, checkpoint, and stop-or-scale rule. A two-to-four-week discovery or prototype sprint is a useful decision cadence for a bounded hypothesis. It is not a deadline by which product-market fit must appear. The sprint should end with a sharper decision, not an automatically enlarged backlog.

    A good stop rule is observable and comparative. For example: pause if the new workflow increases support load while failing to produce repeat use, or if simplifying the experience for a new segment lengthens time-to-value for the retained core. You do not need a universal industry threshold. You need a baseline from your own successful cohort and a clear statement of how much deterioration the business is prepared to accept.

    Choose the expansion model that matches the source of pull

    Expansion is often discussed as if every move were the same. It is not. Each model changes different assumptions and should be validated with different evidence.

    Expansion modelWhat changesUse it whenFirst proof to seekMain failure mode
    Deepen the wedgeMore capability for the same customer and jobRetained users repeat the job but still encounter friction or manual stepsFaster value, more completed workflows, or stronger repeat useAdding options that make the core harder to learn
    Adjacent workflowA job immediately before, during, or after the wedgeThe same handoff or workaround appears across retained accountsUsers adopt the adjacency and continue through the combined workflowBuilding a generic suite of loosely connected features
    Team or account expansionMore roles use the product inside the same customerAn individual’s successful output naturally needs to be shared, reviewed, or reusedOrganic invitations, collaboration, and team-level repeat behaviorAdministrative complexity arriving before collaborative value
    ICP or vertical expansionA new customer segment applies the product to a similar jobThe pain and value mechanism remain stable with limited adaptationThe new cohort begins to approach the core cohort’s activation and retention patternRemoving useful specificity until the product fits nobody well
    New product or SKUA distinct job, value promise, or premium momentExisting customers show repeated pull and the business has shared distribution, identity, or data advantagesStandalone activation plus credible cross-adoption from the coreA bundle concealing weak fit in the new product
    Platform or ecosystemPartners, developers, or customers create value for other participantsIntegration and contribution points already behave like growth or retention nodesThird-party creation increases utility, distribution, or switching value for customersShipping APIs without a participant incentive or value flywheel
    Marketplace cell expansionA new geography, category, or supply-demand clusterThe original cell has reliable liquidity, retained supply, and consistent fulfillmentShort time-to-transaction, repeat activity, and maintained service quality in the new cellFragmenting density before either side has enough reliable choice

    Horizontal expansion should follow the customer workflow

    To find a useful adjacency, map what happens immediately before, during, and after the core job. Favor a move that removes an expensive handoff, compounds a data advantage, or makes the successful workflow easier to repeat. This is more reliable than starting with a broad suite vision and searching for features to fill it.

    Make one connection coherent before stacking another. If users must re-enter data, learn unrelated concepts, or navigate a different product language at each step, you have expanded the feature count without expanding the value system. A strong adjacency makes the original wedge feel more complete.

    Vertical expansion requires fresh discovery

    A nearby industry may appear to have the same problem while differing in workflow, terminology, regulation, implementation, buyer authority, or service expectations. Keep the new segment separate in your analytics and discovery. Do not blend its early usage with the retained core and declare success from the average.

    The market type also matters. Entering an established category with a focused wedge calls for a sharp differentiation and a credible switching path. Creating a new category requires education, use-case sequencing, and a distribution story that helps buyers understand why the behavior should change at all. Reusing one go-to-market playbook across those conditions can make a sound product look weak.

    Marketplace expansion resets liquidity locally

    A marketplace that works in one city or category has not automatically solved the next one. Treat each new cell as a constrained cold start. Protect supply quality, responsiveness, price clarity, trust, and time-to-first-transaction before opening another front.

    Use capacity to decide which side to grow. When retained supply is underused, add qualified demand. When supply is constrained or fulfillment quality is deteriorating, deepen supply before accelerating buyers. Category and geographic expansion should improve marketplace health, not merely increase the number of listings or registered users.

    Protect the wedge with a portfolio and stage gates

    Expansion fails as often through resource allocation as through product judgment. The core quietly loses quality while every ambitious initiative is described as strategic. A practical starting allocation is 70% of capacity on core commitments, 20% on accelerants and adjacencies, and 10% on bolder experiments. Treat that as a portfolio prompt, not a universal benchmark. The right mix depends on the health of the wedge and the cost of the bets.

    The same portfolio can be viewed through three horizons. Horizon 1 protects retention, reliability, activation, and speed in the wedge. Horizon 2 validates adjacencies that deepen customer value. Horizon 3 creates options around new products, platforms, or market shifts. Horizon 3 should be time-boxed and stage-gated so an exciting possibility cannot consume the resources needed to maintain current fit.

    Move each expansion through a visible sequence:

    1. Discover demand: Identify repeated workflow boundaries, workarounds, integration patterns, and buying signals among retained customers.
    2. Prove the value moment: Use a prototype or private beta with power users to test whether the new job produces an outcome worth repeating.
    3. Validate a cohort: Measure activation, repeat behavior, willingness to pay, support burden, and retention separately for the target segment.
    4. Prove distribution: Confirm that the product can acquire, onboard, and serve the new cohort without relying indefinitely on founder attention or bespoke sales work.
    5. Scale or stop: Increase investment only when the expansion passes its behavioral and core-protection gates. Otherwise, narrow, redesign, or end it.

    Power users are excellent scouts because they expose advanced workflows, integration needs, reusable templates, and emerging use cases. They are not automatically a representative market. After co-designing with them, test whether a less advanced customer can understand the promise, reach value, and repeat the workflow without adopting the power user’s entire operating system.

    Record the baseline before the beta starts. Your expansion scorecard should show:

    • Time-to-first-value for the target cohort compared with the successful core cohort.
    • Completion of the first meaningful workflow, not account creation or feature clicks.
    • Repeat usage and retention segmented by job-to-be-done.
    • Organic invitations, shared artifacts, integrations, or other product behaviors that can create distribution.
    • Support tax, implementation effort, and sales-assist ratio.
    • Core activation, retention, reliability, and customer experience as explicit guardrails.
    • Evidence that customers will pay for the added value without a discount masking weak adoption.

    Do not let a blended top-line metric make the decision. Growth in a new cohort can conceal deterioration in the original one, while healthy core retention can conceal a failed adjacency. Keep cohort views side by side until the new motion is independently repeatable.

    The product narrative is another diagnostic. Each expansion should read like the next chapter of the same customer story: a clear problem, a visible before-and-after outcome, and a believable connection to the wedge. If sales needs a different explanation for every module, the portfolio may be a collection of products rather than a coherent platform. That can still be a valid strategy, but it requires explicit product, pricing, and go-to-market choices.

    Finally, maintain a watchlist of external assumptions. Platform changes, privacy rules, AI infrastructure, distribution shifts, and ecosystem consolidation can absorb a feature’s value or create a better expansion path. When one of those assumptions changes, revisit the fit contract before defending the existing roadmap.

    Key takeaways

    • Define PMF for a specific customer, job, outcome, and repeat behavior. Company-wide labels are too broad to guide expansion.
    • Expand the mechanism that created retention, not merely the surface area of the product.
    • Choose among wedge depth, workflow adjacency, team adoption, vertical expansion, a new product, a platform, or a marketplace cell based on observed customer behavior.
    • Keep new cohorts separate from the core so aggregate metrics cannot hide weak fit or core deterioration.
    • Agree on core guardrails and stop rules before building. A kill decision made after launch is easy to rationalize away.
    • Scale only after value, retention, delivery, and distribution repeat without extraordinary intervention.

    At your next planning review, take the highest-priority expansion request and complete three artifacts: the fit contract, the expansion-model row, and the scorecard with a baseline and stop rule. If you cannot fill one in, the next roadmap item is not the expansion. It is the smallest experiment that resolves the missing evidence.

    References

    • Shivam.Consulting Blog — Master Modern Entrepreneurship: Build Lean, Start Young, and Obsess Over Customers
    • Shivam.Consulting Blog — From Vertical Focus to Power Users: My Playbook for Product-Market Fit and Founder Mindset
    • Shivam.Consulting Blog — How to Find Your Product Wedge: Battle-Tested SMB SaaS Lessons from Square, Gusto, and My Playbook
    • Shivam.Consulting Blog — Build Platforms, Not Apps: My Playbook to Delight Customers and Scale Product Strategy
    • Shivam.Consulting Blog — How I Build and Scale Winning Marketplaces: Demand, Supply, PMF, and Growth Loops
    • Shivam.Consulting Blog — How I Find—and Keep—Product-Market Fit: Lessons on Conviction, Distribution, and Mergers
    • Shivam.Consulting Blog — Inside Figma’s Product Playbook: Taste, Simplicity, and Storytelling for Extraordinary PMs
  • What Makes or Breaks Executive Hires: My Lessons on Fit, Red Flags, and Measuring Success

    What Makes or Breaks Executive Hires: My Lessons on Fit, Red Flags, and Measuring Success

    Executive hiring is one of those rare decisions that can bend a company’s trajectory. In my role leading product management at a high-growth SaaS company, I’ve seen the difference between a leader who compounds value and one who quietly drains momentum. That’s why I was eager to examine what actually makes (or breaks) these bets, and to share a practical lens you can use to improve executive hiring outcomes.

    I sat down with Eeke de Milliano for a focused conversation on the realities of executive hiring, leadership transitions, and measuring success. We dig into the “buy or build a leader” decision, how to avoid common red flags, and what it takes to set executives up to thrive in hyper-growth environments.

    Eeke de Milliano is the Head of Global Product at Stripe, helping drive innovation and success in the company’s product line. Before this role, she was Head of Product at Retool and co-founded Constellate. Eeke previously spent 6 years as Product Lead at Stripe, working with the company during their hyper-growth era.

    In today’s episode, we discuss how to rigorously assess executive hiring fit, including the challenges companies face when hiring new executives and the most common red flags and pitfalls I see teams miss under time pressure. We also explore practical advice for measuring success, especially when outcomes vs output get muddled in the first 90–180 days.

    A recurring theme for me is that learning your own strengths is an underrated piece of the process. If you don’t understand the leadership leverage you already have on the team, you’ll over-hire for breadth or under-hire for depth. Great executive hiring clarifies the complementary edge you need—then measures it.

    On the buy vs build decision: early signals matter. If you’re “buying” an external leader, pre-align on scope, authority, and what great looks like before day one. If you’re “building” from within, design a clear on-ramp and operating cadence so the leader can scale without drowning. In both cases, my mental model is to instrument leading indicators (team health, decision velocity, stakeholder trust) well before lagging business metrics fully show up.

    Two red flags I always watch for: first, leaders who default to playbooks without interrogating context; second, leaders who cannot articulate how they measure success beyond activity and output. In hyper-growth, pattern-matching is useful—but uncalibrated pattern-matching is dangerous.

    The human dynamics matter just as much as the strategy. What creates dysfunctional exec relationships is often misaligned interfaces: unclear decision rights, overlapping charters, or incentives that reward local maxima. High-functioning executive teams are like parents—a united front in public, with candid debate in private, anchored to shared principles and measurable outcomes.

    Referenced:

    ASML: https://www.asml.com/en

    Claire Hughes Johnson: https://www.linkedin.com/in/claire-hughes-johnson-7058/

    Constellate: https://constellate.team/

    John Collison: https://www.linkedin.com/in/johnbcollison/

    Mike Maples Jr.: https://www.linkedin.com/in/maples/

    Patrick Collison: https://www.linkedin.com/in/patrickcollison/

    Retool: https://retool.com/

    Stripe: https://stripe.com/

    Will Gaybrik: https://www.linkedin.com/in/william-gaybrick-5730347/

    Where to find Eeke:

    LinkedIn: https://www.linkedin.com/in/eeke-de-milliano-3b05a629/

    Timestamps:

    (00:00) Should you ‘buy or build’ a leader

    (03:45) Why do executive hires fail so often?

    (09:35) Why the stakes are so high for leadership hires

    (12:26) The hardest document Eeke ever wrote

    (14:06) Two red flags in a new hire

    (17:27) An example of an outstanding leader

    (21:40) What creates dysfunctional exec relationships

    (22:38) The three steps towards hiring successful leaders

    (30:30) What you should know about outside hires

    (33:12) Eeke’s advice for easing leadership transitions

    (42:06) How to notice success patterns

    (47:21) Why high-functioning executive teams are like parents

    (52:02) The most surprising lesson from Eeke’s first stint at Stripe

    (55:11) The leadership data Eeke wishes we had


    Book a consult png image
  • How to Create a B2B Category and Scale the GTM Motion

    How to Create a B2B Category and Scale the GTM Motion

    You know you have a category problem when prospects understand the product but still place it in the wrong budget, compare it with the wrong alternatives, or evaluate it against criteria that hide its value. Sales asks for a sharper pitch, marketing proposes a new label, and product adds comparison features. None of those moves fixes the missing buying logic.

    Your job is not to make a new noun famous. It is to help a specific buyer recognize an important change, adopt a better way of working, experience credible proof, and pay for the organizational capabilities that make the new practice safe at scale. The sequence matters: problem clarity before category language, practitioner value before enterprise packaging, and repeatable proof before GTM headcount.

    First decide whether you need a category or better positioning

    Category creation is expensive because you must teach the buyer what changed, why the old approach is inadequate, how the new approach works, and why your product is a credible way to adopt it. A positioning change is narrower. The buyer already understands the problem and budget; you need to show why your approach is the better choice.

    Do not choose category creation because the existing market feels crowded. Choose it only when the existing buying frame actively distorts the value of the product.

    Decision areaYou probably have a positioning problemYou may have a category problem
    Buyer languageBuyers consistently use an established term for the problem.Different buyers describe the same underlying problem with unrelated terms.
    Budget and ownershipA known function owns the budget and buying process.The pain crosses functions, and no established budget fully represents the value.
    Evaluation criteriaExisting criteria expose your differentiation.Existing criteria reduce the product to a misleading feature comparison.
    Behavior changeThe product improves a familiar workflow.The product requires a materially different operating practice.
    Market educationYou mainly need to explain why you are better.You first need to explain why the old way has become insufficient.

    If most evidence lands in the positioning column, resist the temptation to invent a category. Attach yourself to the budget and vocabulary buyers already use, then sharpen the value proposition. If the category column dominates, write a category thesis before you spend on campaigns, analysts, events, or a larger sales team.

    A useful category thesis fits on one page and answers six questions:

    1. What changed in the buyer’s world?
    2. What costly problem does that change create or expose?
    3. Why do established tools or practices handle it poorly?
    4. What new operating principle should replace the old one?
    5. What narrow product experience proves that principle?
    6. What additional value appears when a team or enterprise adopts it broadly?

    Write the thesis without your product name first. If it only makes sense after the brand and feature list are inserted, you have a campaign concept, not a durable market thesis. The strongest category narratives can be taught by a practitioner who has never met your marketing team.

    Then test the thesis against recent opportunities. Look for repeated triggers, failed alternatives, unexpected budget owners, and evaluation criteria that force the wrong comparison. Category creation becomes credible when the same market misunderstanding appears across unrelated accounts. One enthusiastic customer using novel language is a clue, not a market.

    Create a practitioner wedge before an executive narrative

    A B2B category becomes real through behavior before it becomes real through branding. Someone must be able to use the product, get a result, and explain the new practice to a colleague. If adoption depends on an executive accepting the whole category thesis before a practitioner can experience value, the education burden will overwhelm the GTM motion.

    dbt Labs built around an opinionated way for analysts and engineers to work, then reinforced that practice through consulting, open source, and community. Its path ran from three companies using the free tool in 2016 to an ecosystem described as having more than 30,000 enterprise users. The important mechanism was not free distribution by itself. Practitioners could adopt a concrete workflow, improve it together, and advocate for a recognizable standard inside their organizations.

    Clay found early traction in WhatsApp groups and Reddit threads, where operators were already exchanging tactics. Reverse demos made the product useful in the prospect’s workflow instead of asking the prospect to admire a polished feature tour. 1Password found its B2B opening in team adoption patterns around a product people already trusted individually. In each case, observable usage carried more information than an abstract category claim.

    Design the wedge as an adoption ladder:

    1. Individual utility: one practitioner can solve a narrow, painful problem without organizational change.
    2. Visible artifact: the work produces something that can be shared, reviewed, reused, or handed to another person.
    3. Team consistency: collaboration creates demand for common workflows, permissions, templates, or quality controls.
    4. Organizational control: scale creates requirements around governance, administration, reliability, security, and support.
    5. Enterprise expansion: more teams, workflows, data, or regions increase value without changing the original reason for adoption.

    The ladder tells product and GTM where each kind of value belongs. The first step should be easy to experience. The middle steps should make collaboration better. The final steps should make broad adoption manageable. If the first meaningful result only appears after procurement, integration, and an executive rollout, you have made the hardest part of the sale precede the proof.

    Treat community as product instrumentation

    A community is useful when it improves the practice around the product and exposes where that practice breaks. A large member count without recurring practitioner exchange is an audience metric, not a category advantage.

    Choose one primary practitioner environment rather than opening several neglected channels. Each week, review a fixed sample of the highest-signal discussions and classify them as:

    • a repeated job the product handles well;
    • a terminology problem that weakens onboarding or positioning;
    • a workaround that may reveal a missing primitive;
    • a team-level requirement emerging from individual adoption;
    • an enterprise blocker involving control, integration, reliability, or support; or
    • a successful workflow that can become a template, tutorial, or proof asset.

    Send those patterns into one shared product and GTM review. Do not let marketing extract only success stories while product sees only requests and support sees only failures. The combined record is the living map of how the category is being understood and adopted.

    Use services as paid discovery, with an exit condition

    Early consulting and implementation work can reveal the customer’s real workflow faster than detached roadmap research. It also creates a dangerous incentive: every account can look strategically important when it is paying for custom work.

    For each engagement, record the customer’s job, existing alternative, required inputs, workflow changes, blockers, successful output, and reusable elements. Productize a pattern only after it appears across unrelated customers and fits the category thesis. Keep truly account-specific work in services, price it transparently, and do not disguise it as a platform capability.

    The exit condition matters. A service should eventually become a repeatable product workflow, a standardized implementation package, or an explicit premium service. If it remains an open-ended collection of exceptions, it is not accelerating category creation; it is concealing the absence of a scalable product.

    Build the GTM system around proof, then price what compounds

    Replace feature demos with customer-specific proof events

    A conventional demo answers a seller’s question: which capabilities should I show? A proof event answers the buyer’s question: can this work in my environment, for my job, with an outcome I recognize?

    Define one primary proof event for the initial ICP. It should specify:

    • the job the buyer is trying to complete;
    • a representative input from the buyer’s real workflow;
    • the person who should operate or validate the product;
    • the observable output that demonstrates value;
    • the success criteria agreed before the session;
    • the assumptions that remain unproven; and
    • the next organizational dependency, such as integration, governance, rollout, or procurement.

    Keep proof honest. A technical result is not automatically a workflow result, and a successful workflow is not automatically an enterprise business case. Use four distinct levels:

    1. Technical proof: the mechanism works with relevant inputs.
    2. Workflow proof: a practitioner can incorporate the result into real work.
    3. Organizational proof: a team can adopt it with acceptable control, reliability, and effort.
    4. Economic proof: the value is important enough to fund, renew, and expand.

    Do not let sales present level one as if level four has been established. Record which layer each opportunity has actually reached. This makes forecast reviews more useful and tells product whether a stalled deal needs a better core experience, a missing enterprise capability, stronger implementation, or a clearer economic case.

    A proof motion is ready to scale when a person who did not invent it can reproduce the result for the same ICP using a documented input checklist, success criteria, and follow-up path. Until then, adding sellers multiplies variation rather than revenue.

    Place the commercial boundary where organizational value starts

    The low-friction product should spread the practice. The paid product should help customers coordinate, control, and extend that practice. This is why capabilities such as permissions, governance, collaboration, administration, and scale can support monetization without crippling practitioner adoption.

    The pricing unit must also match something the customer can understand before receiving the bill. Clay chose a credit model rather than exposing customers directly to raw usage. Credits can make a variable underlying workload easier to budget when customers consume distinct units across several workflows. Seat pricing is clearer when each additional user receives durable value. A platform subscription can fit organization-wide capabilities whose value is not attributable to individual users. Services should be charged separately when the work is genuinely bespoke.

    Answer these questions before choosing or changing the unit:

    • Which adoption behavior must the pricing model preserve?
    • What unit grows when customer value grows?
    • Can the buyer predict and explain the bill before purchase?
    • Does the unit encourage healthy use, or make customers suppress the behavior that creates value?
    • Which enterprise obligations are included in the price?
    • What causes expansion: more people, more workflows, more volume, more control, or a combination?

    Changing a pricing unit after customers build operating processes around it can damage trust and make budgets unpredictable. Before launch, replay the proposed model against representative historical account usage, inspect the outliers, and write the migration policy. A mathematically elegant model is still wrong if customers cannot forecast it or sales cannot explain it.

    Delaying billing can be useful when it is an intentional validation step. 1Password launched its SaaS platform before billing, allowing adoption to generate evidence for pricing and migration decisions. That sequencing only works when you instrument engagement, define the future value boundary, communicate the transition clearly, and preserve a clean migration path. Free usage without a monetization hypothesis is not validation; it is deferred ambiguity.

    Once the proof and price boundary are stable, layer the GTM roles deliberately:

    • Self-serve product: helps practitioners discover the product and reach the first proof event.
    • Solutions engineering: handles technical validation, integrations, and complex environments without turning every request into roadmap work.
    • Sales: establishes the buying process, economic case, stakeholder alignment, and commercial terms.
    • Customer success: turns an initial purchase into adopted workflows, measurable outcomes, and expansion.

    Enterprise sales can amplify a working adoption loop. It cannot manufacture one. If every deal needs founder persuasion, a custom demo, a new integration, and a unique value story, the motion is still discovery.

    Scale upmarket and globally without severing customer signal

    The danger in scaling GTM is not merely higher cost. It is signal distortion. Large opportunities generate urgent requests, sales teams optimize for the current quarter, and the roadmap drifts toward the loudest account. Meanwhile, the practitioner wedge that created the category becomes slower and harder to adopt.

    Clay layered enterprise customers on top of a functioning product-led engine. Braze invested in platform primitives that could support real-time engagement and a global customer base. 1Password had to preserve usability while meeting the security and administrative expectations of businesses. These motions work when enterprise capability strengthens the core adoption path instead of replacing it.

    Run two connected operating lanes:

    • Core product lane: owns the standardized workflow, activation, platform primitives, APIs, reliability, and the capabilities needed across customers.
    • Field learning lane: uses solutions engineers, forward-deployed talent, product leaders, and customer success to solve high-signal complexity in important accounts.

    Every field request should carry the underlying job, affected persona, current workaround, business consequence, frequency across accounts, reusable product principle, and strategic fit. An account name and contract value are not sufficient product requirements. Promote a request into the core roadmap when it solves a repeatable constraint without weakening the category thesis or the standard product experience.

    Separate enterprise readiness into four layers so teams can see what is actually blocking growth:

    1. Product readiness: administration, permissions, provisioning, auditability, data controls, reliability, and integration.
    2. Proof readiness: a credible way to demonstrate the workflow in the customer’s environment.
    3. Buying readiness: security review, procurement, contracting, support expectations, and a clear commercial model.
    4. Adoption readiness: a rollout owner, champion, implementation path, success definition, and expansion trigger.

    This distinction prevents a common failure mode: treating every stalled enterprise deal as a missing feature. A proof may be technically successful while procurement remains unresolved. A contract may close while rollout ownership remains absent. Those are different problems with different owners.

    Treat each geography as another product-market-fit decision

    Global expansion is not the domestic sales motion with a new territory field. Each region can change latency expectations, compliance obligations, localization needs, support coverage, channel structure, and the credibility required from local references.

    Before committing to a region, document the target use case, buyer and practitioner, required architecture, region-specific data and compliance constraints, localization scope, support model, sales motion, implementation ownership, and first reference path. Assign legal, regulatory, security, and contractual questions to qualified owners; a GTM checklist is not legal clearance.

    Enter narrowly enough that you can distinguish a regional product gap from a weak ICP or an unproven sales motion. One repeatable use case with a supported operating model is more informative than broad pipeline created before the product and field teams can deliver.

    Use metrics that identify the broken handoff

    A category dashboard should connect market understanding to product use and commercial expansion. Track the measures as a system rather than searching for one category-creation metric:

    • Problem recognition: the share of qualified conversations in which buyers recognize the target problem and can describe its consequence in their own words.
    • Activation: the rate and median time from entry to the defined practitioner proof event.
    • Proof progression: movement from technical proof to workflow, organizational, and economic proof.
    • Commercial conversion: proof-to-paid conversion, stage duration, loss reasons, and no-decision reasons by ICP.
    • Adoption: repeated use of the core workflow, team participation, and time to the next relevant use case.
    • Expansion: growth across workflows, teams, volume, or organizational capabilities, including Net Recurring Revenue where it fits the model.
    • Signal quality: recurring community questions, field blockers, support patterns, and the share of requests that generalize across accounts.

    The relationships diagnose the system. If recognition rises while activation stays flat, the story is outrunning the product. If activation improves but expansion does not, the organizational value or paid boundary is weak. If proofs succeed but purchases stall, inspect stakeholder alignment, procurement, trust, and pricing. If enterprise revenue grows while core activation deteriorates, bespoke complexity may be consuming the product.

    Compare these measures by ICP, acquisition motion, and cohort. A blended company-wide number can hide a strong practitioner loop beneath a weak enterprise segment, or make one unusually large account look like a repeatable motion.

    Use the next 90 days to earn the right to scale

    Treat the next 90 days as a sequence of decisions, not a category launch calendar. The objective is to discover whether one buyer, one problem, one proof event, and one commercial path can be repeated without founder-level intervention.

    1. Days 1-30: establish the buying problem. Review the last 20 qualified wins, losses, and no-decisions; if you have fewer, use all of them. Extract the trigger, buyer language, incumbent alternative, budget owner, evaluation criteria, proof requested, and reason the process moved or stopped. Observe at least five relevant practitioners doing the target job. Write the one-page category thesis, choose a narrow ICP, state the disqualifying conditions, and define the first proof event. At the end of this phase, decide whether the evidence supports category creation or simply demands clearer positioning.
    2. Days 31-60: make proof repeatable. Run the same proof structure with a small cohort of relevant prospects or active accounts. Keep the target job, required inputs, and success criteria stable enough to compare results. Record where expert intervention is required. Publish one practical artifact that helps practitioners perform the new workflow, then use their questions to improve onboarding and terminology. Test whether the proposed free-to-paid boundary is understandable before changing pricing.
    3. Days 61-90: test transfer and expansion. Have a person who did not design the motion run it for the same ICP. Separate core-product gaps from implementation, enterprise control, procurement, and pricing gaps. Reprice representative account usage under the proposed model. Test one narrow upmarket or regional hypothesis only if the core proof is stable. Choose explicitly among investing in scale, holding the motion at its current level, narrowing the ICP, or returning to discovery.

    The final decision should be evidence-based. Scale when buyers recognize the problem, practitioners reach proof, a second operator can reproduce the motion, the commercial boundary is legible, and expansion follows the same underlying value. Hold when success still depends on exceptional persuasion, custom work, or an account-specific roadmap.

    Key takeaways

    • Create a category only when the established buying frame hides the problem or misrepresents the product’s value.
    • Start with a practitioner wedge that produces a visible result before asking executives to accept a broad market narrative.
    • Turn demos into defined proof events and distinguish technical, workflow, organizational, and economic proof.
    • Preserve low-friction adoption, then monetize coordination, control, scale, and other capabilities that become valuable as usage spreads.
    • Layer enterprise sales and global expansion onto a repeatable product loop; do not use them to compensate for a weak one.
    • Scale GTM only after someone outside the founding motion can reproduce the proof for the same ICP.

    Start tomorrow with the recent opportunities that did not move. Rewrite the category thesis in the buyer’s language, choose one observable proof event, and identify the first handoff that cannot yet be repeated. Fix that handoff before adding another campaign, segment, region, or sales hire. Category authority is the consequence of a working system, not the starting condition.

    References

    • Shivam.Consulting Blog – Inside dbt Labs’ $4.2B ascent: category creation, open source, and monetization playbook
    • Shivam.Consulting Blog – Inside Clay’s $1.25B Playbook: Unconventional GTM, Pricing Strategy, and Enterprise Wins
    • Shivam.Consulting Blog – Inside Braze’s Blitz to $500M CARR: Bold PM Lessons on Going Global and Outsmarting Rivals
    • Shivam.Consulting Blog – From Bootstrapped to $6B: Inside 1Password’s B2B Pivot, GTM Engine, and CEO Playbook