incident management
- AI Strategy (315)
- Generative AI (106)
- IT Leadership (33)
- Leadership (76)
- Product Management (344)
- Product Management Leadership (306)
- Uncategorized (32)
-

Production AI Agent Operations: A Practical Operating Model
A practical operating model for reliable AI agents, covering orchestration, safe retries, evaluations, controlled releases, monitoring, and incidents.
-

Governing AI Security Beyond the Open-Weights Debate
A practical operating model for product leaders to govern jailbreak disclosure, open-weight releases, agent containment, and incident-response readiness.
-

The July 2026 AI Evaluation Lab Leak: A Leader’s Playbook
A practical governance and containment playbook for testing dangerous model capabilities without turning an internal evaluation into a real-world incident.
-

From Customer Signals to Reliable Product Operations
A practical framework for routing support, behavioral, research, and reliability signals into faster response and stronger product decisions.
-
Built for Your Biggest Days: How We Engineer Fair, Reliable Scale Without Downtime
Enterprise teams ask sharper questions about scale—and they should. I share how we handle 150k+ requests/sec, shard our source-of-truth data with Vitess and PlanetScale, and reshape…
-

Stop Blurring the Lines: Clear Product–Engineering Boundaries to Boost Quality and Prevent Burnout
Blurry lines between product and engineering lead to burnout, slow delivery, and quality problems. I share a practical model that restores clarity: product trios own the…
-

Reliable AI Infrastructure: A Product Leader’s Playbook
A practical framework for exposing silent AI failures, hardening runtime paths, controlling releases, and turning SLOs into product decisions.
-

The Safety of Speed: 180 Deploys a Day, 12‑Minute Releases, 99.8%+ Availability
Speed and safety are not opposites—they reinforce each other when you ship in small, frequent batches. By automating our pipeline end-to-end, we move from merge to…
-

How to Build AI-Enabled Cybersecurity Operations Safely
A practical 90-day model for using AI in threat detection and incident response with clear authority limits, metrics, and governance.
-

Inside the Engine Room: How I Drive Scalable Analytics APIs, Reliability, and Performance
I share how I focus on the middleware and compute systems that power analytics at scale so teams can trust their data. I detail how overseeing…
-

Agentic AI for Incident Response: A Practical Operating Model
A practical operating model for scoping, governing, evaluating, and rolling out incident-response agents without surrendering human control.
Weekly digest
One email a week on AI products, enablement and hiring. No fluff.
Browse topics
- AI Strategy (315)
- Generative AI (106)
- IT Leadership (33)
- Leadership (76)
- Product Management (344)
- Product Management Leadership (306)
- Uncategorized (32)
Work with me
45-minute consultation on AI product strategy, GTM and PM hiring — no charge.
