eval-driven development
- AI Strategy (315)
- Generative AI (106)
- IT Leadership (33)
- Leadership (76)
- Product Management (344)
- Product Management Leadership (306)
- Uncategorized (32)
-

How to Operationalize AI Agents as Recurring Employees
A practical model for turning one-off AI tasks into recurring roles with clear context, review gates, escalation, permissions, and ownership.
-

How Product Leaders Should Read Frontier AI Benchmarks
A practical framework for separating benchmark gains from deployable capability and turning frontier model progress into product decisions.
-

Production AI Agent Operations: A Practical Operating Model
A practical operating model for reliable AI agents, covering orchestration, safe retries, evaluations, controlled releases, monitoring, and incidents.
-

How to Give Autonomous Agents Context and Permission to Act
A practical operating model for agents that detect work, use current context, take bounded action, and escalate before consequences outrun control.
-

How to Benchmark AI Models for Cost, Quality, and Risk
A practical framework for choosing AI models by cost per accepted result, workflow reliability, failure severity, and production economics.
-

Persistent Memory for AI Agents: A Product Leader’s Guide
A practical framework for deciding what an AI agent should remember, how it should retrieve and forget memories, and how to test the value safely.
-

Chinese Frontier AI Models: A Product Leader’s Decision Guide
A practical framework for choosing Kimi K3, Qwen 3.8 Max, or DeepSeek V4 Flash by completion reliability, accepted-result cost, and product fit.
-

Context Engineering: How to Build Reliable AI Applications
A practical framework for designing context, diagnosing agent failures, budgeting each turn, and evaluating AI applications before production.
-

How to Design Persistent Memory for Reliable AI Agents
A practical framework for deciding what an AI agent should remember, retrieving it safely, correcting stale facts, and proving that memory helps.
-

AI Builder Maturity: From Fast Demo to Defensible Product
Use a five-level evidence ladder to diagnose your AI product, test platform risk, and invest in advantages that strengthen as models improve.
-

AI Agent Skill Overload: How to Curate What Stays
Learn how to test, approve, fork, and retire AI agent skills before instruction conflicts and hidden assumptions weaken your team’s output.
-

Embodied AI: A Product Leader’s Guide to Generalist Robotics
Use a task-first framework to choose embodied AI workflows, measure useful autonomy, design safety boundaries, and make a sound robotics investment.
-

The July 2026 AI Evaluation Lab Leak: A Leader’s Playbook
A practical governance and containment playbook for testing dangerous model capabilities without turning an internal evaluation into a real-world incident.
-

Visual Search for Messy B2B Catalogs: An Intent-First Playbook
A practical framework for turning ambiguous product images into intent-aware retrieval, trustworthy ranking, and measurable B2B sourcing outcomes.
-

How to Evaluate AI Risks That Emerge After Deployment
A practical framework for testing stateful AI across long user trajectories, monitoring drift after launch, and resisting misleading satisfaction metrics.
-

Persistent AI Agent Memory: A Knowledge Graph Blueprint
A practical blueprint for modeling, governing, retrieving, and evaluating knowledge graph memory that an AI agent can safely use over time.
-

Codex Token Cost Optimization Without Losing Useful Context
A practical model for trimming Codex context, preserving decisions, and proving token savings without creating more retries or review work.
-

A Practical Guide to Metaphors for Understanding Language Models
A practical framework for choosing AI metaphors, exposing their hidden assumptions, and translating human-sounding behavior into testable product claims.
-

How to Benchmark AI Humanizers Without Gaming the Test
A practical framework for testing AI humanizers across detector disagreement, content fidelity, writing quality, latency, and workflow fit.
-

AI Video Generation Advances: A Product Leader’s Playbook
A practical framework for turning gains in video coherence, control, and open-source access into workflows, evaluations, and sound product bets.
Weekly digest
One email a week on AI products, enablement and hiring. No fluff.
Browse topics
- AI Strategy (315)
- Generative AI (106)
- IT Leadership (33)
- Leadership (76)
- Product Management (344)
- Product Management Leadership (306)
- Uncategorized (32)
Work with me
45-minute consultation on AI product strategy, GTM and PM hiring — no charge.
