# Oshri Cohen: AI-Native Chief Product & Technology Officer (full content export) > Oshri Cohen is an AI-Native Chief Product & Technology Officer with 25 years in software (20 in technology leadership) who works as a fractional and interim CTO. Since 2018 he has served 30+ companies, at one point directing 12 engineering teams across 7 countries. He comes in to solve hard problems fast: building AI-native engineering and product organizations, turning around troubled software, leading technical due diligence for private-equity deals, and modernizing legacy systems. He works directly with his clients, hands-on as a changemaker from the inside out, and stays until the problem is solved. He is based in Canada and serves companies across the USA, working remotely on Eastern Time. Email: hello@oshricohen.me Phone: (514) 777-3883 LinkedIn: https://www.linkedin.com/in/oshricohen/ Location: USA · Remote Site: https://www.oshricohen.me ## Home I'm Oshri Cohen — an AI-native CPTO and fractional & interim CTO for $10M–$50M US companies. I rebuild how the product is built and how the org runs so AI is the default, then build the AI-native team to carry it. Hands-on, until it's done. I think from the product and the business first, then rebuild how the company operates so AI is the default — an AI-native transformation, not a feature bolted on, in service of business development. For 25 years I've sat at the seam between the boardroom and the codebase, translating business strategy into engineering execution, and engineering reality back into business decisions. Now that translation is mostly one thing: AI-native transformation. I've delivered 10+ digital products across HealthTech, E-commerce, Manufacturing, Logistics and Finance, hands-on across architecture, CI/CD, UX, and AI, with a sharp current focus on redesigning for the AI-native era. Since 2018 I've served as a Fractional & Interim CTO to 30+ companies, almost all of them US-based, at one point simultaneously directing 12 engineering teams across 7 countries. Before that, executive seats as VP of Engineering at an intelligent-transportation company and CTO of a healthcare EMR platform. My obsession is the experience, for users, customers, and the operators who run the business. Track record: - 25yrs: In SaaS & enterprise software, 20 in tech leadership - 30+: Companies served as Fractional & Interim CTO since 2018 - 12: Engineering teams led across 7 countries & 4 time zones - 10+: Digital products delivered end-to-end across industries Career: - Oct 2025 to Jun 2026 · Chief Product & Technology Officer, EdTech: Led product, engineering and AI-native transformation for an online learning platform. - 2018 to Now · Consulting & Fractional CTO, USA · Remote: Strategic technology partner to 30+ startups and growth-stage companies. - 2016 to 2018 · VP of Software Development, Intelligent Transportation Systems: Bimodal teams → +30% delivery efficiency; CI/CD → −40% deploy times. - 2014 to 2016 · Chief Technology Officer, Healthcare EMR Platform: Modernized the integrated care platform; health-data compliance. - Earlier · Director of Technology, Hospitality Wellness Platform: POS integration strategy across hundreds of restaurant systems. - Earlier · Principal Engineer, No-code Automation Startup: Founding team; task execution engine & workflow orchestration. - Earlier · Director of Engineering, Market Research Firm · 350+ employees: Digital transformation of research & analytics platform. - 2006 · BA, Business Administration, McGill University: FounderFuel mentor · Montreal CTO Meetup co-organizer. Selected impact: - Redesigning the operating model for AI: Redesigning business, product and engineering processes for the AI-native era, rebuilding workflows, delivery and decision-making so the whole organization operates AI-first. - Engineering Performance (Low → Elite): DORA transformation in three months: Took a low-performing org to elite DORA metrics in 90 days, restructuring process, GitOps and CI/CD to make delivery fast and predictable. - Security & Compliance (SOC 2 I & II): HIPAA-grade compliance, delivered: Achieved SOC 2 Type I & II across HealthTech, InsureTech and E-commerce while leading offshore teams under HIPAA protocols. - Modernization (< 1 hr): On-prem to cloud-native, no drama: Re-platformed legacy enterprise systems to Kubernetes cloud-native, Azure, AWS, GCP, with downtime held under an hour. - Turnarounds (6 products): Rescued troubled SaaS products: Inherited codebases with no tests, docs, or team, rebuilt engineering, led product, and turned six products around in 3–6 months. - Logistics · Re-platform (0 downtime): One-year cloud-native migration: Re-architected an on-prem logistics platform to cloud-native over a year with zero downtime, team built and led end-to-end. Method: Every engagement runs on The Business-Down Method™. — Most technology advice starts with the technology. Mine starts with the business: where you make money, where you lose it, what the board is actually asking. From there it derives the architecture, the team, and the AI. Four movements, one spine, applied the same way every time. The method is the discipline; the solution is always current. - 01 · Read — Diagnose, tied to the P&L: A clear-eyed read on product, team, architecture, security and spend, tied to your numbers. The real root cause, not the loudest symptom, plus where AI creates measurable leverage and where it's a distraction. You leave with a 90-day plan you can act on with or without me. - 02 · Direct — Set the AI-native direction: Strategy tied to revenue, an architecture that scales, an AI-native operating model and the org to run it. Build-vs-buy, security and compliance posture, and the hiring plan. Every call derived down from the business. - 03 · Build — Ship the systems, build the team: Hands-on. Production AI systems in the product and the pipeline, the SDLC rebuilt AI-first with security in every pull request, and the AI-native team hired and operating at elite DORA. The part most advisors don't do. I do. - 04 · Operate — Run it, measure it, make it stick: I run the operating model and measure it: DORA for delivery, cost and quality for AI. I re-run evals as models drift, adopt what's genuinely better, then plan the transition to an internal hire or trained team. No permanent dependency. How I help: - Fractional CTO: Senior technology leadership without a full-time hire. I set direction, build the team, and own delivery. - AI-Native Transformation: I rebuild how the product is built and how the org runs so AI is the default, not an add-on, and I implement it, including the hiring. - Temporary CTO: A CTO in the seat full-time for a defined stretch, to cover a gap or carry a transition, then hand off cleanly. Also called an interim CTO. - Short-Term CTO: Full CTO ownership for a stretch measured in months: cover the gap, hit the deadline, hand back clean. - Interim CTO: Full-time leadership in the seat, temporarily, to cover a gap or carry a transition until the permanent hire lands. - Turnaround CTO: Stabilize troubled products and teams, then chart a credible path back to fast, predictable delivery. - Technical Due Diligence: For PE and investors: a clear read on code, architecture, team and risk before you sign, and an AI-native value-creation plan after. - Technical Recruitment: I recruit your engineering team the way a CTO builds his own: org design first, every candidate vetted personally, hires built for the AI era. - Engineering Org & Culture: High-performing teams via DORA, GitOps and DevOps: measurable, predictable, sustainable delivery. ## Services ### AI-Native Development & Operations Leader: /services/ai-native-leader/ Executive who helps companies become genuinely AI-native, rebuilding how products are built and how the organization runs so AI is the default, not an add-on. Bolting a chatbot onto an old operating model isn't AI strategy. I help US companies become genuinely AI-native, rebuilding how products are built and how the organization runs so that AI is the default, not an add-on. Tags: LLM pipelines at scale, AI in the SDLC, Agentic workflows, AI code & security review What is AI-native transformation leadership? An AI-native transformation leader is an executive who rebuilds how a company builds products and runs its operations so AI is the default rather than an add-on. It goes further than advice: the deliverable is a company that works differently, with workflows redesigned AI-first, AI running inside the SDLC, and agentic systems doing real production work. People searching for an AI transformation consultant usually want this outcome and get a deck instead. I lead the whole cycle hands-on: the build, the operating-model change, and the continuous optimization after launch, because AI moves too fast to rest on a launch. I've run LLM pipelines processing roughly 250 million records a month, and shipped autonomous systems like the order-fulfillment agent graph that finds margin on every order while holding a hard floor. The test of an AI-native operating model is simple: does the leverage show up in cost, speed and quality you can measure? If it only shows up in the all-hands demo, it's theater. Also known as: AI transformation consultant, AI-native operating model, AI-native transformation, AI operations leadership. Not AI features. An AI-native operating model. AI bolted on: Same process, new toy: A copilot license and an "AI feature" on the roadmap; Workflows unchanged; AI lives at the edges; Pilots that never reach production; Spend without measurable leverage AI-native: The model is redesigned around AI: Business, product & engineering processes rebuilt AI-first; AI in the SDLC, code, review, testing, ops; Agentic systems doing real production work; Leverage you can see in cost, speed & quality I lead the whole cycle. - 01 · Build · AI-Native Development: Shipping AI deep in the product and the pipeline, not as a demo, but as production infrastructure that holds up at scale. (LLM-powered data & product systems in production; AI woven through the SDLC: generation, static & security review; Agentic & retrieval architectures, with the right guardrails; Eval, observability & cost control for AI systems) - 02 · Operate · AI-Native Operations: Changing how the organization actually works, so teams, decisions and processes are built around AI from the ground up. (Redesigning business & engineering workflows AI-first; Upskilling teams & setting AI usage, safety & data policy; AI for BI, decision support & operational visibility; A roadmap that ties AI investment to revenue, not hype) - 03 · Optimize · Continuous Optimization: AI moves too quickly to rest on a launch. What was state of the art last quarter is table stakes today, so the work never finishes, it compounds. (Tracking model, tooling & cost curves, and adopting what is genuinely better; Re-running evals as models and prompts drift, so quality never quietly regresses; Tightening cost and latency as usage scales and cheaper paths appear; Feeding production signals back into the product and the operating model) "The companies that win the next decade won't be the ones that added AI. They'll be the ones that rebuilt themselves around it, in the product and in the org.": Oshri Cohen · The AI-native thesis From audit to operating system. - AI-Native Audit: A clear-eyed read on where AI creates real leverage in your product and operations, and where it's a distraction. - Build & Ship: Hands-on architecture and delivery of production AI systems, with eval, observability and cost control built in. - Operating-Model Redesign: Rewiring workflows, teams and decision-making so the whole organization runs AI-first, and it sticks. What founders & boards ask. Q: What does "AI-native" mean in practice? A: It means the operating model is redesigned around AI rather than having AI bolted on. Business, product and engineering processes are rebuilt AI-first, AI runs inside the SDLC for code, review, testing and ops, and agentic systems do real production work, with leverage you can measure in cost, speed and quality. Q: We already have a copilot license. Isn't that AI-native? A: Not on its own. A copilot license and an "AI feature" on the roadmap is AI bolted on, workflows stay unchanged and AI lives at the edges. AI-native means the workflows, teams and decision-making themselves are rewired around AI. Q: Is this just consulting, or will you build? A: Both. I work hands-on, architecting and shipping production AI systems with eval, observability and cost control, as well as redesigning the operating model so the change sticks after I leave. Q: How do you keep AI systems reliable and cost-controlled in production? A: Every system ships with evaluation, observability and cost controls built in, plus guardrails on agentic and retrieval architectures. I've run LLM-powered pipelines processing ~250M records a month on a 75-node cluster, so the patterns are battle-tested, not theoretical. Q: Where do you start with a company that's new to this? A: With an AI-Native Audit: a clear-eyed read on where AI creates real leverage in your product and operations, and where it's a distraction, followed by a roadmap that ties AI investment to revenue rather than hype. Related reading & paths. - AI Strategy: Where every transformation starts: the roadmap, architecture and economics, in writing and ready to build against. - What a transformation costs: The staged numbers: strategy fixed-scope, leadership monthly, and the spend the roadmap budgets. - AI Innovations by Industry: The AI applications and agent systems I've designed across industries, some shipped, some blueprints. Ready to become AI-native? Whether you're shipping AI into the product or rebuilding how your teams operate, let's map the path that actually moves the business. ### AI Strategy Consulting: /services/ai-strategy/ AI strategy consulting from a hands-on fractional AI CTO: an AI roadmap tied to the P&L, the architecture and economics behind it, and the leadership to ship it into production. Every board wants an AI strategy. Most of what gets sold under that name is a deck. I build the other kind: a plan tied to your P&L, written by the fractional AI CTO who then stays to build it, hands-on, until the leverage shows up in the numbers. Tags: Tied to the P&L, Fractional AI CTO, Production, not decks, $10M–$50M US companies What an AI strategy actually is. An AI strategy is a plan for where AI makes your business measurably better, and what it takes to get there. Which workflows it should run. What it does to your cost structure. What your architecture, your data, and your team have to become. It's a business document with an engineering spine. Most AI strategies die at the same spot: the handoff. A firm writes the vision, leaves, and the pilots stall in the gap between the deck and the codebase. I close that gap by refusing to hand off. I write the strategy as your fractional AI CTO, then lead the build until the results are in production and on the P&L. That's how systems like my hybrid AI + human workflow platform got shipped: one owner from plan to production. The work is remote-first. I serve US companies from Eastern Time, with same-day overlap across every US time zone. New to the term? Read the plain-English guide. Also known as: AI consulting, AI consulting services, AI consultant, AI strategy consulting, enterprise AI strategy, AI roadmap, AI advisory, fractional AI CTO, AI transformation strategy. You have AI activity. You don't have an AI strategy. The symptoms vary. The missing piece is usually the same: nobody senior owns the plan. - Pilot purgatory: You've run six AI pilots. None reached production, and nobody can say why the seventh will be different. - Board pressure, no plan: Investors ask about your AI strategy every quarter. What you actually have is a copilot license and a slide. - AI theater: There's a chatbot on the website and a demo in the all-hands. The operation still runs the way it did in 2022. - Spend without leverage: Tokens, tools, vendors and headcount are going out the door, and the leverage they were supposed to buy isn't showing up anywhere you can measure. Six answers, in writing. - Opportunity audit: Where AI creates real leverage in your product and operations, and where it's a distraction. Grounded in your data and your margins. - The AI roadmap: A sequenced plan tied to revenue and cost, with the unglamorous prerequisites, data, evals, security, scheduled instead of ignored. - Architecture & models: Build vs buy, which models, agentic or not, and the eval, observability and cost controls that keep it honest in production. - AI economics: Unit costs, token spend, and the cost curves that decide what's viable. A strategy that ignores the economics is a demo. - Team & operating model: Who you hire, who you upskill, and how the workflows change so AI becomes the default way work gets done. - Governance & safety: Usage policy, data boundaries and review gates, so you can move fast without betting the company on an unreviewed prompt. Strategy as a document vs. strategy as an operating change. The deck: Written, presented, shelved: Authored by a firm that leaves before the build starts; Use cases ranked by excitement, not economics; No owner once the engagement ends; Pilots forever; production never The shipped strategy: Written by the person who builds it: One accountable fractional AI CTO from plan to production; Use cases ranked by P&L impact and feasibility; Architecture, evals and cost control specified up front; Leverage you can see in cost, speed and quality "The strategy isn't the deliverable. The company that runs differently a year later is the deliverable. Everything else is theater.": Oshri Cohen What it costs, published. Strategy and delivery are priced separately, so you can start small and scale commitment as the plan proves out. I publish my numbers so you can self-qualify before we talk. - The Diagnostic (From $20,000 · Fixed · 2–4 weeks · yours to keep): The Business-Down read on your system, org, spend and AI leverage, plus a 90-day plan you can act on with or without me. The best first step. - AI-Native Roadmap & Architecture Sprint (From $35,000 · Fixed scope · 4–6 weeks): The full AI strategy: opportunity audit, sequenced roadmap, architecture and model choices, and the economics. Ready to build against. - Fractional AI CTO ($18,000 / mo · ~2 days a week): I stay to lead the build: architecture, hiring, delivery and the operating-model change. Hands-on, until the leverage is real. All prices USD. Ongoing leadership shares tiers with the Fractional CTO engagement — same person, same method. Market context: what an AI strategy costs, explained. Email me ↗ Related reading & paths. - AI-Native Transformation: Where the strategy leads: rebuilding how the product is built and how the org runs so AI is the default. - Fractional CTO: The flagship engagement. Senior, AI-native technology leadership without the full-time hire, with published pricing. - Escaping AI pilot purgatory: Why pilots stall, and what the organizations that actually reach production do differently. What founders & boards ask. Q: What is an AI strategy? A: An AI strategy is a plan for where AI creates measurable business leverage and what it takes to get there: which workflows and products it should run, the architecture and models behind it, the economics, the team, and the governance. A real one is tied to the P&L and sequenced into a roadmap you can build against. If it can't survive contact with your codebase and your cost structure, it's a vision statement, not a strategy. Q: What does an AI strategy consultant do? A: A typical AI strategy consultant audits your business, identifies AI use cases, and delivers a recommendation. I do that work and then stay to build it. As a fractional AI CTO I write the strategy, make the architecture and model choices, hire and upskill the team, and lead delivery until the systems are in production and the leverage is measurable. Q: What is a fractional AI CTO? A: A fractional AI CTO is an experienced technology executive who leads your AI strategy and its execution part-time, instead of as a full-time hire. You get senior judgment on AI architecture, build-vs-buy, team and governance, plus hands-on delivery, without a six-month executive search. It's the fractional CTO model focused on making AI the default in how the company builds and operates. Q: How much does an AI strategy cost? A: My pricing is published. A fixed-fee Diagnostic starts at $20,000 and takes two to four weeks. The full AI-Native Roadmap & Architecture Sprint, the complete strategy with architecture and economics, starts at $35,000. Ongoing leadership to build it runs $18,000 per month at roughly two days a week. All prices are USD. Q: How long does it take to develop an AI strategy? A: The Diagnostic takes two to four weeks. The full roadmap and architecture sprint takes four to six weeks. That's enough to know where the leverage is, what to build first, and what it will cost. Execution is where the real time goes; engagements that continue into delivery usually run six to eighteen months. Q: How is this different from hiring a big consulting firm? A: Two ways. You work with me directly: no bench, no account manager, and the person who writes the strategy is the person who builds it. And the strategy is written to be shipped: it specifies how each system will be evaluated and what it will cost to run, because I'm the one accountable for shipping it. Q: Do we need an AI strategy or an AI-native transformation? A: The strategy is where you start; the transformation is where it leads. The strategy tells you where AI pays and what to build first. An AI-native transformation then rebuilds how the whole company works so AI is the default rather than an add-on. In practice one flows into the other, and I lead both. Need an AI strategy that survives production? Tell me where you're stuck: pilots that stall, board pressure, or a blank page. I'll tell you honestly what a real plan looks like. ### Fractional CTO: /services/fractional-cto/ An AI-native fractional CTO who comes in, solves the hard technology problem, and builds the engineering and product team to carry it, hands-on, until it's done. Senior, AI-native technology leadership for US companies, without a full-time hire. I come in, solve the hard problem, and build the engineering and product team to carry it, hands-on, until it's done. Tags: $10M–$50M US companies, AI-native teams, Hands-on delivery, Direct, no firm What is a fractional CTO? A fractional CTO is an experienced Chief Technology Officer who leads your technology part-time, typically a day or two a week, instead of joining as a full-time executive. You get the judgment, the architecture calls, the hiring and the delivery ownership without the $400K–$750K all-in cost of a full-time hire. My published price for the full fractional engagement is $18,000 a month at roughly two days a week, with advisory from $10,000 a month. The role travels under many names. If you've been searching for a part-time CTO, a virtual CTO, or CTO-as-a-Service, this is the same job. The label changes; the seat doesn't. What separates my version is that I build. I've shipped systems like a pricing-intelligence platform processing 250 million pages a day, and I bring that same hands-on standard to every fractional engagement. Also known as: part-time CTO, virtual CTO, outsourced CTO, on-demand CTO, CTO-as-a-Service (CTOaaS), fractional technology leadership. You're here because something isn't scaling. The symptoms are always different. The root cause rarely is. - The team can't scale: What worked at ten people and ten thousand users is breaking. Delivery is slowing exactly when it needs to speed up. - Technical debt is winning: Every change is risky, nothing is documented, and the backlog of "we'll fix it later" is now the thing holding you back. - A collaboration tax: Broken processes mean every decision costs three meetings. The org is paying a tax on its own communication. - Spend without leverage: Money is going out the door, headcount, tools, cloud, and you can't see the leverage it's supposed to buy. Set direction. Build the team. Own delivery. - Strategy & roadmap: Technology strategy tied to the P&L, an architecture that scales, and a roadmap that picks the right battles. - Build the team: I hire the engineers and product people, and design an AI-native org around them, not a copy of the last company's. - Own the outcome: Hands-on architecture, delivery, security and compliance. I make the calls and stay accountable for the result. In fast. Direct. Until it's solved. - 01 · Read: Diagnose the system and the business. - 02 · Direct: Set the AI-native direction. - 03 · Build: Ship the systems, build the team. - 04 · Operate & Optimize: Run it, measure it, make it stick. A track record, not a pitch. - 30+: Companies served as fractional & interim CTO since 2018 - Low → Elite: DORA delivery performance, in 90 days - 12: Engineering teams led across 7 countries - 25yrs: In software; 20 in technology leadership "The method is the discipline — repeatable and accountable. The solution is never boilerplate, because the state of the art moves every quarter. I solve the problem in front of you with whatever is genuinely best right now.": Oshri Cohen Related reading & paths. - What is a fractional CTO?: The plain-English explainer, what the role is, what it costs in the market, and when you actually need one. - Fractional vs full-time vs interim: An honest breakdown of which kind of technology leadership your company actually needs right now. - AI Strategy: The AI roadmap, architecture and economics, written by the fractional AI CTO who stays to ship it. - AI-Native Transformation: How I rebuild the operating model so AI is the default, the moat beneath every engagement. What it costs, published. A full-time CTO costs $400K–$750K all-in and takes three to six months to ramp. A fractional CTO gives you 60–80% of the value at a fraction of that — starting in the first 30 days. I publish my numbers so you can self-qualify before we talk. - The Diagnostic (From $20,000 · Fixed · 2–4 weeks · yours to keep): The Business-Down read on system, org, spend and AI leverage, plus a 90-day plan you can act on with or without me. The best first step. - Advisory ($10,000 / mo · ~1 day a week): Senior judgment on direction, architecture, hiring and roadmap. For teams that can execute but need the calls made right. - Fractional CTO ($18,000 / mo · ~2 days a week): The full method, running. I set direction, build and hire the team, and own delivery — hands-on. Most engagements run 6–18 months. - Embedded / Interim CTO ($30,000 / mo · 3+ days a week): Near-full-time leadership in the seat — for turnarounds, transitions, or covering the CTO role until the permanent hire lands. All prices USD. Prefer a fixed scope? AI-Native Roadmap & Architecture sprints from $35K, PE technical due diligence from $15K/deal. Want the market context? What a fractional CTO costs, explained. Email me ↗ What founders & boards ask. Q: What is a fractional CTO? A: A fractional CTO is an experienced Chief Technology Officer who leads your technology part-time instead of as a full-time hire. You get senior judgment on strategy, architecture, hiring, security and roadmap without the cost or the six-month search for a full-time executive. Unlike a consultant who advises from the sidelines, I embed, make the calls, and stay accountable for the outcome. Q: How is this different from hiring a firm or a marketplace? A: You work with me directly, no firm, no rotating bench, no account manager between you and the person doing the work. I'm hands-on, so your problem gets solved quickly: I build, I don't just advise. Q: What makes you an "AI-native" CTO? A: I don't bolt AI onto an old operating model. I rebuild how the product is built and how the organization runs so AI is the default, in the SDLC, in operations, and in the team I hire, and I implement it directly. Q: Who do you work with? A: Mostly $10M–$50M companies where software is core to the business — scaling, modernizing, rebuilding around AI, or being bought. Founders and boards who need senior technology judgment without a full-time hire, and the private-equity firms who back them. I work with US companies and remote teams from Eastern Time, with same-day overlap across every US time zone. Q: How long is an engagement, and how is it priced? A: Engagements last until the problem is solved — sometimes a month, sometimes a year — and most Fractional engagements run 6–18 months. Pricing is published and tiered by commitment: a fixed-fee Diagnostic from $20,000, then Advisory ($10K/mo), Fractional ($18K/mo), or Embedded/Interim ($30K/mo). Fixed-scope projects are available too. The best first step is the Diagnostic, or a direct conversation. Q: Who is this not for? A: If you're a non-technical founder looking for someone to cheaply finish a no-code prototype, I'm not your person. This is senior, hands-on problem-solving for companies with real stakes. Got a problem worth solving properly? If the fit is right, I come in fast and stay until it's done. ### Interim CTO: /services/interim-cto/ Full-time technology leadership in the seat, temporarily, to cover a sudden gap or carry a company through a transition until the permanent hire is in place. Full-time technology leadership for US companies, temporarily. I step in to cover a sudden gap or carry the company through a transition, holding delivery steady, leading the team, and handing off cleanly to the permanent hire. Tags: Leadership gap, Day-one impact, Clean handoff, Direct, no firm What is an interim CTO? An interim CTO is a senior Chief Technology Officer who runs your technology function full-time for a temporary stretch: covering a sudden departure, carrying a merger or restructure, or holding the seat until the permanent hire lands. The ownership is real while it lasts. Strategy, delivery, the team, the board conversation, all of it is mine to run, and the engagement ends when the transition does. The pricing is public. In the US market an experienced interim CTO runs $1,500–$2,500 a day, which lands at roughly $25,000–$50,000 a month near-full-time. My published price is $30,000 a month for embedded, near-full-time leadership in the seat, with a fixed-fee Diagnostic from $20,000 if you want the honest read first. If the gap doesn't need full-time coverage, the fractional engagement runs $18,000 a month. The full breakdown is in the interim CTO pricing guide. I've held this kind of seat before, including as CTO and architect for a regional healthcare network, where I shipped two EMR SaaS platforms while live patient care kept running. Most of the work is remote. I serve US companies from Eastern Time, with same-day overlap across the country. Also known as: temporary CTO, short-term CTO, transition CTO, acting CTO, contract CTO, CTO gap coverage. The CTO just left. Now what? Interim leadership is about stability and momentum, not a placeholder. - A sudden departure: Your CTO is gone and the team, the roadmap and the investors all need a steady hand this week, not in six months. - A transition to manage: A merger, a restructure, a pivot. Someone senior has to hold the technology function together while it changes shape. - A bridge to the right hire: You want to hire the permanent CTO carefully, not panic-hire. I keep things moving and help you bring in the right person. Stabilize. Lead. Hand off. - First · Stabilize: Assess delivery, team, risk and the immediate fires; Reassure the team, the board and the customers; Keep shipping, momentum is the priority - Then · Lead: Run the technology function full-time, hands-on; Fix what's breaking; make the calls that were stuck; Bring AI-native practices in where they create leverage - Finally · Hand off: Help define and hire the permanent CTO; Document decisions, architecture and the plan; Leave the function stronger than I found it A track record, not a pitch. - 30+: Companies served as fractional & interim CTO since 2018 - Low → Elite: DORA delivery performance, in 90 days - 12: Engineering teams led across 7 countries - 25yrs: In software; 20 in technology leadership "Interim isn't a placeholder. It's full ownership of the technology function for as long as it takes to make the next chapter someone else's to run well.": Oshri Cohen What boards & founders ask. Q: What's the difference between an interim CTO and a fractional CTO? A: An interim CTO is full-time but temporary, I'm in the seat every day to cover a gap or carry a transition, typically for a defined period. A fractional CTO is ongoing but part-time. Both are hands-on; the difference is the depth and duration of the commitment. Q: How quickly can you start? A: Usually within one to two weeks. Interim work is urgent by nature, so I keep capacity for exactly this situation and give the mandate full-time attention from day one. Q: Will you help hire my permanent CTO? A: Yes. Part of the job is leaving cleanly, that means helping define the role, assessing candidates, and handing off the architecture, decisions and plan so your permanent hire succeeds. Q: How long does an interim engagement last? A: Until the transition is complete and the function is stable, sometimes a few months, sometimes longer. It ends when the problem is solved, not on an arbitrary date. Related reading & paths. - What an interim CTO costs: The pricing guide: US day rates, monthly numbers, and why firms charge more. - Fractional vs full-time vs interim: An honest breakdown of which kind of technology leadership your company actually needs right now. - Fractional CTO: The part-time alternative when the gap doesn't need full-time coverage, with published pricing. Need someone in the seat now? If the seat is empty and the clock is running, let's talk today. ### Short-Term CTO: /services/short-term-cto/ A short-term CTO who takes over the technology function for a stretch measured in months, carries the company through a gap, a deadline or a transition, and hands off clean. The same role people also call a temporary or interim CTO. You don't always need a CTO for the next five years. Sometimes you need one for the next five months. I take over your technology function as a short-term CTO, carry it through the gap or the deadline in front of you, and hand it back in better shape than I found it. Same role people call a temporary or interim CTO. Tags: Months, not years, In the seat fast, Clean handoff, Direct, no firm What is a short-term CTO? A short-term CTO is a senior Chief Technology Officer who runs your technology full-time for a bounded stretch, usually a few months, instead of joining the payroll for good. You skip the six-month executive search. A proven operator is in the seat this week, making the calls a CTO makes, and the engagement ends when the job does. Dated, high-stakes work is familiar ground; I've built systems like an event purchasing platform engineered to survive its single busiest minute. If you've been searching for a temporary CTO or an interim CTO, you've found the same role under a different name. Whatever you call it, the ownership is total while I'm in it: strategy, delivery, the team, the board conversation. Most of it is remote. I work with US companies from Eastern Time, so I overlap the whole country during the business day and there's nothing to relocate or stand up. Also known as: temporary CTO, interim CTO, contract CTO, CTO for hire, stand-in CTO, short-term technology leadership. When short-term is exactly right. The problem is urgent and bounded. The hire shouldn't outlast it. - A gap to cover: Your CTO left, or the search for one is dragging. The team and the roadmap need a senior hand this week, and the board needs to know someone owns it. - A deadline to hit: A raise, a migration, a launch, an audit. Something with a date on it needs full-time senior ownership until it's behind you. - A decision to de-risk: You're not sure the company needs a permanent CTO yet. A short-term one gets the function in order and tells you honestly what the seat actually requires. Take over. Deliver. Hand back. - First · Take over: Assess delivery, team, risk and whatever is burning; Steady the team, the board and the customers; Keep shipping while I learn the terrain - Then · Deliver: Run the technology function full-time, hands-on; Drive the gap, the deadline or the transition to done; Bring AI-native practices in where they earn their keep - Finally · Hand back: Help define and hire the permanent CTO if you need one; Document architecture, decisions and the plan; Leave the function stronger than I found it A track record, not a pitch. - 30+: Companies served as fractional & interim CTO since 2018 - Low → Elite: DORA delivery performance, in 90 days - 12: Engineering teams led across 7 countries - 25yrs: In software; 20 in technology leadership "Short-term is a duration, not a level of commitment. While I'm in the seat, the function is mine to run and the outcome is mine to answer for.": Oshri Cohen What boards & founders ask. Q: Is a short-term CTO the same as a temporary or interim CTO? A: Yes. Short-term, temporary, interim, contract: different words for one role. A senior CTO steps in full-time for a bounded period, owns the technology function completely, then hands off. Pick whichever word you like; the work is identical. Q: How short is short-term? A: Usually a few months, sometimes a couple of quarters. The engagement is scoped to a problem: cover the gap, hit the deadline, land the transition. It ends when that's done, not on a date picked in advance. Q: How is this different from a fractional CTO? A: Duration versus depth. A short-term CTO is full-time for a bounded period. A fractional CTO is part-time on an ongoing basis, typically a day or two a week. If you need someone in the seat every day right now, that's short-term. If you need senior judgment continuously over the long haul, that's fractional. Q: How quickly can you start? A: Fast. Short-term mandates are urgent by nature, so I move quickly and give the role full-time attention from day one. Q: What happens when the engagement ends? A: A clean handoff is part of the job. I help define and hire the permanent CTO if you want one, and I hand over the architecture, the decisions and the plan so whoever takes the seat next starts strong. Related reading & paths. - What a short-term CTO costs: The pricing guide: it's the interim engagement, and it's priced the same way. - Fractional vs full-time vs interim: An honest breakdown of which kind of technology leadership your company actually needs right now. - Interim CTO: The same seat described from the transition side: cover the gap, carry the handoff. Need a CTO for the next few months? If there's a gap to cover or a deadline coming, let's talk today. ### Technical Due Diligence: /services/technical-due-diligence/ Pre-deal technical due diligence and post-acquisition turnaround for private equity, a clear read on code, architecture, team and risk, then an AI-native value-creation plan. Pre-deal technical due diligence for US private equity and investors, and post-acquisition turnaround after. A clear-eyed read on code, architecture, team and risk, then a credible plan to make the asset AI-native and worth more. Tags: Pre-deal DD, Post-close turnaround, AI-native value creation, PE & investors What is technical due diligence? Technical due diligence is the pre-deal assessment of a software company's code, architecture, engineering team, security and true cost of ownership, done so the buyer knows exactly what they're acquiring before the money moves. The same work goes by several names: technology due diligence, software due diligence, tech DD. Whatever you call it, the output should be a decision-ready report written for the partners making the call, not a stack of engineering notes. My reads start at $15,000 per deal, scoped up with the size and complexity of the target, and most run two to three weeks from data-room access to report. Full numbers are in the technical due diligence pricing guide. The read is an operator's, because I've been on the build side of everything I assess. I've taken platforms through SOC 2 Type I & II and HIPAA-grade compliance myself, so when I flag a gap, I can also tell you what closing it takes. Also known as: technology due diligence, software due diligence, tech DD, IT due diligence, technical audit for acquisition. The deck never shows the technical debt. By the time it surfaces, you've already wired the money. - Is the architecture real?: Does it scale, or is it held together with heroics? I tell you what's solid and what's a liability before you commit. - Can the team deliver?: Key-person risk, capability gaps, culture problems. The people are the asset, or the risk, and I read both. - What will it really cost?: The cleanup, the re-platform, the hires. I give you the true cost of ownership, not the optimistic one. - Where's the upside?: An AI-native operating model is often the fastest path to margin. I show where the value-creation actually is. A decision-ready read. - 01 · Pre-deal · Technical Due Diligence: A clear, honest assessment delivered in time for your decision, written for partners, not just engineers. (Code quality, architecture and scalability; Security, compliance and data risk (SOC 2, HIPAA, GDPR); Team, key-person risk and delivery capability; True cost of ownership and a remediation roadmap) - 02 · Post-close · Turnaround & Value Creation: After the deal, I go in and execute, stabilizing the asset and transforming it into something worth more. (Stabilize delivery and stop the bleeding; Rebuild the team and the engineering approach; Redesign the operating model to be AI-native; Tie technology investment to margin and exit value) "Diligence tells you what you're buying. The turnaround is where the return actually comes from, and in most portfolio companies today, that means making them AI-native.": Oshri Cohen What deal teams ask. Q: What does technical due diligence cover? A: Code quality and architecture, scalability, security and compliance exposure, the engineering team and key-person risk, and the true cost of ownership, delivered as a decision-ready report written for partners, not just engineers, with a clear remediation roadmap. Q: Do you also do the post-acquisition work? A: Yes. Pre-deal diligence and post-close turnaround are a continuum. After the deal I go in hands-on to stabilize the asset, rebuild the team, and transform the operating model to be AI-native, which is usually where the return comes from. Q: How fast can you turn around a diligence report? A: Most reads run two to three weeks from data-room access to the report, and I work to the deal timeline. When a process is moving fast, a compressed read is possible; the scope narrows to the questions that can kill or reprice the deal. Q: Who is the buyer for this? A: Private-equity firms and investors evaluating a software asset, and operating partners responsible for value creation across a portfolio. Related reading & paths. - What tech DD costs: The pricing guide: per-deal numbers, what moves the fee, and what the report includes. - Turnaround CTO: When the read says the portfolio company needs an operator, not just a report. - Selected work: The projects behind the judgment: platforms scaled, systems modernized, teams rebuilt. Evaluating a deal or a portfolio company? Bring me in before you sign, or after, to make the asset worth more. ### Technical Recruitment: /services/technical-recruitment/ CTO-led technical recruitment for the AI coding era. I design the engineering org, vet every candidate myself, and deliver a team that works — not a stack of resumes. Recruitment run by a CTO who has spent twenty years hiring for his own teams. I design the org, vet every candidate myself, and hand you engineers built for the AI coding era. The deliverable is a working team. Tags: CTO-led vetting, Full engineering org, First hire in 6–8 weeks, 90-day guarantee What is CTO-led technical recruitment? CTO-led technical recruitment is engineering hiring run by a working CTO instead of a recruiting agency: I design the org, run every technical screen personally, and hand you a team built for how software gets written now. The fee is 30% of first-year salary per successful hire, the first hire typically lands in 6–8 weeks, and every hire carries a 90-day replacement guarantee. People find this page searching for a technical recruiter alternative or an engineering hiring partner. Both fit. The difference is what gets judged: a recruiter is measured on filling the seat, I'm measured on whether the team ships. Hiring is also part of the job when I'm in the seat as a fractional CTO. This service is that same muscle offered on its own: the org design and the vetting, without the ongoing mandate. Also known as: technical recruiter alternative, engineering hiring partner, CTO-led hiring, engineering recruitment services. Hiring engineers is hard. Hiring them right now is harder. Three very different companies end up on this page. The problem underneath is the same: nobody in the room can truly judge the candidate. - You raised, and now you have to hire: You're a founder without a technical co-pilot. Every candidate sounds impressive. You have no way to tell the strong engineer from the confident one, and the cost of guessing wrong is your runway. - The team has to double: You have engineers, but the leaders are stretched thin and interviews are eating the roadmap. Growth is stalling on the one thing you can't shortcut: finding people worth hiring. - You hired for a different era: Your team was built before AI changed how software gets written. The next hires need a different profile, and your interview process has no way to detect it. - The resumes all look perfect: AI writes flawless resumes and rehearsed answers now. The signals hiring used to rely on are gone, and keyword matching just surfaces whoever gamed it best. Design the org. Vet the people. Deliver the team. - Org design first: Before a single job post goes out, I design the team: which roles, what seniority mix, how it operates. A great hire in the wrong org design is still a miss. - CTO-led vetting: I interview every candidate personally. System design, real code, and a working conversation about how they think. Nobody reaches you on a keyword match. - Close and land: I calibrate offers to the market, help you win the candidate, and stay through onboarding so the hire actually sticks. Backed by a 90-day replacement guarantee. The whole team, or a partner inside your process. - End-to-end: The Team Build: I own the entire thing: org design, role definitions, sourcing, vetting, offers, and onboarding. You get a functioning engineering team, not a pile of candidates to figure out. - Embedded: Hiring Partner: Your recruiters run sourcing; I bring the technical judgment. Scorecards, technical screens, final interviews, and offer calibration — a CTO inside your hiring loop. The full engineering org. - Engineers, all levels: Backend, frontend, full-stack, mobile. From strong mid-level builders to the staff and principal engineers who set the technical bar. - Engineering leaders: Engineering managers, directors, VPs of Engineering, and CTOs. I've hired eight CTOs; I know what the seat requires because I've sat in it. - Data, ML & platform: Data engineers, ML engineers, DevOps and platform roles. The people who make the product measurable, reliable, and cheap to run. The old interview tests what AI does for free. Mine doesn't. - 01 · The CTO screen: A real technical conversation with me: architecture, trade-offs, decisions they've owned. Twenty years of hiring for my own teams goes into this hour. - 02 · AI-era work sample: A realistic exercise with AI tools on the table, the way the job actually works. I watch how they direct the tools, what they accept, and what they push back on. - 03 · Critical thinking: Can they tell plausible from correct? I probe how they challenge generated output, find the flaw in a confident answer, and reason when the pattern doesn't match. - 04 · Creative range: The AI age rewards people who reframe problems, connect domains, and find the path nobody scripted. I test for that range directly, because it doesn't show up on a resume. First hire in 6–8 weeks. A team build repeats this loop role by role, with sourcing running in parallel so hires land in sequence, not single file. - Week 1 · Org design & scorecards: Map the team you actually need: roles, seniority mix, sequence; Write role definitions and scorecards worth measuring against; Calibrate compensation to the market you're really in - Weeks 2–3 · Source & screen: Targeted sourcing through my network and channels that work; CTO screen on every serious candidate — no outsourced filters; AI-era work sample for the ones worth your time - Weeks 4–6 · Your interviews: A short slate of vetted finalists, with my written assessment of each; I sit in or run your loop, whichever helps more; Debriefs that turn opinions into a decision - Weeks 6–8 · Offer, close, land: Offer strategy and negotiation support until the yes; A 30-day onboarding plan so the hire ships early; 90-day check-ins, with a replacement guarantee behind them I've been hiring engineers for twenty years. - 75+: Engineers hired across startups and scale-ups - 8: CTOs recruited into the seat - 20yrs: In technology leadership, building my own teams - 90days: Replacement guarantee on every hire "A recruiter is judged on whether the seat gets filled. I'm judged on whether the team ships. Those are different jobs, and I only know how to do the second one — every candidate I send you is someone I'd put on my own team.": Oshri Cohen The thinking behind the service. - The resume is dead: AI writes flawless resumes now. Every signal hiring used to lean on is gone, and most processes are still screening like it's 2019. - The leetcode interview is over: Algorithm puzzles test exactly the work AI does for free. What a technical interview should measure instead. - Hire for judgment: The profile of a great engineer has changed: taste, skepticism toward plausible output, and leverage with agents. - Four engineers, not twelve: AI changed the math of team size. Why the right hire count is smaller than your plan says. - Technical skill is not the point: Raw technical skill is abundant now. What actually predicts performance is harder to test, and worth more. - All essays: The rest of the writing on AI-native teams, engineering leadership, and how software gets built now. One fee, published. You pay per successful hire, so my incentive is the hire that works, not the search that drags. Both engagement models carry the same fee and the same guarantee. - The Team Build (30% of first-year salary · Per hire · end-to-end): Org design, role definitions, sourcing, CTO-led vetting, offers, and onboarding. I run the whole build and hand over a working team. - Embedded Hiring Partner (30% of first-year salary · Per hire · inside your process): Your team sources; I vet. Scorecards, technical screens, final interviews, and offer calibration, with my written read on every finalist. Every hire is backed by a 90-day replacement guarantee: if it doesn't work out, I run the search again at no additional fee. Email me ↗ What founders ask before we start. Q: How is this different from a recruiting agency? A: Agencies are built to run many searches at once, and good ones do that well. I do something narrower: I'm a CTO who has spent 20 years hiring for my own teams, and I bring that judgment to yours. I design the org before sourcing starts, I run every technical screen personally, and I'm accountable for whether the team ships — not just whether the seat gets filled. Q: What does "built for the AI coding era" actually mean? A: Engineers work differently now. AI tools handle much of the routine code, so the engineers worth hiring are the ones with judgment: they direct the tools, catch the plausible-but-wrong output, and reason clearly when the pattern doesn't match. My vetting tests for exactly that — a CTO-led technical screen plus a realistic work sample with AI tools on the table, scored on critical thinking and creative range. Q: What roles do you recruit for? A: The full engineering org: backend, frontend, full-stack and mobile engineers at all levels, staff and principal engineers, engineering managers, directors, VPs of Engineering, and CTOs, plus data, ML, DevOps and platform roles. I've hired 8 CTOs and 75+ engineers over my career. Q: How fast will I get my first hire? A: Typically 6 to 8 weeks from kickoff to a signed offer: week 1 for org design and scorecards, weeks 2–3 for sourcing and CTO-led screening, weeks 4–6 for your interviews, and weeks 6–8 to close. Team builds run roles in parallel, so subsequent hires land faster than the first. Q: What does the fee cover? A: The fee is 30% of the hire's first-year salary, paid per successful hire. It covers org design, role definitions, sourcing, my personal technical vetting, the AI-era work sample, offer strategy, negotiation support, and a 30-day onboarding plan. Every hire carries a 90-day replacement guarantee: if it doesn't work out, I rerun the search at no additional fee. Q: Can you work inside our existing hiring process? A: Yes. In the Embedded Hiring Partner model your team runs sourcing and I bring the technical judgment: scorecards, technical screens, final-round interviews, and offer calibration. Same fee, same guarantee — the difference is who runs the pipeline. Q: Who is this not for? A: High-volume hiring where speed matters more than the individual pick — a staffing agency will serve you better and cost less. This is for companies where each engineering hire genuinely changes the trajectory, and where getting it wrong is expensive. Ready to build a team worth keeping? Tell me what you're building and who you think you need. I'll tell you honestly whether I'm the right person to find them. ### Temporary CTO: /services/temporary-cto/ A temporary CTO who runs technology full-time for a set period, covers a leadership gap or carries a company through a transition, then hands off cleanly. The same role people also call an interim CTO. Sometimes you need a Chief Technology Officer in the seat today, not on the payroll for the next five years. I step in as your temporary CTO, run technology full-time for as long as the situation demands, and hand it off in better shape than I found it. Some people call this an interim CTO. It's the same job. Tags: In the seat fast, Defined period, Clean handoff, Direct, no firm What is a temporary CTO? A temporary CTO is a seasoned Chief Technology Officer who runs your technology full-time for a fixed period instead of joining for good. You get a senior operator in the seat this week, making the strategy and delivery calls a CTO makes, without the long executive search or a permanent line on the payroll. It's the same role most people have in mind when they say interim CTO. The work is hands-on and the ownership is total. I run the function, unblock what's stuck, and stay until the job is done. Migrations are a common trigger; I once moved a live public-transit system to the cloud over a year with zero downtime. Most of it is remote. I work with US companies from Eastern Time and overlap the whole country during the business day, so there's no relocation to arrange and no payroll to stand up. Also known as: interim CTO, contract CTO, CTO for hire, short-term CTO, stand-in CTO, temporary technology leadership. When a temporary CTO is the right call. The need is urgent and real. The job itself doesn't have to be permanent. - The seat went empty: Your CTO is gone. The team, the roadmap and the board all need a steady senior hand this week, not after a six-month search. - A stretch to get through: A raise, a migration, a launch you can't fumble. Something big needs senior ownership for a few months, and then it's behind you. - A bridge to the right hire: You want to hire the permanent CTO well, not in a panic. I keep technology moving while you take the time to find the right person. Step in. Run it. Hand off. - First · Step in: Assess delivery, team, risk and whatever is on fire; Reassure the team, the board and the customers; Keep shipping, momentum matters more than ceremony - Then · Run it: Own the technology function full-time, hands-on; Make the calls that were stuck; fix what's breaking; Bring AI-native practices in where they add real leverage - Finally · Hand off: Help define and hire the permanent CTO; Document architecture, decisions and the plan; Leave the function stronger than I found it A track record, not a pitch. - 30+: Companies served as fractional & interim CTO since 2018 - Low → Elite: DORA delivery performance, in 90 days - 12: Engineering teams led across 7 countries - 25yrs: In software; 20 in technology leadership "Temporary doesn't mean tentative. It's full ownership of the technology function for as long as it takes to make the next chapter someone else's to run well.": Oshri Cohen What boards & founders ask. Q: Is a temporary CTO the same as an interim CTO? A: Yes. Two names, one role: a senior CTO who steps in full-time for a set period, holds technology together, then hands off. Some people say temporary, others say interim, contract or stand-in. The work is identical, full ownership while I'm in the seat. Q: How is this different from a fractional CTO? A: A temporary CTO is full-time but for a set period. A fractional CTO is ongoing but part-time. Both are senior and hands-on; what changes is the depth and the duration. Need someone every day to cover a gap? That's temporary, or interim. Need senior judgment a day or two a week over the long haul? That's fractional. Q: How quickly can you start? A: Temporary mandates are usually urgent, so I move fast and give the role real full-time attention from day one. Q: How long does a temporary engagement last? A: Until the gap is covered or the transition lands, sometimes a few months, sometimes longer. It ends when the problem is solved, not on a date someone picked in advance. Q: Will you help hire my permanent CTO? A: Yes, and it's part of the job. I help define the role, weigh the candidates, and hand over the architecture, the decisions and the plan so your permanent hire walks into a strong position. Related reading & paths. - What a temporary CTO costs: The pricing guide: it's the interim engagement, and it's priced the same way. - Fractional vs full-time vs interim: An honest breakdown of which kind of technology leadership your company actually needs right now. - Interim CTO: The same seat described from the transition side: cover the gap, carry the handoff. Need a CTO in the seat now? If the seat is empty, or you can see a hard stretch coming, let's talk today. ### Turnaround CTO for PE-Acquired Companies: /services/turnaround-cto/ Interim and turnaround CTO for private-equity portfolio companies, stabilizing engineering, de-risking the asset, and protecting the value of the thesis in 90 days. You closed the deal. Then the technical reality showed up, fragile delivery, key-person risk, security gaps, an undocumented codebase. I parachute in as interim CTO, stabilize the engineering organization, and protect the value of the thesis. Tags: Technical due diligence, Stabilization, SOC 2 · HIPAA, Exit-readiness What is a turnaround CTO? A turnaround CTO is an interim technology executive who takes over a struggling engineering organization, usually inside a private-equity portfolio company, and makes it stable, predictable and worth more. The mandate is broader than a normal interim seat: diagnose what's actually broken, stop the bleeding, rebuild the team, and tie the technology back to the investment thesis. The pattern repeats. The deal closes, the technical reality surfaces, and someone has to own the fix with real authority. I've stabilized six troubled SaaS products this way, and I've taken engineering organizations from low to elite DORA performance inside 90 days. Pricing is published too: a fixed-fee Diagnostic from $20,000 to establish the damage, then $30,000 a month embedded until delivery is stable, with the full breakdown in the turnaround CTO pricing guide. The work usually means modernizing a system nobody can afford to take offline. I've done that under harsher constraints than most portfolio companies face, including re-platforming a live public-transit system onto cloud-native infrastructure with zero downtime. Also known as: rescue CTO, crisis CTO, recovery CTO, PE portfolio CTO, engineering turnaround leadership. The diligence said "minor tech debt." Then you opened the hood. Most acquired companies were built to ship features, not to be owned, scaled, or sold again. That gap becomes your risk on day one. - Delivery you can't forecast: Releases are slow, manual and scary. No one can promise a date, which makes the value-creation plan a guess. - Key-person risk everywhere: The system lives in one or two people's heads. No docs, no tests. If they leave, the asset walks out the door. - Security & compliance gaps: SOC 2, HIPAA or GDPR exposure that scares enterprise buyers, and threatens the next round or the exit. - Cost & integration drag: Legacy infrastructure burning EBITDA, plus integration debt across an acquisitive roll-up that never gets paid down. A 90-day turnaround, built for the deal. I've stabilized six troubled SaaS products this way, inheriting codebases with no tests, no docs and no team, and turning them around in three to six months. - Days 0–30 · Diagnose · Find the truth, fast.: Technical & team due diligence; Architecture, security & cost review; ROI review of every active project; Shortlist: what to rethink or cancel; Key-person & delivery risk map; Stop the bleeding on the worst fires - Days 30–60 · Stabilize · Make delivery boring.: CI/CD, GitOps & trunk-based flow; Tests, docs & runbooks where it counts; Rebuild & reorganize the team; Close the urgent security gaps - Days 60–90 · Position · Align tech to the thesis.: Roadmap tied to value-creation plan; Organizational design for the AI age; SOC 2 / HIPAA path underway; Hire or coach the permanent leader; Board-ready reporting & metrics De-risked, and worth more at the next milestone. - Low→Elite: DORA delivery performance, typically within a quarter - <1hr: Downtime on on-prem-to-cloud migrations I design - SOC2I&II: Compliance achieved across HealthTech, InsureTech & e-commerce - 6+: Troubled SaaS products stabilized and turned around "I read the investment thesis before I read the code. The job isn't clean architecture for its own sake, it's protecting and compounding the value of the asset." Three ways in. - Pre-close: Technical Due Diligence: An honest read on the engineering, team and risk before you sign, so there are no surprises after. - Post-close: Interim / Turnaround CTO: I step in hands-on to stabilize, rebuild the team, and execute the first phase of the value-creation plan. - Ongoing: Portfolio Advisor: Fractional technology oversight across multiple portfolio companies, with board-ready reporting. What PE operators ask me. Q: How quickly can you start on a portfolio company? A: Usually within one to two weeks. Pre-close diligence engagements can start in days when a deal is moving, I keep capacity for time-sensitive transactions. Q: Do you work pre-close or only after acquisition? A: Both. Pre-close I run technical and team due diligence so there are no surprises after signing. Post-close I step in hands-on as interim or turnaround CTO to stabilize and execute the value-creation plan. Q: What does a 90-day turnaround actually deliver? A: A diagnosed, de-risked engineering organization: predictable delivery on CI/CD and GitOps, key-person risk mapped and reduced, urgent security and compliance gaps closed, a roadmap tied to the thesis, and a permanent leader hired or coached, with board-ready reporting throughout. Q: Can you cover multiple portfolio companies at once? A: Yes. As a Portfolio Advisor I provide fractional technology oversight across several companies in a fund, with consistent board-ready metrics and a shared playbook. Q: How do you charge for turnaround engagements? A: My pricing is published: a fixed-fee Diagnostic from $20,000 to establish what's actually broken, then $30,000 a month embedded at three or more days a week until delivery is stable and predictable again. Ongoing portfolio-advisory arrangements are also available. All prices USD; the full breakdown is on my turnaround CTO pricing page, and we agree on outcomes and reporting cadence up front. Related reading & paths. - What a turnaround costs: The pricing guide: the diagnostic-first model, the monthly number, and the cost of the alternative. - Low to elite DORA in 90 days: A field report from a real turnaround: what changed, in what order, and what it measured. - Technical Due Diligence: For PE: price the risk before the deal instead of discovering it after close. Got a portfolio company that needs a steady hand? Let's talk through the situation. I'll tell you straight what I'm seeing and what the first 90 days would look like. ## Industries ### Fractional & AI-Native CTO for E-commerce: /industries/ecommerce/ An AI-native fractional & interim CTO for e-commerce: peak-load scale, re-platforming legacy commerce with minimal downtime, and margin-finding AI agents across fulfillment, sourcing and returns. From platform scale to margin-finding AI agents across fulfillment, I build e-commerce technology that holds up on the busiest day of the year and finds profit in every order. Tags: Peak-load scale, Low-downtime re-platform, Margin-finding agents, Fulfillment AI A CTO who plans for the busiest day. A fractional CTO for e-commerce is an experienced Chief Technology Officer who leads an online retail company part-time: the platform architecture, the engineering team, and the systems that decide whether each order makes money. In commerce the judgment is specific. Peak-load scale, re-platforming risk, and thin margins that punish every inefficiency. I've built for the spike and migrated live stores with downtime measured in minutes. When a re-platform or a troubled org needs someone in the seat every day, the same work runs as an interim engagement, see turnaround CTO. My pricing is published in the fractional CTO cost guide. Also known as: online retail CTO, D2C CTO, direct-to-consumer commerce CTO, part-time e-commerce CTO. The busiest day finds every weakness. And thin margins punish every inefficiency the rest of the year. - Peak-load scale: Black Friday doesn't care about your architecture diagram. The platform has to absorb the spike without falling over or over-provisioning all year. - Legacy re-platforming: Migrating a legacy store is high-stakes, every hour of downtime is revenue. It has to be done with minimal disruption, not a big-bang gamble. - Margin compression: In a commodity market the price is fixed, so profit hides in the cost of each order: sourcing, shipping, returns. Finding it is an engineering problem. - Fulfillment complexity: Sourcing, routing, returns and exceptions multiply fast. This is exactly where well-designed AI agents earn their keep. Scale that holds, margin that shows. - Scale & re-platform: Architecture that absorbs peak load economically, and re-platforming legacy commerce with downtime measured in minutes, I've run zero-downtime migrations. - Margin-finding AI agents: A graph of AI agents that finds profit in the cost of each order, sourcing, shipping and returns, while holding a margin floor. - Fulfillment automation: Automate the routing, exceptions and returns that quietly eat operations time, with the right guardrails and observability. What e-commerce operators ask. Q: Can your systems handle peak traffic? A: Yes. Peak events like Black Friday expose every architectural weakness, so the platform has to absorb the spike without falling over, and without over-provisioning the rest of the year. I've built and re-platformed systems specifically to hold up on the busiest day while staying economical on the quiet ones. Q: How can AI improve e-commerce margins? A: In a commodity market where price is fixed, margin hides in the cost of each order. A graph of AI agents can find profit across sourcing, shipping and returns while holding a margin floor, turning a thin-margin operation into a measurably better one. I've designed exactly these fulfillment agent systems. Q: Can you re-platform our legacy store with minimal downtime? A: Yes. Re-platforming is high-stakes because every hour down is revenue. I've led sub-one-hour and zero-downtime migrations of production systems, sequencing the move so customers barely notice rather than betting the business on a big-bang cutover. Q: What does a fractional CTO for e-commerce cost? A: The same published pricing as the rest of my work, in USD: advisory at $10,000 a month (about a day a week), the full fractional engagement at $18,000 a month (about two days), and embedded or interim leadership at $30,000 a month. A fixed-fee Diagnostic starts at $20,000. The full breakdown is in my fractional CTO cost guide. Q: Do you work with D2C and online retail brands? A: Yes. Direct-to-consumer and online retail share the same engineering economics: fixed prices, thin margins, and profit hiding in the cost of each order. My multi-brand commerce work, one multi-tenant engine serving a portfolio of storefronts, comes straight from that world. Q: Can you take over during a turnaround or a migration? A: Yes. When a re-platform or a troubled engineering org needs someone in the seat every day, I work as an interim or turnaround CTO: full-time but temporary, stabilizing delivery, de-risking the migration, then handing back a stronger function. Scaling an e-commerce platform? Whether the problem is peak-day scale, a risky re-platform, or thin margins, I've solved each before. ### Fractional & AI-Native CTO for EdTech: /industries/edtech/ An AI-native fractional & interim CTO for EdTech, backed by lived experience running product and engineering for an online learning platform, scaling under enrollment spikes, content pipelines, learner-data privacy, and AI done responsibly. I've run product and engineering for an online learning platform. I know where EdTech breaks, and how AI changes the classroom, the content pipeline, and the cost structure. Tags: Online learning operator, Content pipelines, Learner-data privacy, Responsible AI A CTO who has run the platform. A fractional CTO for EdTech is an experienced Chief Technology Officer who leads an education technology company part-time: the architecture, the engineering team, the content pipeline, and the privacy obligations that come with learner data. You get senior judgment shaped by the industry's real constraints, term cycles, enrollment spikes, district procurement, without hiring a full-time executive. Mine is operator experience. I served as Chief Product & Technology Officer of an online learning platform, running product, engineering and AI-native transformation. When the need is full-time but temporary, the engagement runs as an interim CTO; my numbers are published in the fractional CTO cost guide. Also known as: education technology CTO, CTO for learning platforms, e-learning CTO, part-time EdTech CTO. The hard parts are predictable. I've lived most of these from inside the platform. - Enrollment-spike scale: Traffic isn't flat, it spikes with enrollment windows and term starts. The system has to hold up on the day that matters most, then not bankrupt you the rest of the term. - Content & curriculum pipelines: Authoring, versioning, and shipping curriculum at scale is its own engineering problem, and it's usually the one nobody resourced properly. - Learner-data privacy: Student data carries real obligations (FERPA, COPPA where relevant). Privacy has to be designed in, not retrofitted before an enterprise or district deal. - AI without harming outcomes: AI tutoring and grading can help or quietly erode learning. The hard part is using it where it lifts outcomes and keeping a human accountable where it counts. Operator experience, applied. - Build the team & roadmap: Set technology direction for the platform, hire the engineering and product people, and protect a roadmap that survives term cycles. - AI-native learning: Rework the content pipeline and the product so AI is the default, tutoring, generation, grading support, with learning outcomes and privacy protected. - Scale & privacy: Architecture that absorbs enrollment spikes economically, and a privacy posture that passes district and enterprise diligence. What EdTech founders ask. Q: Do you have EdTech operating experience? A: Yes, I served as Chief Product & Technology Officer of an online learning platform, running product, engineering and AI-native transformation. This isn't industry I read about; it's industry I've operated in. Q: How can AI help an EdTech product without harming learning outcomes? A: By using AI where it genuinely lifts outcomes, personalization, tutoring support, faster content production, grading assistance, while keeping a human accountable for anything that affects a learner's progress, and measuring outcomes rather than assuming the AI helped. The goal is leverage on learning, not automation for its own sake. Q: Can you handle student-data privacy? A: Yes. Learner data carries real obligations (FERPA and COPPA where they apply), and privacy has to be designed into the architecture rather than bolted on before a district or enterprise deal. I've delivered SOC 2 and HIPAA-grade programs in adjacent regulated spaces and apply the same discipline here. Q: What does a fractional CTO for EdTech cost? A: The same published pricing as the rest of my work, in USD: advisory at $10,000 a month (about a day a week), the full fractional engagement at $18,000 a month (about two days), and embedded or interim leadership at $30,000 a month. A fixed-fee Diagnostic starts at $20,000. The full breakdown is in my fractional CTO cost guide. Q: Can you handle enrollment-spike scale? A: Yes. EdTech traffic spikes with enrollment windows and term starts, so the architecture has to hold up on the day that matters most without over-provisioning the rest of the term. I design learning platforms for the spike, and for the cloud bill that follows it. Q: Do you take interim engagements in EdTech? A: Yes. If the CTO seat is empty or a transition needs carrying, I step in full-time as an interim CTO, stabilize delivery, and hand off cleanly to the permanent hire. Fractional is the default when you need part-time judgment; interim covers the gap when the need is daily. Building in EdTech? If your platform is straining at the seams, or you want it AI-native before your competitors do, let's talk. ### Fractional & AI-Native CTO for HealthTech: /industries/healthtech/ An AI-native fractional & interim CTO for HealthTech, backed by experience as CTO of a healthcare EMR platform: SOC 2 Type I & II, HIPAA-grade compliance, legacy EMR integration, and AI in clinical workflows with a human accountable. Former CTO of a healthcare EMR platform. SOC 2 Type I & II and HIPAA-grade compliance delivered across HealthTech, without slowing the roadmap. Tags: EMR platform CTO, SOC 2 I & II, HIPAA-grade, PHI security A CTO who has shipped in healthcare. A fractional CTO for HealthTech is an experienced Chief Technology Officer who leads a healthcare technology company part-time, bringing the compliance, integration and security judgment the industry demands without the cost of a full-time executive. In healthcare that judgment is specific: HIPAA, SOC 2, PHI security, and the legacy EMR landscape every digital health product has to live with. I've done the job from the inside, as CTO of a healthcare EMR platform, and delivered SOC 2 Type I & II and HIPAA-grade compliance without freezing the roadmap. When the seat needs full-time coverage rather than part-time judgment, the same work runs as an interim CTO engagement. My pricing is published in the fractional CTO cost guide. Also known as: CTO for healthcare startups, healthcare technology CTO, medtech CTO, digital health CTO, part-time HealthTech CTO. Compliance is the gate. And the roadmap can't stop while you clear it. - HIPAA & SOC 2 for sales: Enterprise and hospital sales stall without SOC 2 and HIPAA-grade controls. Getting there without freezing the roadmap is the real trick. - Legacy EMR & integration: Healthcare runs on legacy EMR and a thicket of integrations. Modernizing without breaking clinical workflows takes someone who's done it. - PHI security: Protected health information raises the stakes on every architectural decision, storage, access, audit, vendors. Mistakes here are existential. - AI in clinical workflows: AI can speed documentation, triage and decision support, but a human has to stay accountable for anything that touches care. Compliance and velocity, together. - SOC 2 & HIPAA programs: Compliance programs that pass audits and unlock enterprise and hospital sales, built into delivery, not bolted on at the end. - EMR & integration: Modernize legacy EMR and integration layers without breaking the clinical workflows people depend on. - Responsible clinical AI: Bring AI into documentation, triage and decision support with PHI protected and a clinician accountable at the end. What HealthTech founders ask. Q: Can you get us SOC 2 and HIPAA-ready? A: Yes. I've delivered SOC 2 Type I & II across HealthTech and adjacent regulated industries and led teams under HIPAA protocols. The aim is to build compliance into how the team ships so it unlocks enterprise and hospital sales without freezing the roadmap. Q: Do you have healthcare platform experience? A: Yes, I served as CTO of a healthcare EMR platform, modernizing an integrated care system and meeting its compliance obligations. I know EMR, integration and PHI from the inside, not from a deck. Q: How do you use AI safely with patient data? A: By keeping PHI tightly controlled, using AI where it clearly helps, documentation, triage, decision support, and keeping a clinician accountable for anything that affects care. Evaluation, observability and access controls are built in, not assumed. Q: What does a fractional CTO for HealthTech cost? A: The same published pricing as the rest of my work, in USD: advisory at $10,000 a month (about a day a week), the full fractional engagement at $18,000 a month (about two days), and embedded or interim leadership at $30,000 a month. A fixed-fee Diagnostic starts at $20,000. The full breakdown is in my fractional CTO cost guide. Q: Do you work with early-stage digital health and medtech startups? A: Yes. A healthcare startup rarely needs a full-time CTO before it has real scale, but it does need HIPAA, PHI security and integration decisions made right from the start, because retrofitting compliance costs far more than designing it in. Part-time fractional leadership fits that stage. Q: Can you cover the CTO seat full-time during a transition? A: Yes. When a HealthTech company loses its CTO or is mid-transition, I step in as an interim CTO: full-time but temporary, keeping delivery and the compliance program moving, then handing off cleanly to the permanent hire. Building in HealthTech? If compliance is blocking sales or legacy EMR is slowing you down, I've cleared both before. ### Fractional & AI-Native CTO for SaaS: /industries/saas/ An AI-native fractional & interim CTO for SaaS: 25 years in SaaS and enterprise software, scaling engineering orgs without breaking them, taking delivery from low to elite DORA in 90 days, and rescuing troubled products. Twenty-five years in SaaS and enterprise software. I scale engineering orgs without breaking them, take delivery from low to elite DORA in 90 days, and rescue troubled products. Tags: 25 yrs in SaaS, DORA low → elite, Org scaling, Multi-tenant platforms A CTO who has scaled the org. A fractional CTO for SaaS is an experienced Chief Technology Officer who leads a software-as-a-service company part-time: the architecture, the engineering organization, and the delivery machine that decides whether the roadmap ships. In SaaS the failure mode is predictable. The org scales faster than the trust and process that held it together, and delivery quietly slows to a crawl. I've spent 25 years in SaaS and enterprise software, 20 of them in leadership. I've scaled the orgs, rescued the troubled products, and taken a low-performing team to elite DORA metrics in 90 days. When the seat needs full-time coverage, the same work runs as an interim / turnaround CTO; my pricing is published in the fractional CTO cost guide. Also known as: B2B SaaS CTO, software-as-a-service CTO, CTO for SaaS scale-ups, part-time SaaS CTO. The org outgrows its operating model. The product usually survives growth. The organization is what snaps. - Scaling breaks the org: Adding people adds process, and process grinds down the trust and speed that got you here. Most SaaS orgs break from that, not from headcount. - Delivery slows to a crawl: Deploys get scarier, lead times stretch, and nobody can say exactly why. DORA metrics make the rot measurable, and measurable means fixable. - The troubled product: The platform underdelivers, the team is stuck, and the board wants a straight answer before more money goes in. A rescue, not a rewrite by reflex. - Multi-tenant economics: One engine serving many customers is the whole SaaS promise. Getting tenancy, isolation and cost-to-serve right is an architecture problem. Delivery you can measure. - Scale the engineering org: Grow capacity without grinding down the trust and speed that made the team good. I've scaled SaaS orgs and written down where they break. - DORA transformation: Take delivery from low-performing to elite DORA in 90 days: process, GitOps and CI/CD restructured so shipping is fast and predictable. - AI-native operating model: Rewire how the product is built so AI is the default in the SDLC and the org. For a SaaS company, that edge compounds with every release. What SaaS founders ask. Q: What does a fractional CTO do for a SaaS company? A: Leads the technology function part-time: architecture, the engineering organization, hiring, and the delivery process that determines whether the roadmap ships. For a SaaS scale-up that usually means growing the org without breaking it, making delivery measurable with DORA, and building toward an AI-native operating model. Q: Can you fix our delivery speed? A: Yes. I've taken a low-performing org to elite DORA metrics in 90 days by restructuring process, GitOps and CI/CD so delivery became fast and predictable. The point of DORA is that delivery stops being a feeling and becomes a number you can manage. Q: Have you scaled a SaaS engineering organization before? A: Yes. I've spent 25 years in SaaS and enterprise software, 20 in technology leadership, at one point directing 12 engineering teams across 7 countries, and I've rescued troubled SaaS products along the way. Scaling breaks orgs through added process more than added people; the job is protecting speed and trust on the way up. Q: Do you handle multi-tenant architecture? A: Yes. I've architected multi-tenant platforms, including a commerce and marketing engine serving a portfolio of brands from one codebase, and multi-tenant SaaS in regulated healthcare. Tenancy, isolation and cost-to-serve are architecture decisions that set your margins for years. Q: What does a fractional CTO for SaaS cost? A: The same published pricing as the rest of my work, in USD: advisory at $10,000 a month (about a day a week), the full fractional engagement at $18,000 a month (about two days), and embedded or interim leadership at $30,000 a month. A fixed-fee Diagnostic starts at $20,000. The full breakdown is in my fractional CTO cost guide. Q: Can you take over full-time in a turnaround? A: Yes. When a SaaS product is troubled or the CTO seat is empty, I work as an interim or turnaround CTO: full-time but temporary, stabilizing engineering and delivery, then handing off cleanly, to the permanent hire or back to a stronger team. Scaling a SaaS platform? If the org has stopped scaling or delivery has slowed to a crawl, I've fixed both before. Tell me where it hurts. ## Guides & hubs ### About Oshri Cohen: /about/ Oshri Cohen is an AI-Native Chief Product & Technology Officer and fractional & interim CTO with 25 years in software. Since 2018 he has served 30+ companies, at one point directing 12 engineering teams across 7 countries. He builds AI-native engineering and product organizations, turns around troubled software, and leads technical due diligence for private-equity deals. USA & remote. I'm Oshri Cohen — an AI-Native Chief Product & Technology Officer who works as a fractional and interim CTO. I come in to solve hard problems fast: building AI-native engineering and product organizations, turning around troubled software, and leading technical due diligence for private-equity deals. Hands-on, from the inside out, until the problem is solved. Tags: AI-Native CPTO, Fractional & Interim CTO, 25 years in software, USA · Remote Two and a half decades, measured in outcomes. - 25yrs: In SaaS & enterprise software, 20 of them in technology leadership - 30+: Companies served as fractional & interim CTO since 2018, almost all US-based - 12: Engineering teams directed at once, across 7 countries & 4 time zones - 10+: Digital products delivered end-to-end across five industries A product-minded technologist who translates both ways. I've spent twenty-five years working the seam between the boardroom and the codebase. Most of what I do there is translation: turning business strategy into engineering that ships, then turning engineering reality back into something a board can actually decide on. Lately that work keeps landing on the same request. Companies want to become AI-native, and that's a much bigger job than adding a feature. It means changing how the product gets built and how the organization runs, until AI is just the default way things work instead of a layer stapled on at the end. Along the way I've shipped more than ten digital products in HealthTech, e-commerce, manufacturing, logistics and finance, usually with my hands on the architecture, the pipeline, the UX and now the AI. Since 2018 I've done it as a fractional and interim CTO, for north of thirty companies, nearly all of them in the US. At the busiest stretch I was directing twelve engineering teams scattered across seven countries. Before the consulting years I sat in the full-time chairs too: VP of Engineering at an intelligent-transportation company, then CTO of a healthcare EMR platform. The thread through all of it isn't a favorite technology. It's the order I make decisions in. I start from the product and the business and let the architecture, the team and the tooling fall out of that, rather than picking a stack first and hoping the business catches up. I've turned that habit into a method I can repeat, but the method only earns its keep because I build. I write code, I hire, I run the team. I don't hand over a deck and disappear. And the thing I keep coming back to is whether the work actually holds up for the people using it: customers, sure, but also the operators quietly keeping the business on its feet. I work with companies across the US, remotely, on Eastern Time. My own path ran in the less common direction. I studied business at McGill and learned to code afterward, and that order turned out to be the whole advantage: I can read a cap table and a pull request in the same afternoon and keep both straight. Also known as: AI-native CTO · fractional CTO · interim CTO · Chief Product & Technology Officer · CPTO · technical co-founder for hire. "I think from the product and the business first, then rebuild how the company operates so AI is the default — an AI-native transformation, not a feature bolted on, in service of business development.": Oshri Cohen · The operating philosophy behind every engagement The hard problems I'm brought in to solve. - 01 · AI-native transformation: Rebuilding how the product is built and how the org runs so AI is the default. Production AI in the product and the pipeline, an AI-first SDLC, and the AI-native team to carry it — measured against the P&L, not against demos. - 02 · Fractional & interim CTO: The senior technology seat, part-time or in the gap. Strategy, architecture, hiring and delivery owned end-to-end, sized to what the company actually needs and built to hand off to a permanent leader. - 03 · PE technical due diligence: A credible technical read before a private-equity deal closes, and an AI-native value-creation plan after. Architecture, team, security, spend and AI leverage, assessed against the thesis the investment rests on. - 04 · Turnarounds & troubled software: Troubled products made predictable. I find the real root cause, not the loudest symptom, then stabilize delivery, rebuild trust with the board, and drive the team toward elite DORA performance. - 05 · Legacy modernization: Modernizing systems that have outgrown their architecture — and traditional organizations making their first serious software or AI bet, where there's no legacy stack to fight and the operating model can be designed AI-native from the start. - 06 · Security & compliance: Security and compliance built into the work, not bolted on after. SOC 2, HIPAA and GDPR posture set up so it survives an audit and a scale-up, with security in every pull request. Most advisors stop at the deck. I keep going. The typical advisor: Hands the company a strategy and leaves: Delivers a deck, a workshop and a list of AI use cases, then exits; Talks about AI in the abstract; never ships it into production; Treats the org chart as someone else's problem to solve later; Measures success by the strategy being accepted, not by the outcome; Leaves the hard part — building it — to a team that wasn't there for the thinking How I work: Owns the problem and stays until it's solved: Sets the direction, then builds it: production AI in the product and pipeline; Ships hands-on — architecture, code, security, eval and cost control; Designs the architecture and the org chart together, both derived from the P&L; Measured on the outcome: margin moved, risk retired, delivery made predictable; Hires and runs the AI-native team, then hands off — no permanent dependency The career behind the method. Executive seats, a founding team, and seven years as a fractional CTO. Where the experience comes from. - Oct 2025 — Jun 2026 · Chief Product & Technology Officer · EdTech: Led product, engineering and AI-native transformation for an online learning platform; Owned the full product and technology org from strategy through delivery - 2018 — Now · Consulting & Fractional CTO · USA · Remote: Strategic technology partner to 30+ startups and growth-stage companies; At one point directing 12 engineering teams across 7 countries and 4 time zones; AI-native transformation, turnarounds, PE due diligence and legacy modernization - 2016 — 2018 · VP of Software Development · Intelligent Transportation: Restructured into bimodal teams for a +30% gain in delivery efficiency; Introduced CI/CD that cut deploy times by 40% - 2014 — 2016 · Chief Technology Officer · Healthcare EMR Platform: Modernized the integrated care platform; Owned health-data compliance across the product - Earlier · Director, Principal & founding-team roles: Director of Technology at a hospitality-wellness platform — POS integration across hundreds of restaurant systems; Principal Engineer on the founding team of a no-code automation startup — task execution engine & workflow orchestration; Director of Engineering at a 350+ person market-research firm — digital transformation of the research & analytics platform - 2006 · BA, Business Administration · McGill University: FounderFuel mentor and Montreal CTO Meetup co-organizer; Business first, engineering second — the order that ended up shaping how I work What you can count on. - Business first, always: Every technical decision starts at the P&L — revenue, margin, risk — and is reasoned down to the stack, never the other way around. - Hands-on, not hand-wavy: I build, ship and hire inside the engagement. The part most advisors skip is the part I'm there for. - AI-native by default: AI is evaluated at every turn as a lever and measured against the business, not adopted because it's the trend. - Measured continuously: DORA for delivery, cost and quality for AI systems. Progress stays legible to the board, not a matter of faith. - Both languages, fluently: I read a cap table and a Kubernetes manifest in the same afternoon, and translate each into the other. - Built to hand off: Every engagement ends in a transition — an internal hire, an advisory cadence, or a trained team. No permanent dependency. About me, answered. Q: Who is Oshri Cohen? A: Oshri Cohen is an AI-Native Chief Product & Technology Officer who works as a fractional and interim CTO. He has 25 years in software, 20 of them in technology leadership, and since 2018 has served 30+ companies — almost all US-based — building AI-native engineering and product organizations, turning around troubled software, leading technical due diligence for private-equity deals, and modernizing legacy systems. He works hands-on, from the inside out, and serves companies across the USA remotely on Eastern Time. Q: What does an AI-Native Chief Product & Technology Officer do? A: It's the senior seat that owns both product and technology, run AI-native. Rather than bolting AI on as a feature, Oshri rebuilds how the product is built and how the organization runs so AI is the default: production AI in the product and the pipeline, an AI-first software development lifecycle, and the team and operating model to carry it — all derived down from the business, not from the technology. Q: What kinds of companies does Oshri work with? A: Primarily $10M–$50M software and software-enabled companies past product-market fit, where a technology decision now moves the valuation rather than just the backlog — founders, CEOs and boards, plus private-equity investors who need a technical read before a deal and a value-creation plan after. He also works with traditional organizations making their first serious software or AI bet, where the operating model can be designed AI-native from the start. Q: Fractional, interim, or full CPTO — what's the difference in how he engages? A: Fractional means the senior technology seat part-time, on an ongoing basis. Interim means stepping fully into the gap when a company is between CTOs and needs the role owned now. CPTO means owning product and technology together. In every case the work is sized to what the company actually needs and built to hand off to a permanent leader — never a permanent dependency. Q: Where is Oshri based, and does he work remotely? A: He serves companies across the USA and works remotely on Eastern Time. The practice is remote-first; engagements run distributed, which is also how he has directed as many as 12 engineering teams across 7 countries at once. Q: What is his background? A: A McGill University business graduate (BA, Business Administration, 2006) who learned to code — the wrong-way-round path that lets him translate between the boardroom and the codebase. He has held executive seats as VP of Engineering at an intelligent-transportation company and CTO of a healthcare EMR platform, was on the founding team of a no-code automation startup, and has run engineering at a 350+ person firm. Since 2018 he has worked as a fractional and interim CTO. Q: How does he actually run an engagement? A: Through The Business-Down Method: a repeatable, four-movement approach — Read, Direct, Build, Operate & Optimize — where every technical decision is derived down from the P&L. The discipline is fixed; the solution is built with today's state of the art and bespoke to your numbers every time. Q: How do I get in touch? A: Email hello@oshricohen.me or call (514) 777-3883. The lowest-risk way to start is a focused Read — a diagnosis of system, org, spend and AI leverage that ends in a 90-day plan you keep whether or not we continue. Have a problem worth solving? If the technology and the business are tangled together and someone has to untangle them on purpose, that's the work I'm built for. The fastest way to find out if I can help is a focused Read — yours to keep either way. ### How Much Does an AI-Native Transformation Cost?: /ai-native-transformation-cost/ A plain guide to what becoming genuinely AI-native costs: the strategy, the leadership, and the spend that actually moves, with published numbers for each stage. The honest answer is that it's an operating change, not a project, so it prices as leadership over 6–18 months plus the AI spend itself. Here's the full stack of numbers, published. Tags: Published USD pricing, Staged, with off-ramps, Leverage you can measure Strategy in the tens of thousands. Transformation as a monthly run-rate. An AI-native transformation isn't a fixed-price project, because what's being rebuilt is how the company works: the SDLC, the operating workflows, the team. It prices in three layers. The strategy layer is fixed: a Diagnostic from $20,000 and the full Roadmap & Architecture Sprint from $35,000. The leadership layer is monthly: $18,000 a month fractional, or $30,000 a month embedded, typically for 6–18 months. The third layer is the spend the transformation redirects: model and tooling costs, the agentic systems you build, and the hires. That's not my fee, but a real strategy budgets it honestly, with evals and cost controls attached, because AI spend without measurement is how companies end up paying for leverage they never get. Staged this way, the decision is never a leap. Twenty thousand dollars buys the read. Thirty-five buys the plan. Only when the plan is credible do you commit to the monthly engagement, and the leverage should start showing up in cost, speed and quality inside the first quarter. Related searches: AI transformation cost, AI-native company, cost of adopting AI, enterprise AI transformation pricing. What I charge, exactly. Three stages, each an off-ramp. All prices USD. - The Diagnostic (From $20,000 · Fixed · 2–4 weeks · yours to keep): Where AI creates real leverage in your product and operations, and where it's a distraction. Plus a 90-day plan you can act on with or without me. - AI-Native Roadmap & Architecture Sprint (From $35,000 · Fixed scope · 4–6 weeks): The complete strategy: roadmap, architecture and model choices, economics, team and governance. Ready to build against. - Transformation Leadership ($18,000–$30,000 / mo · Fractional to embedded · 6–18 months): I run the change: ship the systems, rewire the workflows, hire and upskill the team, and hold it to measurable leverage. All prices USD. The full offering is on the AI-Native Transformation page. Email me ↗ What founders & boards ask about price. Q: How much does an AI transformation cost overall? A: For a $10M–$50M company working with me: a Diagnostic from $20,000 and the full strategy sprint from $35,000, both fixed-scope, then $18,000–$30,000 a month in transformation leadership for six to eighteen months, plus the AI spend the roadmap itself budgets (models, tooling, hires). Staged so each step is an off-ramp. All prices USD. Q: Why isn't it a fixed-price project? A: Because the deliverable is a company that operates differently. Workflows, the SDLC and the org chart change over months, and the state of the art moves every quarter. Fixed pricing that pretends otherwise either pads the number heavily or quietly narrows the scope back down to a pilot. Q: What determines whether it's $18K or $30K a month? A: Time in the seat. At roughly two days a week ($18,000 a month) I lead the transformation alongside your existing leadership. Embedded at three or more days ($30,000 a month) fits companies where I'm also covering the CTO role or the change is organization-wide and fast-moving. Q: When does the spend start paying back? A: The leverage should be visible inside the first quarter: cheaper cost-per-outcome in the workflows you rebuilt, faster delivery in the SDLC, and quality you can measure with evals rather than anecdotes. If it isn't showing up in the numbers by then, the plan is wrong and we change the plan. Q: Can we start smaller than a transformation? A: Yes, and you should. The $20,000 Diagnostic is the front door: it tells you where AI actually pays in your business and what a credible first 90 days looks like. Some companies stop there and execute internally. That's a fine outcome; the plan is yours to keep. Related reading & paths. - AI-Native Transformation: The engagement itself: rebuilding how the product is built and how the org runs. - AI strategy cost: Just the strategy layer: what the roadmap and architecture stage costs on its own. - The five phases of AI adoption: Where companies actually stall on the way to AI-native, and what each phase costs them. Pricing an AI-native move for your board? Tell me the shape of the company. I'll give you the staged numbers and what each stage has to prove before the next one. ### How Much Does an AI Strategy Cost?: /ai-strategy-cost/ A plain guide to AI strategy consulting pricing: what a real AI roadmap costs, why big-firm engagements run to six figures, and my published fixed-scope numbers. What AI strategy consulting actually costs in the US market, why the big-firm version runs to six figures, and my published fixed-scope pricing, including what you get for it. Tags: Published USD pricing, Fixed scope, Written by the builder A real one: fixed-scope, from $20K. A big-firm one: six figures. AI strategy consulting spans an absurd price range because the label covers everything from a workshop to a six-month engagement with a staffed team. Big consulting firms routinely price AI strategy work well into six figures. Boutiques and independents usually land between $20,000 and $100,000 depending on depth. My pricing is published and fixed-scope: the Diagnostic from $20,000 establishes where AI creates leverage in your business, and the AI-Native Roadmap & Architecture Sprint from $35,000 delivers the complete strategy — sequenced roadmap, architecture and model choices, and the economics — ready to build against in four to six weeks. One thing the fee should always buy: a strategy the author is willing to build. Mine converts directly into delivery at $18,000 a month as your fractional AI CTO, which keeps the strategy honest. Nobody writes fantasy roadmaps they'll be accountable for shipping. Related searches: AI strategy consulting rates, AI roadmap cost, AI consultant pricing, fractional AI CTO cost. What I charge, exactly. Strategy is fixed-scope; delivery is monthly. All prices USD. - The Diagnostic (From $20,000 · Fixed · 2–4 weeks · yours to keep): The Business-Down read on system, org, spend and AI leverage, plus a 90-day plan you can act on with or without me. The best first step. - AI-Native Roadmap & Architecture Sprint (From $35,000 · Fixed scope · 4–6 weeks): The full AI strategy: opportunity audit, sequenced roadmap, architecture and model choices, and the economics. Ready to build against. - Fractional AI CTO ($18,000 / mo · ~2 days a week): I stay to lead the build: architecture, hiring, delivery and the operating-model change. Hands-on, until the leverage is real. All prices USD. The full offering is on the AI Strategy page. Email me ↗ What founders & boards ask about price. Q: How much does an AI strategy cost? A: My published pricing: a fixed-fee Diagnostic from $20,000 (two to four weeks), and the full AI-Native Roadmap & Architecture Sprint from $35,000 (four to six weeks). In the wider market, boutiques run $20,000–$100,000 and big consulting firms routinely charge six figures for staffed engagements. All prices USD. Q: Why do big-firm AI strategy engagements cost so much more? A: You're paying for a staffed team, a partner margin, and a brand the board recognizes. What you're usually not paying for is the person who will build it. The strategy is priced as a document; execution is a second, larger engagement, often with a different team. Q: What does the $35K sprint actually include? A: The complete strategy: an opportunity audit grounded in your data and margins, a roadmap sequenced by P&L impact and feasibility, architecture and model choices with evals and cost controls specified, the AI economics, and the team and governance plan. It's written to be built against, not presented and shelved. Q: How much does it cost to execute the strategy? A: Ongoing delivery leadership runs $18,000 a month at roughly two days a week as your fractional AI CTO, scaling to $30,000 a month embedded if the transformation needs near-full-time leadership. Engagements that continue into delivery usually run six to eighteen months. Q: Is a cheap AI strategy worth buying? A: A workshop or a template roadmap costs less and is usually worth what it costs. The expensive part of a bad AI strategy isn't the fee; it's the year your team spends building the wrong things while competitors compound. Judge the price against who is accountable for the outcome, not against other decks. Related reading & paths. - AI Strategy consulting: The engagement itself: what's in the strategy and how it converts into delivery. - AI transformation cost: The bigger number: what becoming genuinely AI-native costs end to end. - Fractional CTO cost: The delivery tier: what ongoing senior technology leadership costs per month. Budgeting an AI strategy for this year? Tell me the shape of the business and where AI is on the board agenda. I'll tell you what the right scope costs, in writing. ### What is an AI Strategy?: /ai-strategy/ A plain-English guide to the AI strategy: what it is, what a real one contains, who should own it, and how to tell a working strategy from expensive theater. A plain answer, what a real one contains, and how to tell it from the deck-shaped theater sold under the same name. Written by someone who writes them and then has to build them. Tags: Plain English, No sales pitch, Written by the builder A plan for where AI pays. An AI strategy is a plan for where AI makes your business measurably better and what it takes to get there: which workflows and products it should run, the architecture and models behind it, what it does to your cost structure, and what your team has to become. It's a business document with an engineering spine. The test is simple. A real AI strategy survives contact with your codebase and your P&L. It names the systems to build, sequences them by payback and feasibility, budgets the model and tooling spend, and says who is accountable for shipping each piece. If it can't do that, it isn't a strategy; it's a vision statement with a logo. Ownership matters as much as content. Strategies written by people who will never build them drift toward the impressive; strategies written by the person accountable for delivery stay honest. That's why mine come attached to a builder. Also known as: enterprise AI strategy, AI roadmap, AI adoption strategy, corporate AI strategy, AI transformation plan. Related searches: AI consultant, AI consulting services. What a real one contains. - Opportunity map: Where AI creates leverage in your product and operations, grounded in your data and margins, and where it's a distraction to skip. - Sequenced roadmap: What to build first and why, ordered by payback and feasibility, with the prerequisites (data, evals, security) scheduled. - Architecture & models: Build vs buy, which models, agentic or not, and how quality gets measured once it's live. - The economics: Unit costs, token spend and cost curves. The line between a viable system and an expensive demo is drawn here. - Team & operating model: Who gets hired, who gets upskilled, and how workflows change so AI becomes the default way work gets done. - Governance: Usage policy, data boundaries and review gates, so speed doesn't come at the price of an unreviewed prompt in production. Strategy vs. theater. AI theater: Impressive, then shelved: Use cases ranked by demo appeal; No economics: costs appear after the build starts; Authored by people who leave before delivery; Measured in pilots launched, not leverage shipped A working AI strategy: Specific enough to be wrong: Use cases ranked by P&L impact and feasibility; Unit economics and eval criteria written down up front; Owned by someone accountable for shipping it; Measured in cost, speed and quality you can see The AI strategy, explained. Q: What is an AI strategy in simple terms? A: A plan for where AI makes your business measurably better and what it takes to get there. It covers which workflows and products AI should run, the architecture and models behind it, the costs, the team, and the governance, sequenced into a roadmap you can build against. Q: What should an AI strategy include? A: Six things: an opportunity map grounded in your data and margins, a roadmap sequenced by payback and feasibility, architecture and model choices with evaluation criteria, the unit economics, the team and operating-model plan, and governance. If any of the six is missing, you'll discover it mid-build, at the expensive moment. Q: Who should own the AI strategy? A: Someone senior enough to change the operating model and technical enough to be accountable for delivery — a CTO, or a fractional AI CTO if you don't have one. Strategies that emerge bottom-up from crowdsourced tool adoption almost never scale into transformation; top-down architecture with clear governance is what separates scaled results from expensive experimentation. Q: How is an AI strategy different from digital transformation? A: Today they're converging: the highest-leverage digital transformation available to most companies is an AI transformation. The difference is that a classic digital transformation moved existing processes onto software, while an AI-native move redesigns the processes themselves around what AI can now do. Q: How much does an AI strategy cost? A: Fixed-scope work from an independent runs in the tens of thousands; big-firm engagements routinely reach six figures. My published pricing is a Diagnostic from $20,000 and a full AI-Native Roadmap & Architecture Sprint from $35,000, in USD. Q: When does a company need an AI strategy? A: When AI activity is happening without leverage: pilots that don't reach production, board pressure without a plan, or spend you can't tie to results. If AI could plausibly change your cost structure or your product and nobody owns the plan for that, you're already late. Q: How often should an AI strategy be revisited? A: Quarterly, at minimum. The model and cost landscape moves fast enough that what was state of the art last quarter is table stakes today. The strategy should name owners for tracking model, tooling and cost curves and re-running evaluations as things drift. Related reading & paths. - AI Strategy consulting: The engagement: I write the strategy as your fractional AI CTO, then stay to build it. Published pricing. - What an AI strategy costs: The pricing guide: fixed-scope numbers, what the big firms charge, and what the fee should buy. - The five phases of AI adoption: Where companies actually stall on the way from tools to transformation, and how to tell which phase you're in. Have AI activity but no strategy? Tell me what's running and what the board is asking for. I'll tell you honestly what a real plan looks like for your company. ### How Much Does a Fractional CTO Cost?: /fractional-cto-cost/ A plain guide to fractional CTO pricing: typical US hourly rates and monthly retainers, how they compare to a full-time hire, and my own published numbers. The real numbers: what the US market charges, what moves the price, and what I charge. No "book a call to find out." I publish my pricing so you can self-qualify before we ever talk. Tags: Published USD pricing, No sales call required, 30+ companies since 2018 Roughly $8K–$30K a month, depending on time and seniority. In the US market, a senior fractional CTO typically charges $200–$400 an hour, which works out to roughly $8,000–$30,000 a month depending on how many days a week you need and how senior the person is. Cheaper exists. So does more expensive. Below that band you're usually buying an architect or a coach wearing the title; above it you're paying a firm's margin. My own numbers are published: $10,000 a month for advisory at about a day a week, $18,000 a month for the full fractional engagement at about two days, and $30,000 a month embedded at three or more. A fixed-fee Diagnostic starts at $20,000 and is the best first step. For comparison: a full-time CTO at a $10M–$50M company runs $400K–$750K a year all-in once salary, bonus, equity and benefits are counted, and the search alone takes three to six months. A fractional CTO gives you 60–80% of that value for a fraction of the cost, starting within the first month. Related searches: fractional CTO rates, fractional CTO pricing, part-time CTO cost, CTO-as-a-Service pricing. Why quotes vary so much. - Time commitment: One day a week is advisory. Two is real leadership. Three or more is an executive in the seat. Most of the price is here. - Hands-on vs advisory: Someone who reviews decks costs less than someone who makes the calls, hires the team, and owns delivery. You usually need the second kind. - Scope & stakes: A stable team needing direction costs less than a turnaround, a compliance push, or an AI-native rebuild with the board watching. - Seniority & track record: A CTO who has done it thirty times pattern-matches your problem in week one. That compresses the timeline, which is where the savings actually live. - Direct vs firm: Marketplaces and firms add an account manager and a margin between you and the person doing the work. Direct is cheaper and faster. - Engagement length: Most real engagements run 6–18 months. Anything priced per-hour with no end state is a meter running, not a plan. What I charge, exactly. The same tiers published on the Fractional CTO page. All prices USD. - The Diagnostic (From $20,000 · Fixed · 2–4 weeks · yours to keep): The Business-Down read on system, org, spend and AI leverage, plus a 90-day plan you can act on with or without me. The best first step. - Advisory ($10,000 / mo · ~1 day a week): Senior judgment on direction, architecture, hiring and roadmap. For teams that can execute but need the calls made right. - Fractional CTO ($18,000 / mo · ~2 days a week): The full method, running. I set direction, build and hire the team, and own delivery — hands-on. Most engagements run 6–18 months. - Embedded / Interim CTO ($30,000 / mo · 3+ days a week): Near-full-time leadership in the seat — for turnarounds, transitions, or covering the CTO role until the permanent hire lands. Prefer a fixed scope? AI-Native Roadmap & Architecture sprints from $35K, PE technical due diligence from $15K/deal. Email me ↗ What people ask about price. Q: How much does a fractional CTO cost per month? A: In the US market, roughly $8,000–$30,000 a month depending on days per week and seniority. My published tiers are $10,000 a month for advisory (~1 day a week), $18,000 a month for the full fractional engagement (~2 days), and $30,000 a month embedded (3+ days). All prices USD. Q: What do fractional CTOs charge per hour? A: Senior US fractional CTOs typically charge $200–$400 an hour. I don't bill hourly; I price by monthly commitment tier or fixed scope, because the value is in owning outcomes over months, not in metered hours. Q: How does a fractional CTO compare to a full-time CTO on cost? A: A full-time CTO at a $10M–$50M company costs $400K–$750K a year all-in, plus a three-to-six-month search and ramp. A fractional CTO at $18,000 a month is $216K a year, starts within weeks, and scales down when the need does. You get 60–80% of the value at a fraction of the cost and none of the severance risk. Q: Is a fractional CTO worth the cost? A: It depends on the stakes. If delivery is stalling, technical debt is compounding, or you're facing a raise, a sale, or an AI rebuild without senior technology judgment, the cost of not having a CTO is usually a multiple of the fee. If your team just needs occasional code review, hire a consultant instead. The fixed-fee Diagnostic exists so you can find out for $20,000 which one you are. Q: Are there hidden costs? A: No. My pricing is published and all-in: no recruiting fee, no equity ask by default, no firm margin, no account manager. You pay the monthly tier or the fixed scope, under a standard consulting agreement. Related reading & paths. - Fractional CTO services: The engagement itself: what I actually do, how I work, and the method behind it. - What is a fractional CTO?: The plain-English explainer: what the role is and when you actually need one. - Fractional vs full-time vs interim: Which kind of technology leadership your company actually needs right now. Want the number for your situation? Tell me what's not scaling and how much time it needs. I'll give you a straight answer, including "you don't need me." ### Fractional CTO for Startups: /fractional-cto-for-startups/ An honest qualification guide to the fractional CTO for startups: who gets real value (funded Series A scale-ups, SMB operators), who doesn't (solo founders hoping to finish a prototype cheaply), what a startup actually gets, and when to wait. Senior technology leadership, part-time, for the startups it actually fits. Here's who gets real value from a fractional CTO, who doesn't, and how to tell which you are. Tags: Honest qualification, No pitch, 30+ companies since 2018 The right startups get enormous value. A fractional CTO for startups gives you an experienced technology executive part-time: someone who sets the architecture, hires the engineers, and makes the expensive calls at a stage where a full-time CTO is more commitment than the company needs. I've done the job for 30+ companies since 2018, at one point directing 12 engineering teams across 7 countries, and I explain the role in plain English in the fractional CTO guide. But it is not for every startup, and pretending otherwise wastes everyone's runway. The fit comes down to your stage, your funding, and what you're actually asking the role to do. Related searches: startup CTO, part-time CTO for startups, startup CTO as a service, outsourced CTO for startups. The sweet spot is specific. Two kinds of company get real leverage from this. - Funded Series A scale-ups: The sweet spot. You've raised, the product works, and now you need a real AI-native engineering and product organization: architecture, hiring, process, delivery. A fractional CTO builds it without the six-month executive search. - SMB operators: Established businesses that need senior technology judgment, a re-platform, an AI roadmap, a team that has stopped scaling, without putting a full-time executive on payroll. - Companies bridging to the permanent hire: You'll eventually need a full-time CTO. A fractional CTO builds the org properly first, then helps define and hire the person who inherits it. An honest no. If you're a non-technical solo founder with a no-code prototype and no funding, hoping a fractional CTO will finish the product cheaply: I'm not your answer, and neither is anyone honest. At that stage you need a technical co-founder or a small development shop, not a part-time executive. Hiring senior leadership before there's an organization to lead just burns runway. Not sure which side of the line you're on? I've written about the seven signs you actually need a fractional CTO and what kind of CTO each stage needs. If neither describes you, keep your money. What a startup actually gets. - Architecture that survives growth: A system designed for the next order of magnitude of users, data and team size, not just the demo that raised the round. - The team, hired and led: Org design, engineering and product hiring, and the culture and process that make the people you hire effective. - The expensive calls: Build-vs-buy, vendors, platforms, security. The decisions that cost the most to get wrong, made by someone who has made them before. - An AI-native operating model: The product built and the team run with AI as the default, from the SDLC to operations. The compounding edge for a young company. - Credibility for the raise: When you're raising or selling, someone senior who can stand behind the architecture in front of investors and diligence teams. - Published pricing: No discovery-call theater. My tiers are public, in USD, so you can self-qualify before we ever talk. When to wait, when to move. Wait: You're not there yet: No funding and no revenue to pay for senior time; A prototype that needs builders, not an executive; A couple of engineers and no scaling pain yet; No raise, compliance push or architecture deadline in sight Hire now: The signs are showing: Funded, and the team has stopped scaling; Technical debt is slowing every release; A raise, a sale or an enterprise deal needs technical credibility; You're building an AI-native org and nobody senior owns it Straight answers, up front. Q: Does my startup need a fractional CTO? A: If you're funded and building a real engineering and product organization, probably yes: you need architecture, hiring and delivery decisions made by someone senior, and you don't yet need that person full-time. If you're pre-funding with a prototype and no team, you need builders first. A fractional CTO is leverage on an organization, not a substitute for one. Q: What does a fractional CTO cost a startup? A: My pricing is published, in USD: advisory at $10,000 a month (about a day a week), the full fractional engagement at $18,000 a month (about two days), and embedded or interim leadership at $30,000 a month. A fixed-fee Diagnostic starts at $20,000. The full market context is in my fractional CTO cost guide. Q: Can a fractional CTO act as my technical co-founder? A: No. A co-founder trades equity for years of full-time commitment; I work for a fee under a standard consulting agreement, with no equity ask by default. What I can do is build the technology organization a co-founder would have built, and help you hire the permanent leader who owns it long-term. Q: Will you finish my no-code prototype cheaply? A: No, and I say so on this page on purpose. A prototype needs builders, not a part-time executive. Once the company is funded and building a real team, that's when a fractional CTO earns the fee. Q: Which stage is the sweet spot? A: Funded Series A scale-ups. You need to build a real, AI-native engineering and product organization, but rarely need a full-time CTO in the seat every day yet. A fractional CTO architects the org, hires it, and operates it, then scales the commitment as you grow. Q: When should a startup hire its first full-time CTO? A: Usually around 25 or more engineers, when technology is the whole company and the role needs someone in the seat every day. Before that, a full-time CTO is often more commitment than the stage warrants. Part of my job is defining and hiring that person when the time comes. Q: How is this different from a CTO-as-a-service firm? A: You work directly with me, the person doing the job, not with a firm that places someone from a bench and adds an account layer and a margin. Direct is cheaper, faster, and you know exactly whose judgment you're buying. Related reading & paths. - Fractional CTO services: The engagement itself: what I actually do, how I work, and the method behind it. - What a fractional CTO costs: The pricing guide: typical US rates, what moves the price, and my published numbers. - What is a fractional CTO?: The plain-English explainer: what the role is and when you actually need one. Are you the right fit? Tell me your stage, your funding, and what's breaking. I'll tell you straight whether a fractional CTO helps, or whether you should keep your money. ### Fractional vs Full-time vs Interim CTO: /fractional-cto-vs-full-time-vs-interim/ An honest breakdown of fractional, full-time, interim, and consultant technology leadership: what each is, how committed it is, and which one a company actually needs at its stage. Which kind of technology leadership do you actually need? An honest breakdown by commitment, accountability, and the stage you're at. Tags: Honest comparison, No pitch, By stage & commitment Commitment and stage, weighed honestly. Full-time CTO: The seat, every day: A full-time, long-term commitment in the seat; An executive search that can take six months; Worth it once you have roughly 25+ engineers; Heavy and slow to unwind before that Fractional / Interim: Senior judgment, now: Senior leadership in the seat in days, not quarters; Scales up or down with the actual need; Cross-industry pattern matching, embedded; Accountable for the outcome, not just advice One cell each, plainly. - 01 · Full-time CTO: Right when you have 25+ engineers and need the seat every day. (A permanent, full-time executive hire; A months-long executive search; Best when technology is the whole company and at scale) - 02 · Fractional CTO: Ongoing, part-time, embedded and accountable. My flagship. (Senior judgment without a full-time hire; Scales with you as the need grows; Best for Series A scale-ups building the org) - 03 · Interim CTO: Full-time but temporary, to cover a gap or carry a transition. (In the seat every day, for a defined period; Stabilizes, leads, and hands off cleanly; Best when the CTO just left or you're mid-transition) - 04 · Consultant: Advice from the sidelines, then it's back to you. (Delivers a recommendation, not the outcome; No accountability for whether it works; A fractional CTO makes the call and owns the result) Same jobs, different labels. The interim model travels under several names. A short-term CTO and a temporary CTO are the same job: full-time in the seat for a defined stretch, then a clean handoff. A turnaround CTO is the interim model with a harder mandate, stabilizing a troubled engineering org, often for a private-equity portfolio company. Price follows commitment. Fractional is part-time and starts lower; interim is near-full-time, so it costs more per month but runs for a shorter stretch. I publish my numbers in both guides: the fractional CTO cost guide ($10K–$18K a month on my ladder, with a fixed-fee Diagnostic from $20,000) and the interim CTO cost guide ($30,000 a month embedded). Which one is right? Q: When is a full-time CTO worth it? A: When technology is the core of the business and you have enough engineers, usually around 25 or more, that the role needs someone in the seat every day, owning the function full-time and for the long term. Below that scale, a full-time CTO is often more commitment than the stage warrants. Q: What's the difference between fractional and interim? A: A fractional CTO is ongoing but part-time, leading your technology alongside other commitments. An interim CTO is full-time but temporary, in the seat every day to cover a gap or carry a transition until the permanent hire lands. Both are hands-on and accountable; the difference is the depth and duration of the commitment. Q: Can a fractional engagement convert to interim or full-time? A: Often, yes. Many engagements start fractional and scale up to interim (full-time, temporary) when the need intensifies, or help define and hire the eventual full-time CTO. The point is to match the commitment to the problem, and change it as the problem changes. Q: Which is right for a Series A company? A: For most funded Series A scale-ups, a fractional CTO is the sweet spot: you need to build a real, AI-native engineering and product organization, but rarely need a full-time CTO in the seat every day yet. A fractional CTO architects the org, hires it, and operates it, then scales the commitment as you grow. Still not sure which you need? Tell me your stage and what's not working. I'll give you a straight answer, even if the answer is "not me." ### What is a Fractional CTO?: /fractional-cto/ A plain-English guide to the fractional CTO: what the role is, what it does, and how to tell whether you need a fractional, interim, or full-time CTO, written by someone who has done the job for 30+ companies since 2018. A plain answer, and how to tell whether you need one, an interim CTO, or a consultant. Written by someone who has done the job for 30+ companies since 2018. Tags: Plain English, No sales pitch, 30+ companies since 2018 Senior CTO judgment, part-time. A fractional CTO is an experienced Chief Technology Officer who leads your technology part-time instead of as a full-time hire. You get the same senior judgment on strategy, architecture, hiring, security and roadmap, without the six-month search for a full-time executive. The good ones don't advise from the sidelines; they embed, make the calls, and stay accountable for the outcome. A fractional CTO is part of your leadership, not a vendor sending a deck. Most of this work is remote-first. I serve US companies from Eastern Time, with same-day overlap across every US time zone — no executive search, no relocation, no payroll complexity. Also known as: part-time CTO, virtual CTO, outsourced CTO, on-demand CTO, CTO-as-a-Service (CTOaaS), fractional technology leadership. What a fractional CTO does. - Strategy & roadmap: Technology strategy tied to the business and the P&L, and a roadmap that picks the right battles instead of all of them. - Architecture & scale: An architecture that survives the next order of magnitude of users, data and team size, not just the current one. - Hiring & team: Defining the org, hiring the engineers and product people, and setting the culture and process that make them effective. - Security & compliance: SOC 2, HIPAA and GDPR programs that pass audits and unlock enterprise sales, built in rather than bolted on. - Build-vs-buy & vendors: The expensive calls: what to build, what to buy, which platforms and vendors to bet on, and when to walk away. - AI-native operating model: Rewiring how the product is built and how the team works so AI is the default. The modern fractional CTO's real edge. Fractional vs the full-time hire. Full-time CTO: Right at scale, heavy before it: The right call once technology is the whole company; A search that can take six months to land; A single perspective, learning your stack from scratch; Hard to unwind if the stage or need changes Fractional CTO: Senior judgment, now: Senior leadership in the seat in days, not quarters; Cross-industry pattern matching from many companies; Scales up or down with the actual need; Embedded and accountable for the outcome The fractional CTO, explained. Q: What is a fractional CTO? A: A fractional CTO is an experienced Chief Technology Officer who leads your technology part-time instead of as a full-time hire. You get senior judgment on strategy, architecture, hiring, security and roadmap without the months-long search for a full-time executive. The effective ones embed, make decisions, and stay accountable for the result rather than advising from the sidelines. Q: How is a fractional CTO different from an interim CTO? A: A fractional CTO is ongoing but part-time, leading your technology alongside other commitments. An interim CTO is full-time but temporary, in the seat every day to cover a gap or carry a transition until a permanent hire lands. Both are hands-on; the difference is how much time, and for how long. Q: How is a fractional CTO different from a consultant? A: A consultant advises from the outside and hands you a recommendation. A fractional CTO is part of your leadership: they make the call, own the roadmap, hire the team, and stay accountable for whether it actually works. Q: How is a fractional CTO priced? A: It varies widely with scope, seniority and time commitment. In the US market, my own pricing is published in USD and tiered by commitment: a fixed-fee Diagnostic from $20,000, then Advisory ($10K/mo, ~1 day a week), Fractional ($18K/mo, ~2 days), or Embedded/Interim ($30K/mo, 3+ days). Fixed-scope projects are available too. The full breakdown, including what the wider US market charges, is in my fractional CTO cost guide, or start with a direct conversation about what you actually need. Q: Do fractional CTOs work remotely with US companies? A: Yes — remote-first is the norm, and it's how I work. I serve US companies from Eastern Time, with same-day overlap across every US time zone, under a standard consulting agreement: no search firm, no relocation, no payroll complexity. Almost all of the 30+ companies I've served since 2018 are US-based. Q: When should a company hire a fractional CTO? A: When you need senior technology judgment but can't yet justify, or can't wait six months for, a full-time CTO. Common triggers: the team has stopped scaling, technical debt is slowing delivery, you're raising or selling, or you need to build an AI-native engineering and product org and don't have someone senior enough to architect it. Q: What is an AI-native fractional CTO? A: An AI-native fractional CTO doesn't bolt AI onto an old operating model. They rebuild how the product is built and how the organization runs so AI is the default, in the SDLC, in operations, and in the team they hire, and they implement it directly. In a landscape that moves this fast, that's the difference that compounds. Q: Can a fractional CTO build and hire the team, or just advise? A: The best ones do both. Oshri, for example, designs the solution and implements it, including doing the hiring and building the AI-native org around the people he brings in. A fractional CTO who only advises is really a consultant. Related reading & paths. - What a fractional CTO costs: The pricing guide: typical US hourly rates and retainers, what moves the price, and my published numbers. - Fractional CTO services: The engagement itself: what I actually do, how I work, and the method behind it. - Fractional vs full-time vs interim: Which kind of technology leadership your company actually needs right now. Think you might need a fractional CTO? Tell me what's not scaling. I'll tell you honestly whether a fractional CTO is the right answer, and whether I'm the right one. ### Industries: /industries/ AI-native fractional & interim CTO work across SaaS, EdTech, HealthTech, and E-commerce, backed by lived operating experience in each. The pattern matching that makes a fractional CTO valuable comes from having done the job. Here's where I've run product and engineering myself, and how AI changes each. Tags: SaaS, EdTech, HealthTech, E-commerce Four industries I know from the inside. - SaaS: Twenty-five years in SaaS and enterprise software: scaling engineering orgs without breaking them, rescuing troubled products, and taking delivery from low to elite DORA in 90 days. - EdTech: I've run product and engineering for an online learning platform. I know where EdTech breaks, and how AI changes the classroom, the content pipeline, and the cost structure. - HealthTech: Former CTO of a healthcare EMR platform. SOC 2 Type I & II and HIPAA-grade compliance delivered across HealthTech, without slowing the roadmap. - E-commerce: From platform scale to margin-finding AI agents across fulfillment, commerce technology that holds up on the busiest day of the year. "Industry experience isn't trivia. It's the difference between learning your problem on your dime and recognizing it on day one.": Oshri Cohen Working in one of these industries? If your problem looks like one I've solved before, I can come in fast and skip the ramp. ### How Much Does an Interim CTO Cost?: /interim-cto-cost/ A plain guide to interim and temporary CTO pricing: typical US day rates, what a full-time-but-temporary executive costs monthly, and my published numbers. What a full-time-but-temporary CTO actually costs in the US market, whether the search-firm route is worth its fee, and what I charge. The same numbers apply to a temporary CTO; it's the same job. Tags: Published USD pricing, Interim = temporary CTO, In the seat in days Figure $25K–$50K a month for a real one. In the US market, an experienced interim CTO typically runs $1,500–$2,500 a day, which lands at roughly $25,000–$50,000 a month near-full-time. Go through an interim-executive firm and you'll pay toward the top of that band, because a placement margin is stacked on the executive's rate. My published price is $30,000 a month for embedded, near-full-time leadership in the seat: turnarounds, transitions, or covering the CTO role until the permanent hire lands. Direct, no firm, no placement fee. A temporary CTO costs the same, because it's the same engagement under a different name. The distinction that actually changes the price is interim (full-time, temporary) versus fractional (part-time, ongoing), which starts much lower. Related searches: interim CTO rates, temporary CTO cost, short-term CTO cost, interim CTO day rate, CTO gap coverage. What I charge, exactly. The interim tier sits on the same published ladder as my fractional work. All prices USD. - The Diagnostic (From $20,000 · Fixed · 2–4 weeks · yours to keep): The Business-Down read on system, org, spend and AI leverage, plus a 90-day plan. If you're not sure you need someone in the seat, start here. - Embedded / Interim CTO ($30,000 / mo · 3+ days a week): Near-full-time leadership in the seat — turnarounds, transitions, or covering the role until the permanent hire lands. Clean handoff included. - Fractional CTO ($18,000 / mo · ~2 days a week): If the gap doesn't need full-time coverage, the fractional engagement delivers the leadership at part-time commitment. All prices USD. Full tier detail on the Fractional CTO page. Email me ↗ What people ask about price. Q: How much does an interim CTO cost per month? A: In the US market, roughly $25,000–$50,000 a month near-full-time, or $1,500–$2,500 a day. My published price is $30,000 a month for embedded, near-full-time leadership at three or more days a week, working directly with me rather than through a firm. Q: How much does a temporary or short-term CTO cost? A: The same as an interim CTO, because it's the same engagement: a CTO in the seat full-time for a defined stretch. My published price is $30,000 a month. The real pricing distinction is interim (full-time, temporary) versus fractional (part-time, ongoing), which starts at $10,000–$18,000 a month. Q: Why do interim-executive firms cost more? A: A firm places an executive from its bench and adds a placement margin, often 25–40% on top of the executive's rate, plus an account layer between you and the person doing the work. Hiring directly removes both. You talk to the person who does the job. Q: Is an interim CTO cheaper than rushing a permanent hire? A: Usually, yes. A rushed executive search still takes months, typically costs a recruiting fee of 25–35% of first-year compensation, and a mis-hire at the CTO level costs a year and often the roadmap. An interim CTO covers the seat immediately, stabilizes delivery, and buys you the time to run the permanent search properly. Q: How long do interim engagements run? A: Sometimes a few months, sometimes longer: long enough to stabilize the org and either hand off to the permanent hire or work myself out of the job. The engagement ends with a clean handoff, documented and deliberate, not a cliff. Related reading & paths. - Interim CTO services: The engagement itself: covering the gap, carrying the transition, handing off clean. - Fractional CTO cost: The part-time alternative, from $10K a month. Often the better fit when the gap isn't full-time. - Fractional vs full-time vs interim: Which kind of technology leadership your company actually needs right now. Need the seat covered now? Tell me about the gap. I'll tell you honestly whether it needs full-time coverage or two good days a week. ### What is an Interim CTO?: /interim-cto/ A plain-English guide to the interim CTO: a full-time but temporary technology executive who covers a gap or carries a transition, then hands off to the permanent hire. How it differs from a fractional CTO, what it costs, and how fast one can start. A plain answer: full-time technology leadership, temporarily. How the role works, what it costs, and how to tell it from a fractional CTO. Written by someone who has covered the seat. Tags: Plain English, No sales pitch, 30+ companies since 2018 A full-time CTO, for a while. An interim CTO is an experienced Chief Technology Officer who takes over your technology function full-time, but temporarily. The seat is empty, or about to be, and someone senior has to run engineering every day: keep delivery moving, steady the team, and make the calls that can't wait for a six-month executive search. The engagement has a built-in ending. An interim CTO stabilizes the function, leads it through the gap or the transition, then helps define and hire the permanent CTO and hands off cleanly. Done well, the permanent hire inherits a stronger org than the one the last CTO left behind. I've done this work across the 30+ companies I've served since 2018, and because interim work is usually urgent, I can typically start within one to two weeks. Most of it is remote-first: I serve US companies from Eastern Time, with same-day overlap across every US time zone. Also known as: temporary CTO, short-term CTO, transition CTO, acting CTO. Interim vs fractional. Fractional CTO: Part-time, ongoing: Leads your technology a day or two a week; An ongoing relationship that scales with the need; Best when you need senior judgment, not daily coverage; $10K–$18K a month on my published ladder Interim CTO: Full-time, temporary: In the seat every day, for a defined stretch; Covers a sudden gap or carries a transition; Ends with a clean handoff to the permanent hire; $30K a month embedded, on the same ladder The interim CTO, explained. Q: What is an interim CTO? A: An interim CTO is an experienced Chief Technology Officer who runs your technology function full-time for a temporary period, covering a sudden leadership gap or carrying the company through a transition. The role ends deliberately: the interim CTO stabilizes delivery, leads the team day to day, then hands off cleanly to the permanent hire. Q: How is an interim CTO different from a fractional CTO? A: An interim CTO is full-time but temporary, in the seat every day to cover a gap or carry a transition until the permanent hire lands. A fractional CTO is ongoing but part-time, leading your technology a day or two a week alongside other commitments. Both are hands-on; the difference is how much time, and for how long. Q: Is a temporary CTO the same as an interim CTO? A: Yes. Temporary CTO, short-term CTO, transition CTO and acting CTO are all the same job under different names: a full-time Chief Technology Officer engaged for a defined stretch rather than permanently. The distinction that actually matters is interim (full-time, temporary) versus fractional (part-time, ongoing). Q: How long does an interim CTO engagement last? A: Until the transition is complete and the function is stable, sometimes a few months, sometimes longer. It ends when the problem is solved and the permanent hire is in place, not on an arbitrary date, and it ends with a documented, deliberate handoff rather than a cliff. Q: How much does an interim CTO cost? A: In the US market, roughly $1,500–$2,500 a day, which lands at $25,000–$50,000 a month near-full-time. My published price is $30,000 a month for embedded, near-full-time leadership in the seat, working directly with me rather than through a firm. The full breakdown is in my interim CTO cost guide. Q: Will you help hire the permanent CTO? A: Yes. Leaving cleanly is part of the job: helping define the role, assessing candidates, and handing off the architecture, decisions and plan so the permanent hire succeeds. An interim CTO who makes himself permanent by stealth isn't doing interim work. Q: How fast can an interim CTO start? A: Fast, because the situation usually demands it. I can typically start within one to two weeks, against the three to six months a permanent executive search takes. That speed is most of the point: the team, the roadmap and the investors need a steady hand this week. Q: When should a company hire an interim CTO instead of rushing a permanent hire? A: When the seat is empty and waiting isn't an option. A rushed executive search still takes months and a mis-hire at the CTO level costs a year, often the roadmap with it. An interim CTO covers the seat immediately and buys you the time to run the permanent search properly. Related reading & paths. - Interim CTO services: The engagement itself: covering the gap, carrying the transition, handing off clean. - What an interim CTO costs: The pricing guide: US day rates, monthly numbers, and why firms charge more. - Fractional vs full-time vs interim: Which kind of technology leadership your company actually needs right now. Is the seat empty and the clock running? Tell me what happened and what has to keep shipping. I'll tell you honestly whether you need full-time interim coverage or two good days a week. ### The Business-Down Method: /method/ Oshri Cohen's signature AI-native CTO methodology: Read, Direct, Build, Operate & Optimize. Engineering decisions derived down from the P&L, for $10M–$50M companies. The discipline is fixed; the solution is always current. Most technology advice starts with the technology. Mine starts with the business: where you make money, where you lose it, what the board is actually asking. From there it derives the architecture, the team, and the AI. It's the same method every time. The answer never is. Tags: Read, Direct, Build, Operate & Optimize "The method is how I engage: disciplined, repeatable, accountable. What I build inside it is never boilerplate, because in the AI-native world the state of the art moves too fast for canned solutions. The spine stays; the solution is always current.": Oshri Cohen · The principle What “business-down” actually means. Technology-up is the default in our industry. You start from the stack (the framework someone likes, the cloud you already pay for, the model that's trending right now) and you reason upward toward the business, hoping the architecture you happened to choose lines up with where the company actually makes money. Sometimes it does. Often it doesn't, and you find out eighteen months and a few million dollars later. Business-down inverts that. I start at the P&L (revenue, margin, cost of delivery, the risk on the board's mind, the number the company is being valued on) and derive every technical decision down from there. The architecture, the build-vs-buy call, the org chart, the security posture, and where AI earns its place all fall out of one question asked relentlessly: what does this do for the business? It sounds obvious. It is almost never how technology gets decided. The reason is that business-down is harder: it requires someone who can read a cap table and a Kubernetes manifest in the same afternoon, hold both in their head, and translate each into the other. That seam, between the boardroom and the codebase, is where I've spent twenty-five years, and it's the only place this method works from. The pay-off is that the method is repeatable without being canned. The four movements (Read, Direct, Build, Operate & Optimize) run the same way every engagement, so the work is disciplined, legible, and accountable. But because each call is derived down from your numbers and built with today's state of the art, the output is bespoke every time. The discipline is fixed. The solution is always current. This is also why it's an AI-native method rather than an AI-flavoured one. AI isn't bolted onto the end as a feature; it's evaluated at every movement as a lever: sometimes the biggest one available, sometimes a distraction dressed up as strategy. Business-down is what tells the difference, because it measures AI the same way it measures everything else: against the P&L. Also known as: the business-first method · P&L-down engineering · top-down technology strategy · the Read-Direct-Build-Operate method. Technology-up vs. business-down. Technology-up · the default: Start at the stack, hope it reaches the business: Begins with frameworks, clouds and models: the tech that's interesting or already on the bill; AI adopted because it's the trend, measured by demos, not dollars; Roadmap is a feature backlog; the board can't see the money in it; Architecture optimised for engineering taste, not unit economics; Hiring fills seats by title; the org chart drifts from the strategy; Success is shipping; nobody can tie it back to margin or risk Business-down · the method: Start at the P&L, derive the technology down: Begins with revenue, margin, cost of delivery and the board's real question; AI evaluated at every movement as a lever, measured against the P&L; Roadmap reads as a business case; each item maps to money or risk; Architecture optimised for unit economics, scale and cost control; Hiring derived from the strategy; the org is built to run the model; Success is the outcome: margin moved, risk retired, delivery made predictable The reason it works: Conway's Law. In 1967 a programmer named Melvin Conway noticed something that has held up ever since: any organization that designs a system will produce a design whose structure mirrors the organization's own communication structure. Put plainly, your software architecture ends up looking like your org chart. Four teams that barely talk will build four services that barely talk. A monolith built by one room of people stays a monolith long after it should have been split. This isn't a metaphor; it's a force, and it acts whether or not anyone is paying attention. Most companies meet Conway's Law by accident. They pick an architecture for technical reasons, staff it with whatever org they already have, and then spend years puzzled about why the system keeps re-growing the shape of the teams instead of the shape they drew on the whiteboard. Technology-up engineering is especially exposed here. It decides the architecture first and treats the org chart as an HR problem to be solved later, so the two get set by different people, at different times, for different reasons. Conway's Law then quietly overrules the architecture diagram and hands you the system your communication structure was always going to produce. Business-down does the opposite, on purpose. In the Direct movement I draw the architecture and the org chart together, both derived down from the same P&L, so the team boundaries and the system boundaries are designed to match. Engineers call this the inverse Conway maneuver: shape the teams to get the architecture you actually want, instead of letting an accidental org produce an accidental system. Because business-down owns both decisions in one movement, it can use Conway's Law as a tool rather than be ambushed by it. This is the single biggest reason AI projects fail, and it's almost never the reason that gets named. The models work. The pilots demo well. Then implementation meets an organization that was never designed for AI-native operations, and Conway's Law does the rest: teams, hand-offs and decision rights built for a pre-AI company can only produce a pre-AI system, so the AI ends up bolted to the side of a structure that quietly rejects it. The post-mortem blames the model, the data or the vendor. The real cause was the org chart. The company asked an organization that wasn't built to run AI to suddenly run it, and it faltered exactly when it was time to implement. That's why an AI-native transformation has to redesign the organization, not just adopt the tools. The org now includes agents, pipelines and the people who supervise them, and the communication paths between them become part of the architecture. Derive that operating model down from the business and Conway's Law works for you: the system you ship is the system the business needed. Bolt AI onto a structure nobody designed, and the law hands you the mess your org chart was quietly describing all along. Business-down treats that redesign as the work, which is why it ships AI that survives contact with the company instead of stalling on it. Conway's Law (Melvin Conway, 1967): organizations build systems that mirror their own communication structure. It's the main reason AI initiatives stall at implementation, and the inverse Conway maneuver is how the method turns it into a design tool instead. Read → Direct → Build → Operate. - 01 · Read · Diagnose the system and the business: A clear-eyed read on product, team, architecture, security and spend, tied to your P&L. The real root cause, not the loudest symptom, plus where AI creates measurable leverage and where it's a distraction. You leave with a 90-day plan you can act on immediately, with or without me. (A system & architecture read; A team & delivery read; A spend & risk read; A 90-day plan you keep either way) - 02 · Direct · Set the AI-native direction: Technology strategy tied to revenue. An architecture that scales. An AI-native operating model and the org to run it. Build-vs-buy, security and compliance posture, and the hiring plan. Every call derived down from the business. (A roadmap tied to revenue; An architecture blueprint; An org & hiring plan; An AI governance & usage policy) - 03 · Build · Ship the systems, build the team: Hands-on. Production AI systems in the product and the pipeline: agentic and retrieval architectures with the right guardrails, eval, observability and cost control built in. The SDLC rebuilt AI-first, security in every pull request. And the AI-native team, hired and operating, delivering at elite DORA. This is the part most “AI advisors” don't do. I do. (Production AI in the product; An AI-first SDLC & pipeline; Security in every pull request; An AI-native team, hired & running) - 04 · Operate & Optimize · Run it, measure it, make it stick: AI moves too fast to rest on a launch, so the work compounds. I run the operating model and measure it: DORA for delivery, cost and quality for AI systems. I re-run evals as models drift, adopt what's genuinely better, and feed production signals back into the product. Then I plan the transition: to an internal hire, an advisory cadence, or a trained team. No permanent dependency. (DORA & AI cost/quality metrics; Eval re-runs as models drift; Production signals fed back in; A clean transition plan) A typical engagement, movement by movement. The shape is consistent; the depth flexes with the company. A representative path from first read to a team that no longer needs me. - Movement 01 · Read · Diagnose: Interview leadership, engineering and the operators who run the business; Trace revenue and cost of delivery to the systems behind them; Audit architecture, security, compliance and cloud spend; Map where AI is a real lever and where it's theatre; Deliver a written diagnostic and a 90-day plan you keep - Movement 02 · Direct · Decide: Set a roadmap tied to revenue and the board's question; Choose the architecture and the build-vs-buy calls; Design the AI-native operating model and governance; Draw the org chart and the hiring plan to run it; Align founders, board and team on one direction - Movement 03 · Build · Ship: Ship production AI into the product and the pipeline; Rebuild the SDLC AI-first with eval, observability and cost control; Put security and compliance into every pull request; Hire and onboard the AI-native team, hands-on; Drive delivery toward elite DORA performance - Movement 04 · Operate & Optimize · Compound & transition: Run the operating model and report on it monthly; Re-run evals as models drift; adopt what's genuinely better; Feed production signals back into product decisions; Coach the internal leader who will take the seat; Plan a clean exit: internal hire, advisory cadence, or trained team Discipline you can measure. - 90day: Actionable plan from the Read, yours to keep with or without me - Low → Elite: DORA trajectory the Build movement is engineered to reach - 1: Method, run the same way every engagement, disciplined and accountable - 0: Permanent dependency; every engagement is built to hand off What holds every time. - Derived down from the P&L: Every technical decision starts at the business (revenue, margin, risk) and is reasoned downward to the stack, never the other way around. - One method, many answers: The four movements are fixed and repeatable. What they produce is bespoke to your numbers and built with today's state of the art. The discipline is the constant. - AI-native by default: AI isn't a feature bolted on at the end. It's evaluated at every movement as a lever, and measured against the P&L like everything else. - Hands-on, not hand-wavy: I build, ship and hire inside the engagement, not a deck and a goodbye. The Build movement is the part most advisors skip. It's the point. - Measured continuously: DORA for delivery, cost and quality for AI systems. The method reports on itself, so progress is legible to the board, not a matter of faith. - Built to hand off: The method ends in a transition: an internal hire, an advisory cadence, or a trained team. Success is the company carrying it without me. Who it's for, and when. It's built for $10M–$50M software and software-enabled companies past product-market fit, with real revenue and real stakes, where a technology decision now moves the valuation, not just the backlog. Founders, CEOs and boards who need the engineering to serve the business, and private-equity investors who need a credible read before they sign and an AI-native value-creation plan after. It's just as much for traditional organizations that run on little or no software: companies where the work still lives in spreadsheets, email, phone calls and people's heads, and where AI is the first serious technology bet the business has ever made. Business-down is arguably a better fit there than anywhere, because there's no legacy stack to reason up from and no engineering culture defending it. You start where the value and the cost actually sit, in the operation itself, and derive the software, the automation and the AI-native operating model down from the business for the very first time, with the org designed for it from the start rather than retrofitted later. It fits best when the question is genuinely business-shaped: margins are being eaten by cost of delivery, the board is asking what AI means for the company, a troubled product needs to become predictable, or a transformation needs an owner who will actually build it. In those moments the bottleneck isn't more hands; it's a single person who can derive the right calls down from the P&L and then execute them. It's not the fit for everything, and I'll tell you so. A pre-revenue idea that needs a first prototype, a pure staff-augmentation ask, or a mandate that's already been decided and just needs typing: those don't need a method, they need a different kind of help. Business-down earns its keep where the technology and the business are tangled together and someone has to untangle them on purpose. The lowest-risk way to find out is the Read: a focused version of Movement 01 that ends in a 90-day plan you own whether or not we continue. It's the method's opening move, and it's designed so you can act on it immediately, with me or without me. The method, answered. Q: What is The Business-Down Method? A: It's my signature methodology for running an AI-native CTO engagement. Instead of starting from the technology and reasoning toward the business, it starts at the P&L (revenue, margin, cost of delivery, board-level risk) and derives the architecture, the team, the security posture and the AI down from there. It runs in four movements: Read, Direct, Build, and Operate & Optimize. Q: Why “business-down” instead of technology-up? A: Because technology-up decisions are bets that the stack you happened to pick will line up with where the company makes money, and you find out whether the bet paid off eighteen months later. Business-down starts from the money and the risk and reasons downward, so every technical call is tied to an outcome the board can see. It's harder to do, because it needs someone fluent in both the boardroom and the codebase, but it's the only version that survives contact with a P&L. Q: What are the four movements? A: Read: diagnose the system and the business, tied to your P&L, ending in a 90-day plan you keep. Direct: set the AI-native direction, covering strategy, architecture, operating model, org and hiring plan. Build: hands-on, ship production AI systems, rebuild the SDLC AI-first, and hire and run the team. Operate & Optimize: run it, measure it with DORA and AI cost/quality metrics, re-run evals as models drift, and plan a clean transition. Q: If the method is always the same, how is the solution not boilerplate? A: The method is the discipline: fixed, repeatable, accountable. The solution is derived down from your specific numbers and built with today's state of the art, which in the AI-native world moves every quarter. So the spine stays the same across engagements while the output is bespoke every time. The discipline is the constant; the solution is always current. Q: How is this different from a typical AI consultant? A: Most AI advice stops at strategy (a deck, a workshop, a list of use cases) and leaves the building to someone else. The Build movement is the centre of this method: I ship production AI into the product and pipeline, rebuild the SDLC, and hire and run the team, hands-on. And because every call is derived from the P&L, AI gets measured against the business, not against demos. Q: What does Conway's Law have to do with it, and why do AI projects fail? A: Conway's Law (Melvin Conway, 1967) says any organization that designs a system produces a design that mirrors the organization's own communication structure: your architecture ends up shaped like your org chart. It's the primary reason AI projects fail. The organization was never designed for AI-native operations, so when it's time to implement, the system mirrors the old org and the AI falters, no matter how good the model is. Technology-up engineering makes it worse by picking the architecture first and setting the org separately, so the law quietly overrules the diagram. Business-down draws the architecture and the org chart together in the Direct movement, both derived from the same P&L, so team boundaries and system boundaries match by design. That deliberate use of the law is the inverse Conway maneuver, and it's a core reason the method works where bolt-on AI stalls. Q: Who is it for? A: $10M–$50M software and software-enabled companies past product-market fit, where a technology decision now moves the valuation. Founders, CEOs and boards who need engineering in service of the business, and private-equity investors who need a clear technical read before a deal and an AI-native value-creation plan after. It also fits traditional organizations that run on little or no software and are making their first serious AI or digital bet: with no legacy stack to fight, business-down can derive the technology and the AI-native operating model down from the operation itself. It's not built for pre-revenue prototypes or pure staff augmentation. Q: How do we start, and what's the lowest-risk way in? A: Start with the Read movement: a focused diagnosis of system, org, spend and AI leverage, ending in a 90-day plan you can act on with or without me. It's the best first step: it tells us both whether a deeper engagement makes sense, and you keep the plan either way. Q: How does the engagement end? A: By design, in a transition, never a permanent dependency. The Operate & Optimize movement includes coaching the internal leader who will take the seat and planning a clean exit: an internal hire, a lighter advisory cadence, or a trained team that carries the operating model on its own. Success is the company running it without me. Want the read first? The fastest way to know if I can help is the Read: a focused version of Movement 01 that ends in a 90-day plan. Yours to keep whether or not we continue. ### Work: /projects/ A selection of digital products Oshri Cohen has delivered end-to-end, EMR and healthcare BI, AI agent systems, large-scale data platforms, IoT, and multi-brand commerce, built hands-on, across industries. Across 25 years I've delivered digital products end-to-end, architecture, delivery, and the team to run them. A selection is below, anonymized: the problem, what I built, and what it changed. Tags: End-to-end delivery, Hands-on, Across industries, Anonymized by design Built end-to-end. - EMR SaaS Platforms: Two electronic medical records SaaS builds, integrated care, interoperability, and HIPAA-grade compliance. - Healthcare BI Platform: A business-intelligence product turning clinical and operational data into decisions clinicians and operators trust. - AI Order-Fulfillment Agents: Autonomous fulfillment agents on AWS AgentCore that find margin in the cost of every order. - Hybrid AI + Human Workflow: A platform that mixes AI workloads with human-verified work across multiple industries, with the human accountable. - Event Purchasing Platform: A complex event purchasing and management platform built to hold up under spiky, high-concurrency demand. - Cloud-Native TMS + IoT: A transportation management system rebuilt cloud-native, with IoT hardware deployed on public buses. - Marketing Automation (Twilio): Automatic lead qualification using dynamic phone numbers through Twilio, with call tracking and attribution. - Pricing Intelligence at Scale: A massively scalable pricing and product-intelligence architecture processing over 250 million data pages a day. - CMS Data as a Database + MCP: The full CMS public-data universe auto-scraped into one relational database, with aggregate APIs and an MCP so AI can answer questions that used to be too expensive to ask. - Multi-Brand Commerce: A shared marketing and e-commerce architecture for a company operating a portfolio of brands. Delivered, at scale. - 10+: Digital products delivered end-to-end across industries - 250M/day: Data pages processed by the pricing-intelligence platform - SOC 2 I & II: Delivered under HIPAA across HealthTech and beyond - 0: Downtime on a year-long cloud-native re-platform "I don't hand over a deck and leave. I architect the thing, build it, and stay accountable for whether it actually works in production.": Oshri Cohen Have something that needs building? If your problem looks like one of these, I've shipped it before, and I can do it again, hands-on. ### Services: /services/ AI-native fractional & interim CTO services: fractional CTO, temporary CTO, short-term CTO, AI strategy, AI-native transformation, interim CTO, turnaround CTO, and technical due diligence for founders, boards and private equity. I work directly with you as an AI-native fractional and interim CTO for $10M–$50M US companies — no firm, no bench, no account manager. I don't just tell you what needs to be done: this is white-glove, bespoke work that evolves you, your team, and then the whole organization into the AI-native world. One method, The Business-Down Method™, run by the person you actually hire. Below are the shapes that work usually takes; pricing is published on the Fractional CTO page. Tags: Direct, no firm, White-glove & bespoke, A changemaker from the inside out A CTO in the seat, however you need one. - Fractional CTO: Senior, AI-native technology leadership without a full-time hire. I set direction, build the team, and own delivery. My flagship. - Temporary CTO: A CTO in the seat full-time for a defined stretch, to cover a gap or carry a transition, then hand off cleanly. Also called an interim CTO. - Short-Term CTO: Full CTO ownership for a stretch measured in months: cover the gap, hit the deadline, hand back clean. - Interim CTO: Full-time leadership in the seat, temporarily, to cover a gap or carry a transition until the permanent hire lands. - Turnaround CTO: Stabilize troubled products and teams, then chart a credible path back to fast, predictable delivery. - Technical Due Diligence: For PE and investors: a clear read on code, architecture, team and risk before you sign, and an AI-native value-creation plan after. - Technical Recruitment: A CTO who builds your engineering team, not a resume pipeline: org design, personal vetting, and hires built for the AI coding era. Becoming AI-native, for real. - AI Strategy: An AI roadmap tied to the P&L, with the architecture and economics behind it, written by the fractional AI CTO who stays to ship it. - AI-Native Transformation: I rebuild how the product is built and how the org runs so AI is the default, not an add-on, and I implement it, including the hiring. - AI Innovations by Industry: The AI applications and agent systems I've designed across industries, some shipped, some blueprints. "The method is the discipline. The solution is never boilerplate — in the AI-native world the ground moves too fast for canned answers. I solve the problem in front of you with whatever is genuinely best right now.": Oshri Cohen By industry. - EdTech: I've run product and engineering for an online learning platform, I know where EdTech breaks and how AI changes it. - HealthTech: Former CTO of a healthcare EMR platform. SOC 2 Type I & II and HIPAA-grade compliance without slowing the roadmap. - E-commerce: Platform scale on the busiest day of the year, and margin-finding AI agents across fulfillment. Before we talk. Q: What services does Oshri Cohen offer? A: Oshri works as an AI-native fractional and interim CTO. The core services are fractional CTO (his flagship), AI-native transformation, interim CTO, turnaround CTO, and technical due diligence for private equity. All are hands-on and senior. He doesn't just say what needs to be done — he works as a changemaker from the inside out, evolving the founder, the team, and then the organization into the AI-native world, designing the solution and implementing it, including the hiring. Q: How do I know which one I need? A: If you need ongoing senior judgment without a full-time hire, that's fractional. If you need someone full-time but temporary to cover a gap, that's interim. If delivery is broken, that's a turnaround. If you're evaluating an acquisition, that's technical due diligence. The fastest way to find out is a direct conversation. Q: How is pricing structured? A: It's published and tiered. A fixed-fee Diagnostic from $20,000 is the front door; ongoing leadership is Advisory ($10K/mo, ~1 day a week), Fractional ($18K/mo, ~2 days), or Embedded/Interim ($30K/mo, 3+ days). Fixed-scope projects — roadmap sprints, PE due diligence — are quoted to scope. Full detail is on the Fractional CTO page. Not sure which one you need? Tell me what's not scaling. I'll tell you honestly whether, and how, I can help. ### How Much Does Technical Due Diligence Cost?: /technical-due-diligence-cost/ A plain guide to technical due diligence pricing for PE and investors: typical market ranges, what drives the fee, and my published per-deal pricing. For PE firms and investors buying software companies: what a real technical read costs, what moves the fee, and my published per-deal pricing. Cheap DD is the most expensive kind. Tags: From $15K per deal, Written for partners, An operator's read From $15K a deal, scoped to the target. In the US market, technical due diligence on a software acquisition typically runs from the low tens of thousands for a focused read on a mid-market target to well into six figures when a big advisory firm staffs a team on a large platform. My published pricing starts at $15,000 per deal, scoped up with the size and complexity of the target. What you get is a decision-ready report written for partners, not just engineers: code quality and architecture, scalability, security and compliance exposure, the team and key-person risk, and the true cost of ownership, with a clear remediation roadmap you can price into the deal. The fee is small against what it protects. A missed compliance gap, a rewrite-in-disguise, or a key-person dependency discovered after close doesn't cost $15K; it costs a chunk of the thesis. And because I operate as a turnaround and fractional CTO after deals close, the report reads like a plan, not a caveat list. Related searches: tech DD cost, software due diligence pricing, IT due diligence fees, technical audit for acquisition. What I charge, exactly. Fixed per-deal pricing, scoped before we start. All prices USD. - Technical Due Diligence (From $15,000 · Per deal · fixed · scoped up with target size): Code, architecture, scalability, security, team and cost of ownership, delivered as a decision-ready report for partners with a remediation roadmap. - Post-Close Diagnostic (From $20,000 · Fixed · 2–4 weeks): After the deal: the deeper Business-Down read of the acquired company, turned into a 90-day value-creation plan. - Value-Creation Leadership ($18,000–$30,000 / mo · Fractional to embedded): If the plan needs an operator, I stay on as fractional or interim CTO to execute it, including AI-native modernization. All prices USD. Full tier detail on the Fractional CTO page. Email me ↗ What investors ask about price. Q: How much does technical due diligence cost? A: My published pricing starts at $15,000 per deal for a mid-market software target, scoped up with size and complexity. Market-wide, focused independent reads run in the low tens of thousands, while large advisory firms staffing teams on big platforms charge well into six figures. All prices USD. Q: What does the fee include? A: A decision-ready report covering code quality and architecture, scalability, security and compliance exposure, the engineering team and key-person risk, and true cost of ownership, plus a remediation roadmap priced for the deal model. Written for the partners making the decision. Q: How long does technical due diligence take? A: Most reads run around two to three weeks from data-room access to report, and I work to the deal timeline. Compressed reads are possible when a process is moving fast; the scope narrows to the questions that can kill or reprice the deal. Q: Why not have the big advisory firm do it? A: Sometimes you should, on very large platforms. But a firm sends a team and a template; I send the person who has run engineering organizations and fixed these companies after close. The read is operational, the report says what I'd actually do, and the fee doesn't carry a partner-leverage margin. Q: What happens after the deal closes? A: That's the point of the model. The DD report converts into a post-close Diagnostic and 90-day plan, and if the thesis needs an operator I stay on as fractional or interim CTO to execute it. The person who priced the risk is the person accountable for retiring it. Related reading & paths. - Technical Due Diligence: The service itself: what the read covers and how it converts into a value-creation plan. - Turnaround CTO cost: When the portfolio company needs more than a report: pricing for post-close stabilization. - Fractional CTO cost: The ongoing leadership tier: what part-time senior technology leadership costs. A deal on the table and a codebase you can't see into? Send me the timeline. I'll scope the read to the questions that can kill or reprice the deal. ### How Much Does a Turnaround CTO Cost?: /turnaround-cto-cost/ A plain guide to turnaround CTO pricing: what stabilizing a troubled engineering organization costs, why it's priced like interim leadership, and my published numbers. The honest frame: a turnaround is priced like interim leadership, and the number that matters more is what the broken state is already costing you every month. Here are both. Tags: Published USD pricing, Diagnostic first, For founders, boards & PE A diagnostic first, then $30K a month in the seat. A turnaround CTO is an interim executive with a harder job, so it's priced like interim leadership. In the US market that means roughly $25,000–$50,000 a month near-full-time. My published numbers: a fixed-fee Diagnostic from $20,000 to establish what's actually broken, then $30,000 a month embedded until delivery is stable and predictable again. The right first step is the Diagnostic. Nobody should pay for a rescue before an honest read of the damage, and you keep the findings and the 90-day plan whether or not I stay to run it. The comparison that matters isn't my fee against a cheaper consultant. It's my fee against the burn: a stalled engineering org at a $10M–$50M company costs hundreds of thousands a month in payroll that isn't shipping, missed roadmap, and customers quietly evaluating alternatives. Most turnarounds pay for themselves by showing up in DORA metrics within the first quarter. Related searches: turnaround CTO pricing, engineering turnaround cost, rescue CTO, PE portfolio company CTO. What I charge, exactly. Turnarounds run on the embedded tier of my published ladder. All prices USD. - The Diagnostic (From $20,000 · Fixed · 2–4 weeks · yours to keep): The honest read: system, org, spend, and what's actually broken versus what's just loud. Plus a 90-day stabilization plan. The best first step. - Embedded Turnaround CTO ($30,000 / mo · 3+ days a week): In the seat, running the stabilization: delivery, team, architecture and the hard calls. Until the org is fast and predictable again. - Fractional follow-through ($18,000 / mo · ~2 days a week): Once stable, most companies step down to fractional leadership to hold the gains and keep building. Same person, lower commitment. All prices USD. Full tier detail on the Fractional CTO page. Email me ↗ What founders & boards ask about price. Q: How much does a turnaround CTO cost? A: It's priced like interim leadership: roughly $25,000–$50,000 a month near-full-time in the US market. My published pricing is a fixed-fee Diagnostic from $20,000, then $30,000 a month embedded at three or more days a week until delivery is stable. All prices USD. Q: How long does an engineering turnaround take? A: The Diagnostic takes two to four weeks. Visible stabilization, delivery cadence, fewer incidents, a believable roadmap, typically shows inside the first 90 days; I've taken organizations from low to elite DORA performance in three months. Full turnarounds usually run three to six months, then step down to fractional leadership to hold the gains. Q: Why not just hire a cheaper consultant to fix it? A: Because a turnaround isn't advice, it's authority. Someone has to make the unpopular calls: stop the death-march project, restructure the team, freeze the rewrite. A consultant recommends; a turnaround CTO decides and owns the consequences. That's what the price buys. Q: What does a failed turnaround cost by comparison? A: The downside case is the whole company. A stalled org burns its payroll without shipping, the roadmap slips a year, and in PE situations the hold period runs out with the value-creation plan unexecuted. Against that, $30,000 a month with measurable 90-day checkpoints is the cheap option. Q: Do you work with private-equity portfolio companies? A: Yes, that's a core use case: post-acquisition stabilization, carve-outs, and portfolio companies where delivery has stalled. PE firms also engage me for technical due diligence before the deal, from $15K per deal, so problems are priced in rather than discovered. Related reading & paths. - Turnaround CTO services: The engagement itself: stabilize delivery, fix the org, chart the path back to fast. - Low to elite DORA in 90 days: A field report from a real turnaround: what changed, in what order, and what it measured. - Interim CTO cost: The pricing frame turnarounds inherit: full-time, temporary leadership in the seat. Delivery stalled and the board is asking? Start with the Diagnostic. In two to four weeks you'll know what's broken and what it takes to fix it, with or without me. ### Web Properties: /web-properties/ Products Oshri Cohen owns and operates himself, built end-to-end with the same AI-native methods he brings to clients. Each property is a live business: the product, why it exists, and what running it proves. Client work lives under Work. This is the other shelf: products I own outright and run in production, from the first line of code to the support inbox. Each one exists because a real problem annoyed me enough, and each one is living proof of the AI-native way of working I bring to clients. Tags: Built end-to-end, Run in production, AI-native by design, My own P&L Live and in production. - Watchly · watchlyplayer.tv: YouTube parental controls done right: parents approve every video, kids watch ad-free with no algorithm, and time limits enforce themselves. Born at home, run daily. The consultant who ships his own. Most advisors have opinions about how software should be built. Fewer are willing to bet their own money and evenings on those opinions. My web properties are that bet. I do the product thinking, write the code, run the servers, watch the analytics and answer the support email myself, using AI as leverage the whole way through. They serve two purposes. They solve problems I actually have, which keeps the products honest. And they keep my consulting honest too: when I tell a client that a small AI-native team can ship what used to take a department, I'm not quoting a study. I'm describing my own operating reality, and I can show the receipts. What people ask about this. Q: What are Oshri Cohen's web properties? A: Web properties are software products Oshri Cohen owns and operates himself, separate from his client work. Each one is a live production business he built end-to-end and runs day to day. The first is Watchly (watchlyplayer.tv), a YouTube parental-controls product where parents approve every video their kids can watch. Q: How is this different from the client work under Work? A: The projects listed under Work were delivered for clients, and the clients own them. Web properties are Oshri's own: he owns the code, the infrastructure, the revenue and the roadmap, and he operates them in production himself. Q: Why does a fractional CTO run his own products? A: Running his own products keeps his advice grounded in first-hand, current practice. The AI-native methods Oshri Cohen installs in client organizations are the same ones his web properties are built and operated with, so clients get guidance that has been tested on a real production business, at his own risk. Want your team to ship like this? The methods behind these products are the same ones I install in client organizations. If you want to see what one AI-native operator can carry, let's talk. ## Projects (case studies) ### Cloud-Native TMS with IoT on Public Buses: /projects/cloud-native-tms-iot/ I re-architected a legacy on-prem Transportation Management System into a cloud-native platform and put IoT devices on a live public-bus fleet, with zero downtime over a year-long migration. A public-transit operator was running its Transportation Management System on aging on-prem hardware, with no live view of where its buses actually were. I rebuilt the platform cloud-native, designed and deployed IoT devices across the fleet, and stood up a real-time telemetry pipeline, all without taking daily service offline. Tags: Cloud-native rebuild, Kubernetes, Fleet IoT & firmware, Real-time telemetry An on-prem TMS, and buses you couldn't see. The system ran the operation, but it ran it blind. Hardware was aging out, scaling meant buying servers, and nobody could say where a given bus was right now without calling someone. - Legacy on-prem core: The TMS sat on physical servers in one location, brittle to scale, expensive to maintain, and a single point of failure for a service the public depends on. - No live fleet visibility: Vehicle position was estimated from schedules and radio, not measured. Dispatch and planning ran on stale, partial data. - Hardware on every bus: Putting telemetry on the road meant real devices, on real buses, surviving heat, vibration and patchy cellular, not a slide deck. - Service can't stop: This is public transit. There was no maintenance window where the buses simply stop running while we cut over. From the device on the bus to the cloud that runs it. - Cloud-native re-platform: I re-architected the legacy on-prem TMS into containerized services on Kubernetes, designed to run across AWS, Azure or GCP rather than a single rack. Scaling became a config change, not a purchase order. - IoT devices on the fleet: I designed and deployed the on-bus hardware and its firmware, ruggedized for vibration, temperature and intermittent connectivity, then rolled it out across a live public-bus fleet. - Real-time telemetry pipeline: GPS and vehicle telemetry stream off every bus into an ingestion pipeline built to absorb bursts, tolerate dropouts, and deliver position data in real time instead of after the fact. - Fleet visibility, finally: Dispatch and planning got a live picture of the whole fleet, where every bus is, right now, turning guesswork into operational data the operator could actually run on. - Zero-downtime migration: I moved the operator off on-prem and onto the cloud-native platform over roughly a year, in stages, without disrupting daily transit service. The public never saw the cutover. - Built to be owned: Infrastructure as code, observability and runbooks so the operator's own team could run and extend the platform, the same discipline I bring as a fractional CTO. Modernized under load, without a service gap. - 0downtime: Daily transit service stayed live through the entire migration - ~1yr: End-to-end re-platform from on-prem legacy to cloud-native - Real-time: Live GPS and telemetry across the public-bus fleet - Device→cloud: One system spanning on-bus firmware up to the cloud platform "A transit operator can't take a maintenance window, the buses run whether your migration is ready or not. So you modernize underneath a moving service, not on top of a paused one." What people ask about this build. Q: How do you migrate a TMS without stopping the buses? A: In stages, with the old and new systems running side by side. I moved workloads onto the cloud-native platform piece by piece, validated each one against the live system, and only cut traffic over once it was proven, so daily transit service never went dark during the year-long migration. Q: Why did this involve hardware and not just software? A: Real-time fleet visibility starts on the bus. There's no telemetry to ingest until a device is physically capturing GPS and vehicle data and surviving the road. I designed and deployed the IoT devices and firmware as well as the cloud platform, so the work spanned from the metal on the bus up to Kubernetes. Q: Why cloud-native instead of replacing the on-prem servers? A: Buying more servers solves today and recreates the same problem next year. Re-architecting onto containers and Kubernetes meant the operator could scale on demand, run across AWS, Azure or GCP, and stop carrying a single physical point of failure for a public service. Got a legacy platform that needs to go cloud-native? Whether it's a transit TMS, an IoT fleet, or an on-prem system you can't afford to take offline, let's talk through what a clean migration looks like. ### CMS Healthcare Data, Turned Into Answers: /projects/cms-data-mcp/ The U.S. government publishes a staggering amount of healthcare data, and almost nobody can get a straight answer out of it. I built the thing that does: ask a question in plain English, get a clear answer in seconds, so clients build dashboards and make decisions that used to be impossible. The government publishes an enormous amount of public healthcare data, and almost none of it is usable. Getting one real answer meant hiring analysts and waiting weeks, so most questions never got asked. I fixed that. Now you ask a question the way you'd say it out loud, and you get a clear answer back in seconds. Tags: One place for all of it, Ask in plain English, Answers in seconds, No data team needed The data is public. The answers weren't. All the information is right there, free, for anyone. Actually getting an answer out of it is where everyone gets stuck. - Scattered everywhere: The piece you need is split across hundreds of separate places that were never meant to work together. Answering one question first means hunting down and lining up half a dozen of them. - Nothing lines up: The same hospital, doctor or drug is labeled three different ways in three different places. Before you can compare anything, someone has to make it all agree, and that someone is usually expensive. - Every question started from zero: Want to compare quality against cost across a region? That was a project: scope it, staff it, wait weeks, pay for it, for a single answer. So most questions simply never got asked. Gather it once, then just ask. - All of it in one place: I brought the whole sprawling collection together into one place, with the same hospital, doctor and drug finally lined up so they can be compared instead of just collected. - Always current: It keeps itself up to date on its own. When the government publishes something new, it's there, so you're always working from today's picture, not a snapshot someone pulled a year ago. - Ask in plain English: No special skills, no query language, no ticket to a data team. You ask the question the way you'd say it in a meeting, and the answer comes back. - Answers, not raw data: You get the finished answer, the ranking, the comparison, the number you were after, rather than a giant file you'd still have to make sense of yourself. - Built for the questions people actually ask: It's tuned around the real questions clients bring, quality versus cost, who's growing, where the outliers are, so the answers land on the decision, not a chart nobody reads. - Powers your own dashboards: Teams plug it into their own dashboards and tools, turning a pile of public data they could never use into a living source of answers they own. What changed about getting an answer. Before: A project per question: Track down the right sources by hand; Pay specialists to make them line up; Build something custom for this one question; Wait weeks and pay for the analyst time; Most questions never get asked at all After: Ask, and get an answer: Everything is already gathered and lined up; It stays current on its own; Ask in plain English, the way you'd say it; Get a clear answer in seconds, not weeks; Cheap enough to ask the questions you used to skip Questions that were once too expensive to ask. - Allof it: The whole public healthcare data collection, gathered and lined up in one place - 50Qs: Real questions it answers out of the box, asked in plain English - Weeks→sec: Answers that took weeks of analyst time now come back on demand "The data was always public. What changed is the cost of asking a question of it. When the answer drops from a three-week project to a sentence, people finally ask the questions that move the business.": Oshri Cohen · On making data useful What people ask about this. Q: What can clients actually do with it? A: They ask questions and get answers, without a data team in the loop. Instead of commissioning an analysis and waiting, someone types the question the way they'd say it in a meeting and gets a clear answer back, a ranking, a comparison, the number they needed. That's what makes the dashboards and reports possible: the hard part is handled once, up front, so every question after that is quick and cheap. Q: How do you keep it current? A: It updates itself. When the government publishes new information, it gets pulled in automatically, so you're always working from the latest picture rather than a stale extract someone downloaded once and forgot. Nobody has to babysit a download folder. Q: Why is this so much cheaper than before? A: Before, every question was its own project: find the sources, make them line up, build something custom, wait weeks, pay for the time. Most questions never got asked because it wasn't worth it. Here all of that hard work is done once, up front. After that, asking a brand-new question costs almost nothing, so the answers people used to skip are suddenly within reach. Sitting on public data you can't actually use? Whether it's this healthcare data, another government data dump, or your own scattered systems, I build the thing that turns it into answers anyone can ask for. ### AI Order-Fulfillment Agents on AWS AgentCore: /projects/ecommerce-fulfillment-agents/ I built an order-fulfillment system for an e-commerce operator as a graph of autonomous agents on AWS AgentCore that finds margin in the cost of each order while holding a hard floor. For an e-commerce operator in a commodity market, the price was fixed. The only place left to win was the cost of each order. I built the fulfillment system as a graph of autonomous agents on AWS AgentCore, each one narrow, each one accountable, all of them holding a hard margin floor. Tags: Agent-graph design, AWS AgentCore, Guardrails & margin floor, Eval & observability When you can't move the price, you move the cost. In a commodity market, every competitor charges roughly the same. The order isn't where you make money. The cost of fulfilling it is. - Price is fixed: The market sets it. Discounting is a race to the bottom, and there's no premium tier to climb into. The top line was effectively capped. - Margin hides in the order: Sourcing, shipping and returns all have choices behind them, and each choice moves the per-order cost by a few points. At volume, a few points is the business. - Humans can't price every order: Re-evaluating sourcing, carrier and routing on every single order, in real time, is not a job a human can do at scale. It is exactly a job for a system. A graph of narrow agents, not one big swarm. - Specialist agents: Each agent owns one decision, sourcing, carrier selection, returns handling, and does it well. Narrow scope means I can reason about, test and trust each one, instead of debugging an opaque mega-agent. - A directed graph, not a free-for-all: Agents are wired into an explicit graph with defined hand-offs, so work flows along paths I designed. No agent improvises outside its lane; the orchestration is the product, not an afterthought. - A hard margin floor: The system optimizes cost per order aggressively, but a guardrail enforces a margin floor no decision can breach. It finds margin where it exists and refuses orders that would lose money. - Hosted on AWS AgentCore: The agent graph runs on AWS AgentCore, so hosting, scaling and the runtime are managed infrastructure. I spent the effort on the decision logic and guardrails, not on babysitting servers. - Traceable & accountable: Every order carries a trace: which agent decided what, on what input, and why. When a margin looks wrong, I can follow it back to the exact step instead of guessing. - Built for e-commerce reality: Sourcing, shipping and returns are modeled as the real, messy levers they are. It's the kind of operational AI I describe across e-commerce work, see the broader picture on the e-commerce page. Optimization with brakes. - 01 · Guardrails · The floor that doesn't move: Cost optimization without a floor will happily chase volume into a loss. The margin floor is enforced as a hard constraint, not a suggestion an agent can talk itself out of. (A hard margin floor enforced on every order; Agents constrained to their lane in the graph; Refuse-the-order paths when no choice clears the floor; Decisions bounded by explicit rules, not vibes) - 02 · Evaluation & observability · Watching what the agents actually do: Autonomous agents drift. The only way to trust them in production is to measure them continuously and trace every decision back to its inputs. (Per-order traces across the whole agent graph; Evals that catch quality and margin regressions; Observability into where cost is won and lost; Signals fed back into the agents that make the calls) - 03 · Cost control · The agents pay for themselves: An agent system that costs more to run than it saves is a science project. I kept the model and orchestration spend in proportion to the per-order margin it captures. (Right-sized models per agent, not one expensive model everywhere; Orchestration cost tracked against margin captured; AgentCore hosting tuned for scale, not vanity; Cheaper paths adopted as they become available) Margin where there wasn't any. - Marginfloor: A hard floor no order can breach - Perorder: Margin found in sourcing, shipping & returns - Agentgraph: Narrow specialists, traceable end to end "When the price is fixed, the only honest place to compete is the cost of the order. The agents don't chase volume, they hold the floor and find the margin.": Oshri Cohen · On the fulfillment system What operators ask about this. Q: Why a graph of narrow agents instead of one capable agent? A: Because narrow agents are accountable. Each one owns a single decision, sourcing, carrier, returns, so I can test it, reason about it and trace it. One big agent doing everything is an opaque box: when a margin comes out wrong, you can't tell which part of its reasoning failed. A directed graph of specialists keeps the system legible at scale. Q: How do you stop cost optimization from losing money? A: A hard margin floor, enforced as a constraint the agents can't override. The system optimizes the cost of each order aggressively across sourcing, shipping and returns, but no decision is allowed to breach the floor. When no available choice clears it, the system refuses the order rather than fulfilling it at a loss. Q: Why AWS AgentCore? A: AgentCore handles hosting, scaling and the agent runtime as managed infrastructure, so I could spend the engineering effort on the decision logic, guardrails and observability, the parts that actually capture margin, instead of operating servers. It's the hosting platform for the agent graph, not a constraint on the design. Need agents that hold the line? If your margin is hiding in operations and you want autonomous systems you can actually trust, let's talk about what to build. ### EMR SaaS Platforms: /projects/emr-saas-platforms/ Two electronic medical records SaaS platforms delivered as CTO and architect, integrated care, HL7/FHIR interoperability, PHI security and audit, SOC 2 Type I & II, and HIPAA-grade and regional health-data compliance, without breaking live clinical workflows. As CTO and architect for a regional healthcare provider network, I designed and shipped two electronic medical records SaaS platforms, integrated care across providers, clinical data flowing cleanly between systems, and the security and audit posture that lets a healthcare buyer say yes. All of it delivered while live patient care kept running. Tags: EMR / EHR, HL7 & FHIR, PHI security, SOC 2 Type I & II, HIPAA-grade In healthcare, you can't stop the line. Clinicians are seeing patients while you rebuild the system under them. - Data trapped in silos: Each provider held its own records in its own format. Care suffered because nobody saw the whole patient. The platforms had to make data move. - PHI is not normal data: Protected health information carries legal weight. Access control, encryption and a defensible audit trail aren't features, they're the price of being allowed to operate. - A legacy platform mid-flight: One system was already in clinical use. Modernizing it meant changing the engine while the plane was carrying passengers, no downtime patients or clinicians would feel. - Two compliance regimes at once: Provincial health-data rules and HIPAA-grade controls applied together. The architecture had to satisfy both without bolting on compliance after the fact. Two platforms, one operating standard. - 01 · Interoperability · Clinical data that moves: An integration layer that let records flow between providers and external systems instead of dying in silos, so a clinician sees the whole patient. (HL7/FHIR-style integration between disparate provider systems; A normalized clinical data model under the messaging; Integrated care across providers, not isolated record-keeping; Built for HealthTech's interoperability and compliance reality) - 02 · Security & audit · PHI handled the right way: Security and access control designed into the platform from the data model up, with an audit trail that survives scrutiny. (Role- and context-based access control over PHI; Encryption in transit and at rest; Immutable audit trails on every record touch; Least-privilege boundaries between services and roles) - 03 · Modernization · Legacy platform, rebuilt live: Modernized an existing integrated-care platform in place, re-architecting it without an outage clinicians or patients would notice. (Incremental migration off legacy components; No interruption to live clinical workflows; A maintainable architecture the team could keep extending; Capacity to grow with the provider network) - 04 · Compliance · Built to pass: Compliance treated as an architectural property, not paperwork, which is how the platforms cleared formal audit. (SOC 2 Type I & II delivered; HIPAA-grade compliance posture; Provincial health-data compliance; Controls evidenced, not asserted) What the platforms cleared. - 2platforms: EMR SaaS platforms delivered as CTO and architect - SOC 2Type I & II: Delivered across the platforms - HIPAAgrade: Compliance posture for PHI - Provincial: Health-data compliance met "In healthcare, the hard part isn't the feature list. It's earning the right to hold the data, and rebuilding the system without ever stopping care.": Oshri Cohen What buyers ask about this work. Q: What was your role on these platforms? A: CTO and architect for a regional healthcare provider network. I owned the technical direction and the architecture for two EMR SaaS platforms, interoperability, PHI security and access control, audit, modernization of a legacy integrated-care platform, and the compliance posture behind them. Q: How did you modernize a legacy platform without downtime? A: Incrementally. I migrated off legacy components piece by piece behind stable interfaces, so the platform kept serving live clinical workflows throughout. Clinicians and patients never experienced an outage from the rebuild. Q: What compliance did the platforms achieve? A: SOC 2 Type I & II were delivered, with a HIPAA-grade compliance posture for protected health information and regional health-data compliance. Compliance was designed into the architecture, access control, encryption and immutable audit trails, rather than added afterward, which is why the controls held up under audit. Building or fixing a clinical platform? Bring me in to architect the data, the security and the compliance, without stopping care. ### Event Purchasing & Management Platform: /projects/event-purchasing-platform/ I designed and built a high-concurrency event purchasing and management platform: a checkout that holds together during on-sale surges, inventory that never double-sells, and an organizer side for scheduling and capacity. I built a purchasing and management platform for an events company. The demand is spiky: an on-sale opens and tens of thousands of people hit the same finite inventory in the same minute. I designed it so the busiest minute is the spec, not the surprise, no double-sells, no collapse at peak. Tags: High concurrency, Inventory under contention, Queueing & back-pressure, Payments Most of the day it's quiet. Then an on-sale opens. Event commerce isn't a steady stream of orders. Traffic sits flat, then spikes by orders of magnitude the second tickets drop. Everything that's fine at average load breaks at peak, and peak is exactly when correctness matters most. - Everyone wants the same seats: Thousands of buyers contend for the same finite inventory in the same instant. Naive reads and writes either oversell or grind the database to a halt under lock contention. - Double-sells are unforgivable: Selling the same seat twice isn't a glitch, it's a refund, an angry customer, and a reputation hit. Allocation has to be correct under contention, not just usually correct. - The surge is the spec: If the architecture only works at average load, it doesn't work. The on-sale minute, with its thundering herd, is the real workload the system exists to survive. - Payments can't be the bottleneck: Checkout holds inventory while money moves through a third party with its own latency and failures. Hold too long and you starve other buyers; release too soon and you oversell. Correctness under load, then everything around it. - Queue & back-pressure: A virtual waiting room absorbs the on-sale spike. Buyers are admitted at a rate the core can actually serve, so load is shaped instead of dropped, and the system degrades gracefully instead of falling over. - Inventory under contention: Allocation uses short, well-scoped holds and atomic claims so a seat can be reserved by exactly one buyer at a time. No oversell, no double-sell, even when thousands hit the same inventory at once. - High-throughput checkout: The checkout path is kept narrow and fast: hold, pay, confirm. Slow and failure-prone work is pushed off the critical path so the hot loop stays fast when it's under the most pressure. - Payments without oversell: Holds are time-boxed against payment latency and expire cleanly. A failed or abandoned payment returns inventory to the pool automatically, so money and seats never drift out of sync. - The management side: Organizers create and schedule events, define capacity and allocations, set what goes on sale and when, and watch it sell in real time. The buying and managing sides share one source of truth. - Built like an owned asset: Tested, observable, and documented so it can be operated and grown, not just shipped once. This is the same standard I bring to my fractional CTO work: build for the team that inherits it. The decisions that actually matter at peak. - 01 · Shape the load before it hits the core: You can't make a surge disappear, but you can decide how it arrives. Queueing and admission control turn an uncontrolled spike into a steady, serveable rate. (Virtual waiting room at the edge; Admission paced to real capacity; Fair ordering, no silent drops; Predictable behavior at the worst moment) - 02 · Make allocation atomic: Correctness lives in how a seat moves from available to held to sold. I kept that transition atomic and short so contention resolves cleanly instead of corrupting state. (Single-owner holds; Time-boxed reservations; Automatic release on expiry or failure; One source of truth for inventory) - 03 · Keep the hot path thin: Everything non-essential moves off the checkout critical path, so the part that runs tens of thousands of times in a minute does the least possible work. (Async work for anything slow; Idempotent operations on retries; Observability on the surge path; Graceful degradation over failure) "In event commerce the architecture isn't judged on the average day. It's judged on the one minute everyone shows up at once, and it has to be correct, not just alive." What people ask about this build. Q: How do you stop the same seat from being sold twice? A: Allocation is atomic. Moving a seat from available to held is a single-owner operation, so exactly one buyer can claim a given seat at a time. Holds are time-boxed and release automatically if payment fails or is abandoned, so inventory and payments never drift apart, even under heavy contention. Q: What happens when an on-sale opens and traffic spikes? A: A virtual waiting room absorbs the spike and admits buyers at a rate the core can actually serve. Instead of letting an uncontrolled surge overwhelm the database, the system shapes the load and degrades gracefully, so the busy minute behaves predictably rather than collapsing. Q: Does the platform also handle the organizer side? A: Yes. Beyond purchasing, organizers create and schedule events, define capacity and allocations, control what goes on sale and when, and watch sales in real time. The buying and managing sides share one source of truth so capacity and inventory stay consistent. Got a system that has to survive its busiest minute? High-concurrency commerce, inventory under contention, checkout that can't go down at peak, that's the kind of problem I like. Tell me what you're building. ### Healthcare Business Intelligence Platform: /projects/healthcare-bi-platform/ I built a healthcare business intelligence platform, data pipelines, a warehouse, governed metrics and dashboards, that clinicians and operators actually trust, with HIPAA-grade handling of sensitive data. A healthcare provider organization had clinical and operational data scattered across systems, with no trustworthy way to see what was happening. I built the pipelines, the warehouse, the governed metrics and the dashboards, and made the numbers something clinicians and operators rely on. Tags: Data pipelines, Warehouse, Governed metrics, HIPAA-grade handling Lots of data. No visibility. Reports existed. Nobody believed them, so nobody acted on them. - Data was scattered: Clinical systems, scheduling, billing and operations each held a piece of the truth, and none of them agreed. Pulling a single number meant a week of manual reconciliation. - Numbers couldn't be trusted: Two reports on the same question gave two answers. When leaders can't trust the dashboard, they fall back on gut, and the data investment is wasted. - Sensitive data, real exposure: This is patient-adjacent data. Moving it into a warehouse without rigorous handling and access control isn't a reporting problem, it's a compliance and trust problem. - Charts, not decisions: Even where dashboards existed, they showed activity, not answers. Nobody could tell you what to do differently on Monday morning. The platform, end to end. - Data pipelines: Reliable ingestion from the clinical, scheduling, billing and operational systems, with validation built in so bad data is caught at the door, not discovered in a board deck. - A warehouse with one truth: A central warehouse modeled so every team draws from the same definitions. One number, one source, the end of duelling spreadsheets. - Governed metrics: A defined, versioned metrics layer. 'Utilization' or 'no-show rate' means exactly one thing across every dashboard, agreed with the people who use it. - Dashboards for two audiences: Clinical views and operational views built for how each group actually works, not a generic BI template bolted onto a database. - Access control on sensitive data: HIPAA-grade handling throughout: least-privilege access, scoped views, and a clear line on who can see what, so the platform is safe to put in front of real users. - Reporting that drives decisions: I shipped reporting tied to the operating questions leaders actually ask, so the output is a decision, not just another chart. More on the domain on my HealthTech work. Trusted, safe, and used. - HIPAA-grade: Data handling and access control on sensitive, patient-adjacent data - Onesource: A single governed set of metrics every team draws from - Trusted: Reporting clinicians and operators rely on to make decisions "A dashboard nobody trusts is worse than no dashboard, it just adds an argument. The hard part isn't the chart; it's earning the trust that lets people act on the number.": Oshri Cohen What people ask about this. Q: How did you handle sensitive patient-adjacent data? A: HIPAA-grade handling was a design constraint from day one, not a bolt-on. That meant rigorous controls on how data moved into the warehouse, least-privilege access, scoped views by role, and a clear, defensible answer to who can see what. You can't ask clinicians and operators to trust a platform that isn't safe. Q: Why did clinicians and operators actually trust the numbers? A: Because the metrics were governed and agreed with the people who use them. Every key metric had one definition, one source, and a clear lineage back to the system it came from. When two reports always agree, and you can explain where a number comes from, trust follows, and that's what turns reporting into decisions instead of arguments. Q: What made this more than a dashboarding project? A: Dashboards were the visible part, but the value was underneath: dependable pipelines, a warehouse with a single source of truth, a governed metrics layer, and access control on sensitive data. I built it so the reporting answered the operating questions leaders ask, so the output is a decision, not just a chart. Sitting on data you can't actually use? I build data platforms and reporting that people trust enough to act on, in healthcare and beyond. ### Hybrid AI + Human-Verified Workflow Platform: /projects/hybrid-ai-human-workflow/ A workflow platform that routes work between AI and people, gating low-confidence output through human review so teams get AI throughput without giving up accountability. I built a workflow platform that mixes AI workloads with human-verified ones across several industries. The hard part was never getting AI to do the work. It was knowing when not to trust it, and proving the answer was right. Tags: AI + human routing, Confidence thresholds, Review queues, Audit trails AI is fast. Fast and wrong is worse than slow. Teams across these industries wanted AI to take the volume off their people. What they couldn't accept was AI quietly making mistakes that nobody caught until a customer, a regulator or an auditor did. - Volume they couldn't staff: Work arrived faster than people could process it, but it was too consequential to hand to a model and walk away. - No idea when AI was guessing: A confident-sounding answer and a correct answer look identical until someone checks. Nothing flagged the difference. - No paper trail: When an output was challenged, no one could say whether AI or a person made the call, what the input was, or who signed off. A platform that routes work, then proves it. - AI-and-human routing: Every unit of work enters one pipeline. The platform decides what AI can handle alone and what needs a person, so the two run as one system instead of bolted-together silos. - Confidence thresholds: Each AI result carries a confidence signal. Above the line it proceeds; below it, the work is automatically pulled out and sent to a human. The threshold is a dial the business controls, not a black box. - Human-in-the-loop review queues: Low-confidence and high-stakes work lands in a review queue built for speed: the AI's draft, its reasoning and the source side by side, so a reviewer verifies in seconds instead of starting over. - Audit trails on everything: Every decision records who or what made it, the input, the confidence, and the reviewer who signed off. When an output is questioned, the answer is one query away. - Agents on a tight leash: I treat agents like very stupid employees: narrow scopes, explicit guardrails, and no authority to act outside their lane. They're fast and tireless, never trusted to improvise. - Quality you can see: Human verdicts feed back as a continuous measure of AI accuracy, so quality never silently regresses as inputs and models drift. The same discipline I bring to AI-native transformation work. Three things that keep it honest. - 01 · Route · The right worker for the job: The platform's job is allocation. AI takes the volume it can handle confidently; people take the rest. Neither is the default, the work decides. (One pipeline for AI and human workloads; Confidence thresholds tuned per workflow and industry; Automatic escalation when the model isn't sure; No silent hand-offs, every route is logged) - 02 · Verify · Humans where they matter: People aren't there to rubber-stamp. They're there for the cases AI shouldn't decide alone, with everything they need to judge fast in front of them. (Review queues prioritized by risk and confidence; AI draft, reasoning and source shown together; Reviewer decisions captured as ground truth; Agents kept to tight scopes with hard guardrails) - 03 · Account · Proof, not vibes: Throughput is worthless if you can't defend the output. Audit trails and accuracy measurement make the whole system answerable. (Full audit trail on every AI and human decision; Accuracy tracked continuously from review verdicts; Drift caught before it becomes a customer problem; Clear answer to who decided what, and why) "Anyone can wire an LLM into a workflow. The work that matters is deciding when a human must look, and being able to prove, later, that the right one did.": Oshri Cohen · On hybrid AI systems What teams ask before they trust this. Q: How does the platform decide when a human has to step in? A: Every AI result carries a confidence signal. Each workflow has a threshold the business sets; above it, the AI's output proceeds automatically, below it the work is pulled out and routed to a human review queue. High-stakes categories can be sent to a person regardless of confidence. The threshold is an explicit dial, not a hidden heuristic, so the trade-off between throughput and verification stays in the team's hands. Q: How do you stop AI accuracy from quietly degrading over time? A: Human reviewers are the ground truth. Every verdict they give is captured and fed back as a continuous measure of how often the AI was right. If accuracy starts drifting as inputs change or a model is updated, the numbers move before a customer notices, and the confidence threshold can be tightened in response. Quality is measured on an ongoing basis, not assumed. Q: How do you keep the AI agents from doing something they shouldn't? A: I treat agents like very stupid employees: they get a narrow scope, explicit guardrails, and no authority to act outside their lane. They're fast and tireless, which is exactly why they can't be trusted to improvise. Combined with audit trails on every action, that keeps the system fast without making it reckless. It's the same discipline I bring to AI-native transformation engagements. Need AI throughput you can stand behind? If you want AI doing real volume without surrendering accuracy or accountability, this is the pattern. Let's talk about your workflow. ### Marketing Automation with Dynamic Twilio Numbers: /projects/marketing-automation-twilio/ A marketing automation platform that qualifies leads in real time using dynamic phone numbers provisioned through Twilio, tying every inbound call back to the campaign that produced it. I built a platform for a marketing and lead-generation company that qualifies inbound phone leads automatically and ties every call back to the campaign that paid for it. Dynamic Twilio numbers turn the phone, the hardest channel to attribute, into clean, closed-loop data. Tags: Dynamic number pools, Call tracking & attribution, Real-time lead scoring, CRM integration Phone leads were a black box. The company spent across dozens of campaigns and channels, but the moment a prospect picked up the phone instead of filling out a form, the trail went cold. - No campaign attribution: A call came in on one shared business number. Nobody could say which ad, landing page or source actually drove it, so budget decisions were guesswork. - Manual qualification: Every call had to be triaged by a person before it reached sales. Good leads waited in a queue while reps burned time on calls that were never going to close. - Disconnected systems: Marketing spend lived in ad platforms, calls lived in a phone system, and deals lived in the CRM. None of it joined up, so the real cost per qualified lead was unknowable. A closed-loop attribution engine. - Dynamic number pools via Twilio: I provision and recycle pools of phone numbers programmatically through Twilio's API and assign a unique number to each campaign, source or visitor session, so the number a prospect dials is itself the attribution key. - Call tracking & attribution: Every inbound call is matched to the campaign, channel and landing page that surfaced its number. Marketing finally sees which spend produces phone leads, not just form fills. - Real-time lead scoring: Calls are qualified as they happen using signals from the call and the originating campaign, so a lead is scored and ranked before the conversation is even over. - Routing to sales: Qualified leads are routed straight to the right rep while intent is hot; unqualified traffic is filtered out so the sales team only spends time on calls worth having. - CRM integration: Scored calls, their attribution and the full context flow into the CRM automatically, creating and enriching lead records without anyone retyping anything. - Engineering leadership behind it: The platform was built and shipped the way I run delivery as a fractional CTO, pragmatic architecture, owned end to end. See my fractional CTO services. "The win wasn't a smarter dialer. It was closed-loop attribution, finally knowing which marketing dollar produced which qualified phone lead.": Oshri Cohen What people ask about this build. Q: Why use dynamic phone numbers instead of one tracking number? A: A single number tells you the phone rang; it can't tell you why. Provisioning a pool of numbers through Twilio and assigning them per campaign, source or session makes the dialed number itself the attribution key, so every call is traced back to the exact spend that produced it. Numbers are recycled from the pool, so it scales without buying one number per ad forever. Q: How does the platform qualify a lead in real time? A: Each call carries context from the campaign and source that generated its number, plus signals from the call itself. The platform scores that context as the call happens, ranks the lead, and decides whether it's worth a salesperson's time before the conversation ends, instead of leaving qualification to a manual triage queue. Q: Does it fit into an existing CRM and ad stack? A: Yes. It's built to sit on top of what the company already runs. Scored, attributed calls flow into the CRM automatically as enriched lead records, and the attribution data reconciles against the ad platforms so the team can see true cost per qualified phone lead without changing how they buy media. Have a channel you can't attribute? If marketing spend is going somewhere you can't measure, that's a system problem with a system fix. Let's talk about what it would take to close the loop. ### Multi-Brand Marketing & Commerce Architecture: /projects/multi-brand-commerce-architecture/ A shared, multi-tenant marketing and e-commerce platform I architected for a holding company with a portfolio of brands, build once, run many, without flattening what makes each brand distinct. A portfolio holding company owned a growing set of brands, each running its own stack, its own checkout, its own marketing tooling. Every brand reinvented the same plumbing, and none of them did it well. I architected a shared, multi-tenant marketing and commerce platform: build once, run many, without making every brand look and behave the same. Tags: Multi-tenant, Per-brand theming, Shared catalog & payments, Brand isolation Every brand had rebuilt the same store, badly. When a holding company grows by acquiring and launching brands, each one shows up with its own stack. Left alone, that means N teams maintaining N copies of the same commerce plumbing, and N places for it to break. - Duplicated everything: Each brand had its own checkout, catalog model, and marketing integrations. The same bug got fixed three different ways, three different times. - No shared view of the customer: Order, product, and campaign data lived in silos. Nobody could see a customer across brands, or move a winning play from one storefront to another. - Marketing reinvented per brand: Email, promotions, and analytics were wired up brand by brand. Every new launch started the marketing stack from scratch. - Scaling meant copy-paste: Adding a brand meant cloning a codebase and praying. Cost and risk grew linearly with the portfolio instead of being amortized across it. A shared platform that serves many storefronts. - Multi-tenant core: One codebase, one deployment, many tenants. A brand is configuration, not a fork. Shared services handle the work that no brand should be reimplementing on its own. - Per-brand storefronts: Each brand gets its own theme, domain, and storefront experience layered on the shared services, so the customer never sees the platform underneath, only the brand. - Shared catalog & data: A common product and order model with a single source of truth, so the company can finally see catalog and customers across the whole portfolio. - Shared payments: One payment integration and one set of compliance controls, hardened once and reused everywhere, instead of each brand maintaining its own fragile checkout. - Shared marketing tooling: Email, promotions, and analytics built into the platform, so a new brand inherits a working marketing stack on day one and proven plays travel between brands. - Governance & isolation: Tenant isolation, access control, and clear boundaries so brands share infrastructure without leaking into each other's data, traffic, or blast radius. "The whole game is the seam: shared where it's leverage, separate where it's identity. Get that line wrong and you either ship one bland store wearing different hats, or you rebuild the same engine N times. The architecture exists to hold that line." Shared efficiency vs. per-brand autonomy. - Build once, run many: Commerce, payments, data, and marketing are built and maintained in one place. The cost of running another brand drops toward configuration, not a new project. The portfolio gets leverage from every improvement. - Room to be different: Where a brand genuinely needs to differ, its look, its storefront flow, its promotions, the platform gives it real control instead of forcing it into one mold. Autonomy is a designed-in feature, not an escape hatch. What operators ask me about going multi-brand. Q: Why a shared platform instead of letting each brand keep its own stack? A: Because the expensive, risky parts of commerce, payments, the catalog and order model, marketing tooling, compliance, are the same for every brand. Maintaining N copies multiplies cost and risk with no upside. A shared, multi-tenant platform builds those once and lets every brand benefit from every fix, while the brand-facing experience stays its own. Q: Won't a shared platform make all the brands look the same? A: Only if you design it lazily. I treat per-brand identity as a first-class part of the architecture: each brand gets its own theme, domain, storefront experience, and merchandising on top of the shared services. Customers see the brand, never the platform underneath. The shared layer handles the plumbing, not the personality. Q: How do you keep the brands isolated from each other? A: Tenant isolation is built into the data model, access control, and runtime boundaries. Brands share infrastructure but not data or blast radius, one brand's traffic spike, bad deploy, or sensitive customer data stays contained to that tenant. Governance is a designed-in property, not something bolted on later. Running a portfolio of brands on too many stacks? If you're maintaining the same commerce plumbing five times over, there's a better architecture. Tell me what you're running and I'll tell you straight what I'd consolidate first. ### Pricing & Product Intelligence at 250M Pages/Day: /projects/pricing-intelligence-scraping/ I built the distributed scraping and data platform behind a pricing-intelligence company, processing over 250 million product pages a day into clean, deduplicated pricing intelligence. I designed and built the distributed scraping and data platform behind a pricing-intelligence company, turning hundreds of millions of raw product pages a day into clean, normalized pricing intelligence. The hard part wasn't reading one page. It was reading 250 million of them, correctly, cheaply, every day. Tags: Distributed crawling, Anti-bot resilience, Parsing & dedup pipeline, Cost control at scale Scale breaks everything that worked at small volume. A scraper that works on a thousand pages is a script. At hundreds of millions a day, every assumption you made quietly becomes the bottleneck, the bill, or the outage. - Targets fight back: Sites rotate layouts, throttle traffic, fingerprint clients and deploy anti-bot defenses. At this volume you hit all of them, all the time, and a 1% failure rate is millions of lost pages. - Raw pages aren't data: A captured page is noise: stale prices, duplicates, malformed markup, currency and unit mismatches. Pricing intelligence only exists after parsing, normalization and dedup, done reliably at the same scale. - Economics is the real constraint: At 250M pages a day, a few cents per thousand pages is the difference between a healthy margin and a platform that bankrupts the company. Correctness and cost are the same engineering problem. A platform engineered for correctness and economics. - Distributed crawling: A horizontally scaled crawl fleet with scheduling and back-pressure, so the system pulls hundreds of millions of pages a day without overwhelming targets or itself, and recovers cleanly from failure. - Resilience by design: Defenses against anti-bot systems, rotation, and layout drift. Sites change constantly; parsers and crawl strategies are built to detect breakage and degrade gracefully instead of silently emitting garbage. - Parsing & normalization: A pipeline that turns raw HTML into structured records: parsed, normalized to common product and currency shapes, and validated, so downstream pricing intelligence is comparable across sources. - Dedup & freshness: Hundreds of millions of pages collapse into the unique products and price points that matter, with freshness tracking so customers see current prices, not yesterday's snapshot. - Warehouse & scheduling: A data warehouse sized for this throughput, fed by schedulers that decide what to crawl, how often, and in what order, balancing coverage, freshness and cost against finite capacity. - Cost control at volume: Per-page economics tracked and tuned end to end, compute, bandwidth and storage, because at this scale efficiency is the product. This work sits at the core of how I think about AI and data at scale. Throughput that holds up every single day. - 250M+/day: Product data pages processed into clean pricing intelligence - 24/7: Continuous crawling with scheduling, back-pressure and failure recovery - Per-page: Cost tracked and tuned end to end so the unit economics actually work "At 250 million pages a day, correctness and cost are the same problem. A platform that reads the web cleanly but can't pay for itself is a science project, not a business.": Oshri Cohen · On data platforms at scale What teams ask about this. Q: How do you keep scraping reliable when sites are actively trying to block you? A: You assume breakage is the normal state, not the exception. The crawl fleet rotates and adapts, parsers are validated against expected shapes so layout changes are detected rather than silently passed through, and the pipeline degrades gracefully and recovers instead of failing whole runs. At 250M pages a day a small failure rate is millions of lost pages, so resilience is a first-class design goal, not a retry loop bolted on at the end. Q: How do raw pages become actual pricing intelligence? A: Through a pipeline: capture, parse, normalize and dedup. Raw HTML is parsed into structured records, normalized to common product and currency shapes, validated, then deduplicated down to the unique products and price points that matter, with freshness tracking so customers see current prices. The captured page is just noise until that pipeline runs reliably at the same scale as the crawl. Q: How do you control cost at that volume? A: By treating per-page economics as a core metric and tuning compute, bandwidth and storage end to end. At 250M pages a day, a few cents per thousand pages decides whether the platform has a healthy margin or quietly bankrupts the company, so I engineer for correctness and economics together rather than optimizing one and discovering the other later. Need a data platform that holds up at real scale? Whether it's pricing intelligence, large-scale crawling, or any pipeline that has to be correct and cheap at extreme volume, let's map what it actually takes. ## Web properties (products Oshri owns & operates) ### Watchly: /web-properties/watchly/ Watchly (watchlyplayer.tv) is a YouTube player for kids where parents approve every video before it can be watched. No algorithm, no recommendations, no ads, PIN-protected profiles and time limits that enforce themselves. Built, owned and operated by Oshri Cohen. Watchly is a YouTube player for kids where nothing plays unless a parent approved it first. You hand-pick videos or import whole playlists, your kids watch inside a clean, ad-free player, and when their time is up it simply stops. There is no recommendation engine anywhere in it. I built it, I run it, and my own family uses it every day. Tags: Parent-approved library, No algorithm, Ad-free, Works with free YouTube YouTube was never built for your kids. The content is fine. The delivery system is the problem: a recommendation engine tuned for watch time, pointed at a seven-year-old. - The algorithm has its own agenda: YouTube's recommendations exist to maximize watch time. Your kid starts on an innocent video and forty minutes later autoplay has drifted somewhere you'd never have chosen. That isn't a bug. It's the system working as designed. - YouTube Kids doesn't fix it: Filtering billions of videos with automation means weird and inappropriate content slips through constantly. Every parent who has actually used it knows this. A filter that's mostly right is still a gamble you take every single day. - Supervision doesn't scale: You can't sit next to them for every video, and the moment you look away the recommendations take over. Then comes the time-limit fight, because the app itself is engineered to make stopping feel like punishment. You are the algorithm. - Approve every video: Hand-pick single videos or import entire playlists and channels. Nothing is watchable until you've approved it, so the library is 100% parent-curated. - No recommendations, ever: There is no algorithm inside Watchly. When a video ends, kids pick the next one from their library. The rabbit hole simply doesn't exist. - Ad-free on free YouTube: Videos play without ads and without a YouTube Premium subscription. Kids never see a pre-roll, a mid-roll, or a thumbnail designed to bait them. - Time limits that hold: Set daily screen time per kid and the player stops itself when time is up. No negotiating with a child, no being the bad guy every night. - PIN-protected profiles: Each kid gets their own profile with their own library and limits. Settings and approvals sit behind a parent PIN, so nothing changes without you. - One subscription, whole family: Every kid, every device, one plan. Watchly runs in the browser and as an app, and there's a 7-day free trial to see if it sticks in your house. Built at home, out of necessity. Watchly started the way most honest products do: with a problem in my own house that I got tired of losing to. Put on one harmless video, walk away for ten minutes, and come back to find the algorithm has taken over. The content drifts. First it's slightly off, then it's loud junk engineered for clicks, and eventually it's something no parent would have chosen. I watched that loop repeat until the pattern was impossible to ignore. Here's the uncomfortable part: the algorithm isn't broken. It's doing exactly what it was built to do, which is maximize watch time for an advertising business. A recommendation engine trained on billions of hours of attention data, aimed at a child, is not a fair fight. The kid never had a chance, and neither does the parent standing behind the couch trying to referee it. The industry's answer has been filtering: take the infinite feed and try to block the bad parts. YouTube Kids is the biggest attempt, and it still lets strange content through every day, because filtering billions of videos with automation is a losing game at the margins, and the margins are where kids live. After enough evenings watching filters fail, I landed on a different conclusion. Don't filter the feed. Replace it. Start from zero and let the parent add what's allowed, one video or one playlist at a time. The parent becomes the algorithm. That one decision drives everything in the product. Approval before playback, so the library is entirely parent-built. No recommendations of any kind, so a finished video leads back to the library instead of down a hole. Ads stripped out, because the attention economy has no business inside a kid's player. Time limits enforced by the player itself, so the screen goes off without a nightly standoff. PIN-protected profiles, so every kid's world stays their own and every setting stays yours. There's a professional reason too. I spend my working life telling companies that a small AI-native team can carry a product from idea to production and keep it running. Watchly is me taking my own medicine. One person owns the product design, the engineering, the infrastructure, the analytics and the support, with AI in the loop at every stage. It is the method I sell, running live, with real families depending on it, and it keeps my advice honest in a way no slide deck can. Watchly is live today at watchlyplayer.tv, in the browser and as an app, with a 7-day free trial. It runs daily in my own home, which means every rough edge gets found at my house before it gets found at yours. Filter the firehose, or build the library. The filter model: Everything, minus the bad parts: Starts with an infinite feed of everything; Automation guesses what's kid-safe at scale; Weird content slips through the margins daily; Autoplay and recommendations keep pulling; Parents audit the damage after the fact The Watchly model: Nothing, plus what you approve: Starts with an empty, per-kid library; A parent approves every video and playlist; Nothing unapproved can ever play; No recommendations, no autoplay, no ads; Parents decide up front, then relax The numbers that actually matter. - 100%: Of every kid's library is parent-approved before it can play - 0: Ads, recommendations, and autoplay surprises inside the player - 1plan: Covers the whole family, every kid and every device - 7days: Free trial to find out if the screen-time fights actually stop "Nobody loves your kid like you do, so nobody should curate for your kid but you. Watchly just makes that job take minutes instead of requiring you in the room.": Oshri Cohen · Founder, Watchly What parents ask about Watchly. Q: What is Watchly? A: Watchly is a YouTube player for kids where parents approve every video before it can be watched. Parents hand-pick videos or import playlists into a per-kid library, and kids watch inside an ad-free player with no algorithm, no recommendations, and no autoplay. It lives at watchlyplayer.tv and works with free YouTube. Q: How is Watchly different from YouTube Kids? A: YouTube Kids filters an infinite feed with automation, so inappropriate and strange content regularly slips through, and it still runs on recommendations. Watchly uses a whitelist instead of a filter: kids can only watch videos a parent explicitly approved, and there are no recommendations at all. Nothing unapproved can ever play. Q: Do I need YouTube Premium for it to be ad-free? A: No. Watchly plays videos without ads on free YouTube. There is no separate Premium subscription required, and kids never see pre-roll or mid-roll ads inside the player. Q: How do the time limits work? A: Parents set daily screen-time limits per kid, and the player enforces them itself: when time is up, playback stops. Kids' profiles are PIN-protected, so limits, libraries, and settings can only be changed by a parent. Q: How many kids and devices does one subscription cover? A: One subscription covers the whole family: every kid gets their own PIN-protected profile with their own library and limits, on every device. Watchly runs in the browser and as an app, and offers a 7-day free trial. Q: Why did Oshri Cohen build Watchly? A: Oshri Cohen built Watchly after watching YouTube's recommendation algorithm repeatedly steer kids' viewing toward content no parent would choose, and concluding that filtering an infinite feed can't work. Watchly replaces the algorithm with the parent. It is also a working proof of his AI-native operating thesis: one person building and running a production consumer product end-to-end. Tired of refereeing the algorithm? Try Watchly free for 7 days. And if you're a founder who wants to see how one person ships and runs a product like this, I'm happy to talk shop. ## Media appearances (podcasts & video) ### Technical Recruiting Is Broken — Here's How to Fix It: /media/technical-recruiting-is-broken/ Published: 2025-10-15 · Video: https://www.youtube.com/watch?v=bbTXexV5Pls Why technical hiring keeps failing teams — and how to combine recruiter instincts with real engineering judgment. Technical recruiting is broken, and Oshri Cohen has strong opinions about why. In this conversation he argues that most hiring processes screen for the wrong things — and that fixing them means pairing a recruiter's instinct for people with an engineer's ability to actually evaluate the work. He draws on building and leading engineering teams across countries and time zones to lay out a more honest, more effective way to hire developers: assess the real signal, respect candidates' time, and stop optimizing for interview theater. Practical guidance for founders and hiring managers tired of expensive mis-hires. Takeaways: - Why most technical interview processes screen for the wrong things - Combining recruiter instinct with real engineering judgment - How to find genuine signal in a hiring process - Reducing the cost of expensive mis-hires Topics: Technical recruiting, Hiring, Engineering teams, Team building ### Why Most Startups Don't Need a CTO: /media/why-most-startups-dont-need-a-cto/ Show: Software Without Borders · Episode 30 · Published: 2025-05-15 · Video: https://www.youtube.com/watch?v=ZM2M8F1wWDU A hacker-turned-CTO on why most startups don't need a full-time CTO — and what to do instead. On Software Without Borders (Episode 30), Oshri Cohen makes a deliberately provocative case: most startups don't need a full-time CTO. Drawing on a career spent rescuing struggling tech teams across industries, he explains what founders actually need at each stage — and how to buy it efficiently. The conversation covers the difference between a builder and an executive, the real cost of hiring the wrong leader too early, and how a fractional model gives small companies access to senior judgment without the senior salary. Equal parts contrarian and practical — a good listen for any founder about to post a "CTO wanted" ad. Takeaways: - Why a full-time CTO is often the wrong first hire - Builder vs. executive — and which one you need now - The cost of hiring senior leadership too early - How fractional leadership closes the gap Topics: Fractional CTO, Startups, Hiring, Engineering leadership ### Navigating the Tech Landscape: The Role of a Fractional CTO: /media/navigating-the-tech-landscape/ Show: Think Big with Dan and Qasim · Published: 2025-04-15 · Video: https://www.youtube.com/watch?v=oOeZxnzQUBw Oshri Cohen joins Think Big with Dan and Qasim to explain what a fractional CTO really does — and when founders should bring one in. On Think Big with Dan and Qasim, Oshri Cohen maps the modern technology landscape for founders and explains exactly where a fractional CTO fits: senior technical leadership, on demand, without a full-time hire. He covers how he comes in to solve the hard problem and build the team to carry it, the signals that tell a founder it's finally time for real technical leadership, and how to get outsized leverage from one or two days a week of the right person. A clear, founder-friendly primer on a model more and more companies are reaching for as they scale. Takeaways: - What a fractional CTO actually does day to day - The signals that it's time to bring in technical leadership - How to get leverage from a part-time senior leader - Solving the hard problem, then building the team to carry it Topics: Fractional CTO, Founders, Technology strategy, Scaling ### Becoming an AI-Native Organization: /media/becoming-an-ai-native-organization/ Published: 2025-03-01 · Video: https://www.youtube.com/watch?v=ROLhtE2Il8Y Oshri Cohen on rebuilding how a company builds — so AI is the default, not an add-on. AI-native isn't a feature you bolt on — it's how the whole organization operates. In this appearance, Oshri Cohen describes rebuilding business, product and engineering processes for the AI-native era, so workflows, delivery and decision-making are AI-first by default. He talks through what that concretely changes — not just the tooling, but how teams are structured and how the org hires — drawing on his current focus redesigning operating models for AI. Takeaways: - Why AI-native is an operating model, not a feature - Rewiring workflows, delivery and decision-making - What changes in how teams are structured and hired Topics: AI-Native, Transformation, Operating model, Engineering leadership ### What a Fractional CTO Actually Does: /media/what-a-fractional-cto-does/ Published: 2025-02-01 · Video: https://www.youtube.com/watch?v=Dvj4Mcov0-c A quick take on the fractional CTO model — senior technology leadership without a full-time hire. A short-form take: what does a fractional CTO actually do? Oshri Cohen sums up the model — senior technology leadership on demand, without a full-time hire. He comes in to solve the hard problem, sets technical direction, and builds the team to carry it — hands-on, until it's done. Sixty seconds only gets you the shape of the answer. The fuller version: the role carries the same accountability as a full-time CTO, on a fraction of the hours and at a fraction of the cost. The fractional CTO explainer walks through how an engagement actually works. Still weighing whether that's the right shape of leadership for your company? Start with the comparison of fractional, full-time and interim CTOs. Takeaways: - Senior technology leadership on demand, without a full-time hire - Solve the hard problem, set technical direction, build the team to carry it - The same accountability as a full-time CTO, on a fraction of the hours Topics: Fractional CTO, Technology leadership ### Engineering in Service of the Business: /media/engineering-in-service-of-the-business/ Published: 2024-11-01 · Video: https://www.youtube.com/watch?v=tif6tWt8oZc Oshri Cohen on his core operating philosophy: product and business first, with technology and engineering in service of it. Oshri Cohen leads from a simple operating philosophy: think from a product and business point of view first, and let everything derive down to technology and engineering in service of business development. In this conversation he unpacks what that means in practice. He talks through translating business strategy into engineering execution — and engineering reality back into business decisions — the work he's done for 25 years at the seam between the boardroom and the codebase, and the lens he brings to every fractional CTO engagement. Takeaways: - Product and business first, technology in service of it - Translating business strategy into engineering execution - Bringing engineering reality back into business decisions Topics: Engineering leadership, Business strategy, Product, Operating model ### When to Bring In a Fractional CTO: /media/when-to-bring-in-a-fractional-cto/ Published: 2024-09-01 · Video: https://www.youtube.com/watch?v=DseAdeLcJCM Oshri Cohen on the signals that tell a founder it's time for senior technical leadership — and what changes once it arrives. In this appearance, Oshri Cohen — an AI-native Chief Product & Technology Officer who works as a fractional and interim CTO — talks through when a growing company should bring in senior technical leadership, and what actually changes once they do. Drawing on 25 years in software and 30+ fractional engagements since 2018, he covers the signals that it's time, the problems he tends to solve first, and how founders get outsized leverage from one or two days a week of the right person. Takeaways: - The signals that tell a founder it's time for senior technical leadership - The problems a fractional CTO tends to solve first - What actually changes once technical leadership arrives - Getting leverage from one or two days a week of the right person Topics: Fractional CTO, Founders, Scaling, Technology strategy ### Technical Due Diligence for Investors: /media/technical-due-diligence-for-investors/ Published: 2024-07-01 · Video: https://www.youtube.com/watch?v=kTVD_qjRosw Oshri Cohen on reading code, architecture, team and risk before a deal — and building the value-creation plan after. Before an investor signs, someone has to give a clear, honest read on the technology. In this conversation, Oshri Cohen talks about technical due diligence for private-equity deals — code quality, architecture, scalability, security and compliance, team and key-person risk. He explains how he writes it for partners rather than just engineers, and how the diligence becomes an AI-native value-creation plan once the deal closes. Takeaways: - What technical due diligence actually covers - Writing the report for partners, not just engineers - Turning diligence into a value-creation plan Topics: Technical due diligence, Private equity, Risk, AI-Native ### Why Every Tech Founder Needs a Fractional CTO by Their Side: /media/why-every-founder-needs-a-fractional-cto/ Show: The Jeff Bullas Show · Episode 211 · Published: 2024-06-15 · Video: https://www.youtube.com/watch?v=MhmA2rLDgyk Oshri Cohen on The Jeff Bullas Show: why a fractional CTO can be the highest-leverage hire a founder makes. On The Jeff Bullas Show (Episode 211), Oshri Cohen explains why so many founders are turning to fractional technology leadership — and how the right CTO "by your side" changes what a small team can realistically build. He breaks down the economics — founders can save roughly $400k a year in direct technical-leadership cost — the kinds of problems a fractional CTO solves fastest, and the signals that tell you your business is ready for one. A founder-focused conversation about getting senior technical judgment without committing to a senior full-time salary. Takeaways: - The economics: ~$400k/year saved on direct technical leadership - The problems a fractional CTO solves fastest - How to tell when your company is ready for one - Why "a CTO by your side" beats no CTO at all Topics: Fractional CTO, Founders, Tech leadership economics, Scaling ### The CTO's Role at Each Stage of a Company: /media/the-cto-role-at-each-stage/ Published: 2024-04-01 · Video: https://www.youtube.com/watch?v=kG5kNQyUdBo Oshri Cohen on how the CTO job changes from seed to scale — and why the right leader at one stage is the wrong one at another. The CTO job at seed looks nothing like the CTO job at scale. In this appearance, Oshri Cohen walks through how the role changes stage by stage — from the hands-on builder to the strategic executive — and why the right leader at one stage is often the wrong one at the next. It's a practical guide for founders and boards trying to match the technology leader to where the company actually is, rather than where they hope it's going — and for anyone weighing whether a fractional CTO fits their current stage. Takeaways: - How the CTO job changes from seed to scale - From hands-on builder to strategic executive - Why the right leader at one stage is often wrong at the next - Matching the technology leader to where the company actually is Topics: The CTO role, Startup leadership, Scaling, Hiring ### Managing Global Dev Teams: Cultural Nuances & Leadership Challenges: /media/managing-global-dev-teams/ Published: 2024-02-15 · Video: https://www.youtube.com/watch?v=oAazwAaqKkA Lessons from leading 12 engineering teams across 7 countries: managing distributed, multicultural development teams. Oshri Cohen has, at one point, directed twelve engineering teams across seven countries and four time zones — so the challenges of managing global, multicultural development teams are familiar territory. In this conversation he digs into the cultural nuances, communication norms, and leadership habits that make distributed engineering actually work: building trust across borders, keeping delivery aligned across time zones, and turning a scattered group of teams into one organization that ships. Useful for any leader scaling engineering beyond a single office or country. Takeaways: - Building trust across borders and cultures - Keeping delivery aligned across time zones - Communication norms that make distributed teams work - Turning many teams into one organization Topics: Distributed teams, Engineering leadership, Global teams, Org & culture ### Fractional vs. Full-Time vs. Interim CTO: /media/fractional-vs-full-time-vs-interim-cto/ Published: 2024-01-01 · Video: https://www.youtube.com/watch?v=T9hEQYcYMgU Oshri Cohen on the difference between fractional, interim, and full-time technology leadership — and which one a company actually needs. "Fractional," "interim," and "full-time" CTO get used interchangeably, but they're different jobs. In this conversation Oshri Cohen draws the distinctions — and helps founders figure out which one fits their stage and their problem. (See also the full comparison.) He explains where a part-time fractional leader creates the most leverage, when a full-time interim CTO in the seat is the right call, and how to avoid the expensive mistake of hiring the wrong shape of leader too early. Takeaways: - How fractional, interim and full-time CTO roles actually differ - Where a part-time fractional leader creates the most leverage - When a full-time interim CTO in the seat is the right call - Avoiding the wrong shape of leader too early Topics: Fractional CTO, Interim CTO, Hiring, Technology leadership ### Should You Hire a Fractional CTO?: /media/should-you-hire-a-fractional-cto/ Show: That Tech Pod · Published: 2023-11-21 · Audio: https://www.buzzsprout.com/1738696/episodes/13947930 Oshri Cohen joins Laura Milstein and Kevin Albert on That Tech Pod to unpack what a fractional CTO actually is, how the job compares to a full-time seat, and whether a C-suite role can be vended out at all. On That Tech Pod, hosts Laura Milstein and Kevin Albert put the question directly to Oshri Cohen: should you hire a fractional CTO? It's a fair challenge. The model still sounds odd to a lot of founders, so the episode starts at the beginning and defines what the role actually is. Oshri lays out how a fractional CTO differs from a full-time one. The accountability is the same; the hours and the price tag are a fraction. From there the conversation gets into the harder question of whether you can vend out a C-suite seat at all, and where doing so makes sense for a company that needs senior technology leadership before it can justify the full-time hire. A useful listen for anyone weighing that hire right now. Takeaways: - What a fractional CTO actually is, and how the role differs from a full-time CTO - Whether a company can, or should, vend out its C-suite - How to think about the fractional model when hiring technology leadership Topics: Fractional CTO, Technology Leadership, Hiring ### Scaling a SaaS Engineering Organization: /media/scaling-a-saas-engineering-org/ Published: 2023-11-01 · Video: https://www.youtube.com/watch?v=xY3ycWDLEi4 Oshri Cohen on scaling engineering from a handful of developers to many teams without losing speed. Growth breaks the things that worked when you were small. In this conversation, Oshri Cohen talks about scaling a SaaS engineering organization — the structure, process and leadership that keep delivery fast and predictable as headcount climbs. He draws on directing twelve engineering teams across seven countries and four time zones at one point, and on building organizations that stay measurable and aligned while they grow. Takeaways: - The structure and process that keep delivery fast as headcount grows - Lessons from directing twelve teams across seven countries - Keeping a growing org measurable, predictable and aligned Topics: Scaling, SaaS, Engineering org, Delivery ### Programming, Business, and Becoming the Fractional CTO: /media/becoming-the-fractional-cto/ Show: Exponential Growth · Episode 44 · Published: 2023-10-17 · Video: https://www.youtube.com/watch?v=kLgrvcgKeus From self-taught developer to fractional CTO: the path through programming and business that led to the work Oshri Cohen does today. On Exponential Growth (Episode 44), Oshri Cohen tells the long-form story behind the title: a self-taught developer who worked through nearly every role in a technical organization on the way to CTO — and ultimately to building a fractional and interim CTO practice. He talks through the mindset shift from writing code to driving business outcomes, why he leads with "product and business first, technology in service of it," and how that lens reshapes the way he builds teams and rescues troubled software. A useful listen for engineers wondering how a technical career compounds into leadership — and for founders trying to understand how a fractional CTO actually thinks. Takeaways: - How a self-taught developer grows into a CTO, one role at a time - The shift from writing code to owning business outcomes - "Product and business first, technology in service of it" - Why the fractional model fits the way Oshri likes to work Topics: Fractional CTO, Career journey, Programming, Business strategy ### Unlocking Startup Success: Mastering Tech Leadership: /media/unlocking-startup-success/ Published: 2023-09-15 · Video: https://www.youtube.com/watch?v=bsmja06tyzE Oshri Cohen on the technology-leadership habits that actually move a startup forward. What does it take to lead technology at a startup without slowing it down? In this conversation, Oshri Cohen shares the leadership patterns behind the products and teams he's built — from setting clear technical direction to keeping delivery fast and predictable. He connects engineering decisions back to the business at every turn, and talks through how founders can get senior technical leadership working for them early, drawing on 25 years across HealthTech, e-commerce, logistics and more. A grounded look at the habits that separate teams that ship from teams that stall. Takeaways: - Setting technical direction a team can actually follow - Keeping delivery fast and predictable as you grow - Tying every engineering decision back to the business - Getting senior technical leadership early Topics: Tech leadership, Startups, Engineering management, Delivery ### The Role of a CTO in Compliance: /media/the-role-of-a-cto-in-compliance/ Show: Innovation in Compliance · Compliance Podcast Network · Published: 2023-08-29 · Video: https://www.youtube.com/watch?v=lPB4mQ5sHT8 Oshri Cohen joins Thomas Fox on Innovation in Compliance to unpack how the CTO's role in data strategy and governance shifts with company size. On Innovation in Compliance with Thomas Fox, Oshri Cohen explores what a Chief Technology Officer actually owns when it comes to compliance and data governance — and how that mandate changes dramatically with the size of the company. In larger organizations the CTO is a strategic planner; in smaller ones the CTO is often the head engineer writing the code. Oshri and Tom dig into data strategy, security, and the working partnership between the CTO and the Chief Compliance Officer that keeps a growing company both fast and safe. Drawing on Oshri's work across HealthTech and other regulated industries — including SOC 2, HIPAA and GDPR programs — it's a grounded look at building compliance into engineering rather than bolting it on afterward. Takeaways: - How the CTO's compliance role scales from head engineer to strategist - Why data strategy and governance are a CTO responsibility, not an afterthought - Where the CTO and Chief Compliance Officer have to partner - Lessons from SOC 2, HIPAA and GDPR programs in regulated industries Topics: Compliance & governance, Data strategy, SOC 2 · HIPAA · GDPR, The CTO role, Security ### Technology Management in the Age of AI: /media/technology-management-in-the-age-of-ai/ Show: In Systems We Trust · Episode 71 · Published: 2023-07-18 · Video: https://www.youtube.com/watch?v=j1z0RAb9SoM Oshri Cohen and Marquis Murray on managing technology — and technology teams — as AI reshapes how software gets built. On In Systems We Trust, host Marquis Murray sits down with Oshri Cohen to talk about managing technology in the age of AI — how leaders should think about tooling, teams, and decision-making as AI moves from novelty to default. The conversation ranges across building and leading engineering organizations, where AI genuinely creates leverage versus where it's just noise, and how a technology leader keeps both the systems and the people trustworthy while the ground keeps shifting. It's an early, prescient take on the AI-native operating model Oshri now builds with clients — rewiring how the product is built and how the organization runs so AI is the default rather than an add-on. Takeaways: - Where AI creates real leverage in an engineering org — and where it doesn't - How to manage teams while the tooling changes underneath them - Keeping systems and decisions trustworthy in an AI-first workflow - The seeds of an AI-native operating model Topics: AI-Native, Technology management, Engineering leadership, Systems ### Modernizing Legacy Systems: /media/modernizing-legacy-systems/ Published: 2023-06-01 · Video: https://www.youtube.com/watch?v=vtCwbj58VLs Oshri Cohen on re-platforming legacy systems to cloud-native — with downtime measured in minutes, not days. Moving critical software off legacy infrastructure is where many modernization efforts stall. In this appearance, Oshri Cohen talks about re-platforming on-prem systems to cloud-native architectures on Kubernetes, AWS, Azure and GCP — with downtime held under an hour. He shares lessons from a year-long, zero-downtime logistics migration and other re-architectures, including how to sequence the work so the business keeps running throughout. Takeaways: - Sequencing a migration so the business keeps running - Holding downtime under an hour on critical systems - Re-platforming to Kubernetes across AWS, Azure and GCP Topics: Cloud-native, Modernization, Migration, Architecture ### Rescuing Troubled Software and Teams: /media/rescuing-troubled-software/ Published: 2023-05-01 · Video: https://www.youtube.com/watch?v=MgwQSAvjuJ4 Oshri Cohen on turnarounds: inheriting codebases with no tests, docs, or team — and turning them around in months. Some of Oshri Cohen's hardest engagements start with a codebase that has no tests, no documentation, and no team. In this appearance he talks about turnarounds — stabilizing troubled software and the people around it, then charting a credible path back to fast, predictable delivery. He shares how he diagnoses the real root cause rather than the loudest symptom, rebuilds engineering from the ground up, and has turned six products around in three to six months. Takeaways: - Diagnosing the real root cause, not the loudest symptom - Stabilizing a troubled product and team - Charting a credible path back to predictable delivery Topics: Turnarounds, Engineering leadership, Delivery, Rescues ### Building High-Performing Engineering Teams: /media/building-high-performing-engineering-teams/ Published: 2023-03-01 · Video: https://www.youtube.com/watch?v=yU_UTmukXQI Oshri Cohen on turning engineering orgs into measurable, predictable, high-performing teams with DORA, GitOps and DevOps. What separates a team that ships from one that stalls? In this appearance, Oshri Cohen talks about building high-performing engineering organizations — measurable, predictable, and sustainable — using DORA metrics, GitOps and modern DevOps practice. He draws on turnarounds where he took a low-performing org to elite DORA metrics in 90 days by restructuring process, CI/CD and delivery, and on leading teams across countries and time zones. Takeaways: - Why DORA metrics are the scoreboard for delivery - Restructuring process, GitOps and CI/CD for speed - Making delivery predictable and sustainable as you grow Topics: DORA & DevOps, Engineering culture, Delivery, Team building ### From Self-Taught Developer to Fractional CTO: /media/from-self-taught-to-fractional-cto/ Show: Software Developer's Journey · Episode 233 · Published: 2022-12-20 · Audio: https://devjourney.info/Guests/233-OshriCohen.html On episode 233 of Software Developer's Journey, Oshri Cohen tells Tim Bourguignon how a self-taught engineer worked through every technical role on the way to CTO, and why he chose to go fractional. Tim Bourguignon's Software Developer's Journey asks developers how they became who they are. In episode 233 it's Oshri Cohen's turn, and his answer starts without a classroom: he taught himself to code, then learned the rest of the craft by doing the work. The conversation follows the whole arc. Oshri came up through every technical role a software company has, from developer onward, and each seat shaped how he leads engineering organizations today. Eventually the path reached CTO. Then it kept going, and Tim digs into why Oshri traded the single full-time seat for fractional work across several companies at once. If the model is new to you, the fractional CTO explainer covers what the job actually is. If you're a self-taught engineer wondering how far the road goes, this episode is the long answer. Takeaways: - How a self-taught developer built a career without the traditional credential - The path through every technical role on the way to CTO - Why Oshri left the single full-time seat and went fractional Topics: Career Journey, Fractional CTO, Engineering Leadership ### From Developer to CTO: /media/from-developer-to-cto/ Published: 2022-09-01 · Video: https://www.youtube.com/watch?v=clbt4PrNW0Y Oshri Cohen on the path from self-taught developer through every technical role to CTO — and what he learned along the way. Oshri Cohen started coding at 13 and worked through nearly every role in a technical organization on the way to CTO. In this conversation he traces that path — and the mindset shifts that turn a strong engineer into a technology leader. From principal engineer and director of engineering to VP and CTO seats, he reflects on what compounds over a career and how it shaped the fractional and interim CTO practice he runs today. Takeaways: - Working through nearly every technical role on the way to CTO - The mindset shifts that turn a strong engineer into a technology leader - What compounds over a career, from principal engineer to CTO seats - How the path shaped a fractional and interim CTO practice Topics: Career journey, Engineering leadership, Programming, Mentorship ### What Kind of CTO Do You Need?: /media/what-kind-of-cto-do-you-need/ Show: Tiny DevOps · Episode 48 · Published: 2022-05-31 · Video: https://www.youtube.com/watch?v=pcgbHSpGaoY On Tiny DevOps, Oshri Cohen breaks the CTO role into four distinct phases — and explains why one person rarely fits all of them. The title "CTO" hides at least four very different jobs. On Tiny DevOps with host Jonathan Hall, Oshri Cohen pulls the role apart into its phases — from the hands-on founding builder to the strategic executive who runs a large organization — and explains why the same person rarely satisfies all of them as a company grows. Oshri makes the case that many early-stage startups don't need a full-time CTO at all, and walks through what a founder could do with the roughly $150k saved by bringing in a fractional CTO instead — buying senior technical judgment exactly when it's needed, without the full-time cost. It's a practical, myth-busting conversation for any founder trying to work out who should own technology at their stage of the journey — and what to actually hire for. Takeaways: - The four phases of the CTO role — and why one person rarely covers them all - Why most early-stage startups don't need a full-time CTO - What a founder could do with the ~$150k a fractional CTO saves - How to match the technology leader to the company's stage Topics: Fractional CTO, The CTO role, Startup leadership, Hiring ### AI Innovations in Industry: /ai-innovations/ AI applications and agent systems Oshri Cohen has designed across industries, some shipped to production, others architected as blueprints, from accounting and tax to HealthTech, logistics and manufacturing. A portfolio of AI applications and agent architectures I've designed across industries, some shipped to production, others mapped out as blueprints. Organized by the business problem they solve, not by the model that happened to be in fashion. Tags: AI agent architecture, Graph, not swarm, Human-in-the-loop, From problem to system "The interesting question was never "where can we add AI?" It's which problem in this industry is really a reading, reasoning and form-filling problem in disguise, and then designing the system that solves it.": Oshri Cohen What I've built, and what I'd build. - 01 · Accounting & Tax · An agent team for tax & filings: A graph of narrow AI specialists, tax-law, accounting-data, deductions, compilation, a critical thinker, an auditor and an orchestrator, that reads the data, computes against the rules, and drafts the filings, with the accountant accountable at the end. (One tax-law agent per section: personal, corporate, trusts; A deductions agent that hunts for legitimate savings; A graph, not a swarm, so every conclusion is traceable; Human signs off on a summary of the reasoning, not raw data) - 02 · E-commerce & Fulfillment · Margin-finding agents for fulfillment: In a commodity market the price is fixed, so profit lives in the cost of fulfilling each order. A graph of agents sources every line, across owned inventory and drop-ship suppliers, picks the cheapest carrier that still meets the promise, and holds every order above a margin floor. (Sourcing agents across owned stock and drop-ship suppliers; Carrier-rate agents pick the cheapest path that meets the promise; A landed-cost agent computes true margin per fulfillment plan; A margin guardrail blocks any order that would ship at a loss) - 03 · HealthTech & Public Data · Turning public healthcare data into answers: The government publishes an enormous amount of healthcare data that almost nobody can actually use. I built the thing that makes it usable: ask a question the way you'd say it out loud, and get a clear answer in seconds, the kind that used to mean hiring analysts and waiting weeks. It's the same instinct as the agent work, do the hard part once, up front, so every answer after that is quick and cheap. (All of it gathered into one place, finally lined up so it can be compared; Stays current on its own as new data is published; Ask in plain English, no specialists or special skills; Powers clients' own dashboards and reports) From problem to system. - Start from the problem: I start from the business outcome, not the model. The first question is which work in an industry is really reading, reasoning and computation in disguise, and where AI is just a distraction. - Design the agent system: Most of these are agent architectures: narrow specialists organized under a hierarchy you can trace, designed so the output is something a professional can defend, not a confident guess. - Ship where it counts: Where a design goes to production, it ships with evaluation, observability and cost control, and a human kept at the point where judgment and accountability belong. What teams ask about this work. Q: Have you shipped all of these, or are some designs? A: Both. Some are systems running in production; others are agent architectures I've designed for a specific problem, blueprints ready to build. Each entry is clear about which it is. Q: Can you adapt one of these to my industry? A: Usually, yes. The pattern, narrow specialist agents under a traceable hierarchy with a human accountable at the end, transfers across domains. The work is mapping it to your data, your rules, and your regulatory reality. Q: Do you build, or only design? A: Both. I work hands-on, architecting and, where it makes sense, building and shipping the system with evaluation, observability and cost control. Some engagements stop at a design and a roadmap your team executes. Q: How do you keep these reliable in a regulated industry? A: By organizing agents as a graph rather than a swarm, so every conclusion is traceable, and by keeping a human at the end of the process to review the reasoning, not just the result. Have a problem that AI could solve? Tell me the industry and the outcome you're after. I'll tell you straight whether AI is the right tool, and what it would take to design and ship it. ## Philosophy Twenty-five years at the seam between the boardroom and the codebase taught me a handful of principles I don't compromise on. They shape every engagement, from the first board conversation to the last pull request. - 01 Business before backlog: I start from the product and the P&L, then build the engineering to match. Technology exists in service of business development, never the other way around. - 02 Read the thesis before the code: Architecture is a means, not an end. Before I judge a system I learn what the business is trying to become, then protect and compound the value of the asset. - 03 Measure the system, not the people: Metrics that rank individuals get gamed. Point the measurement at flow, stability and friction, and you improve the machine everyone works inside. - 04 Make delivery boring: Elite performance isn't heroics, it's a system that makes the safe path the easy path. Small batches, automation and predictability over drama. - 05 Obsess over the operator: Customer experience wins the deal; operator experience decides whether you can afford to keep it. I design for the people who run the business, not only the people who buy it. - 06 AI-native by default: Don't bolt AI onto an old operating model. Rebuild how products are built and how the organization runs so that AI is the default, not an add-on. "I think from a product and business point of view first, and everything derives down to technology and engineering in service of business development." ## Essays (Thoughts and Musings) ### A2A Explained: How AI Agents Delegate to Each Other: /blog/a2a-how-agents-talk-to-each-other/ MCP gives one agent hands. A2A is what happens when agents need coworkers: discovery, delegation, and a task lifecycle that survives the messy middle of real work. A single agent with good tools is an employee. Useful, bounded, and easy to reason about. The interesting problems start when one agent can't finish the job alone and needs to hand part of it to another agent, possibly one built by another team, or by another company entirely. That handoff is what A2A exists for. This is the second essay in a series on how agents communicate. The first covered MCP, the agent-to-tool layer. The third covers ACP, the rival protocol that merged into this one, and the comparison piece ties it all together. The short version of where A2A sits: MCP connects an agent to capabilities, A2A connects an agent to peers. Why tools aren't enough You can get surprisingly far pretending other agents are tools. Wrap an agent in an API, expose it through MCP, call it like a function. I've done it. It works right up until the work stops being a function call, and real work usually does. A tool call is synchronous and stateless: send arguments, get a result. Delegated work isn't like that. It drags on for minutes or days. Sometimes it fails halfway and somebody has to decide what happens next; sometimes it stops with a question. A hiring agent asked to source candidates might need to ask the requesting agent whether remote is acceptable before it can continue. A function signature has no room for "hang on, I need more information from you." A task does. "A tool call is a question with one answer. A task is a relationship with a lifecycle." How A2A actually works Google published A2A in 2025 with a long list of launch partners, then handed it to the Linux Foundation, which is exactly what you want to see in a protocol you might bet on: no single vendor holding the pen. Mechanically it's built on boring, proven parts, HTTP and JSON-RPC and server-sent events, and that's a compliment. The interesting design is in three ideas. Discovery through Agent Cards. Every A2A agent publishes a card, a JSON description at a well-known URL that says what the agent can do, what skills it offers, and how to authenticate with it. When an agent can't complete a task, it doesn't need a hardcoded integration to a specific peer. It reads cards, finds an agent whose declared skills match the need, and delegates. It's a resume plus a front door, machine-readable. Tasks with a real lifecycle. Delegation creates a task object with states: submitted, working, completed, failed, and the one that makes the whole protocol worth having, input-required. When the remote agent hits something it can't resolve alone, it pauses the task in that state and loops back to the requester with a question. The requester answers, the task resumes. Long-running work streams progress over SSE, and results come back as structured artifacts rather than a blob of prose. Opacity by design. This is the piece people miss. A2A agents don't share memory, tools, or internal state. Each one is a black box that accepts tasks and returns results. Your agent doesn't get to see how the other agent works, and that's a feature. It means the two sides can be built on different frameworks, run by different companies, and neither has to trust the other with anything more than the task itself. Two businesses work the same way. They exchange commitments, and neither hands the other a login to its systems. The org-design lens I keep telling teams to design each agent like a very stupid employee with exactly one job. A2A is what turns those employees into an organization. One agent owns travel booking, another owns expense policy, another owns calendar negotiation, and they coordinate through explicit, inspectable task handoffs rather than one bloated genius agent trying to hold everything in its head. The parallel goes further than people expect. Everything we know about human organizations starts applying: clear ownership beats overlapping mandates, and the interface between teams matters more than what's inside them. A handoff without a status update is where work goes to die. A2A essentially forces good org design onto your agents, because the protocol only carries tasks, statuses, and artifacts. If you can't express the collaboration in those terms, you haven't defined the jobs clearly enough. An honest word about timing Here's the part vendors won't tell you: most companies don't need A2A yet. If all your agents live in one codebase, run on one framework, and answer to one team, a shared orchestrator with plain function calls is simpler and easier to debug. A2A earns its complexity at boundaries: between departments that ship independently, or between your company and a vendor's agent. No boundary, no need for a boundary protocol. The other cost is verification. Two autonomous, non-deterministic systems negotiating with each other multiplies the ways a workflow can go subtly wrong, and your evals have to grow to match. I wrote about that problem in testing non-deterministic agents: every new component in the pipeline is a new place for failure to hide. Multi-agent handoffs are exactly such a component. Budget for grading the coordination, not just each agent's individual output. So my playbook is unglamorous. Get MCP right first, because tool access is where the immediate ROI lives. Design narrow agents with clean ownership. Then, when a real boundary shows up, and it will, adopt A2A at that boundary instead of inventing a bespoke agent-messaging scheme you'll regret. The standard exists precisely so that this seam is not your problem to design. The bottom line A2A is the org chart layer of the agent stack: discovery through Agent Cards, delegation through tasks, collaboration between black boxes that never have to trust each other's internals. It complements MCP rather than competing with it, one protocol for hands, one for coworkers. If you're deciding how those layers fit your own stack, the field guide to all three protocols is the map. And if you're staring at a multi-agent architecture decision right now and want a second opinion from someone who has shipped these systems, let's talk → ### ACP: The Agent Protocol That Merged Into A2A, and Why That's Good News: /blog/acp-the-protocol-that-merged-into-a2a/ IBM's Agent Communication Protocol took a REST-first swing at agent-to-agent communication, then folded into A2A. The merge is the most instructive part of the story. Most essays about protocols are about winners. This one is about a protocol that folded, and I'm writing it anyway, because the way ACP ended tells you more about how to bet on agent standards than any launch announcement ever will. This is the third essay in a series on agent communication. The first covered MCP, the agent-to-tool layer, the second covered A2A, the agent-to-agent layer, and the field guide compares all three. ACP was A2A's direct competitor. Same problem, different philosophy, and for a while in 2025 you genuinely had to pick. What ACP was The Agent Communication Protocol came out of IBM's BeeAI work, and its pitch was refreshingly plain: agent-to-agent communication should just be REST. No special SDK required. Any agent exposes ordinary HTTP endpoints, and anything that can speak HTTP, which is everything, can call it. You could drive an ACP agent with curl. For a lot of engineers that was the whole seduction. Nothing new to learn, just endpoints. Discovery worked through an Agent Manifest, a description of the agent's capabilities that peers could read to find the right collaborator, much like A2A's Agent Card. Communication was flexible about time: call an agent synchronously when you needed a fast answer, or go async and stream results over server-sent events when the task ran long. Async was the default posture, which was the honest choice, because real agent work is rarely quick. Put next to A2A, the differences were real but not philosophical chasms. A2A leaned on JSON-RPC and a richer task-lifecycle model. ACP leaned on plain REST conventions and a lower barrier to entry. Both had discovery documents, both handled sync and streaming, both wanted to be the lingua franca between agents built on different frameworks. Two standards, one job. "Two standards doing one job is not competition. It's a tax on everyone who has to choose." The merge, and why it was the right call In 2025 the story resolved cleanly: Google contributed A2A to the Linux Foundation, IBM brought ACP into the same effort, and ACP as a separate track wound down. The ecosystem consolidated on A2A, with ACP's REST-first instincts absorbed into the surviving standard rather than thrown away. I want to be clear that this outcome was good, and not just for the winner. A fragmented agent ecosystem, where every framework speaks its own dialect, recreates exactly the integration hell that MCP had just eliminated at the tool layer. The entire value of an interoperability protocol is that there's one of it. The moment serious vendors lined up behind A2A under neutral governance, continuing ACP would have served IBM's ego at the ecosystem's expense. They folded it instead. That's maturity you don't often see in standards fights, which historically run for a decade out of pure stubbornness. What this teaches you about betting on standards If you built on ACP, you weren't wrong. That's the lesson people miss. The teams that adopted it early got their agents talking over clean HTTP interfaces with explicit manifests, and when the merge came, migrating to A2A was a port, not a rewrite, because the concepts mapped almost one to one. The teams that got burned were the ones who used the standards war as an excuse to build a proprietary agent-messaging layer, and who now own a dialect nobody else will ever speak. So here's the rule I give boards and CTOs when a young standard is still contested. Don't bet on which spec wins. Bet on the shape both specs share. When every contender has capability discovery, task delegation, and streamed results, those concepts are the actual standard, and the winning wire format is an implementation detail you can swap. Keep your agents behind thin adapters, and consolidation becomes a chore instead of a crisis. It's the same instinct as any good architecture bet: read the thesis, not just the code. The consolidation also simplified the pitch for everyone. The stack now has two layers with clear jobs: MCP for an agent's tools, A2A for an agent's peers. One protocol war, settled early, with the losing side's best ideas kept. If only every infrastructure decision resolved this well. The bottom line ACP mattered because it pushed the simplest possible answer, agents as plain REST services, hard enough that the surviving standard had to absorb it. Its afterlife inside A2A is a better legacy than limping along as a fragment. And the episode is a compact case study in how to adopt standards during consolidation: commit to the concepts and stay loose on the wire format, with a thin adapter between your agents and whatever the industry finally agrees on. The full comparison shows how the settled map looks now. Wrestling with a build-on-emerging-standard decision of your own? I've helped a lot of companies place those bets without getting locked in. Let's talk → ### MCP Explained: How AI Agents Actually Use Tools: /blog/mcp-how-agents-use-tools/ The Model Context Protocol is the reason your agent can read a database, file a ticket, or query your analytics without a custom integration for each one. Here's how it works and where teams get it wrong. Every agent demo hits the same wall. The model is smart, the conversation is impressive, and then someone asks the only question that matters: can it actually touch my systems? Can it read the CRM? Can it file a ticket in the tracker the support team actually uses? A model that can't act on real data is a very expensive chat window. The Model Context Protocol is the answer the industry settled on, and it settled fast. Anthropic published MCP in late 2024. Within a year it had stopped being an Anthropic thing and become the de facto standard for connecting agents to tools, the way USB-C became the default port. This essay is the first in a series on how agents communicate. The next two cover A2A, the agent-to-agent protocol, and ACP, the one that didn't survive, and the field guide puts the whole map together. The problem MCP solves Before MCP, connecting a model to a tool meant writing a custom integration. Your agent needed Salesforce, Postgres, and Slack? Three integrations. You switch model providers? Write them all again, because every provider had its own function-calling format. With M models and N tools you were staring at M×N integration projects, each one bespoke, each one rotting quietly as APIs changed underneath it. MCP collapses that to M+N. The tool side implements one MCP server, once. The model side implements one MCP client, once. Anything on one side of the protocol can now talk to anything on the other. That's the entire pitch, and it's why adoption was so quick: nobody loves writing glue code, and MCP made most of it somebody else's problem. "MCP turned M×N integration projects into M+N. That one change is why it won." How a tool call actually flows The protocol has three roles, and keeping them straight clears up most of the confusion. - Host — the application the user is actually in: Claude Code, an IDE, your internal agent app. It owns the conversation and the model. - Client — a connector embedded in the host. One client per server connection. It speaks the protocol so the host doesn't have to. - Server — the thing wrapping a capability: a database, a SaaS API, a file system. It advertises what it can do and executes requests. A request runs end to end like this. You ask the agent something that needs outside information. The model decides a tool is required and names the one it wants, with arguments. The client formats that as a JSON-RPC call and routes it to the right server. The server does the real work, calling the API or running the query, and returns a structured result. The model reads the result and keeps reasoning. From your side it looks like the agent just knew the answer. Under the hood it was a round trip you can log and put permissions around. Servers can expose more than tools. The spec also covers resources, which are pieces of context the host can pull in, like a file or a record, and prompts, reusable templates the server offers. Most of the ecosystem's energy is in tools, but the other two matter once you build serious internal servers. What it feels like in practice I run a lot of my work through agents, and MCP is the plumbing behind almost all of it. Analytics, GitHub, email, media tools: each is just a server my session connects to. When I want an agent to reach a new system, I don't scope an integration project. I find or write a server, add it to the config, and the tools show up in the agent's hands the same afternoon. That speed compounds. It's a big part of why a handful of agents on one laptop can cover ground that used to take a platform team. Writing a server for your own systems is also less work than people expect. If your internal API already exists, an MCP server is mostly a thin layer that describes each endpoint well enough for a model to use it responsibly. The code is the easy part. Deciding what the agent should be allowed to do takes longer, and that's a governance question you were going to have to answer anyway. Where teams get it wrong Three mistakes come up constantly. Treating MCP as agent-to-agent communication. It isn't. MCP connects an agent to capabilities. The moment you want two autonomous agents negotiating with each other and handing work back and forth, you've left MCP's job description and entered A2A territory. You can wrap an agent inside an MCP server and call it a tool, and sometimes that's fine, but you lose the task lifecycle and the back-and-forth that real delegation needs. Connecting everything and drowning the model. Every tool you expose is another thing the model can misuse or get distracted by. Forty servers with three hundred tools doesn't make an agent powerful. It makes it scattered. The fix is the same discipline I argue for in designing agents like very stupid employees: give each agent the few tools its one job requires, and nothing else. Ignoring the trust boundary. An MCP server executes real actions with real credentials, and tool results flow straight back into the model's context. A malicious or sloppy server is an injection vector. Treat servers like you treat third-party dependencies: vet them, pin them, and give them the narrowest credentials that still work. Log every call. Your security team already preaches this supply-chain hygiene; servers are just a new place to apply it. The bottom line MCP is the least glamorous layer of the agent stack and the most important one. It's how capability gets into an agent's hands, and it's already the standard, so adoption stopped being a real decision a while ago. What's left to decide is which of your systems deserve a server first, and what each agent is allowed to touch. Get that right and everything above it, including multi-agent delegation, gets much easier. If you're mapping out which internal systems to expose to agents and how to keep that safe, that's work I do with companies every week. Let's talk → ### MCP vs A2A vs ACP: A Field Guide to How AI Agents Communicate: /blog/mcp-vs-a2a-vs-acp/ Three protocols, two jobs, one settled map. How MCP and A2A divide the agent stack, where ACP went, and what to actually adopt at each stage of your rollout. One agent with good tools is useful. Several agents that can hand work to each other stop being a chatbot and start being an organization. Underneath either setup sits the same unglamorous question: how does everything talk to everything else? Three protocol names dominate that conversation: MCP, A2A, and ACP. They get compared as if they were three competitors, and that framing is wrong in a useful way. Two of them never competed at all, and the two that did have already merged. This essay is the map. Each protocol also has its own deep dive in this series: MCP, A2A, and ACP. Two different jobs Start with the distinction that dissolves most of the confusion. An agent has two kinds of conversations, and they need different plumbing. - Agent to tool. The agent needs a capability: run this query, send this email. The other side isn't intelligent. It executes and returns. This is MCP's job. - Agent to agent. The agent needs a colleague: another autonomous system that plans, works over time, and might come back with questions. This is A2A's job, and it was ACP's job too, which is exactly why one of them had to go. Comparing MCP to A2A is comparing a screwdriver to a phone. The real comparison was A2A versus ACP, and the market resolved it. MCP: the hands The Model Context Protocol, from Anthropic, standardizes how a host application connects its model to capabilities. The host embeds an MCP client, the tool side runs an MCP server, and tool calls flow between them as structured JSON-RPC: model picks a tool, client routes the call, server executes and returns a result the model can reason over. Before MCP, every model-tool pairing was a custom integration and the ecosystem was drowning in glue code. After it, tools are written once and work everywhere, which is why it became the layer everyone standardized on almost immediately. The mistakes teams make with it are about discipline rather than protocol: too many tools per agent, too little attention to the trust boundary. The MCP deep dive covers both. A2A: the coworkers A2A, started by Google and now under the Linux Foundation, standardizes delegation between autonomous agents. Discovery happens through Agent Cards published at well-known URLs, so an agent that can't finish a task can find a peer whose declared skills fit. Work travels as tasks with a real lifecycle, and the state that justifies the whole protocol is input-required: the remote agent can pause mid-task, ask the requester a question, and resume. Agents stay opaque to each other, no shared memory, no shared tools, which is what lets systems from different vendors and frameworks collaborate without trusting each other's internals. It's org design as a wire protocol, and the same rules apply: it shines at boundaries and adds overhead where none exist. The A2A deep dive goes deeper, including when not to bother yet. ACP: the one that merged ACP, from IBM's BeeAI effort, attacked the same agent-to-agent problem with a REST-first philosophy: agents as plain HTTP services, discovery through an Agent Manifest, sync responses for fast tasks and async SSE streams for long ones. It was a credible, simpler rival to A2A for much of 2025, and then IBM did the mature thing and folded it into the A2A effort under the Linux Foundation. The concepts survived inside the winner; the separate spec did not. If you're evaluating protocols today, ACP is history, but instructive history: the ACP essay pulls out what the merge teaches about betting on young standards without getting burned. "MCP gives an agent hands. A2A gives it coworkers. ACP gave A2A its best ideas and got out of the way." How they compose in production In a real deployment the two surviving protocols stack. Picture a procurement workflow. Your purchasing agent uses MCP to read inventory and pull contract terms: tool calls, all of it. Then it needs quotes evaluated against a supplier's own agent, an autonomous system your company doesn't control. That handoff goes over A2A. Your agent finds the supplier agent's card, hands over the evaluation, answers whatever input-required questions come back, and gets a structured result. That agent, in turn, is using its own MCP servers on its side of the fence. Every agent is an MCP client downward and an A2A peer sideways. The composition is also why the layering matters for safety and evals. Tool calls are auditable actions you can permission per agent, which pairs naturally with narrow, one-job agent design. Delegations are commitments between non-deterministic systems, which your evaluation suite has to grade as coordination, not just output, a problem I dig into in testing non-deterministic agents. What to adopt, in what order - Now, for everyone: MCP. Tool access is where agents earn their keep, the ecosystem is mature, and the cost is low. Pick the two or three internal systems with the most agent leverage and stand up servers for them. - When a boundary appears: A2A. The trigger is organizational, not technical: agents owned by different teams, companies, or vendors that must collaborate. Inside one team, an orchestrator and function calls remain simpler. - Never: a proprietary agent-messaging layer. The standards war is over. Building your own dialect now buys you nothing except a migration project later. - If you're on ACP: port to A2A. The concepts map nearly one to one, so the migration is mostly mechanical. The bottom line The agent communication map is settled enough to act on. MCP is the standard for giving agents capabilities, A2A is the standard for letting them delegate, and ACP's merge means you don't have to hedge between rivals at the agent-to-agent layer. At most companies the protocol choice turns out to be the easy part. The hard questions are which workflows deserve agents at all, what each agent may touch, and how you'll know the system is working. Those are strategy and operating-model questions wearing a technical costume. That's the layer I work at: helping leadership teams turn agent plumbing into an actual AI strategy with owners, boundaries, and measurable outcomes. If you're drawing this map for your own company, let's talk → ### How to Hire AI-Native Engineers: Screen for Judgment: /blog/hire-for-judgment/ The profile of a great engineer has changed. What to look for now: taste, skepticism toward plausible output, and leverage with AI agents. Two engineers, same seniority, same salary band. Give each of them the same AI tools and the same ticket. One ships a clean, tested change in an afternoon. The other ships three thousand lines of confident-looking code that costs the team a week to untangle. Same tools. Wildly different outcomes. That gap is the entire story of engineering hiring right now. The tools are a multiplier, and a multiplier is only as good as the number you feed it. Judgment is the number. Everything else on the resume is increasingly decoration. What judgment actually means "Hire for judgment" risks becoming one of those phrases everyone nods at and nobody can act on, so let me break it into the specific behaviors I screen for. After vetting engineers for my own teams for twenty years, and now doing it for clients, I've found judgment shows up in four observable habits. They treat plausible as a warning sign. Generated code fails differently than human code. It rarely looks wrong. It compiles, the names are sensible, the structure is textbook, and the bug is a quiet assumption three layers down. Engineers with judgment have recalibrated their suspicion: the cleaner the output looks, the harder they check the assumptions underneath it. Engineers without judgment read clean code as done code. They know what not to build. When producing code was expensive, restraint was enforced by cost. Now that generating another abstraction layer is free, restraint has to come from taste. The best engineers I've hired lately are notable for how little code they ship. They delete the speculative flexibility the model added. They ask why the feature exists before improving how it works. They decompose before they delegate. Working with agents is a management skill. You break the problem into pieces with clear contracts, hand off the pieces the tools do well, and keep the pieces where the risk lives. Watch a strong engineer run an agent and it looks like a good tech lead running a sprint. Watch a weak one and it looks like a wish. They can descend the stack when it matters. Abstraction is wonderful until something breaks beneath it. The engineers worth hiring can drop below the generated layer, read what's actually happening, and come back up with the fix. This is where fundamentals still matter, just differently: less recall of algorithms, more the deep model of how systems behave. "The cleaner the output looks, the harder they check the assumptions underneath it." The seniority illusion Here's the uncomfortable part: years of experience predict this profile far less than you'd hope. I've screened fifteen-year veterans who use AI tools like a slot machine, pulling the lever until something passes the tests. I've screened engineers four years in who orchestrate agents with the discipline of a staff engineer, because they built their working habits in this era instead of retrofitting them. Experience still matters. Someone has to have seen a database fall over on a holiday weekend to truly respect a migration. But the correlation between tenure and AI-era effectiveness is loose enough that title-based hiring is now genuinely risky. You have to test the actual behaviors, on actual work, or you're hiring the label on the tin. Curiosity is a hiring criterion now One more trait, and I've stopped treating it as a nice-to-have: genuine curiosity. The tools change quarterly. The engineer who treats their workflow as finished, whatever that workflow is, gets a little more outdated every month without noticing. The engineer who keeps poking at the new model, the new agent pattern, the new way to structure context, compounds instead. In interviews this is easy to surface. I ask what they've changed about how they work in the last six months. People with real curiosity light up and get specific: they'll tell you what they tried, what failed, what stuck. People without it give you a tool name and a shrug. Six months from now, both answers will have compounded. How to actually screen for this None of these traits show up in a resume, and only weak echoes of them survive a conversational interview. You need to watch candidates work. My vetting for clients runs a realistic exercise with AI tools expected, and grades the process: what they questioned, what they caught, what they refused to ship, how they explained their reasoning afterward. The exercise is small. The signal is enormous. And be honest with yourself about the bar. A team of people with this profile is smaller, more expensive per head, and dramatically cheaper per outcome. Which changes how many people you should hire in the first place, but that's its own essay. If you want engineers vetted this way, by a CTO who has hired 75+ of them for his own teams, this is the service. Or just reach out → ### Is the Leetcode Interview Dead? What to Use Instead: /blog/the-leetcode-interview-is-over/ Algorithm puzzles test exactly the work AI now does for free. Here's what a technical interview should measure instead, and what mine looks like. The classic technical interview asks a candidate to invert a binary tree on a whiteboard, from memory, under pressure, while a stranger watches. We ran this ritual for twenty years. It was never a great test, and everyone quietly knew it, but it had one honest defense: writing correct algorithmic code under constraints was at least adjacent to the job. That defense is gone. The exact skill leetcode measures, producing known solutions to well-defined puzzles quickly, is now the single most automated part of software engineering. A model does it faster than your best candidate on their best day. When your interview tests the one thing the machine fully absorbed, passing it tells you almost nothing about whether someone can do the job that's left. How the puzzle interview actually broke It broke twice, in opposite directions, and the combination is fatal. First, it broke as a measurement. The job changed. An engineer's day is now spent directing AI tools, reviewing generated code, and making judgment calls about architecture, trade-offs, and what not to build. Recalling the optimal solution to a graph problem correlates with that work about as well as spelling bees correlate with writing novels. You're grading a skill the role no longer exercises. Second, it broke as a filter. Remote interviews plus capable models mean the puzzle can be solved by something other than the person on the call, and detection is a losing arms race I have no interest in fighting. But notice what the cheating panic obscures: even a completely honest leetcode pass has stopped predicting performance. The dishonest pass is just a louder version of the same emptiness. "When your interview tests the one thing the machine fully absorbed, passing it tells you nothing about the job that's left." There's a quieter cost, too. The engineers I most want to hire, the ones with a decade of shipped systems and taste to show for it, increasingly refuse to grind puzzle prep for the privilege of interviewing. The process filters them out before it ever sees them, and keeps the people with the most spare evenings. That is precisely backwards. What the interview needs to measure now Strip the job to what humans still uniquely contribute and the list is short: judgment about what to build, the ability to evaluate work they didn't type, and clear reasoning when the problem doesn't match any pattern in the training data. So test those. Directly. - Can they review? Reviewing code you didn't write used to be maybe a tenth of the job. It's now closer to half. Almost no interview process tests it at all. - Can they direct? Given powerful tools, do they decompose the problem well, give the tools the right context, and notice when the output is confidently wrong? - Can they decide? Trade-offs, sequencing, what to leave out. The expensive mistakes in software were never syntax errors. - Can they go deep? When something breaks two layers below the abstraction, do they have the foundations to descend, or do they only operate at the level the tools hand them? The work sample I actually run When I vet engineers for a client, the technical exercise is boring by design. No puzzles, no tricks. I take a small, realistic problem, the kind of thing that would be a two-day ticket, and I ask the candidate to make real progress on it in an hour, with AI tools not just allowed but expected. Then I watch how they work, because the process is the product. The strong candidates all do a version of the same dance. They interrogate the problem before touching the tools. They give the model context in deliberate, structured pieces. When it produces something plausible, they slow down instead of speeding up, because plausible is exactly when the traps appear. One candidate recently caught a generated migration that would have silently dropped rows on a table with a null foreign key. Nothing in the code looked wrong. He just knew where that class of bug likes to live. That single moment told me more than any whiteboard session in my career. The weak candidates ship the first answer that runs. They aren't stupid and they aren't lazy. They've simply never been asked to be responsible for output at this volume, and it shows immediately. You cannot detect this from a resume, a puzzle score, or a pleasant conversation. You have to watch the work. The part that isn't technical After the exercise we talk about it, and this conversation carries as much weight as the code. Why this approach? What would break at ten times the load? What did the model get wrong, and how did you know? What would you do with another day? Candidates who own their reasoning can defend it, adjust it, and tell me where they were guessing. That honesty about the edge of their own knowledge is the strongest hiring signal I know, and it's completely invisible to an autograder. Companies keep running leetcode because it scales, it feels objective, and changing an interview loop is organizational surgery nobody volunteers for. I understand all three reasons. They're also how you end up with a team optimized for an era that ended. The interview is a specification of what you value; engineers read it and sort themselves accordingly. Specify puzzle recall and watch who shows up. Specify judgment and taste, and watch who shows up instead. I run this vetting as a service: CTO-led screens and AI-era work samples, so every candidate who reaches you has already shown judgment on real work. Here's how it works, or let's talk → ### Do I Need a Fractional CTO? Seven Signs You Do: /blog/seven-signs-you-need-a-fractional-cto/ Patterns I see over and over in founder conversations. Recognize yourself in two or more and it's worth a call. I can usually tell in the first ten minutes of a call whether a founder needs a fractional CTO. The story arrives with a different name each time, but the shape of it repeats. After a few hundred of these conversations, the patterns are hard to miss. The common thread is simple: you need a fractional CTO when technical decisions start outrunning the technical judgment in the room. The seven signs: you're a non-technical founder who got burned, your tech lead is drowning, investors are asking questions you can't answer, you're scaling and technology is the bottleneck, you're confused about AI, you can't afford a full-time CTO yet, and what you need is experience rather than hours. None of them is a diagnosis on its own. But if you read two or more and feel a little seen, we should probably talk. "A fractional CTO is a senior technology executive you hire for a day or two a week: the judgment of a CTO without the full-time salary." 1. You're a non-technical founder who got burned This is the most common one, and it usually costs the most to learn. You had a real idea. You didn't write code, so you hired an agency or stitched together a few freelancers, and you assumed that hiring people meant the work would get handled. It got built. You have no way to tell if what you got is any good. I've talked to a founder who spent well into six figures on a shop that shipped an app with sharp edges everywhere. It fell over the moment a dozen people logged in at once. Then the same shop billed her to fix the bugs it had created. She ended up sitting on a codebase nobody could maintain and a launch that was half a year late. The details vary. Maybe the agency vanished halfway through. Maybe what they delivered sort of works and you can't tell whether it'll survive real traffic. Maybe you brought developers in-house to fix it and now you suspect they're padding estimates, but you can't prove it because you don't speak the language. That last part is what actually eats at people. Not the money. The helplessness. "The money isn't what traumatizes founders. It's not being able to tell whether they're being told the truth." What you need is someone in your corner whose paycheck doesn't grow with the hour count. Someone who can read your code and tell you the plain truth, whether that's "this is solid, keep going" or "this needs to be rebuilt, here's the plan." Someone who speaks both business and engineering and can sit between you and your developers so nothing gets lost in translation. That's the job. Not to type. To be the technical judgment you're missing. 2. Your tech lead is drowning You promoted your best engineer. Maybe you handed them the CTO title because the cap table needed one, or because they were your first technical hire and it felt right. They know the system cold. Nobody writes better code. And they're sinking. They were never trained to manage people, and half the time they don't want to. They can't push back on a timeline they know is fantasy. They freeze when an investor asks about architecture. They're stuck in meetings all day and the code, the one thing they're great at, is getting worse. They're on their way to burning out. None of that is a failure on their part. The skills that make a great engineer are simply different from the ones that make a great technology executive. Some people grow into both. Plenty don't, and there's nothing wrong with not wanting to be a manager. A fractional CTO can take the executive load, the investor conversations, the strategic calls, the architecture decisions that ripple across the whole company, and let your tech lead do the work they're actually good at. Sometimes that split is permanent. Sometimes it's a bridge while they grow into the bigger role, which does happen, it just takes longer than anyone wants. Either way it lifts a weight off someone you were quietly asking to do a job they were never set up for. 3. Investors are asking questions you can't answer There's a board meeting on the calendar. Or a Series A pitch. Or diligence on an acquisition. At some point someone is going to ask about your technology strategy, and "we're building good software" is not an answer. Neither is waving your hands at AI. They want specifics. How does the architecture scale? Where's the roadmap headed and what are the risks along it? How do you know the team is productive? What happens the day your lead developer quits? The red flags they're hunting for aren't obvious from the inside. They're patterns an experienced technologist spots in a few minutes. I've sat on both sides of a technical due-diligence table. I know what the people writing the check are looking for, because I've been the one doing the looking. A fractional CTO can help you get ahead of it: audit the technology, surface the problems before an outside reviewer does, help you tell a credible story, and be in the room when the questions get technical so your team looks like it has an adult supervising the stack. More often than not, this is the exact moment a founder finally reaches out. Fundraising is close, and the gap suddenly feels real. 4. You're scaling and technology is the bottleneck The company is growing. You're hiring, revenue is up, and somehow every feature takes longer than the last one. Each sprint feels heavier than the one before it. That's not a sign you hired wrong. It's a sign the system around your team has stopped keeping up. The code three developers wrote doesn't hold when there are ten hands in it. The architecture that shrugged off a thousand users groans under fifty thousand. The informal way things got decided when everyone shared a room falls apart the minute you're remote and growing. This is normal, and it's also the thing that quietly caps a company. The bottleneck is almost never what founders assume. It isn't that the developers are slow. It's that friction has crept into everything around them: - Technical debt compounding faster than anyone pays it down. - Organizational debt, unclear ownership, decisions with no home. - Communication overhead that grows faster than headcount. - Testing gaps that turn every release into a gamble. Someone who has scaled a team from five to fifty knows where these show up before they show up. A fractional CTO can find what's actually slowing you down and fix it while it's still cheap, instead of after scale has made it expensive. 5. You're confused about AI Everyone tells you to be using AI. Your competitors talk about it. Investors ask for your AI strategy. You know you're supposed to be doing something and you have no idea what. Integrate ChatGPT? Train a custom model? Roll out copilots? Which parts are real and which are theater, and how would you even measure whether any of it helped? I've lived through a few hype cycles. I was around when Java applets were going to replace desktop software. They didn't. I was around when every company "needed" a mobile app. Some did, plenty didn't. AI is real in a way those weren't, but that doesn't mean every AI feature is worth building. I've watched startups pour six months into AI that added nothing, and I've watched a small, boring integration save engineers hours a day. "The difference is never the model. It's whether you knew which problem you were solving before you reached for it." A rough map of where the value tends to sit, and where it tends to evaporate: Usually hype - An "AI-powered" sticker slapped on a feature that already existed. - A custom ML model built for a problem a plain API would solve. - A chatbot that frustrates users more than a good search box would. Usually real - Copilot tools that give each developer back a couple of hours a day. - Claude or ChatGPT APIs doing the text-shaped work nobody wanted to. - Proven tools threaded into workflows people already use every day. Most of my work right now is exactly this: helping companies decide what to try, what to ignore, and how to get developers to adopt something new instead of quietly resisting it. That's what an AI-native transformation is really about, and it's a lot more about the operating model than the model. 6. You can't afford a full-time CTO yet Run the numbers and they don't work. A strong CTO in the US costs real money, salary well into the mid six figures, plus equity, plus benefits. Then add the three to six months of searching and the cost of having nobody in the seat while you look. At your stage, spending that on one executive can feel like setting runway on fire when the same money buys you two more engineers. Here's the part nobody says out loud: at your stage there probably isn't forty hours a week of CTO work to do. You need someone to make the architecture calls that are hard to reverse, help with the key hires, hold the investor conversations, and be there when something breaks. That's more like ten or fifteen hours a week of genuinely senior time. A fractional CTO gives you that judgment at a price that fits the stage you're actually in, a day or two a week of experienced leadership instead of a full-time salary you'd feel every month. And when you're truly ready for a full-time CTO, with enough work and budget to justify one, a good fractional helps you make that hire and hands off cleanly. The goal was never to make you dependent. It's to get you to the point where you don't need me anymore. 7. You need experience, not hours This one runs underneath all the others, so let me say it flat. You need experienced leadership, not more hours. With a fractional CTO you get a very experienced one for less than the cost of an inexperienced one who's also writing your code. You already have people to write code. You have someone to manage tickets. What you don't have is someone who has made these decisions before and paid for the wrong ones. Someone who can look at your situation and say, "I've seen this exact pattern, here's how it usually plays out, here's what I'd do," and be right often enough to matter. Experience compounds in a way hours never do. A CTO who has built five companies makes a hard architecture call in ten minutes that a green one agonizes over for weeks and still gets wrong. Not because they're smarter. Because they've already made that mistake once and know what it costs. That's the thing you're actually buying. Not time. Pattern recognition, and all the expensive lessons that came free with it. So what now? If two or more of these felt uncomfortably familiar, a fractional CTO is probably worth a real conversation. Not definitely, every company is its own thing, but probably. The fastest way to know is to talk it through. In thirty minutes I can usually tell you what makes sense. Sometimes the answer is "hire a fractional CTO." Sometimes it's "you're not ready yet, and here's what to fix first." Sometimes it's "you actually just need to unblock your existing tech lead." I'll tell you the truth either way, even when the truth is that you don't need me. If you recognized yourself up there, let's figure out what you actually need. Here's how I run a fractional CTO engagement and what it costs. Book a call → ### Is the Resume Dead? What to Screen For in the AI Era: /blog/the-resume-is-dead/ AI writes flawless resumes now. Every signal hiring used to lean on is gone, and most interview processes are still screening like it's 2019. I read a resume last month that was perfect. Quantified impact on every line, verbs in the active voice, a career arc that built to exactly the role I was hiring for. It was one of the best resumes I've seen in twenty years of hiring. So were the other forty in the pile. That's the whole problem in two sentences. The resume used to carry signal because writing a good one took effort, and the effort itself told you something. A candidate who could describe their work crisply, quantify their impact, and tailor the story to your role had already demonstrated a kind of competence. Now a model does all of that in eight seconds, for everyone, for free. The polish is still there. The signal is gone. What actually died Let me be precise, because "the resume is dead" gets said a lot and usually means nothing. The document still exists. People still send it. What died is every inference you used to make from it. - Writing quality used to proxy for clarity of thought. Now it proxies for having an internet connection. - Tailoring used to show genuine interest in your company. Now a model tailors two hundred applications an hour. - Keyword fit used to be a rough relevance filter. Now it's the easiest thing in the world to game, and the people gaming it hardest are often the weakest candidates, because they need to. - The cover letter used to be a work sample of communication. It's now the least trustworthy document in the whole process. None of this makes candidates dishonest. Most of them are doing the sensible thing with the tools available, and honestly, an engineer who refuses to use AI for a writing task in 2026 worries me more than one who does. The point isn't that candidates are cheating. The point is that the artifact no longer measures anything. "The polish is still there. The signal is gone." The screening stack built on sand Here's what makes this dangerous rather than just annoying. Most companies' hiring funnels are a stack of filters, and the resume is the bottom layer. An ATS scores it for keywords. A recruiter skims it for pedigree. A hiring manager decides in ninety seconds whether the phone screen happens. By the time a human has a real conversation with the candidate, three decisions have already been made based on a document a machine wrote. When the bottom layer of a stack stops carrying information, everything above it inherits the noise. You're not filtering for engineering ability anymore. You're filtering for prompt quality, and you're doing it with a straight face. I've watched this play out from the inside. A founder I work with ran a search for a senior backend role and got nine hundred applications in a week. The ATS shortlist looked immaculate. The first five phone screens were a massacre: candidates who couldn't explain a single line of the experience their resume described, because they hadn't written the description and, in one memorable case, seemed to be reading their answers off a second screen with a familiar latency. What still carries signal The good news is that real signal didn't disappear. It just moved. It moved out of documents and into anything a candidate has to do live, unrehearsed, with their actual judgment on display. A real conversation about a real decision still works. Not "tell me about a time you showed leadership," which has a rehearsed answer, but "walk me through the worst technical decision you were part of, and what you'd do differently." People who lived the work can go infinitely deep. People who resumed their way in run out of floor in two follow-up questions. The follow-up question is the entire game now. Any first answer can be manufactured. The third layer can't. Work samples still work, with one big change: you have to let the candidate use AI, because forbidding it tests a job that no longer exists. Watch how they direct the tools. Watch what they accept and what they push back on. An engineer who ships the model's first plausible answer is telling you exactly how they'll perform on your codebase. And references, the most unfashionable tool in hiring, quietly became one of the most valuable. A fifteen-minute call with someone who actually worked beside the candidate is the one channel AI hasn't polished, at least for now. What I do instead When I recruit for a client, the resume gets me a name and a rough shape of a career. That's all I let it do. Every candidate who goes further talks to me, a CTO who has hired 75+ engineers for his own teams, in a real technical conversation with real follow-ups. The ones worth a client's time then do a working exercise with AI tools on the table, scored on judgment: what they questioned, what they caught, what they chose not to build. It's slower per candidate than keyword filtering. It's also the only version of screening that still measures the thing you're buying. The resume was a shortcut, and it was a good one for fifty years. It's dead now. The companies that keep screening on it aren't saving time, they're just making their bad hires more efficiently. If you're hiring engineers and the pile of perfect resumes is telling you nothing, that's exactly the problem I solve. Here's how I recruit engineering teams, or just get in touch → ### Why Technical Skill No Longer Decides Who to Hire: /blog/technical-skill-is-not-the-point/ For thirty years we hired engineers for what they could type. AI ended that. What actually predicts performance now is harder to test, and worth more. I've spent most of my career evaluating engineers on technical skill. How fluently they wrote code, how many frameworks they'd internalized, how quickly they could produce a working solution. For thirty years that was rational, because technical skill was the scarce input that everything else depended on. It isn't scarce anymore. Competent code is now the cheapest ingredient in software. And when an input stops being scarce, hiring for it stops making sense, no matter how deep the habit runs. This essay is the why; the how, what to screen for instead, is in its companion piece on hiring for judgment. What "technical skill" was actually buying Be precise about what changed. When we tested for syntax fluency, framework recall, and speed, we were buying production capacity: the ability to convert a decision into working code. That conversion used to be the bottleneck of the whole industry. Entire org charts, interview loops, and salary bands were built around finding the people who did it fastest. AI tools now handle most of that conversion. Not perfectly, and not unsupervised, but well enough that the marginal value of one more fast typist has collapsed. What did not collapse, what actually exploded in value, is everything wrapped around the conversion: deciding what to build, noticing what's wrong, imagining the option nobody listed. The parts of the job that were always quietly the hard parts are now openly the whole job. "The parts of the job that were always quietly the hard parts are now openly the whole job." This is not "engineers don't need to know how computers work." Foundations matter more than ever, because someone has to know when the confident output is nonsense, and that requires a real model of the system underneath. What's devalued is skill as recall and speed. What's revalued is skill as understanding. The two capacities that predict performance now When I vet engineers for clients, two things separate the people who multiply with AI from the people who merely operate it. Neither appears on a resume, and the standard interview loop tests neither. Critical thinking, by which I mean something specific: the ability to distinguish plausible from correct. A model's failure mode is confident, fluent, well-structured wrongness. The engineer's job has become an ongoing act of epistemic hygiene: what is this output assuming, where would it break, does this answer actually address my situation or just a situation shaped like mine? People with this capacity treat every generated artifact as a claim to verify. People without it treat fluency as evidence, and fluency is exactly the thing the machine has infinite amounts of. Creative thinking, which in engineering doesn't mean artistic flair. It means reframing. The model is a phenomenal interpolator: ask it a well-posed question and it gives the consensus answer. It's much weaker at noticing the question is wrong. The engineers earning their salaries now are the ones who look at a ticket and say "we don't need this queue at all if we change the contract upstream," or who connect a pattern from a different domain because they've been curious about more than one thing in their lives. Every reframe like that is worth more than a week of generated code, because it changes what needs to be generated at all. Why your interview can't see any of this The standard loop was engineered, carefully, over decades, to measure production capacity. Algorithm rounds measure recall under pressure. Take-homes measure unsupervised output. System design comes closest to testing thinking, but it usually rewards the memorized reference architecture rather than live reasoning. So companies keep running loops that grade the abundant thing and stay blind to the scarce thing. The candidates who ace them are often genuinely skilled in exactly the dimension that matters least. This is how you assemble a team that looks stellar on paper and drowns in its own confident, machine-generated code. Testing the scarce thing is possible, it's just more work. Give candidates a real problem with AI tools on the table and watch where their attention goes. Ask them to review a plausible, subtly broken piece of generated work and see what they catch, and just as revealing, what they praise. Ask for the second and third way they'd solve the problem, and watch whether the alternatives are real or decorative. Ask what they'd refuse to build and why. An hour of this tells you more than a full day of the old loop. The uncomfortable conclusion for hiring If technical skill is abundant and judgment is scarce, then the whole apparatus of technical hiring, keyword filters, framework checklists, years-of-experience gates, puzzle scores, is optimized for the wrong scarcity. Not slightly miscalibrated. Pointed at the wrong target. It also means the evaluator matters more than it used to. Grading syntax was nearly mechanical; any competent engineer could do it. Grading judgment takes judgment. You need someone who has made these calls at production scale, been wrong, paid for it, and calibrated. That's why I vet every candidate myself when I recruit for a client: after twenty years of building my own teams, plausible-but-wrong sets off an alarm in my head that no rubric replicates. The engineers you want are still out there, and ironically they're easier to spot than ever, because the contrast between operators and thinkers has never been sharper. You just need a process, and a person, actually looking for the right thing. That's the service. Get in touch → ### How Big Should an Engineering Team Be? Four Beats Twelve: /blog/four-engineers-not-twelve/ AI changed the math of team size. Why the right hire count is smaller than your plan says, and what that does to roles, budgets, and recruiting. A founder showed me his post-raise hiring plan a few months ago. Twelve engineers in twelve months: two squads, a platform team, an engineering manager layer. It was a good plan. I'd have approved it myself in 2021. I told him to hire four people and keep the difference. Not because the roadmap shrank. Because the math under every headcount plan quietly changed, and most hiring plans are still running the old constants. The math that changed Team size was always a trade between production and coordination. Every engineer adds output; every engineer also adds communication paths, and those grow faster than headcount. The old equilibrium landed where it did because one person could only produce so much, so you ate the coordination tax to get the production. AI tools moved one side of that trade and left the other alone. A strong engineer directing agents now carries the throughput that used to justify a pod of three or four. But the coordination cost of a twelve-person org didn't drop a cent. Standups, handoffs, interface negotiations, the pull-request queue, the roadmap meeting about the roadmap meeting: all still priced in human hours. When production per person triples and coordination per person doesn't move, the optimal team gets smaller. That's not a philosophy. It's arithmetic. "When production per person triples and coordination per person doesn't move, the optimal team gets smaller." And a smaller team of stronger people compounds in ways the spreadsheet undersells. Fewer handoffs means fewer places for context to die. Whole-system ownership means the person debugging the API also wrote the schema and remembers why. The velocity difference between a tight four and a coordinated twelve isn't twenty percent. On the teams I've run, it's routinely two to three times, in the small team's favor. What the twelve-person plan actually buys you I want to be fair to the big plan, because it isn't stupid, it's nostalgic. Headcount used to be the honest measure of seriousness. Investors read team size as traction. Managers read span of control as career progress. And redundancy was real risk management when every engineer held irreplaceable context in their head. But look at what the extra eight hires cost in the AI era, beyond salary. You now need a management layer, so you hire managers, so you need alignment rituals, so your best builders spend mornings in meetings. Onboarding twelve people through a codebase moving at AI speed is its own project. And averaged-down talent doesn't average: on a modern team, output that has to be rewritten is negative work, and the engineer who ships confident, unreviewed, agent-generated code is a cost center with a good attitude. The failure mode I keep getting called into is exactly this: a company that hired to its 2021 org chart, is paying 2026 salaries for it, and can't understand why a team of fourteen ships less than the founding three did. The answer is almost never effort. It's structure built for a constraint that no longer exists. Role design changes too Smaller teams don't just mean fewer of the same roles. The roles themselves reshape. - The narrow specialist gets rarer. When agents cover breadth, a person who only does one slice of the stack creates handoffs a four-person team can't afford. You hire product-minded generalists with a deep spike, and rent true specialization when you hit it. - Every senior hire is a force multiplier or a mistake. There's no crowd to hide in. One low-judgment engineer on a four-person team is a quarter of your company's output. - The first leadership hire comes later, and matters more. Four strong owners barely need a manager. What they eventually need is a leader who shapes problems and sets appetite, which is a different job description than meeting-runner. - Junior hiring becomes deliberate instead of volumetric. You still hire juniors, but as an investment you actively develop, with real mentorship, not as cheap capacity, because capacity is no longer what's scarce. What this means for how you recruit Here's the consequence nobody puts in the pitch deck: hiring four instead of twelve makes recruiting harder, not easier. Every single pick carries triple the weight. A tolerable-miss process, where one mediocre hire in ten washes out in the noise, becomes intolerable when the miss is 25% of the team. The bar goes up exactly as the volume goes down. That's why I tell founders to spend on vetting depth what they used to spend on pipeline width. Fewer searches, run properly: org design first, so you know which four roles actually compose into a team; a real technical screen run by someone who has built these teams; a work sample that shows judgment with AI tools on real work. The savings from the eight hires you didn't make will fund the most rigorous search you've ever run, several times over. This is the work I do: I design the small, sharp team and then I go find it, vetting every candidate personally. Here's how the team build works, or let's talk about your hiring plan → ### How to Use Claude Plugins for Knowledge Work: Staff Claude Like a Department: /blog/claude-knowledge-work-plugins/ Almost everyone runs Claude as one freelancer with amnesia: open a chat, paste a task, close the tab, re-explain everything next time. Anthropic quietly open-sourced the repo that turns it into a department instead. Here is what almost everyone does with Claude. They open a chat. They paste a task. They get an answer. They close the tab. Next time they start from zero and re-explain the whole situation again, what the company does, who the customer is, what good output even looks like. That isn't an assistant. That's one freelancer with amnesia. Useful, occasionally impressive, and permanently small. There is a completely different way to run this, and the gap between the two has nothing to do with the model. Anthropic quietly open-sourced a repo that turns Claude into a set of specialized office roles, a sales rep, a marketer, a financial analyst, a legal reviewer, a data analyst, each one pre-loaded with the workflows, the domain knowledge, and the tool connections that role actually needs. You stop prompting from scratch. You start hiring a department. What the repo actually is It's called knowledge-work-plugins, and it's a free, open-source marketplace of role-based plugins for Claude. Each plugin turns Claude into one narrow specialist, and inside every one there are three things doing the work. - Skills, the domain knowledge and best practices for that role. Claude pulls them automatically when they're relevant. You don't invoke them, you don't even think about them, they're just the part where the worker already knows how the job is done. - Commands, ready-made workflows you trigger with a slash, like /sales:call-prep or /data:write-query. These are the repeatable jobs that role does every day, packaged so you don't rebuild them each time. - Connections, the tools that role plugs into. The sales plugin reaches for your CRM. The finance plugin reaches for your data warehouse. The marketing plugin reaches for your analytics and your design tool. The detail that should make you sit up: this is the same foundation Anthropic built Claude for Legal and Claude for Financial Services on top of. You're getting the base layer those paid products are made from, in the open, for free. The expensive version is the same skills with a support contract wrapped around them. "You're not prompting from scratch anymore. You're hiring a department." The roles you can hire The repo ships with a full org chart, and each role is one command to install. You don't take all of them. You build the team your work actually needs. - Productivity, tasks, calendars, daily routine, personal context. Plugs into Slack, Notion, Asana, Linear, Jira, ClickUp, Microsoft 365. - Sales, account research, call prep, pipeline, cold outreach, competitive analysis. Plugs into HubSpot, Close, Clay, ZoomInfo, Fireflies. - Marketing, content, campaigns, brand voice, competitor sweeps, channel reporting, SEO audits. Plugs into Canva, Figma, HubSpot, Klaviyo, Ahrefs, SimilarWeb. - Customer support, ticket triage, reply templates, escalations, turning solved tickets into help-center articles. Plugs into Intercom, HubSpot, Guru. - Product management, specs, roadmaps, user research synthesis, stakeholder updates. Plugs into Linear, Figma, Amplitude, Pendo. - Finance, journal entries, reconciliations, statements, variance analysis, month-end close, audit support. Plugs into Snowflake, Databricks, BigQuery. - Legal, contract review, NDA triage, risk assessment, templated responses. Plugs into Box, Egnyte, Microsoft 365. - Data, queries, SQL, stats, dashboards, sanity-checking results before you publish them. Plugs into Snowflake, Databricks, BigQuery, Hex. - Enterprise search, one search across your email, chat, docs, and internal wikis. Step 1: Get the desktop app These plugins are built for Cowork, Anthropic's agentic desktop app, though they also run in Claude Code. Download Claude Desktop from claude.com/download and open the Cowork tab. This is the moment Claude stops being a chat window and starts touching real files, real tools, and real workflows. Everything below assumes you're working in there. Step 2: Add the marketplace Cowork has a terminal. Point Claude at the full catalog of roles with one command, which you only ever run once: claude plugin marketplace add anthropics/knowledge-work-plugins Step 3: Hire your first worker Install the role you need most. Say it's a sales rep: claude plugin install sales@knowledge-work-plugins Swap sales for any role, marketing, finance, legal, data, product-management, customer-support, productivity. The plugin activates the moment it's installed. Start with one. Get a feel for how it changes the work before you build out the whole floor. Step 4: Put it to work standalone Here's the part people miss: every plugin works on day one without connecting a single outside tool. You just hand it the raw material. Trigger a workflow with a slash command, /sales:call-prep takes a company name and hands back a full pre-call brief, /data:write-query takes a plain-English question and hands back the SQL, /marketing:seo-audit takes a page and hands back keyword gaps and fixes. Paste your notes, upload a CSV, describe the situation. The skills behind the plugin already know how that role does the job, so you skip the entire part where you explain what good looks like. That's the part that used to eat the whole session. "Standalone is the intern. Connected is the senior hire." Step 5: Connect its tools This is where the worker goes from competent to genuinely useful. Each plugin has tool connections built in. Connect the sales plugin to your CRM and it stops asking you to paste pipeline data and starts pulling it. Connect finance to your data warehouse and it reconciles against the real numbers, not the ones you typed in. Connect marketing to your analytics and the reports build themselves. In Cowork, open Connectors and authorize the tools that role uses. The standalone version is the intern who needs everything handed to them; the connected version is the senior hire who already has the logins. Step 6: Build the rest of the team Now repeat Step 3 for every role your work actually requires. And here's the part that turns a clever trick into an operating model: once they're installed, they work together in the same session. Your data worker pulls the numbers, your finance worker reconciles them, your marketing worker turns the result into a report, in one place, in one pass. One operator, a full cross-functional team, no payroll, no scheduling, no handoff lost in a Slack thread. This is the same instinct I keep coming back to, designing each agent like a narrow employee with a clear job, rather than expecting one genius generalist to do everything. Step 7: Make them yours The default plugins are a strong starting point and a generic one. The real edge is customizing them for how you actually work. The repo ships a meta-tool, the cowork-plugin-management plugin, built for exactly this: tell it your tools, your terminology, your process, and it reshapes a plugin to fit. And because plugins are just markdown files, you can edit them directly, fork the repo, and keep your own private versions. This is the same move as writing the context once so the system prompts itself, the difference between Claude that knows how a generic sales rep works and Claude that knows how your company sells. "The model didn't change. The setup did. And the setup is exactly what almost nobody bothers to do." What you have after seven steps Before this, Claude is a chatbot you ask questions, one at a time, starting over every session. After this, Claude is a building full of specialists. A sales rep who preps every call. A marketer who runs the campaign. A data analyst who writes the queries. A finance lead who closes the month. All pulling from your real tools, all working in one place, all running off a free open-source repo. Same subscription. Completely different operation. I've written before that when building gets cheap the scarce skill moves to shaping the problem, and this is the same lesson wearing office clothes. The capability was sitting there the whole time. What separates the people getting leverage from the people still pasting tasks into a chat box isn't access to a better model, it's that one group did the setup and the other didn't. Most people will read all seven steps and install nothing. If that's not going to be you, the next move is simple, run the first command. And if you want help turning this into how your company actually operates rather than a weekend experiment, that's the work I do. Let's talk → ### How Engineering Leaders Create Their Own Luck: 9 Habits: /blog/create-your-own-luck/ The luckiest people I've worked with weren't lucky. They built a surface area for luck to land on, then put in the work that made it stick. Someone sent me one of those lists that does the rounds on every feed: 9 ways to create more luck for yourself. I almost scrolled past it. Then I read it twice and admitted it was the most honest career advice I'd seen all month. I've spent 25 years around people who got lucky — the engineer who happened to be in the room, the founder who happened to meet the right investor, the operator who landed the role that changed everything for them. From the outside it looks like a coin flip. Get close enough and almost none of it was luck. It was a surface they'd been building for years, and the break finally had somewhere to land. That's what the word hides. Luck is real, but it isn't random — it's manufactured. You don't get to decide whether the break shows up. You get to decide how big a target you are when it does. Here's how the people I coach actually do it. 1. Work your ass off There's no version of this that skips the work. Every lucky person I know was also the person doing the unglamorous reps long before anyone was watching. Effort doesn't guarantee the break, but it's what makes you ready to catch it. The opportunity that changes your career almost always arrives disguised as more work than everyone else wanted to do. 2. Add value Stop asking what a role gives you and start asking what you leave behind. The people who get pulled upward are the ones who made the team, the product, or the customer measurably better, and did it visibly enough that someone noticed. Value is the only currency that compounds across jobs. Your title resets when you switch companies. A reputation for making things better follows you. "Luck is real, but it's not random. It's manufactured. You can't control whether the break comes, only how big a target you are when it does." 3. Get good at sales Engineers flinch at this one, and they shouldn't. Sales isn't manipulation. It's the work of getting someone else to see what you already see. You're doing it constantly whether you admit it or not — pitching an idea in planning, pitching yourself in an interview, walking a skeptical board through a roadmap. I've watched genuinely great work die in silence because nobody in the room could say out loud why it mattered. Learning to sell is just giving your good work a voice. 4. Listen more, talk less This is the cheat code, and almost nobody uses it. The fastest way to become valuable is to understand the problem better than everyone else in the room, and you don't do that by talking. You do it by listening until you hear the thing under the thing, the real constraint, the unspoken fear, the leaky faucet nobody's named yet. Most people are waiting for their turn to speak. The ones who actually listen end up holding information nobody else has. 5. Find a mentor and learn You can buy back years by borrowing someone else's. A good mentor has already paid the tuition on mistakes you haven't made yet, and a single honest conversation can save you a season of flailing. The trick is to be worth mentoring: come with specific questions, do the reading first, and actually act on the advice. Mentors invest in people who move. I've spent a lot of my career on the other side of this, and the ones who grew fastest were always the ones who made it easy to help them. 6. Have faith Not in the cosmic sense, necessarily, though whatever gets you there. I mean the working belief that effort compounds even when you can't see the curve yet. Most meaningful careers go through a long flat stretch where nothing seems to be landing and the temptation to quit is loudest right before it breaks. Faith is what carries you across the gap between the work and the result, when the only evidence you have is your own conviction. 7. Stop listening to the wrong people Your inputs become your ceiling. If the loudest voices around you are cynical, risk-averse, or quietly invested in you staying where you are, their fear becomes your fear. Audit your inputs the way you'd audit a budget. The people who told me my biggest bets were reckless were almost never the people who'd taken big bets themselves. Take counsel from people in the arena, not the ones narrating it from the stands. "Your inputs become your ceiling. Take counsel from people in the arena, not the ones narrating from the stands." 8. Give without expecting anything This is the one that looks naive and turns out to be the most strategic thing on the list. When you help people with no scoreboard running, you build a web of goodwill that you can't predict and can't engineer. Years later a door opens, and you have no idea it traces back to an hour you gave someone who never forgot it. Generosity is the highest-yield, longest-dated investment in a career. The catch is that it only works if you genuinely don't keep score. 9. Always do the right thing, no matter what Everything above compounds slowly. Your integrity is the one thing that can vanish in a single decision. The right call is often the expensive one in the moment, the harder conversation, the deal you walk away from, the credit you give away. But your reputation is the asset that opens every other door, and it's built one unglamorous right choice at a time. People remember how you behaved when it cost you something. What the list is really about Read the nine again and notice that none of them are actually about luck. They're about becoming the kind of person luck keeps happening to. You do the reps so you're ready, you add value so you're worth a bet, you listen so you understand the problem better than anyone, you give so the network has roots, you stay honest so the door stays open. Do that for a decade and people will call you lucky. Let them. The alternative is explaining ten years of quiet work nobody was around to see. None of this is fast, which is exactly why most people skip it. They want the break without building the thing it lands on. So build the thing. The luck will find it. If you're working on this, on becoming the leader your next break is waiting on, that's the kind of thing I coach people through. Let's talk → ### What Is Go Fever? How Groupthink Ships the Wrong Product: /blog/go-fever-and-groupthink/ Go Fever and groupthink don't just ship broken products. They ship the wrong ones: the feature the room fell in love with and the user never asked for, defended past every signal that it didn't fit. The launch went out on time. It demoed beautifully. Six weeks later, almost nobody was using it. The strangest part wasn't the silence in the metrics, it was the silence that came before it: half the room had a quiet feeling the thing wasn't what users actually wanted, and not one of us said so out loud. We didn't ship something broken. We shipped something nobody asked for, exactly on schedule. That failure has a name, and the name isn't bad luck. Go Fever is the collective excitement to hit a goal that quietly overrides every signal telling you to stop. It's a phrase engineers first reached for in the wreckage of an aerospace program, then watched repeat itself for decades. Its partner is groupthink, the term Irving Janis gave in 1972 to a group trading accuracy for harmony, swallowing the doubt that would break the consensus. Go Fever is the momentum to ship. Groupthink is the silence that lets a team mistake its own enthusiasm for the user's. Inside a product organization, the two together rarely produce an outage. They produce something quieter and far more expensive: the feature the team fell in love with, the redesign everyone knew users wanted, the roadmap bet defended long past the evidence. All of it shipped on time, to a shrug. The crater you can't see Software's Go Fever leaves no mark you can point to. No fireball, no inquiry, no front page. There's just adoption that never crosses the line you drew, retention that dips and gets blamed on seasonality, a support queue filling with confusion that you decide is an onboarding problem rather than a fit problem. None of it lands on launch day, so none of it gets traced back to the afternoon the team stopped asking whether the thing fit and started asking only when it would ship. Aerospace was forced to reckon with Go Fever because its failures were undeniable. Product teams get to keep it, because their failures are deniable: diffuse, delayed, and easy to re-attribute. So the team reorganizes, picks a fresh bet, and runs the same play with new faces. The disease that forces a reckoning in every industry where the stakes are visible gets to quietly become a culture in ours. "A product launch with Go Fever doesn't explode. It under-performs, and then gets explained away." The tell isn't the data. It's the silence. You won't catch this in a dashboard. You catch it in the room. Janis's symptoms map almost perfectly onto a launch review. There's the illusion of unanimity, where nobody objects so everybody assumes everybody agrees. There are the self-appointed mindguards, the lead who quietly handles the skeptical designer before the meeting so the room stays smooth. There's direct pressure, usually delivered as a joke: you're not seriously going to be the one to push the date? And there's self-censorship: the product manager who has seen the research, suspects it doesn't fit, and decides the doubt isn't worth the political cost of saying so. None of it looks like dysfunction in the moment. It looks like a team getting along. Strip away the politeness and the signs are specific: - The date was set before anyone validated that users want the thing. - "Users will get it once they're in" is doing an enormous amount of work. - The only research that survives the meeting is the research that agrees with the plan. - The loudest case for shipping is excitement, not evidence. - Someone says "we're too far in to change course now," and it ends the conversation. - The real doubts are frequent and specific, and only ever spoken in DMs. Normalization of deviance, product edition Underneath Go Fever sits a deeper mechanic with the least catchy and most useful name of the three: normalization of deviance, coined by the sociologist Diane Vaughan. A warning sign appears. Nothing visibly breaks. So the warning sign gets quietly reclassified from problem to normal, and the next one has a lower bar to clear. Each acceptance moves the line a little further. Product teams run this exact loop on fit signals. The usability test where three of five users got lost becomes "small sample." The beta cohort that didn't come back becomes "wrong audience." The activation metric that looked ugly becomes the metric you quietly stop putting on the dashboard. No single dismissal is irrational, and none of them kills the launch on its own. Together, they are the runway Go Fever needs. By the time you reach ship day, you've spent weeks rehearsing how to ignore the exact signals that were trying to tell you the thing doesn't fit. "Go Fever doesn't start at launch. It starts the first time you explain away a user who didn't get it." Make "this doesn't fit" cheap to say This is a culture problem, which means you can't fix it with a poster that says speak up. Smart, committed, courageous people stayed silent in every case study you've ever read, because in the moment, silence was cheaper than speaking. The only durable fix is to change the price. You make dissent structural, so that saying this doesn't fit the user becomes an ordinary move anyone can make without spending a reserve of personal courage they may not have on a given day. A few things that actually work: - Give dissent a seat. Name a person whose explicit job in the review is to argue that users don't want this. Doubt that is assigned isn't disloyalty; it's a role, and a role can't be punished for doing its job. - Run a pre-mortem. Before you commit, have the team write the story of the launch that flopped and explain why users shrugged. It launders private unease into shared evidence, on the record, before the momentum starts. - Separate the decider from the champion. The person who can say stop should not be the same person who is in love with the idea and whose name is on the date. - Tie the launch to a fit signal, not a date. Decide in advance, while everyone is calm, what evidence would prove users want this and what result would mean you pull it. A date defended past the evidence is just Go Fever with a calendar. - Protect the user who didn't get it. The outlier in the test is the cheapest warning you will ever receive. Make it expensive to wave away and cheap to take seriously, not the other way around. Notice what every one of these does. It turns "no, this isn't right for the user" from an act of individual bravery into a normal step in how the team decides. That is the whole game, and it's the opposite of becoming the Department of No. The goal isn't a team that says no more often. It's a team where the truth about the user is cheap to say out loud, in the room, while there's still time to act on it. A delay is better than a launch nobody wanted Safety cultures have a line they repeat until it's reflex: a delay is better than a catastrophe. The product translation is less dramatic and just as true: a delay is better than a quarter spent shipping the wrong thing beautifully. The reason it's so hard to act on is a trick of perspective. The wasted quarter is abstract and lives in the future, while the delay is concrete and sits right in front of you, in a room that is already excited and has already promised the board a date. Your culture isn't the values on the wall. It's what happens in the thirty seconds after someone says I don't think users want this. If that sentence is cheap to say and gets taken seriously, Go Fever never builds up enough speed to matter. If it's expensive, no amount of process will save you, because the process will be run by people who have already learned to stay quiet. Leadership's real job here was never to have the best instinct in the room. It's to make sure the room can hear its own doubts before the user has to deliver them for free. The strongest product cultures I've worked inside weren't the ones with the best ideas. They were the ones where it was safe, and cheap, and slightly boring to say the idea didn't fit, before it shipped. If you're trying to build a culture like that and you can feel how much harder it is than buying a tool, let's talk → ### How to Roll Out Claude Code Across an Engineering Team: /blog/build-a-system-that-prompts-itself/ You're not supposed to prompt Claude. You're supposed to build a system that prompts itself. The leverage was never in the wording. It's in the wiring. If you have used Claude for more than a month and never left the chat window, you have been using one agent. When you could be running a team of them. There is a difference between how most people use Claude and how the teams who build Claude Code actually run it, and it is not a difference of prompt-writing skill. The people getting the most out of these tools stopped trying to write the perfect message a long time ago. They build the thing that sends the messages for them. The mindset shift takes one sentence to state and a while to live by: you are not supposed to prompt Claude. You are supposed to build a system that prompts itself. "The leverage was never in the wording. It's in the wiring." The chat window is the manual setting A single chat window is the hand-cranked version of this work. You type, it answers, you read, you correct, you paste, you type again. You are the loop. Every turn is fed by hand, and every piece of context the model needs has to be carried in by you, in the moment, from memory. That is fine for exploration. It is a terrible way to run anything you do more than once. The chat window makes you the bottleneck and then hides that fact behind how good the answers feel. One sharp agent in a text box still feels clever right up until you realize you have been doing all the lifting around it. The bloat that cripples you before you type a word Here is the part nobody mentions when they teach prompting: most of your context is spent before you have written a single useful word. Stale files, leftover history, half the repository dragged into the window for no reason: that clutter crowds out the model's attention and degrades every answer that follows. A system treats context as an input to be curated, not an accident of whatever was lying around. It decides what the agent should see for this task and nothing more. When you are designing the environment instead of typing into it, trimming context stops being a chore and becomes part of the architecture. What you get back is not a nicer-sounding reply but a cheaper and more repeatable one. Stop prompting. Start wiring. The automation most people never discover is the whole point. Routines that fire on a schedule, daily task pipelines that run without anyone touching the keyboard, a standing /goal that the system pursues on its own while it breaks the work down, does it, and surfaces only what needs a human. This is the move that compounds. The first time you watch a pipeline open a branch, make the change, and leave you a result to review, the chat window stops looking like the product and starts looking like a debugger. The real work was the workflow. The prompt was just one deterministic step inside it, codified once and reused forever instead of retyped every morning. - Repetitive work becomes a routine that runs on a schedule, not a task you remember to do - Multi-step work becomes a pipeline with defined stages, not a marathon chat session - A standing goal pursues itself between sessions and escalates only the exceptions - Your judgment gets reserved for the decisions that actually need it Run a team, not an agent Once you accept that you are building a system, the natural next question is: why only one worker? The unlock the experienced teams reach for first is Git worktrees, running several Claude Code instances at once, each in its own checkout, so they never collide on the same files or stomp on each other's branch. This is quietly a distributed-systems problem wearing an AI costume. The moment you have more than one agent against the same repository, you need boundaries, and CLAUDE.md stops being a notes file and starts functioning as a contract. It is the shared agreement every parallel agent reads before it touches anything: how this codebase works, what each agent owns, and what it should leave alone. Once you treat that context as a contract the agents have to honor, parallel work stops fighting itself. "A single agent feels clever. A worktree full of them turns into an org chart you actually run." The honest caveat It would be dishonest to pretend the system comes first. It does not. The messy, manual prompting is where the discovery happens, where you find the pattern that is actually worth automating. The system is what you build after you have found a win, to make that win repeatable. It codifies the work. It does not replace the part where you figure out what the work even is. So the sequence matters. Prompt by hand until something clearly works. Then, the moment you catch yourself doing it a second or third time, stop prompting it and wire it. The skill is not avoiding the chat window. It is knowing when you have learned enough in it to graduate the task out of it. What actually compounds Models learn from the open internet. What they cannot easily learn is why your team rejected a deal, escalated a ticket, or changed a policy two years ago. That institutional memory, encoded into how your systems prompt themselves, is the asset that compounds. The teams that capture it pull away from the teams that just keep buying a slightly better model. Prompting is labor. You do it, you get an output, and the value stops the moment you stop typing. A system keeps producing while you are asleep, and every workflow you encode makes the next one cheaper to build. So one of these scales only as far as your own effort reaches, and the other keeps going without you. That is the whole reason to leave the chat window. If you are trying to make this shift, from prompting individual agents to designing the environment a team of them runs in, that is exactly the work I help organizations build. Let's talk → ### Why AI Pilots Fail: Purgatory Is an Orchestration Problem: /blog/escaping-ai-pilot-purgatory/ AI got funded, piloted, and then stalled. Escaping pilot purgatory has little to do with the model or the tooling. It comes down to orchestration, and that architecture decision belongs in the boardroom. Most organizations I speak to are stuck in the same place. AI has been funded, pilots have launched, use cases identified. Then it stalls. AI gets pushed into the CIO organization, progress fragments across teams, and shadow AI starts showing up everywhere. Six months later the board asks an uncomfortable question: why aren't we seeing real outcomes? This is what I call AI pilot purgatory: AI pilots fail because the technology works but the organization never learns to operationalize it. That gap isn't about the model or the tooling. It's about orchestration. Most companies are past the point where tinkering is enough, but their strategy hasn't caught up yet. The ones pulling ahead have stopped treating AI as a collection of tools and started running it as an operating system, which is what I mean by AI orchestration. This is not a technical paper. It is a business call to arms. What orchestration actually is, and what it is not The term is being used to sell everything from chatbots to enterprise transformation platforms. The vendor noise is deafening, so here is the version I use. AI orchestration is the coordination of multiple AI models, agents, data sources, tools, and humans to complete real, end-to-end business work, reliably, at scale, and under governance. What orchestration is not - A chatbot on your website - An AI that summarizes documents - A single model accessed through an API - "We use a consumer AI tool internally" What orchestration is - Coordinated AI agents completing multi-step workflows autonomously - Decisions being made, actions being taken, exceptions being escalated - Memory, context, and learning shared intelligently across processes - Humans engaged at the right moments, not all moments The standalone chatbot has run its course. Siloed AI tools force the employee to act as the middleware: copying data out of an ERP, pasting it into a prompt, taking the output, dropping it into an email. Nobody should mistake that for transformation. It's just a faster hamster wheel. "Siloed AI tools force the employee to be the middleware. That isn't transformation. It's a faster hamster wheel." This is not a future state. It is how competitors are already operating. The global orchestration market sits at roughly fourteen billion dollars today and is projected to reach sixty billion within the decade. Multi-agent deployments have grown more than 300% in recent months as enterprises moved from pilots to production. - $14B: Orchestration market today - $60B: Projected within the decade - 300%+: Growth in multi-agent deployments This is a management problem, not a technology problem When you decide how to break a business problem into AI agents, how many of them, in what sequence, with what decision authority, you are making a strategy decision. You are defining how your organization processes information, makes choices, and acts. That isn't a technology question, it's a leadership one. In most enterprises today, AI sits in IT or Data. Use cases are scoped like projects. Success is measured in pilots deployed. Governance is bolted on afterward. Teams build in isolation. The result: twenty different tools solving similar problems, no shared memory, no consistent cost model, no clear ownership of outcomes, and growing shadow AI risk. Once AI starts making decisions, executing workflows, shaping customer outcomes, and driving cost structures, it can no longer sit inside IT. It's now a financial issue, a governance and risk issue, a question of how you design your workforce and where you compete. I have yet to see it work the other way: organizations that let AI strategy emerge organically from the bottom up, crowdsourcing tools and use cases across teams, almost never achieve transformation. Top-down strategic architecture, with clear governance from day one, is what separates scaled transformation from expensive experimentation. "The architecture decision must be made at the C-suite level. Not delegated to IT." Build vs. buy vs. co-create: ask the right question Every leadership team eventually faces this question, and almost every one of them underestimates it. Most ask: should we build or buy? That is the wrong question. The right question is: where do we want to create competitive advantage, and where will the market commoditize? The practical answer for most enterprises: - Buy what will commoditize: base models, agent frameworks, connectors, generic workflow engines - Co-create where speed and domain fit matter most - Build where orchestration logic becomes genuine advantage: your decision logic, cost routing, institutional memory, human-AI collaboration design Enterprise deployments typically run between half a million and two million dollars including platform, integration, and team capability. But the visible costs are not the dangerous ones. The dangerous costs are lock-in risk, talent dependency, governance debt, and opportunity cost. Every sprint your team spends building orchestration infrastructure is a sprint not spent on your actual business problem. So the question I'd actually put to a leadership team is what your competitive advantage is. If it's orchestrating AI agents, build. If it's healthcare, financial services, logistics, or any other domain, buy the orchestration layer and build the intelligence on top of it. Don't pay VP rates for entry-level work This is happening in boardrooms right now. A CFO opens a cloud invoice. The number is wrong, not wrong as in a mistake, but wrong as in nothing we planned for. Welcome to the token economy. A token is roughly three-quarters of a word. Every time a model processes your data it consumes tokens, and you get billed for them. Most providers have now moved to consumption-based pricing, so the meter is always running. And token cost is only twenty to forty percent of what a deployment actually costs. The rest is integration overhead, human review time, compliance infrastructure, and retry waste. Companies that budgeted fifty thousand dollars for an agent deployment regularly find the real number closer to three hundred and eighty thousand once the full stack is visible. You wouldn't hand basic, repetitive work to your vice presidents. Their cost per hour is too high, and it's the wrong use of what they're good at. AI works the same way. Route every task to the most powerful, most expensive frontier model and costs explode until the economics fall apart. Good orchestration routes tasks by what they actually need: - High-volume, low-complexity tasks → efficient, specialized local models, at fractional pennies per task - Context-heavy analysis → intermediate models with optimized retrieval, at a moderate, predictable cost - Strategic reasoning → frontier models deployed with precision, a premium investment, not a default The CFO question to ask today: do we have a fully loaded cost model for every AI workflow we are running? Not just the token invoice, the full stack. If the answer is no, you are flying blind on one of your fastest-growing cost lines. Human in the loop vs. human on the loop This distinction, in versus on, matters more than most people realize, and most organizations are getting it dangerously wrong. Human in the loop means a person is part of every transaction. The AI stops and waits. Scale is limited by human capacity. Add enough checkpoints and your middle management drowns in what I call guardrail fatigue, the ROI of your AI deployment plummets to zero, and you have simply built more expensive approval queues. Human on the loop means AI executes within defined boundaries. Humans monitor outcomes. Intervention happens only on exceptions. This is the architecture that enables real scale: route ninety percent of standard transactions autonomously, with automated guardrails flagging only the anomalies, high-value decisions, or borderline compliance risks to human executives. Reserve genuine human judgment for the cases where it cannot be delegated. "Are we using human oversight to teach our AI to improve, or to compensate for AI we don't trust?" If it's the latter, what you've built isn't an AI capability. It's an expensive human review function with an AI front end. Measure what actually matters Traditional productivity measurement was built for a world where humans did discrete tasks in measurable time. Orchestration breaks that model entirely. A finance team might process ten times the volume of contracts at the same headcount. A customer service operation might resolve eighty percent of cases without human involvement. Yet most operational dashboards would flag fifteen AI iterations as "inefficiency." The metrics that matter now: - Outcome velocity: how fast decisions reach resolution, not how many steps were taken - Value per token: business output generated per dollar of AI compute consumed - Exception rate: what share of workflows required human intervention, and why - Cycle-time compression: days to hours, weeks to days - Judgment yield: the share of human time spent on high-value decisions versus low-value drafting The board question: what are we actually measuring, and does it reflect the value AI is creating, or destroying, in our operations? If your AI KPIs are adoption rates and user-satisfaction scores, you are measuring inputs, not outcomes. Governance is your fiduciary duty, not a checkbox Regulators and investors have shifted their stance. Board oversight of AI has jumped sharply in public-company disclosures this year. Major asset managers are now factoring AI governance maturity into company valuations. Full enforcement provisions for high-risk AI systems are coming into force. AI agents are not passive tools. They hold API access, execute transactions, read sensitive data, write to production systems, and make decisions with real consequences, at machine speed and machine scale, around the clock. That makes them an entirely new class of security target. Prompt injection, privilege escalation, data exfiltration, and cascading failures across multi-agent systems are live threats, not theoretical ones. Industry analysts warn that more than forty percent of agentic AI projects will be cancelled within a few years due to runaway costs, unclear business value, or governance failures. The organizations that escape pilot purgatory treat AI agents as accountable infrastructure with measurable KPIs from day one, not experiments with open-ended timelines. Every agent in your organization needs: - A sovereign identity with role-based access permissions - The minimum access necessary to complete its task - Every action logged and auditable, with an unalterable trail - A kill switch that actually works "Can we demonstrate, in writing, how every AI agent makes decisions, who is accountable, and what happens when they go wrong?" If the answer is no, that is a material governance gap, and increasingly, a material disclosure obligation. Shadow AI: the risk already inside your organization Your employees are already using AI, and they're doing it without any of the controls you're carefully designing at the enterprise level. They upload customer data to free tools, draft client communications in consumer models, and wire up personal automation workflows that touch company systems in ways nobody has reviewed. Shadow AI isn't a future risk you can plan around. It's already happening, in virtually every organization that employs knowledge workers. The answer is not prohibition. It is governance through enablement. Organizations that try to ban AI find that employees simply get better at hiding it. The principle: your people will use AI. The only question is whether they do it inside your governance framework or outside it. Make the inside option attractive enough that the outside option becomes unnecessary. The gap is widening right now The organizations winning with orchestration are not winning because they have better technology. They are winning because they started earlier, governed better, and built faster feedback loops. And every month that passes, their advantage compounds. Each workflow they automate generates data, which trains better models, which produce better decisions, which generate more value, which funds more investment. The flywheel keeps speeding up. A recent global study of CEOs found that those who have systematically folded proprietary data and IP into custom AI models and agents expect a materially larger share of their end-of-decade revenue to come from products and services they don't offer today. That's not a marginal edge. It's the kind of gap competitors don't close by trying harder. Meanwhile, the organizations still in the pilot phase, still debating architecture, still waiting for the technology to mature, are watching that gap open. The technology is mature enough, the platforms are production-ready, the ROI is proven. What's missing in every company that isn't moving isn't capability. It's a decision. The decision in front of you AI orchestration isn't a technology decision. It's a decision about strategy, governance, culture, and ultimately about what kind of organization you intend to be at the end of this decade. The leaders who look back on this year with satisfaction won't be the ones who moved fastest or spent the most. They'll be the ones who moved deliberately, with a clear architecture, costs they could account for, a real plan for their people, and the discipline to turn all of it into business value. "Judgment has never been more important. The decision is yours." Boardroom discussion starters - The cost audit: do we have a fully loaded cost model for every AI workflow, and dynamic model routing in place to optimize that spend? - The governance audit: can we demonstrate how every AI agent makes decisions, who is accountable, and what happens when they go wrong? - The shadow AI audit: how are we tracking unsanctioned AI usage, and what is our roadmap for governed enablement that makes compliance the path of least resistance? - The workflow blueprint: which core processes are ready to move from human-intermediated tasks to fully orchestrated, governed multi-agent workflows? - The ownership question: who at the C-suite owns our orchestration strategy, with the cross-functional authority, business credibility, and AI fluency to drive it? If you are trying to answer those questions, the hardest part won't be the tools. That's the work I do in AI strategy engagements, rebuilding the operating model rather than bolting a model onto the org chart. Let's talk → ### AI Adoption Phase 5, Transform: The 6% Who Redesigned the Org: /blog/phase-five-transform/ Only 6% of companies capture real EBIT impact from AI, and they're nearly three times more likely to have redesigned how they work. The last phase isn't technical. It's the operating model, and it's the one only leadership can change. Every essay in this series has been about a phase you can delegate. An executive can commission the audit, fund the pilots, hire the engineers who build and scale the agents. Transform is different. Transform is the phase where the executive is the work. Per the funnel I laid out in the series opener, the research finds that only about 6% of organizations qualify as AI high performers, companies that attribute 5% or more of EBIT to AI and report significant value from it. And here's the detail that should keep boards up at night: those high performers are 2.8 times more likely than everyone else to have fundamentally redesigned their workflows, 55% of them have, against roughly 20% of the rest. The other 94% have plenty of spend and pilots to show, and still struggle to find the impact on the P&L. The money was never in the tools or even the agents. It was always in the redesign. Why the last transition is the hardest There's a reason the funnel collapses from 23% to 6% at exactly this point, and it isn't technical difficulty. Every previous transition could be framed as an addition: more tools, more pilots, more agents, a bigger ops function. Additions are politically cheap. Transform is the first change that requires subtraction, ripping out process and collapsing roles and approval chains, and subtraction always has a constituency against it. Worse, the people who must lead the redesign are the people the current design made successful. The org chart, the planning cadence, the approval matrix, those aren't neutral artifacts. They're the accumulated answers to the question "how do we coordinate expensive humans doing slow work?" Every executive in the building earned their position by mastering those answers. Asking them to redesign the model is asking them to devalue their own expertise. That, not model quality, is why 94% stall. The limiting reagent of AI transformation is leadership courage, not technology. "The limiting reagent of AI transformation is leadership courage, not technology." What actually gets redesigned "Redesign the organization" sounds abstract until you list what the 6% actually changed. Four things, consistently. The unit of work. Pre-AI organizations move work in big batches, quarterly roadmaps, projects, epics, because coordination overhead made small batches uneconomical. When building gets cheap, the economics invert: the winning unit of work is the small, shaped bet, a problem framed as an outcome, given an appetite, and handed to a team that owns the result. I wrote the full argument in when building gets cheap, shaping becomes the job; Transform is where it stops being a personal habit and becomes the official operating system: you shape problems instead of writing specs, set an appetite instead of estimating, place bets instead of grooming a backlog, and measure the outcome rather than the output. The shape of teams. The handoff chain, product writes, design draws, engineering builds, QA checks, was a coping mechanism for specialization scarcity. With AI collapsing the cost of each specialty's mechanical layer, small autonomous teams own outcomes end to end, and the boundary between product and engineering thins toward vanishing. Roles change with it: every engineer is a manager now, directing fleets of agents and reviewing their output, and the leverage of every senior person is measured by what they shape and judge, not what they type. The decision system. Approval chains exist to ration expensive, slow execution. When execution is cheap and fast, the chain itself becomes the most expensive component, a committee spending three weeks deciding whether to attempt two days of work. The 6% replace permission with boundaries: appetites, error budgets, and guardrails that make small failures survivable, then they let teams move. Failure handling is the real tell of the culture. A well-run failed bet gets treated as the system working, because the alternative is a culture where misses are punished, and that quietly reinstates every approval chain you tore out, only now it does it informally. The economics. Transform-phase companies re-cost their assumptions from scratch. Things that were premium become default, every customer gets the white-glove onboarding, every deal gets the custom analysis, because the marginal cost collapsed. Capacity freed by agents gets reinvested in the leaky faucets and ambitions that never made the old roadmap, instead of being banked as headcount reduction and handed to a spreadsheet. The companies that treat AI purely as a cost-cutting tool get a smaller version of their old company. The 6% get a different company. "Cost-cutters get a smaller version of the old company. The 6% get a different company." You can't pilot your way here, but you can slice your way here The trap executives fall into at this point is concluding that Transform requires a big-bang reorg, the all-hands, the new org chart, the consulting deck with the word "transformation" in 60-point type. It doesn't, and the big bang usually fails, because you can't PowerPoint an operating model into existence. What works is the vertical slice from the series opener, applied at full depth: take one value stream, one, and run it the Transform way end to end. An autonomous team works shaped bets, agents do the mechanical labor, you measure outcomes instead of output, and you treat every failure as information rather than a crime. Protect it from the old process like an organ transplant from rejection. That slice does two jobs no memo can. It generates evidence: cycle times, costs, and customer outcomes you can hold up against the old model's baseline. And it generates defectors: people who've worked the new way and won't go back, who become the most credible advocates the change will ever have. The redesign then spreads the only way operating models ever spread, by envy, one team at a time, with leadership clearing the path and retiring the old machinery behind it. What this looks like when I do it with you Transform is the reason my AI-native engagement is structured the way it is. The audit, the shipped pilots, the agent patterns and the ops discipline are also quietly building the credibility and the evidence base this phase spends. By the time we're redesigning the operating model, we're not arguing from a whitepaper. We're arguing from your own numbers, on your own workflows, with your own people as the proof. The white-glove part here is the most personal in the whole arc, because this phase runs on trust, not deliverables. I work directly with the executive team on the redesign itself: which value stream goes first, how to shape and bet instead of plan and approve, what each leader's role becomes when their function stops being a handoff station, how to handle the manager whose job the new model genuinely eliminates, and how to talk about failure so the new boundaries hold. I stay through the messy middle, the quarter where the old process is dying and the new one isn't trusted yet, because that's the exact moment companies lose their nerve and snap back. And then the engagement is designed to end: the operating model is yours, run by your people, or it was never a transformation at all. Five phases. Most companies will keep circling the first two, equipping and experimenting, because it feels like motion and nobody has to change. The returns sit in the last phase precisely because it costs something the early ones don't: leadership willing to redesign the machine it sits on top of. If you're ready to be in the 6%, let's talk → ### How to Say No to Feature Requests Without Killing Momentum: /blog/saying-no-to-good-ideas/ Focus doesn't mean rejecting bad ideas. Those reject themselves. Part 4 of my product design philosophy: kill projects that miss the bar, prototype with working software, and restart when it isn't right. When Jobs said he was as proud of the things Apple didn't do as the things it did, and that focus means saying no to 1,000 things, most people filed it away as a productivity tip. It's the most expensive discipline in product work, because the thousand things you say no to are good ideas, and good ideas are the hardest ones to refuse. This is the final part of a four-part series on my product design philosophy. The rest: Part 1: If Your Product Needs a Manual, the Design Already Failed, Part 2: Start With the Experience, and Part 3: Nobody Asks for the Future. Bad ideas reject themselves Nobody needs a philosophy to say no to bad ideas. Bad ideas die in the meeting where they're proposed. The dangerous ideas are the good ones: a feature a real customer is asking for, an adjacent market that really is sitting there, a partner with a working integration on the table. Each one passes every test except the only test that matters, whether it serves the one thing you're trying to be exceptional at. "Bad ideas reject themselves. Good ideas are the ones that dilute you." Do a few things exceptionally well rather than many things adequately. The math behind that sentence is brutal and simple: exceptional is a threshold, not a gradient. A product that's adequate at six things loses to a product that's exceptional at one, in every market I've ever operated in, because users don't buy averages. They buy the one thing you're for. I've argued before that prioritization is the job; this series is why. Every yes is funded by quality withdrawn from everything you already said yes to. Kill what misses the bar Saying no at the idea stage is the cheap version. The expensive version is killing a project that's already alive: months invested, a team attached, a roadmap commitment made, and the honest assessment is that it will ship as adequate, not exceptional. Most organizations ship it anyway, because the sunk cost has a constituency and the standard doesn't. Jobs killed products days before launch. I'm not romantic about the brutality, I've been in those rooms and they're miserable, but I am certain about the economics. The cost of shipping adequate isn't the one mediocre feature; it's the permanent tax of maintaining it, the support load, the surface area in every future redesign, and, worst, the message to the team that the bar is negotiable under deadline. Killing it says the opposite, and the team you keep hears it loudest. A high standard you enforce once buys you a hundred arguments you never have to have. Prototype with real working models Make real working models, not just drawings. Of all the principles in this series, this is the one where the ground has shifted most since Jobs said it, and it's shifted in his direction. A mockup is a drawing of an opinion. It cannot be slow, it cannot have an awkward empty state, it cannot reveal that step three is infuriating, because it cannot be used. Decisions made from mockups are decisions made from the most flattering possible version of the idea. Working software is the only honest critic on the team: it pushes back, it embarrasses the idea early, it tells you things nobody in the design review knew to look for. "A slide deck never pushes back. Working software argues with you." In 1997, demanding working models was an expensive eccentricity only Apple could afford. In 2026 it's nearly free. When building gets cheap, a working prototype costs an afternoon with a coding agent. Which means deciding from static mockups is no longer a budget constraint, it's a choice, and it's the wrong one. The new rule on my teams: if we're debating it in a document for the second meeting, someone builds it instead. The prototype ends arguments that decks can only prolong. Refine until it feels right, restart when it doesn't Keep refining until it feels absolutely right, and don't be afraid to restart if it isn't. "Feels right" sounds unrigorous next to metrics and acceptance criteria, but it's the most demanding bar on the list, because feel is the sum of every detail from Part 2 and you can't unit-test the sum. The restart is the part teams fear, and the fear is outdated. A restart was a catastrophe when the build took a year. When the build takes days, a restart is just iteration with honesty about how far back the problem goes. The second build is always faster and always better, because the first build was the real spec, the working model that taught you what the thing actually needed to be. Version one of anything important I've shipped was thrown away at least once. That isn't waste, it's how the work gets done. The whole philosophy on one page Four essays, one position. Simplicity is the ultimate sophistication, and a product that needs a manual has failed. Start with the experience and work backwards to the technology, with quality going all the way through, including the parts nobody sees. Build what customers can't ask for, and own the layers that carry the vision. Then protect all of it with ferocious focus: say no to the good ideas, kill what misses the bar, prototype for real, and restart without fear. None of it is new. That's rather the point. The philosophy has been sitting in plain sight for decades while building was too expensive for most teams to live by it. Cheap building just removed the last excuse. If your roadmap has a thousand yeses on it and you know which essay you skipped to get here, let's talk → ### Why Customer Research Fails for New Products: /blog/nobody-asks-for-the-future/ Customers describe their problems in the vocabulary of today's solutions. Part 3 of my product design philosophy: faster horses, seeing what others can't, and owning the parts of the stack that carry your vision. If you ask customers what they want, they'll say faster horses. The line is attributed to Henry Ford, popularized by Jobs, and misunderstood by almost everyone who quotes it. It does not mean ignore your customers. It means understand what they're actually telling you. This is Part 3 of a four-part series on my product design philosophy. The rest: Part 1: If Your Product Needs a Manual, the Design Already Failed, Part 2: Start With the Experience, and Part 4: Saying No to 1,000 Good Ideas. Customers are experts in problems, amateurs in solutions "Faster horses" is actually a perfect piece of customer feedback, if you listen to it correctly. The customer is telling you the problem with total precision: getting places takes too long. What they can't tell you is the solution, because they can only describe the future using the parts of the present. Nobody who had never seen a car could request one. Nobody asked for the iPhone. Nobody filed a feature request for the spreadsheet, the search engine, or the chat-based coding agent. "Customers are experts in their problems and amateurs in your solutions." So the job splits cleanly. Take the problem statement literally, it's gold. Take the solution request as a clue, never as a spec. When a customer asks for a faster horse, the roadmap-driven team breeds horses. The team that's listening hears "speed matters more than anything" and starts questioning the horse. This is why I'm cold on market research as a source of direction. Research is a rear-view mirror; it describes the world the last generation of products created. It can validate an idea, sharpen your priorities, or kill a bad assumption, and that's worth a lot. What it cannot do is show you the future, because the future isn't in the data yet. Create products people don't know they need. By the time they know they need it, someone has already built it. Innovation is seeing, then having the nerve to act True innovation means seeing what others can't, and the romantic version of that sentence stops there, as if vision were a gift. It isn't. In my experience the "seeing" is mostly proximity: the people who spot the future early are the ones closest to the friction, watching real users do real work and noticing the workaround everyone else has normalized. The insight is lying on the workshop floor. Most people step over it daily. The rare part isn't the seeing. It's the nerve. Acting on something only you can see means committing resources to a thing no customer asked for and no analyst put in a slide. There's no competitor doing it yet to make you feel safe. It will look wrong to smart people, by definition, because if it looked right to everyone it wouldn't be early. That's why conviction is a design input, not a personality flaw. "Think different" was never decoration; it's the willingness to be temporarily wrong in public on the way to being right. Own what carries the vision Jobs' refusal to build on other people's components, the obsessive integration of hardware and software, gets dismissed as control-freakery. It was strategy. Great experiences come from controlling the stack that delivers them, because every layer you don't control is a place your vision gets negotiated down to someone else's roadmap. Translated out of Cupertino and into the companies I work with: you don't need to build everything, you need to own the layers where your differentiation lives. Buy commodity, build identity. Your auth, your billing, your email delivery, rent them happily. But the workflow that is your product, the data model that encodes how you see the customer's world, the experience from Part 2 that you worked backwards from? Outsource those and you've outsourced the company. I've written about owning the whole stack organizationally; this is the product version of the same conviction. Everything must work together seamlessly, and seamless is precisely the property you cannot buy from three vendors and glue. "Every layer you don't control is a place your vision gets negotiated down to someone else's roadmap." The modern test case is AI. Companies bolting a vendor's chatbot onto an unchanged product are integrating someone else's component and calling it innovation, faster horses with a plugin. The companies getting it right are rethinking what their product even is now that software can act, and they're keeping that rethinking in-house, because it's the new core. Conviction has a price, and it's Part 4 Betting on what you see instead of what's requested means you will sometimes be wrong, and you'll be wrong with real money. The only thing that makes this philosophy survivable is the discipline that pairs with it: ferocious focus, small bets, working prototypes, and the willingness to kill what isn't great. That discipline is the finale of this series: saying no to 1,000 good ideas. If you're sitting on something only you can see and you need help finding the nerve, or the plan, let's talk → ### On-Device AI Economics: When a $20K Box Beats the Cloud: /blog/on-device-ai-economics/ If frontier AI lands at $1K to $5K per employee per month, a $20K private AI server per employee stops sounding crazy and starts sounding like procurement. The hardware has mostly caught up; the math hasn't, not quite. Here's a back-of-the-napkin exercise every CFO is going to run within the next two years. Take a heavy AI user on your team, an engineer running agents all day, an analyst with a research pipeline, anyone whose job has quietly become "directing models." At frontier pricing as of June 2026, that person can burn $1,000 to $5,000 a month in tokens. I wrote about Fable 5's pricing when it landed: ten dollars per million input tokens, fifty per million output. Run serious agentic workloads against those numbers and a four-figure monthly bill per employee isn't an edge case. It's the median. Now ask the question the CFO will ask: if that employee costs $3,000 a month in inference, why don't we just buy them a $20,000 private AI server and be done with it? It's a good question. The answer today is "because you can't." What's worth paying attention to is how quickly that answer is falling apart. The rent-versus-own math is brutal Do the arithmetic without flinching. An employee spending $3,000 a month on frontier inference costs $36,000 a year, $108,000 over three years, and the meter never stops. A $20,000 box amortized over the same three years is about $550 a month, plus electricity. Even at the bottom of the range, $1,000 a month in tokens, the box pays for itself before the second year is out. At the top of the range it pays for itself in four months. We have seen this exact movie before. Companies rented time on mainframes until the PC made compute cheap enough to put on every desk. They rented data center racks until the cloud made elasticity worth a premium, and now the same companies are repatriating steady-state workloads because renting a constant is the most expensive way to own it. Inference for a knowledge worker is about as steady-state as workloads get: the same person, hammering models, eight hours a day, every working day. That is precisely the shape of demand that capex was invented for. "Renting a constant is the most expensive way to own it. A knowledge worker's daily inference is about as constant as workloads get." And the per-token bill is only half the argument. The box on the desk doesn't rate-limit you, and it doesn't ship your code, your contracts, or your customer data off to someone else's data center. Nobody quietly changes its pricing, deprecates your model, or hands you an outage in the middle of a product launch. For regulated industries (healthcare, defense, finance, legal), "the data never leaves the building" can be the entire reason the purchase order gets signed. Walk into Micro Center and look at the shelf Here's what makes this more than a thought experiment: the hardware is no longer hypothetical. Nvidia is already selling the early versions of the box, at consumer and prosumer prices. - $1,999: GeForce RTX 5090: 32GB GDDR7, ~1.8 TB/s bandwidth. Runs ~30B-parameter models quantized on a gaming card. - $3,999: DGX Spark at launch: GB10 Grace Blackwell, 128GB unified memory, a petaflop of FP4. Runs models up to ~200B parameters on a desk. - $8,565: RTX PRO 6000 Blackwell: 96GB GDDR7 ECC. ~70B parameters at full precision, ~180B at FP4, in a workstation. Read that middle line again. Nvidia ships a book-sized machine with 128GB of unified memory that runs 200-billion-parameter models, locally, for four grand. Wrap an RTX PRO 6000 in a workstation and you're at roughly $12,000 to $15,000, comfortably inside the $20K envelope, serving 70B-class models at full precision with room for long context. The category the CFO's question assumes already exists and is sitting on a shelf at Micro Center. Demand is already outrunning supply, which tells you something about where this is headed. As of June 2026, the DGX Spark's street price has crept from $3,999 toward $4,700, largely on memory supply constraints, and 5090s have spent most of their life trading above MSRP. The market is voting. So why can't you do it yet? Because of the gap nobody puts on the spec sheet: the model that's worth $3,000 a month doesn't fit in the box, and the models that fit in the box aren't worth $3,000 a month. The frontier models behind that per-employee bill are vastly larger than anything 128GB of memory can hold, and it wouldn't matter if they weren't, because the labs don't sell the weights. What you can actually run locally is the open-weights tier: excellent 30B-to-180B models that land roughly where the frontier was a year or so ago. For a lot of tasks, that's genuinely fine. For the work that justifies a four-figure token bill (the long-horizon agents, the hard reasoning, the stuff that replaced a workflow rather than autocompleting one), the capability gap is exactly the part you were paying for. That's the honest state of on-device AI in 2026. The cost math is begging you to buy the box, while the capability you actually need is still telling you to keep renting. The unit economics are not there yet, and "yet" is the word I'd underline. Two curves, one inflection point What makes this a timing question rather than a fantasy is that two curves are converging from opposite directions, and neither shows signs of stopping. Models are being optimized downward. The cost of inference at a fixed capability level has been falling roughly tenfold per year, a thousandfold over three years, driven by distillation, quantization, mixture-of-experts sparsity, and plain algorithmic progress that cuts the compute needed for a given score by multiples annually. The practical version of that statistic: today's 70B open-weights model does what a frontier model did twelve to eighteen months ago. The frontier keeps moving, but the copy of it that fits in a box keeps arriving about a year later, at a hundredth of the price. Hardware is climbing upward. Consumer cards went from 24GB to 32GB in a generation. The unified-memory desk machines went from zero to 128GB in one product cycle, with native FP4 turning every gigabyte into three. Each hardware generation raises the ceiling on what "fits locally" by roughly a model class, and Nvidia, Apple, and AMD are all racing to own the category, because they can read the same CFO napkin everyone else can. "The frontier keeps moving. But the copy of it that fits in a box keeps arriving a year later, at a hundredth of the price." The inflection point is where those curves cross: the moment a locally-runnable model handles, say, 80% of a given employee's actual workload at quality indistinguishable from the API. It doesn't have to be 100%, and chasing that number is the wrong bar to set. The desk box doesn't need to beat the frontier; it needs to be good enough that the cloud becomes the exception instead of the default, with local handling the daily grind and the frontier API reserved for the 20% of problems that genuinely need the best model on Earth. The day that hybrid math works, the $20K box stops being an enthusiast purchase and becomes an IT line item, which is the same path the PC and the workstation took as compute kept migrating toward the user whenever the economics finally allowed it. What to do about it now Whatever you do, don't rush out and buy the hardware. That's the trap. A $20K box purchased today is a depreciating bet on this year's model class, and both curves are moving fast enough that next year's box will embarrass it. Early hardware hedges almost always lose to one more year of renting. - Architect for portability. Every AI workflow you build should sit behind a routing layer that doesn't care whether the model is an API in Virginia or a card under the desk. If your agents are hard-wired to one provider's endpoint, the inflection point will arrive and you won't be able to take the discount. - Route by difficulty, not by habit. Start splitting workloads now into "needs the frontier" and "a strong open-weights model would do." That's the same discipline that saves you money on tokens today, and it's the exact seam where local inference will slot in later. If you already route by difficulty, adopting local models is mostly a config change; if you don't, you'll be rewriting workflows under time pressure. - Run the napkin math per role. Token spend per employee is about to be a number your CFO knows. Know it first, and know which roles are at $200 a month and which are at $4,000, because the $4,000 roles are where on-device lands first. - Let privacy lead the business case. If you're in a regulated industry, the inflection point arrives early for you, because the comparison isn't box-versus-API price. It's box-versus-not-being-allowed-to-use-AI-at-all. On-device AI is the future for the same reason the PC was: compute migrates toward the user the moment the economics permit it, and the economics are converging from both directions at once. The labs optimize down, the silicon climbs up, and somewhere in the next few years they cross. Winning that crossing has very little to do with who buys hardware earliest. It comes down to whose architecture, routing, and cost discipline are ready on the day the box finally makes sense. If you're trying to figure out where your organization's inflection point is, what to route where, what your real per-employee AI economics look like, and how to build for the switch without betting early, that's the work I do. Let's talk → ### AI Adoption Phase 4, Industrialize: Scale Agents, Not Chaos: /blog/phase-four-industrialize/ Fewer than a quarter of companies have scaled AI agents beyond the first win. Scale is where AI stops being a project and becomes infrastructure, and infrastructure has rules most AI teams haven't learned yet. There's a moment in every successful AI program when the question flips. For the first few agents the question is "can we make this work?" Then one quarter the agents are handling real volume across three functions, the API bill has a comma in a new place, a model deprecation notice lands in someone's inbox, and the question becomes "can we keep all of this working, at this price, while everything underneath us keeps moving?" That's the Industrialize phase. Per the funnel in my series opener, the research finds just 23% of organizations scaling an agentic system anywhere in the enterprise, and the ones that get here discover that scale changes the nature of the work entirely. One agent is a project you ship. A fleet of twenty is something you operate, and operating it pulls in the same rules that govern databases and payment systems, applied to a layer most teams are still treating like a science experiment. What breaks at scale The failure modes of this phase are nothing like the phases before it. Nobody here is wondering whether AI works. They're drowning in the consequences of it working: - Cost stops being a rounding error. At pilot volume, nobody reads the API bill. At production volume, unit economics decide whether the agent is a margin story or a margin leak, and most teams can't tell you their cost per resolved case to within an order of magnitude. - Model churn becomes weather. Providers ship better-cheaper-different models every few months and deprecate the ones you built on. Every agent you run is built on ground that moves, and "we'll stay on the old model" is a strategy with an expiration date. - Quality drifts silently. Prompts get edited, retrieval corpora grow stale, traffic shifts toward inputs you never evaluated. Without continuous measurement, an agent degrades the way a bridge rusts: invisibly, then suddenly. - The portfolio loses legibility. With twenty agents owned by five teams, nobody can answer "what is our AI doing right now, what is it costing, and which of these things still earn their keep?" Notice that every one of these is an operations problem, not an intelligence problem. The model is the least of your worries in this phase. The worries are the same ones every ops discipline eventually codified: visibility, budgets, regression safety, lifecycle management. "Building one agent is a project. Running twenty is infrastructure, and infrastructure has rules." The three instruments of a scaled AI operation Companies that industrialize well converge on the same three instruments, whatever tools they use to implement them. First: evals as the regression suite. The eval harness you built in the Operationalize phase (or should have) graduates into the central nervous system of the whole operation. Every prompt change, every retrieval tweak, every model upgrade runs the suite before it ships, exactly like a CI pipeline, because that's what it is. This is what makes model churn survivable: when a new model drops, you don't convene a committee, you run the evals, read the diff in quality and cost, and decide in an afternoon. Second: observability with cost attached. Every agent interaction traced, prompt, context, output, latency, tokens, dollars, score, so quality and spend are queryable in one place. I've written about the concrete stack in LLM ops with Langfuse and Finout, but the tools matter less than the discipline: cost per case and quality per case on the same dashboard, per agent, per week. The instant an agent's unit economics become visible, "is the AI worth it" stops being a matter of belief and turns into a number you can read off the dashboard. Third: a portfolio review with teeth. A standing rhythm, monthly is right for most, where every production agent defends its existence with three numbers: volume handled, quality against budget, cost against the human baseline. Agents that no longer earn their keep get retired without sentiment. This sounds obvious and almost nobody does it, because nobody assigns an owner to the portfolio as a whole. Individual agents have owners; the fleet has none. Fix that and half of this phase fixes itself. Scale is a flywheel, not a checklist The companies that do this well treat those three instruments as the engine of compounding rather than overhead. Because every interaction is traced and scored, production becomes a continuous source of new eval cases. Because evals are cheap to run, model upgrades get adopted in days, which keeps cost falling and quality rising. Because cost is visible per case, workflows that were marginal last quarter become viable this quarter. The optimization never reaches a finish line; each turn of the loop makes the next one pay off more. I've run this loop on LLM pipelines processing roughly 250 million records a month across a 75-node cluster, and the honest lesson is that the loop, not any individual agent, is the asset. - 250M/mo: Records processed by LLM-powered pipelines - 90%: Reduction in data-processing time - 75: Kubernetes nodes orchestrating the fleet "The loop, not any individual agent, is the asset." The ceiling of Industrialize And yet, this phase has a ceiling, and it's worth naming honestly because it sets up the final essay in this series. You can run a flawless agent fleet inside an organization whose processes, roles, and decision-making were designed for a pre-AI world, and what you get is a beautifully optimized version of the old company. The agents accelerate the existing workflows. They don't question them. Nobody asks whether the workflow should exist at all, whether the department boundary it crosses still makes sense, or what the organization should do with capacity that suddenly costs a tenth of what it did. Those are operating-model questions, and no amount of LLM ops answers them. That's the transition into the Transform phase, the one only 6% make. What this looks like when I do it with you Industrialize maps to the third movement of my AI-native engagement: continuous optimization. The white-glove version is that I build the three instruments into your organization rather than describing them to it, standing up the eval-gated deploy pipeline, wiring cost and quality into one pane of glass, installing the portfolio review and chairing the first few so the standard is set by demonstration, not memo. I also handle the unglamorous calls this phase runs on: which model migrations are worth taking now versus next quarter, where caching and routing cut cost without cutting quality, which agents to retire even though someone loves them. The deliverable isn't a fleet that works this quarter. It's an operation that keeps getting cheaper and better without me in the room, because the loop is owned, instrumented, and reviewed on a rhythm your team runs. Final essay in the series: the organizational redesign that only 6% attempt, Transform: the 6% who redesigned the organization. And if your API bill just grew a comma and nobody can say what it bought, let's talk → ### What Is Product Engineering? It Was Always One Job: /blog/product-engineering-one-job/ We split "what to build" from "how to build it" because building was expensive. Now the split itself is the expensive part. Product engineering is what's left when you remove the seam. Walk into almost any software company and you'll find the same wall. On one side, people who decide what to build. On the other, people who build it. Between them: tickets, specs, refinement meetings, estimation rituals, and a translation layer so thick we invented entire job families just to maintain it. I've spent twenty-five years on both sides of that wall, and I'll say plainly what most people in this industry already suspect: the wall was never a principle. It was a workaround. We split the job because building software was so expensive and so specialized that the person who understood the customer couldn't afford to also be the person who understood the code. The split was rational. It was also a compromise from day one, and anyone who's watched a good idea die in a handoff knows it. AI just removed the reason for the compromise. Building got cheap. And when building gets cheap, the most expensive thing left in your delivery pipeline is the seam itself, the lossy, slow, accountability-diffusing handoff between the person who knows why and the person who knows how. The discipline that emerges when you remove the seam needs a name, and it has one: product engineering. It isn't product management with a code editor, and it isn't engineering with a few extra meetings bolted on. It's one person, or one small team, owning the loop from customer problem to shipped outcome, with nothing handed off in between. Why we split the job in the first place The split made sense when we made it, and it's worth remembering why. In a world where a feature cost a quarter of engineering time, a wrong bet was catastrophic. So you put a careful, customer-facing person in front of the engineers to pre-filter the work. And in a world where engineering skill took a decade to build and was scarce on the market, you protected engineers from anything that wasn't engineering: customer calls, pricing arguments, the support queue. Keep them building, because building was the bottleneck. Every artifact of modern product development is downstream of that economics. The spec exists because the builder wasn't in the room when the problem was understood. The estimate exists because the decider couldn't judge the cost themselves. The backlog exists because demand for building always exceeded the supply of builders. The sprint demo exists because the people who asked for the thing had to wait weeks to see it. Now run the numbers in 2026. A competent engineer with AI leverage ships in days what used to take a quarter. The cost of a wrong bet has collapsed alongside the cost of a right one, which means the elaborate pre-filtering apparatus is protecting an asset that's no longer scarce. Meanwhile the seam, the handoff, hasn't gotten one minute faster. Specs still take days to write and still lose the nuance. Refinement meetings still burn whole teams' afternoons. The question still travels from engineer to PM to customer and back over a week, when the build itself would have taken two. "When building gets cheap, the handoff becomes the most expensive thing in the building." What a product engineer actually is The title needs pinning down, because it's already being diluted into "full-stack engineer who attends roadmap meetings." A product engineer owns a customer problem end to end. They talk to the people who feel the problem, directly, not through a research summary. They decide what's worth building and how much it's worth, because they can judge both the value and the cost in the same head. They build it, with AI doing most of the typing. They ship it, watch the metrics, sit in on the support tickets, and decide whether it actually worked. For them, done isn't the moment the code merges; it's the moment the problem stops happening. Notice what's absent from that loop: a handoff. There's no moment where context gets serialized into a document, transmitted across a wall, and deserialized with loss on the other side. The person who heard the customer sigh is the person who ships the fix. Almost everything good about the model traces back to that one fact. It helps to be just as clear about what the role is not. A product engineer is not a PM who learned to prompt. They carry real engineering judgment, they're accountable for the system staying coherent, secure, and operable, and AI doesn't hand them that judgment so much as multiply whatever they already had. They're also not an engineer who got handed a backlog column. They carry real product judgment: they say no, they kill their own features, they understand that owning a feature costs far more than building it. The role is demanding precisely because it refuses to let you hide in either half of the old split. The skills are different from either parent Product engineering isn't the union of two job descriptions. It selects for things neither parent role hired for: - Problem fluency over solution fluency. The scarce skill isn't writing code or writing specs, it's holding a customer problem in your head at the right altitude and noticing when the thing you're building has quietly stopped solving it. - Appetite judgment. Deciding how much an outcome is worth before deciding how to build it, and scoping the build to fit the worth rather than gold-plating to fit the ego. - Taste under speed. When you can ship anything in days, the discipline isn't shipping, it's choosing. A product engineer's portfolio is judged by coherence, not throughput. - Comfort with the whole consequence. There's no one to pin a bad requirement on and no one to pin a bad implementation on, because you were both of those people. You decided, you built, and you live with what happened next. If that list sounds like what we used to expect from founders, that's not an accident. Product engineering is founder-mode made into an organizational discipline. Every early-stage founder already works this way: talks to a customer in the morning, ships the fix in the afternoon. Then most companies spend their entire scaling journey engineering that loop out of the org, one specialized role at a time. The AI-era opportunity is to scale without severing the loop. What it does to the org chart I won't pretend this is a painless reshuffle. It isn't. Two roles feel it directly. PMs bifurcate. The ones whose value was the translation layer itself, gathering requirements, grooming backlogs, running ceremonies, are translating between two sides that no longer need an interpreter. The ones whose value is editorial judgment move up: fewer of them, covering more surface, acting as the keepers of product coherence across many fast-moving product engineers. I've written about that inversion before; the short version is that the PM becomes the editor, and the product engineers become the writers. Engineers face the more uncomfortable question, because the comfortable version of the engineering job, take a well-specified ticket, implement it well, hand it back, is exactly the part AI absorbed first. The engineers who thrive are the ones who move toward the customer, not away. That transition is real work. Most engineers were actively trained out of product judgment, told for years that the requirements were someone else's job. You don't undo that with a memo. You undo it by handing people small, real outcomes, an appetite, and the safety to miss. Which is why this is an operating-model change, not a re-titling exercise. Teams get smaller, two or three product engineers instead of eight specialists in a relay. Cycles get shorter. The unit of work stops being the feature and becomes the outcome. And the management layer stops assigning tasks and starts shaping problems and betting on them. "Product engineering is founder-mode made into an organizational discipline." How to start without blowing anything up If you run a product or engineering org and this resonates, don't announce a reorg. Run an experiment. Pick one real customer problem, small enough to be safe, real enough to matter. Pick one engineer with latent product instincts, you already know who they are, they're the one who keeps asking why in refinement. Give them the problem, not a spec. Give them direct access to the customers who feel it. Give them an appetite, say, two weeks, and explicitly take estimation, tickets, and the ceremony layer off their plate for the duration. Then judge the result by one question: did the problem stop happening? Run that three or four times and you'll learn more about whether product engineering works in your organization than any framework comparison will tell you. You'll also learn who your product engineers actually are, and it will not perfectly match your current seniority chart. Expect that, and treat it as information rather than a problem. The companies that get this right won't be the ones with the best AI tooling. Tooling is available to everyone. They'll be the ones that noticed the wall was a workaround, took it down deliberately, and rebuilt their teams around the loop instead of the seam. That rebuild, the roles, the bets, the culture that makes owning an outcome safe, is the core of an AI-native transformation, and it's the work I do. If you're somewhere in the middle of taking the wall down, or deciding whether to, I'd genuinely like to compare notes. Let's talk → ### Product Maturity Model: Is Your Product Ready to Scale?: /blog/product-maturity-model/ Feature flags, a real testing environment, an experimentation subsystem. The unglamorous infrastructure that decides whether your product can absorb more building, or just ship chaos faster. Every scaling conversation I get pulled into starts the same way: "We want to ship faster. Should we hire more engineers, or roll out AI tooling, or both?" And almost every time, the honest answer is one nobody booked the meeting to hear: your product can't absorb the output you already have. Deploys happen twice a month and everyone holds their breath. There's a staging environment, technically, but it drifted from production two years ago and the test data in it is a folk legend. Releases are all-or-nothing: the feature goes out to every customer at once, and if it breaks, the rollback is a war room. Nobody can tell you whether last quarter's big launch actually moved a number, because nothing measures that. Pour more building into that system and you don't get more innovation. You get more incidents, and a team that has learned to be afraid of its own deploys, because every bad release teaches the organization to bolt on another approval step and slow down. The constraint was never how much you could build. It's how much change your product can safely take in. So I've started making the argument explicitly, as a maturity model: a product has to reach a certain operational maturity before scaling development is anything other than scaling risk. Maturity isn't about how old the product is or how clean the code is. It's about whether three specific capabilities exist: you can change the product safely, you can verify changes before customers see them, and you can learn from what you ship. Feature flags, a real testing environment, an experimentation subsystem. Miss any of them and there's a ceiling on your velocity that no amount of hiring will raise. The ladder The model has four levels, and it's deliberately blunt. At each one the question is the same: what happens when you double the volume of change? - Level 0: Fragile. Deploys are events. Testing happens in production, mostly by accident. Releases are coupled to deploys, so every push is a customer-facing gamble. Doubling change volume here doubles your incident count. - Level 1: Repeatable. CI runs a real test suite, deploys are boring and frequent, and there's an environment that genuinely resembles production where changes can be verified first. Doubling change volume is survivable but still scary. - Level 2: Safe. Deploy and release are decoupled by feature flags. New code ships dark, turns on for 2% of users, and turns off with a switch instead of a rollback. Doubling change volume is fine, because each change has a controlled blast radius. - Level 3: Learning. An experimentation subsystem (GrowthBook, or something like it) sits on top of the flags. Every meaningful release is a question with a measured answer. Doubling change volume doubles how fast you learn. The rule that falls out of this: your safe rate of change is set by your maturity level, not your headcount. Hiring and AI tooling raise the rate at which you produce change. Only maturity raises the rate at which you can absorb it. When production outruns absorption, the surplus turns into risk, and the organization responds the only way fear knows how: with process. It freezes releases, stands up a change advisory board, adds another signature to the sign-off chain. You hired to go faster and ended up slower. "Your safe rate of change is set by your maturity level, not your headcount." Level 1: an environment you can trust The unglamorous foundation is a testing environment that tells the truth. Most companies have something they call staging. Far fewer have one where a green result actually predicts production behavior, same infrastructure shape, same configuration path, data that's realistic in volume and weirdness, and integrations that behave like the real ones instead of returning a hard-coded 200. A staging environment that lies is worse than none, because it launders risk. Teams do the responsible thing, verify in staging, ship, and get burned anyway, and after a few rounds of that they stop believing in verification entirely. Then the real testing environment becomes production, you just don't call it that. What I push for: staging built from the same infrastructure-as-code as production so it can't silently drift, a seeded dataset that's regenerated on demand and includes the ugly edge cases from real usage, and ephemeral preview environments per change so engineers aren't queuing for the one shared environment and quietly overwriting each other's state. None of this is exotic. All of it is the difference between "the tests passed" meaning something and meaning nothing. Level 2: deploy is not release Feature flags are the single highest-leverage piece of product infrastructure I know, and they're routinely dismissed as a nice-to-have. What a flag actually does is separate two decisions that have no business being coupled: the engineering decision to ship code, and the product decision to expose behavior. Once those are separate, everything about scaling gets easier. Engineers merge small and merge often, because unfinished work can ship dark behind a flag instead of rotting on a long-lived branch. Releases stop being cliff edges: a feature goes to internal users, then 2%, then 25%, then everyone, and the moment a metric twitches you turn it off. No rollback to coordinate, no war room, nobody shipping a fix at 2 a.m. The blast radius of any mistake stops being "the customer base" and becomes "the cohort we chose." That containment is what makes a high rate of change survivable, and it's why flags are the gate between Level 1 and everything above it. Two honest caveats. Flags need observability next to them, the switch is only useful if a dashboard tells you when to flip it, so error rates and key product metrics have to be visible per-flag, not just globally. And flags rot: an expired flag is dead code with a control panel. Mature teams treat flag cleanup as part of the definition of done, not as someday-debt. Level 3: shipping becomes asking Here's the part that turns operational plumbing into an innovation engine. Once every feature already ships behind a flag to a chosen percentage of users, you are one step away from experimentation: assign the cohorts randomly, attach metrics, and let the math decide. That's all an A/B test is, a feature flag with a hypothesis and a scoreboard. This is why I tell clients to stop treating experimentation as a distant someday and deploy a real subsystem for it. GrowthBook is my usual recommendation, since it's open source, self-hostable, and runs flags and experiments in one place. Once the flag infrastructure exists, the marginal cost of running real experiments collapses. And the cultural effect is bigger than the technical one: roadmap arguments that used to be settled by seniority get settled by data. "I think users want X" becomes "we ran X for two weeks against a holdout and it moved activation 4%." The loudest voice in the room loses its monopoly. "An A/B test is just a feature flag with a hypothesis and a scoreboard." A Level 3 product changes what "shipping" means. At Level 0, shipping is a risk you take. At Level 2, it's a routine you trust. At Level 3, it's a question you ask, and every release comes back with an answer. That's the actual innovation flywheel: not more output, more answered questions per quarter. AI just made this urgent This model used to be a five-year conversation. AI compressed it into a now conversation, because building got cheap, and the volume of change a small team can produce went up by an order of magnitude. Every one of those changes still has to cross the same bridge: verified somewhere honest, released with a contained blast radius, measured against a real metric. An immature product hit with AI-accelerated output doesn't innovate faster. It floods. The deploy queue backs up, staging becomes a rumor, releases get batched into bigger and scarier bundles, exactly the wrong direction. I've watched teams adopt agentic coding tools on a Level 0 product and conclude, three incidents later, that "AI writes bad code." The code was fine. The product had no way to absorb it. The same applies to product engineers who own the loop from problem to outcome: that loop only closes if the product gives them flags to release with and experiments to learn from. You can't own an outcome you can't measure. Maturity work is product work The reason most companies are stuck at Level 1 isn't ignorance, it's framing. Flags and staging and experimentation all get filed under "tech debt" or "platform," which means they lose every prioritization fight against a feature with a customer's name on it. That filing is wrong. This is product work, it determines how fast every future feature ships and whether you ever find out if it worked, and it deserves a place at the table on those terms. The pitch I make to boards is one sentence: this investment raises the safe rate of change for everything we build afterward. Concretely, the sequencing I push is: honest staging and boring deploys first, flags with per-flag observability second, experimentation third, and only then pour in the headcount and the AI tooling. Each step typically pays for itself within a quarter or two, and the DORA metrics will show it: deployment frequency up, change-failure rate down, recovery measured in minutes because recovery is a flag flip. So before you approve the hiring plan or the AI rollout, ask the three questions: Can we change the product safely? Can we verify before customers see it? Do we learn from what we ship? If any answer is no, that's the work, and it comes first. If you want help finding where your product sits on this ladder, and what the shortest path up looks like, let's talk → ### Claude Fable 5 Pricing: The Price Went Up, the Knobs Came Off: /blog/fable-5-price-went-up-knobs-came-off/ Anthropic just shipped a model tier above Opus. The price doubled and the dials disappeared, and both of those facts tell you how to run an AI-native organization. Anthropic just put a new rung on top of the ladder. It's called Claude Fable 5, it sits above Opus, and it costs twice as much: as of June 2026, ten dollars per million input tokens and fifty per million output, against Opus 4.8's five and twenty-five. Most of the coverage will be about capability, and fair enough. It's billed as the most intelligent model they've shipped. But I think the two most instructive things about Fable 5 have nothing to do with benchmarks. The first is the price. The second is what they removed from the API. Both are signals about how this technology wants to be managed, and most organizations are reading neither. The price is the announcement For the last few years, the implicit deal was that frontier intelligence got cheaper. Every release either raised the ceiling at the same price or held the ceiling and dropped the floor. You could budget AI the way you budget bandwidth: assume the unit cost only goes down, and don't think too hard about it. Fable 5 breaks that pattern. For the first time in a while, the ladder grew upward instead of the floor dropping. As of June 2026 the lineup runs Haiku at a dollar in, Sonnet at three, Opus at five, and Fable at ten. A real spread, with a meaningful price jump at the top. - 2×: the per-token price of Opus 4.8 - 1Mtokens: of context window - 128Ktokens: of maximum output That spread changes the nature of the decision. When the top model costs roughly the same as the one below it, "just use the best one" is a defensible default. When the top model costs double, model choice becomes a portfolio decision, and a portfolio decision is a management decision. Somebody in your organization now has to be able to answer: which of our problems are actually worth frontier pricing? "When the best model costs double, "just use the best one" stops being a default and becomes a decision." I've written before about the failure mode where token consumption becomes a status symbol, tokenmaxxing. Fable 5 raises the cost of that vanity. The right question was never "are we using the most powerful model," it was "what does it cost us to be wrong here?" A migration that corrupts data, a security review that misses the hole, an overnight agent run that has to be redone by a human in the morning: those are worth ten-dollar tokens. Reformatting a support ticket is not. They took away the knobs Here's the part that fascinates me more than the price. Fable 5 doesn't have a temperature setting. No top-p, no top-k, no fixed thinking budget measured in tokens. Those parameters don't exist on this model. Send them and the API rejects the request. You can't even explicitly switch its reasoning off; the model decides for itself when a problem deserves thought and how much. What you get instead is a much smaller set of controls, and look at what they are. An effort level: how hard should this attempt try, from quick-and-cheap up to spare-no-expense. A task budget: here's roughly how many tokens this whole job is worth, and the model watches its own countdown and prioritizes accordingly. That's the entire interface. If that vocabulary sounds familiar, it should. It's how you delegate to a senior person. You don't regulate a principal engineer's brain chemistry, and you don't hand them a step-by-step script. You tell them what the outcome is, how much it matters, and how much of their time it's worth, and then you let them work. The API now asks for the same three things a manager asks for: what you want done, how hard to push, and how much it's worth spending. "You don't set its temperature anymore. You set its appetite." I find this genuinely clarifying, because it settles an argument I keep having. There's a persistent instinct in engineering organizations to treat the model as a component to be tuned: find the magic parameters, wrap it in enough scaffolding, micromanage every step. Each generation of these models punishes that instinct a little more. The guidance that comes with this tier is the same advice I give about delegating to people: give the full specification up front, state what done looks like, set the budget, and get out of the way. Vague asks dribbled out over many small corrections burn tokens and produce worse work. Sound familiar? Route work like a manager, not a fan So what do you actually do with a four-tier lineup? You staff it the way you staff a team. You don't put your most senior engineer on every ticket, not because they couldn't do the work, but because it's a waste of the scarcest thing you have. The same logic now applies, with a price list attached. - Haiku-class work: high-volume, mechanical, cheap to verify. Classification, extraction, routing. Being occasionally wrong is recoverable and obvious. - Sonnet-class work: the everyday middle. Most coding tasks, most drafting, most internal tooling. Fast feedback loops catch the mistakes. - Opus-class work: hard problems with real stakes. Long agentic runs, serious refactors, analysis your team will act on. - Fable-class work: the small set of problems where the cost of being wrong dwarfs the cost of the tokens, and where nobody is watching closely enough to catch a subtle miss. Notice that the routing criterion is never "how impressive is the task." It's the cost of error and the cost of verification. Work that's cheap to check can run on cheap intelligence, because your checks are the safety net. Work that's expensive to check (overnight runs, subtle judgment calls, anything reviewed by a tired human at 9 a.m.) is exactly where paying double for fewer mistakes is the bargain of the year. I learned this lesson the hard way running fleets of coding agents: the bottleneck was never the agents, it was my capacity to verify what they produced. What this means for your organization First, AI spend is graduating from a rounding error to a line item with structure, and that's healthy. A budget with tiers forces the conversation that a flat budget lets you skip: which problems are worth what. If your organization can't answer that question for tokens, I'd gently suggest it couldn't answer it for engineering hours either, and the tokens are just making an old problem visible. Second, the skill that compounds is the one I keep coming back to: shaping. A model that takes a full, well-specified goal and runs with it for hours rewards exactly the discipline most organizations lack: deciding what the outcome is, what it's worth, and what's explicitly out of scope before the work starts. The organizations getting the most out of these models aren't the ones with the cleverest prompts. They're the ones whose leaders can shape a problem crisply enough to hand it to anyone, human or model, and bet on the result. And third, stop waiting for this to settle. The ladder will keep growing rungs, top and bottom, and the prices will keep forking. The durable capability isn't familiarity with any one model. It's an operating model that can route work by value, verify cheaply, and treat intelligence, artificial or otherwise, as a portfolio to be managed rather than a status symbol to be maxed. That operating model is the actual work of an AI-native transformation, and it's the work I do. If your AI bill is growing faster than your confidence in what it's buying, let's talk → ### AI Adoption Phase 3, Operationalize: Agent to Employee: /blog/phase-three-operationalize/ A third of companies have an agent doing real work in production. Most of them built a heroic one-off: brilliant, fragile, and understood by exactly one engineer. That's not a capability. It's a liability with good PR. This is the phase where I stop teasing companies, because getting here is real. Per the funnel in my series opener, the research finds only about a third of organizations have begun to scale AI beyond experimentation, even though 62% are experimenting with agents. That gap, between trying agents and having one do actual work in production, is this phase's entry fee. If that's you, something genuinely shipped: a system reads the ticket and resolves it, or clears the invoice without a human in the loop. Money or time changes hands because of a model's decisions. And yet the Operationalize phase has its own trap, and most companies who reach it are already in it. The agent works, but nobody can say how well it works. It was built by one talented engineer in a burst of enthusiasm, it lives in a repo nobody else touches, it has no evals, and the only guardrail is hope. Nobody has written down what happens when it goes wrong. It is, in other words, a heroic artifact. And heroic artifacts don't compound, they accumulate risk while everyone congratulates themselves. The one-off trap Here's the test I run in phase-three companies. I ask three questions about their proudest agent. One: what's its error rate this week, and is that better or worse than last week? Two: if its model provider upgrades the underlying model tomorrow, how will you know whether the agent got better or worse? Three: if the engineer who built it resigns on Friday, who runs it on Monday? Silence on all three, almost every time. The agent is in production but not operated. It's the difference between owning a truck and running a logistics company. And the silence has a consequence that compounds quietly: because nobody trusts the agent in a way they can defend with numbers, nobody dares to extend it, replicate the pattern, or point it at anything more important. The organization's one success becomes its ceiling. "An agent in production but not operated is the difference between owning a truck and running a logistics company." Manage agents like employees, because that's what they are The mental model that gets companies out of this trap is one I've argued before: design agents like stupid employees. Not stupid as an insult, stupid as a design constraint. A new hire who is fast, tireless, occasionally brilliant, and prone to confidently making things up. You would never onboard that person by giving them root access and walking away. You'd give them a narrow job description, examples of done-well, boundaries on what they can decide alone, and an escalation path for everything else. Then you'd review their work, frequently at first, then by sampling. Every piece of that maps directly onto agent architecture: - Job description → a narrow, explicit scope. Agents fail in proportion to the vagueness of their mandate. "Handle refunds under $200 matching these policies" works; "handle customer issues" is an incident waiting for a timestamp. - Probation review → an eval suite. A library of real cases with known-good outcomes, run continuously, so quality is a number on a chart instead of a vibe in a standup. - Decision boundaries → guardrails in code. What the agent may do autonomously, what requires human sign-off, and what it must never touch, enforced by the system, not the prompt. - Escalation path → an explicit uncertainty route. The single highest-leverage design feature of any agent is knowing when to say "this one's not for me" and hand off cleanly to a human. - Manager → a named owner in the org chart. Not "the AI team." A person whose job includes this agent's output, the same way a team lead owns a junior's output. When companies tell me their agent "isn't reliable enough for production," the problem is almost never the model. It's that they built an employee with no job description, no manager, and no performance review, and they're surprised it behaves like one. Testing the non-deterministic The engineering discipline underneath all of this is evaluation, and it deserves its own emphasis because it's the thing phase-three companies most consistently skip. Traditional software testing assumes determinism: same input, same output, green check. Agents broke that assumption, and most teams responded by quietly giving up on testing rather than changing how they test. The answer, which I detailed in testing non-deterministic agents, is statistical: golden datasets, scored rubrics, pass-rate thresholds, regression runs on every prompt change and every model bump. Evals are not a quality nicety. They are the asset that makes everything after this phase possible. Your eval suite is what lets you upgrade models the week they ship instead of the quarter after, swap providers when the price drops, and widen the agent's scope because the numbers say you can, not because someone felt brave. Companies in the Industrialize phase, the ones scaling, are running on eval suites the way traditional companies run on accounting. No books, no business. "Your eval suite is the asset. The agent is just this quarter's expression of it." Exit criteria: from artifact to pattern You've exited the Operationalize phase when the first agent stops being a story and becomes a template. Concretely: the second agent took a fraction of the time of the first because it inherited the scaffolding, evals, guardrails, escalation, observability, instead of re-inventing it. There's a named owner for every agent. Error rates are charted, reviewed, and tied to an error budget someone agreed to. And, the most telling sign, the line organization is asking for the next agent, with a workflow already in mind, rather than the AI team hunting for use cases. What this looks like when I do it with you Operationalize is the most hands-on stretch of an AI-native engagement, and it's where the white-glove part is most literal: I'm in the architecture, in the code review, and in the eval design. The work is converting your heroic artifact into a managed pattern, building the eval harness it never had, wrapping it in real guardrails and escalation, putting observability on it, and then extracting the scaffolding so the second and third agents are weeks, not quarters. The other half of the work is organizational, and it's the half nobody budgets for: deciding who manages the agents. Someone in your line organization is about to become a manager of non-human workers, reviewing samples of their output, tuning their boundaries, owning their mistakes. I've argued that every engineer is a manager now, and this phase is where that stops being a metaphor and starts being someone's actual Tuesday. Choosing those people, training them, and adjusting what success looks like for them, that's operating-model work disguised as an engineering project, and it's exactly the seam where I work. Next in the series: what happens when the agents multiply and the real costs show up, Industrialize: scaling agents without scaling the chaos. And if your proudest agent just failed my three questions, let's talk → ### How to Work Backwards From the Customer Experience: /blog/start-with-the-experience/ Design is not how it looks, it's how it works. Part 2 of my product design philosophy: the best interface is no interface, and the parts users can't see should be as beautiful as the parts they can. Most products are built inside-out. The team starts from the database schema, the vendor's API, or whatever the framework makes easy, then decorates the result with an interface and calls the decoration "design." You can always tell. The product works the way the system is organized, not the way the user thinks. This is Part 2 of a four-part series on my product design philosophy. The rest: Part 1: If Your Product Needs a Manual, the Design Already Failed, Part 3: Nobody Asks for the Future, and Part 4: Saying No to 1,000 Good Ideas. The fix is a sentence Jobs said to developers in 1997 and most teams still haven't absorbed: start with the customer experience and work backwards to the technology. Decide what the user should feel at every step, then bend the architecture until it delivers that. The experience is the spec. The technology is the implementation detail. Design is how it works "Design is not just what it looks like and feels like. Design is how it works." Everyone quotes this line; almost nobody staffs like they believe it. If design is how it works, then your designers belong in the architecture conversation and your engineers own the user experience. Splitting "what to build" from "how to build it" produces exactly the inside-out products I just described, which is why I argue product and engineering were always one job. In practice, working backwards looks like this: write the experience first, in plain words. "The operator pastes a tracking number and the order is reconciled before they can switch tabs." That sentence is a design document. It sets a ceiling on latency, dictates how errors get handled, and decides how much of your vendor's clunky API you're allowed to expose. Every technical decision after it either serves that sentence or gets rejected. The best interface is no interface The highest form of this philosophy: make the technology invisible. Every screen, button, and confirmation dialog is a cost the user pays to get what they actually wanted. The best interface isn't a cleaner version of that cost. It's the step that disappears entirely. The file that syncs without a sync button. The payment that happens when you walk out. The report that arrives before anyone asks for it. Users don't describe these as good interfaces; they describe them as the product "just working," which is the entire point. The best parts of a product are the ones nobody thinks to praise, because they never had to think about them at all. "Every screen you don't build is a screen nobody has to learn." This is also where AI genuinely changes the design vocabulary. For thirty years, "make it invisible" was aspirational because software couldn't infer intent; it had to ask. Now it can read context, draft answers, and complete whole steps on its own. Which means interfaces that interrogate users with forms aren't just dated, they're a choice. The new design question isn't "how should this screen look," it's "why does this screen exist?" Magic is engineered, not sprinkled "Every interaction should feel magical and delightful" sounds like marketing fluff until you reverse-engineer what magic is made of. It's a response under a hundred milliseconds, and defaults so well chosen the user never opens settings. It's the empty state that teaches without a tutorial, the error message that says what to do next, the form that remembers what it asked you last time. None of that is delight-as-decoration. All of it is engineering. Perfection in details matters, obsess over every pixel and every transition, not because users consciously notice each one, but because they unconsciously notice all of them at once. The sum has a name: the product feels either crafted or assembled. Users can't tell you which details produced the feeling, but they can tell you the feeling within ten seconds. The parts you can't see Jobs got this from his father, who refused to build a fence with an ugly back even though no one would see it. The software translation: the parts users can't see should be as beautiful as the parts they can. The schema, the service boundaries, the naming, the test suite, the deploy pipeline. Quality must go all the way through. I hold this position for a thoroughly unsentimental reason: in software, the back of the fence doesn't stay hidden. It leaks. The messy schema becomes the slow query becomes the spinner the user stares at. The tangled service boundary becomes the feature that takes a quarter instead of a week. The codebase nobody is proud of becomes the team that stops arguing for the user, because people don't fight for craft in a place that has none. "Users can't see your architecture, but they can feel it." You cannot bolt a magical experience onto an ugly core. The experience is the part of the iceberg above the water; its shape is determined by everything underneath. Next in the series: why customers can't ask for the future, and why you have to build it anyway. And if your product was built inside-out and you can feel it, let's talk → ### How to Run Ten Coding Agents in Parallel on One Laptop: /blog/ten-coding-agents-one-laptop/ Running six to ten coding agents at once was never a model problem. It's an environment problem, and the moment you solve it, the constraint moves to the one thing you can't refactor: the RAM on your desk. Right now, as I write this, there are ten coding agents running on my laptop. Not ten browser tabs with a chatbot in each, ten agents, each on its own branch, each building a different slice of the same product, each able to run that product and prove its own work before it hands it back. You'd think the interesting part is the models, or some clever orchestration trick, but it really isn't. The agents were never the hard part. When people hear "run six to ten agents at once," they picture the easy version: a row of chat windows, each spitting out code. That part is genuinely easy. Ten agents editing files is a solved problem. The wall you hit, the one nobody shows you in the demo, is that an agent editing code is useless until it can run the thing it changed. And the moment two agents try to run the same product on the same machine at the same time, the whole illusion of parallelism falls apart. Parallel agents, one shared everything Most setups break in the same spot. You point every agent at one local environment: one dev server, one database, one queue, one set of ports. Agent A runs a migration to test its feature, and Agent B's tests start failing against a schema it never asked for. Agent C boots the web app and grabs port 3000, and the other nine line up behind it. Two agents seed the same database with conflicting fixtures, and both of their test runs lie to them. The work fanned out, but the environment didn't. So you aren't really running ten agents in parallel. You're running one environment, single-file, with ten agents elbowing each other for a turn, and paying for it in race conditions you now have to debug. It's actually worse than running them one after another, because now you also have the coordination overhead on top. "Ten agents editing code is easy. Ten agents that can each run the product is the whole problem." The unit of isolation is the worktree The fix starts with git worktree. A worktree gives each branch its own working directory on disk, same repository, separate checkout, so ten agents can hold ten different states of the code at once without touching each other's files. I run all of this inside Conductor, a desktop app that spins up worktrees and keeps them organized, so I'm orchestrating agents instead of hand-managing a tangle of directories and branches. But a worktree only isolates the code. The running product, the web app, the API, the database, the queue, the workers, still has to live somewhere, and by default every worktree's stack wants the same ports and the same database as every other one. Isolating the files gets you part of the way and then leaves you stuck. You have to isolate the runtime too. Make the environment ephemeral, and parameterized So I made the environment as disposable as the branch. The product already ran in Docker Compose; the change was to stop treating Compose as a single fixed thing and make it parameterized. I wrap docker compose in a shell script that derives, from the worktree's branch name, the public ports it binds and the names of every container it starts. Branch feature/billing gets its own ports and its own named stack; feature/search gets a different set. Nothing collides, because nothing shares a name or a number. Then I told Conductor, for this one project, to run that script automatically whenever it creates a new worktree. I start a new piece of work and a complete, isolated, running copy of the entire stack comes up with it, no checklist, no "wait, which port was this one." The environment is born with the branch and dies with it. - Public ports derived from the branch, so two stacks never fight over the same one. - Container names matched to the worktree branch, so each stack is legible at a glance and nothing clobbers anything. - Isolated data, so no two worktrees ever share a database or a queue. - The whole stack per worktree (web app, API, database, queue, workers), not a stripped-down stand-in. The hard part was the tests (and I was making dinner) The dev server was the easy half. The genuinely hard part, and the part that decides whether any of this is worth doing, was getting the full test suite to run inside that same isolated Compose. Both layers: the fast unit tests and the slow, stateful end-to-end ones. A green checkmark is supposed to mean the code works, but that depends entirely on what it ran against. A green run against a database nine other agents are mutating means nothing. A green run against this worktree's own database, its own queue, its own services that no other process can touch, is a result you can actually trust. How it got built is the whole point. I didn't sit there hand-wiring test harnesses into Compose. I shaped the problem, every agent verifies its own work against its own stack, no shared state, unit and end-to-end both, handed it to an agent, and went to make dinner. It did the fiddly, unglamorous plumbing while I cooked. That's what the job looks like now. You describe the outcome precisely enough that you can stand behind it, and then the building happens while you're off doing something else. "A green test run only means something if it ran against an environment no one else could touch." Isolation is what makes the parallelism honest With that in place, the parallelism actually does what it claims to. Each agent owns a worktree, a full running stack, and a test suite that proves its slice works in isolation. They stopped being ten autocomplete windows and became ten teammates who can each build something, run it, test it, and hand back work that has actually been checked. This is the part that never makes the highlight reel. The demos show the swarm of agents and the code flying by, and they never show the environment underneath, because it isn't photogenic. But it's what decides whether running many agents is real engineering or an expensive way to generate merge conflicts. It's the unglamorous layer beneath every engineer becoming a manager of agents, and beneath the "staff specialized sub-agents" and "build the entire testing pyramid" pillars of professional vibe coding. The agents get all the attention, but the environment is where the actual work lives. Then the constraint moves, to the one thing you can't refactor And here is where it gets funny, in the way this whole era keeps being funny. Every time you knock down a constraint, the bottleneck doesn't go away, it just moves somewhere else. Building got cheap, so shaping the problem became the scarce skill, and once change got cheap the thinking moved back upstream. I knocked down the environment constraint, and the bottleneck sailed straight past every layer I know how to refactor and landed on physics. - 6-10: coding agents in parallel - ~4GB: memory per worktree stack - 16 / 32GB: used on an M1 Pro Each one of these per-worktree stacks costs me about four gigabytes of memory. The web app, the API, the database, the queue, the workers, multiply that by every branch I want alive at once and the number climbs fast. I'm sitting at sixteen gigabytes used and I can feel the ceiling getting close. My machine is an M1 Pro with thirty-two gigabytes, which until recently felt like plenty. The thing throttling how many agents I can run is no longer the model, the tooling, or my process. It's the silicon soldered to the board in front of me. "You can shape a problem. You can't shape a memory chip you already bought." Where this goes next I haven't solved this part yet, so I'll tell you what I'm actually weighing rather than pretend I have a tidy answer. - Shrink the footprint per worktree. Does every branch really need all five services running, or can the heavy, stateless ones be shared while only the parts a change actually touches spin up fresh? - Share the stateful pieces, namespace the data. One database and one queue across worktrees, with a schema or namespace per branch, trading a little isolation for a lot of memory. - Get the environments off the laptop. Keep the worktree local but run the heavy stack on a remote or cloud dev environment, the way ephemeral preview environments already work for deploys. - Buy more RAM. The least clever option, and sometimes the right one. When the constraint is hardware, the cheapest fix is occasionally just hardware. Each of those trades something away, isolation, simplicity, or money, and I don't yet know which trade I'll make. That's genuinely where I am. Ask me in a month. But notice the shape of the thing, because that's the real lesson and it has nothing to do with Docker. Going AI-native never removes the bottleneck. It keeps moving it, and the job is to keep finding where it went. For me, today, it's the memory in my laptop. Next month it'll be something else, somewhere I'm not looking yet. Chasing that down is the actual work, far more than the agents or the tooling ever were, and it's the work I do as an AI-native leader. If you're running agents in parallel too, I want to know where your constraint moved. Let's talk → ### Product Design Principles: If It Needs a Manual, It Failed: /blog/design-failed-if-it-needs-a-manual/ Simplicity is the ultimate sophistication. Part 1 of my product design philosophy: eliminate ruthlessly, question every assumption, and treat every "how do I…" as a bug report. I didn't invent my product design philosophy. The bones of it are Steve Jobs', and before him Leonardo's: simplicity is the ultimate sophistication. What I can claim is twenty-five years of testing it against actual products, teams, and budgets, and it has held up. This series is that philosophy, written down the way I actually use it. This is Part 1 of a four-part series. The rest: Part 2, Start With the Experience, Part 3, Nobody Asks for the Future, and Part 4, Saying No to 1,000 Good Ideas. Simplicity is a destination, not a starting point People hear "keep it simple" and think it means doing less work. It's the opposite. Anyone can ship complicated; complicated is what naturally happens when a team builds under deadline. Every first draft of every product I've ever seen was complicated, including mine. Simplicity is what's left after you've understood the problem deeply enough to throw most of your solution away. That's why I treat simplicity as an engineering outcome, not an aesthetic preference. A simple product means somebody did the hard thinking so the user doesn't have to. A complicated product means the team shipped its org chart and its unresolved arguments straight into the interface. The manual test Here's the cleanest quality bar I know: if users need a manual, the design has failed. Not "the documentation team has work to do." Failed. A manual is a confession that the product couldn't explain itself. And the test applies everywhere, not just to consumer apps. An internal admin panel that needs a manual has failed the same way a consumer app would. So has the API that needs a two-hour onboarding call, and the build that new engineers can't run without a wiki page. I push this hard because I've watched companies bleed margin through tools their own operators can't use while everyone shrugged and said "it's just internal." "Every "how do I…" message is a bug report filed against your design." When someone asks how to do something in your product, the instinct is to answer the question. The discipline is to also log it, because the question itself is the defect. Products should be intuitive and obvious. "Intuitive" isn't a compliment users give you; it's the absence of all the questions they never had to ask. Eliminate, don't organize When a product gets confusing, the common response is to organize the complexity: better navigation, better grouping, a settings page with tabs. That's tidying the clutter instead of removing it. The Jobs move, and the right one, is elimination. The unnecessary button or option doesn't get filed under a tidier menu, it gets deleted. My least favorite artifact in software is the settings page that grew for five years. Most settings exist because two people on the team disagreed and nobody had the authority to decide, so the user got handed the argument. Every toggle is a decision you outsourced to someone who paid you specifically so they wouldn't have to make it. Elimination hurts because everything you'd cut is there for a reason. Some customer asked for it. Some deal closed on it. That's exactly why subtraction is a leadership job, not a design task: the designer can see what should go, but only leadership can absorb the cost of removing it. I'll go deeper on saying no in Part 4, because it deserves its own essay. Question every "should" The second half of this philosophy is the part people skip: challenge every assumption about how things "should" work. Every product category carries inherited conventions, and most of them are fossils, somebody else's old constraint that outlived the constraint. Invoices look like invoices because paper was 8.5 by 11. Dashboards open with charts because executives once got printed reports. Signup asks for a company name on step one because a salesperson needed it for routing in 2009. None of these are laws. When I review a product, the question I ask most isn't "does this work?" It's "why is this here at all?" The answer "that's how it's done" is not an answer; it's an invitation. Break from conventional wisdom when necessary, and necessary means whenever the convention serves the category's history instead of your user's next thirty seconds. Why this matters more in 2026 Here's what makes a decades-old philosophy urgent now: building got cheap. When building gets cheap, adding a feature costs almost nothing, which means bloat is now the default failure mode. Engineering cost used to filter out the marginal feature for you. Now that it doesn't, the only thing standing between your product and endless bloat is your own willingness to say no. "When features are free, restraint is the product strategy." Teams that don't internalize this will ship more than they ever have, and most of it won't matter. The ones that win will use cheap building the other way: to afford the three rewrites it takes to make something genuinely simple. Next in the series: start with the experience and work backwards to the technology, including why the parts of your product nobody sees still have to be beautiful. And if your product is failing the manual test right now, let's talk → ### AI Adoption Phase 2, Experiment: From Pilot to a Production Decision: /blog/phase-two-experiment/ Two-thirds of companies are running AI pilots. Most pilots are built to demo, not to ship, and a pilot without a production path is just an expensive way to postpone a decision. Every company stuck in the Experiment phase has the same artifact: a demo that kills in meetings. It reads a contract and drafts a passable reply, or it triages a support ticket on cue. Executives see it and say some version of "this is going to change everything." And then it changes nothing, because eleven months later it is still a demo, kept alive by an innovation team and shown off once a quarter, with not a single real workflow running through it. Welcome to pilot purgatory. Per the numbers in my series opener, the research finds 62% of organizations at least experimenting with AI agents, and the overwhelming majority of those experiments will never ship. MIT's GenAI Divide study put a number on it: roughly 95% of enterprise GenAI pilots deliver no measurable P&L impact. The technology mostly works; the pilot was just never built to do anything beyond impress a room. Pilots are built to demo, not to ship The defining property of a purgatory pilot is what's missing. Look under the hood of the typical one and you find: no evaluation harness, no error budget, no data pipeline that survives contact with production permissions, no owner in the line organization, and, most damning, no definition of done. It runs on a curated sample of inputs, the happy path, demonstrated by the person who built it. None of that is an accident. The pilot was commissioned to answer the question "is this possible?", and demos answer that question beautifully. But "is this possible?" stopped being the interesting question two years ago. With modern models, almost everything in the demo tier is possible. The questions that matter now are: is it reliable at the 95th percentile? What does it cost per case at real volume? Who owns it when it's wrong? A demo answers none of those, and a pilot designed as a demo can't be upgraded into one that does, it has to be rebuilt. ""Is this possible?" stopped being the interesting question two years ago." There's a structural reason this keeps happening: pilots live in innovation teams, and innovation teams are graded on demonstrations, not operations. The line organization, which owns the actual workflow, was never asked whether it wants this thing, never gave up budget for it, and quietly regards it as a threat or a toy. So the pilot has no landing zone. It orbits the org chart indefinitely, fully funded and going nowhere. The MIT research backs this up from the data: the pilots that cross the divide are overwhelmingly the ones driven by line managers who own the workflow, not by a central AI lab. The purgatory economics Purgatory has a seductive economics. Each individual pilot is cheap, a couple of people, a few months, API credits. Killing one is awkward, so nobody does, and continuing all of them costs less per quarter than the political fight of forcing one into production. The portfolio grows. The aggregate spend quietly becomes enormous, and the return is a slide titled "AI Initiatives: 23 Active Pilots" that the board mistakes for momentum. Here's the reframe I push on executives: a pilot is not a project, it's a bet with a kill criterion. The language comes from the operating model I described in when building gets cheap: shape the problem, set an appetite, and decide in advance what evidence would make you ship it and what evidence would make you kill it. A pilot allowed to run past its appetite without a verdict has stopped being research. You're just paying to feel like something is moving. "A pilot allowed to run past its appetite without a verdict is a subscription to the feeling of progress." Anatomy of a pilot that graduates The pilots that escape purgatory look different from day one. After taking enough of them through, the pattern is mechanical: - It starts in the line organization. The team that owns the workflow runs the pilot inside their real daily work, with the innovation function supporting, not owning. - It's evaluated on production data from week one, the messy real inputs and the people actively trying to break it, not a curated golden set. - The eval harness is built before the prompt. You cannot improve, or even honestly describe, what you don't measure, and an eval suite is also the regression net you'll need for every model upgrade afterward. - It has a number: the baseline cost or cycle time of the current process, and the threshold the pilot must beat. "People liked it" is not a number. - It has an appetite and a verdict date. On that date it ships, it's killed, or in rare justified cases it gets one explicit extension. Orbiting is not an outcome. - The production path is designed up front: where it runs, who's paged, what the human-escalation route is, what the rollback is. Notice that a pilot built this way is barely a pilot at all. It's the first iteration of a production system, scoped small. That's the real exit from the Experiment phase: stop building disposable demonstrations and start building small, real things. The distance from "pilot" to "production" should be a deploy, not a rebuild. Kill more, ship more The counterintuitive metric of a healthy pilot portfolio is the kill rate. Purgatory companies kill almost nothing, everything stays "active." Healthy companies kill the majority of their pilots, quickly and without ceremony, because a fast, cheap, well-documented kill is a successful outcome: you bought certainty for the price of a few weeks. The teams that win let themselves fail early and cheaply, inside boundaries that make any single failure easy to absorb. If your AI program has never killed a pilot, it isn't a program, it's a museum. What this looks like when I do it with you The Experiment phase is where my engagement shifts from audit to build and ship. The white-glove part is that I don't hand you a pilot framework and wish you luck, I take the two or three workflows the audit ranked highest and personally drive each through the anatomy above: eval harness first, production data from week one, a kill-or-ship date that I hold everyone to, including the executive sponsor who would rather extend than decide. I also do the quiet political work that determines whether any of this lands: moving the pilot out of the innovation silo and into the line team's hands, negotiating what the workflow owner gets out of it, and making sure the first shipped pilot is visible enough that the next one is pulled by the organization rather than pushed. The technology has almost never been the blocker in these engagements. The blocker is that production is a commitment, and organizations are built to defer commitments. My job is to make deferral more expensive than a decision. Next in the series: what happens after the pilot ships, when you discover that an agent in production is less like software and more like a new hire, Operationalize: you built an agent, now make it an employee. And if you recognize your own pilot portfolio in this essay, let's talk → ### AI Adoption Phase 1, Equip: You Bought Tools, Nothing Changed: /blog/phase-one-equip/ 88% of companies have bought AI tools. Most of them mistook the purchase order for the transformation. Procurement is not adoption, and seat counts are not leverage. There is a meeting that happens in almost every company right now. Someone presents a dashboard: AI tool licenses purchased, seats activated, weekly active users trending up. Everyone nods. The company is "doing AI." The meeting ends. Nobody in that meeting can answer the only question that matters: what changed about how the work gets done? This is the Equip phase, and per the latest enterprise AI surveys, 88% of organizations are in it, using AI in at least one business function, up from 78% a year earlier. It's the most crowded phase of the funnel because it's the easiest to enter. Buying tools requires a budget line and a signature. It requires no decisions about workflows, no arguments about process, no one to change how they spend their Tuesday. That's exactly why it produces nothing on its own. This is the second essay in my five phases of AI adoption series, and Equip deserves the first deep dive because the failure here is the cheapest to fix and the most expensive to ignore. Procurement theater I call what happens in the Equip phase procurement theater. The company performs the motions of adoption, contracts, rollouts, lunch-and-learns, an internal Slack channel called #ai-tips, and mistakes the performance for the thing itself. The tell is what gets measured. Equip-phase companies measure inputs: licenses, activations, prompts per user. None of those numbers connect to a workflow, a cycle time, a cost, or a customer. I've written before about why token consumption is not productivity, and the Equip phase is where that confusion is born. Usage is the easiest thing to measure and the least meaningful. A developer can burn ten thousand tokens a day on autocomplete and ship exactly what they would have shipped anyway, a little faster, with the saved minutes dissolving into the same meetings as before. "Tools don't change how people work. They change how fast people do the work they already feel safe doing." Drop powerful tools into unchanged workflows and you get the old behavior at a slightly higher clock speed. The handoffs and approval chains haven't moved. The report that takes four people three days still takes four people, because the bottleneck was never typing speed, it was the process wrapped around the typing. Why smart companies stall here It's tempting to be smug about the 88%, but companies stall in Equip for reasons that are locally rational. - Buying tools is a decision one executive can make alone. Changing a workflow requires agreement between several, and nobody owns the seam between departments. - Tool rollouts produce immediate, reportable numbers. Workflow change produces awkward questions for two quarters before it produces results. - "Give everyone access and let a thousand flowers bloom" feels democratic and safe. Picking one workflow to transform means someone's territory gets redesigned, and someone has to be accountable if it doesn't work. - The vendors are selling seats, so every piece of collateral the organization consumes equates adoption with seat count. The result is a strategy of horizontal coverage: a thin layer of AI spread across the entire organization, deep nowhere. And horizontal coverage has a nasty property, it generates the feeling of progress at almost exactly the rate it avoids the substance of it. The dashboard goes up and to the right while the operating model stands perfectly still. The hidden cost of standing still Equip feels safe because the spend is modest and nothing breaks. But standing still has a price, and it compounds. Every quarter spent equipping without changing, your best people are learning that "AI initiative" means nothing changes. That's cultural scar tissue, and it makes the real transformation harder later, because by the time leadership gets serious, the organization has already metabolized AI as another CRM rollout. Meanwhile the gap between you and the companies a phase or two ahead isn't static. They are accumulating evals, production patterns, and redesigned workflows, the kind of advantage that compounds and can't be purchased retroactively. You can buy their tools any day of the week. You cannot buy their two years of operating experience. "You can buy the tools any day. You can't buy the two years of operating experience." Exit criteria: how you know you've left Equip Graduating from this phase is not about more tools or better training. It's about converting diffuse access into a concentrated bet. You've genuinely entered the Experiment phase when three things are true: - You can name the specific workflows you're transforming, who owns each one, and what business number each is supposed to move. - You measure outcomes (cycle time, cost per case, error rate, revenue per head) rather than inputs (seats, tokens, prompts). - At least one pilot exists that a line team, not the innovation team, is running inside its real, daily work. Notice that none of these are technical. The exit from Equip is a leadership act: someone with authority picks a narrow front, frames the outcome, and accepts accountability for it. That's the vertical-slice principle from the series opener: rebuilding one workflow end to end gets you further than sprinkling AI across a hundred of them. What this looks like when I do it with you This stage of an engagement is the AI-Native Audit, and it's deliberately unglamorous. I sit inside your actual workflows, sales ops, engineering, support, finance, and map where the hours actually go, where AI creates real leverage, and where it's a distraction wearing a demo's clothes. The output isn't a maturity scorecard. It's a short, ranked list of workflows worth betting on, each with a named owner, a target number, and an appetite, how much time and money that outcome is worth. White glove, at this stage, means I do the part organizations are structurally bad at: saying no. Eighty percent of the "AI opportunities" surfaced in a typical audit aren't worth pursuing yet, and an internal champion can rarely kill them without spending political capital. I can, because I've seen the same twenty ideas at the last ten companies and I know which three pay. You end the audit with fewer initiatives than you started with, and that's the point. The 88% got here by buying for everyone at once. Getting out means committing to a few workflows and ignoring the rest for now. Next in the series: the phase where good ideas go to die politely, Experiment: pilot purgatory. And if your dashboard of seat counts is starting to look like a confession, let's talk → ### The Five Phases of AI Adoption, and Where Companies Stall: /blog/five-phases-of-ai-adoption/ AI adoption moves through five phases: Equip, Experiment, Operationalize, Industrialize, Transform. Each transition fails for a different reason, and almost everyone is stalled in the first two. Ask a hundred companies where they are on AI and nearly all of them will tell you they're "doing it." They're not lying. They're just not measuring the same thing. AI adoption moves through five phases, Equip, Experiment, Operationalize, Industrialize, and Transform, and almost every company is stalled in the first two. The latest large-scale enterprise AI research, surveys covering nearly 2,000 organizations across more than a hundred countries, puts numbers on it. 88% of organizations now use AI in at least one business function. About two-thirds, 62%, are at least experimenting with AI agents. Roughly a third report they've begun to scale AI beyond the experiments. Fewer than a quarter, 23%, are scaling an agentic system anywhere in the enterprise. And about 6%, six in a hundred, qualify as genuine AI high performers: organizations that attribute 5% or more of EBIT to AI, the ones that did the deeper thing and redesigned how they work. The surveys vary on the exact percentages, but the shape never varies. MIT's GenAI Divide study found that roughly 95% of enterprise GenAI pilots deliver no measurable P&L impact. It's a funnel that collapses brutally between "we spent money on AI" and "AI changed how we operate," and almost every company I talk to is somewhere in that collapse, wondering why the spend isn't showing up in the P&L. I think about that funnel as five phases, and I've given each one a name, because naming the phase you're in is the first honest act of any transformation. This essay is the opener of a series that takes each phase apart: what it looks like from the inside, why companies stall there, and what the transition to the next phase actually demands. Because the central finding is uncomfortable: each transition fails for a different reason, and companies keep applying the fix for one phase to the problems of another. The five phases Phase one, Equip. Copilot seats, a ChatGPT Enterprise contract, an AI line item in the budget. This is where 88% of companies live, and most of them mistake procurement for progress. The tools are real, but nothing about how the work gets done has actually moved. I cover this in Equip: you bought the tools and nothing changed. Phase two, Experiment. The innovation team has a demo. It's genuinely impressive. It has been genuinely impressive for eleven months, and it has never touched a production workflow or a customer. Two-thirds of companies are here, in what I call pilot purgatory. Phase three, Operationalize. Something agentic is doing real work in production. This is genuine progress, and it's where a new failure mode appears: the agent works, but it's a fragile one-off that one engineer understands, with no evals, no guardrails, and no plan for what happens when the model underneath it changes. That's Operationalize: you built an agent, now make it an employee. Phase four, Industrialize. Agents are doing meaningful volumes of work across more than one function. Now the problems are operational: evaluation, observability, cost curves, model churn. This is where AI stops being a project and becomes infrastructure, and infrastructure has rules. That's Industrialize: scaling agents without scaling the chaos. Phase five, Transform. The 6%. The companies that stopped asking "where can we add AI to what we do" and started asking "given what AI makes cheap, how should we work?" The processes change, the roles change, the decision rights change, and the underlying economics change with them. That's the phase where the returns live, and it's the subject of the final essay in the series. "Naming the phase you're in is the first honest act of any transformation." Why the funnel collapses Each transition in the funnel fails for a different reason, and that matters, because companies keep applying the Equip-phase fix, buy more, train more, roll out more, to problems that live three phases later. - Equip → Experiment fails because nobody owns the question "what is this for?" Buying is easy; choosing a workflow to change is a decision someone has to be accountable for. - Experiment → Operationalize fails because pilots are built to demo, not to ship. No eval harness, no error budget, no owner in the line organization, no definition of done. - Operationalize → Industrialize fails because the first agent was a heroic one-off, not a repeatable pattern. You can't scale a thing that only worked because one person willed it into existence; you scale the platform underneath it. - Industrialize → Transform fails because it stops being a technology problem entirely. It's an operating-model problem, and the people who must change the model are the people the current model made successful. Notice the progression. The early failures are about focus, the middle ones are about engineering discipline, and the last one is about leadership. That's why "we hired a great ML team" doesn't get you to Transform, and neither does "we bought the best tools." The bottleneck moves as you advance, and by the end it's sitting in the executive suite. Getting that diagnosis right is the heart of any serious AI strategy. The dirty secret: the phases are not a maturity ladder Most maturity models imply you should climb one rung at a time. I want to push back on that, because it's exactly how companies waste two years. You do not need to perfect Equip before you Experiment. You do not need a hundred pilots before you ship one agent. The companies that reach Transform fastest run the phases concurrently on a narrow front: they pick one workflow that matters, take it all the way from tool to pilot to production agent to redesigned process, and let that one vertical slice teach the organization what the new operating model feels like. Then they widen. The companies that stall do the opposite, they go horizontally, rolling tools out to everyone, piloting everything, and finishing nothing. "Go vertical on one workflow that matters, not horizontal across a hundred that don't." This is the same argument I made in when building gets cheap, shaping becomes the job: the scarce skill isn't the building, it's deciding what's worth building and how much it's worth. The phases are a diagnosis, not a curriculum. How I take companies through this This series doubles as a map of the work I do as an AI-native development and operations leader, because the engagement is structured around exactly these transitions. It starts with an AI-Native Audit: a clear-eyed read of which phase you're actually in, function by function, and where AI creates real leverage versus where it's a distraction. Most executives think they're a phase ahead of where they are. The audit hurts a little. It's supposed to. Then it's hands-on. Not a slide deck and a retainer, I architect and ship the production systems with evaluation, observability, and cost control built in, while simultaneously rewiring the workflows and decision-making around them so the change survives my departure. White glove means I'm in the codebase and in the boardroom in the same week, because the phase transitions fail precisely in the gap between those two rooms. Over the next five essays I'll take each phase apart: the traps, the exit criteria, and what the work looks like when it's done properly. If you already know which phase you're stalled in and you'd rather skip ahead, that's what an AI strategy engagement is for. Let's talk → ### Tokenmaxxing: Why Token Usage Is a Bad Productivity Metric: /blog/tokenmaxxing-is-not-productivity/ Counting tokens tells you nothing about whether work got done. But rationing them to save money is the more expensive mistake. The real skill is context hygiene, and the real win is letting people experiment. A few months into every AI rollout, someone in finance discovers the usage dashboard. Then I get the call. "Engineer A burned ten times the tokens of Engineer B last month. Should we be worried?" The honest answer is: worried about what, exactly? Because that number, on its own, tells you almost nothing about either of them. I've started calling the instinct behind that question tokenmaxxing. Tokenmaxxing is the belief that tokens consumed measure work: that a high count means productivity, or that a capped count means savings. Both readings are wrong, because a usage number says nothing about the value of what it produced. It shows up in two opposite-looking forms, and each one follows from the same bad premise. One form treats high token usage as a sign of productivity, the engineer who runs the model all day must be getting more done. The other treats it as a cost to be stamped out, cap the budgets, downgrade everyone to the cheapest model, ration access until the line goes down. They look like opposites. They're the same mistake wearing different clothes: both have confused a usage number with value. We have run this exact play before If "a usage number that has nothing to do with value" sounds familiar, it should. We spent decades counting lines of code and learned, painfully, that the most productive engineer on the team is often the one who deleted ten thousand lines and shipped a smaller, simpler system. Tokens are lines of code with a fresh coat of paint. A high token count can mean deep, sustained work on a genuinely hard problem. It can also mean someone flailing in circles, re-prompting the same broken request twenty times because they never learned to set it up properly. The number doesn't distinguish between those. Neither does it distinguish the engineer who shipped a quarter's worth of value in an afternoon from the one who generated a quarter's worth of plausible nonsense nobody will ever merge. Goodhart's law isn't a theory in software; it's a Tuesday. Make tokens the metric and you'll get more tokens. You will not get more value, and you may well get less, because the people gaming the number are optimizing for the number instead of the work. "Token usage is a bad productivity metric for the same reason lines of code was: it counts motion, not progress." The rationing mistake is the expensive one Of the two failure modes, premature rationing is the one that actually costs you money, which is the irony, because saving money is its entire justification. The reasoning sounds responsible: tokens cost money, the better models cost more, so we'll cap usage and default everyone to the cheapest tier until they prove they need more. Run the arithmetic on that decision and it falls apart. A senior engineer's fully loaded hour costs you somewhere in the range of one to two hundred dollars. The difference between the cheap model and the capable one, across a heavy day of real work, is measured in single-digit dollars. When you force a skilled, expensive person onto a weaker model to save the price of a sandwich, and that model sends them down two extra dead ends before lunch, you didn't save anything. You spent an expensive hour to avoid a trivial cost. Penny-wise, hundred-dollar-foolish. And the damage isn't only on the spreadsheet. There is a real morale cost to handing a craftsperson a deliberately worse tool. People notice when the org's message is "we'd rather you struggle than spend three dollars." It reads as distrust, and it lands hardest on exactly the people you most want to keep, the ones who can tell the difference between a sharp tool and a dull one and resent being made to work with the dull one. Friction you imposed to save money on tokens, you pay back with interest in frustration and attrition. "Forcing an expensive person onto a cheaper model to save a few dollars in tokens is the most expensive saving on the books." Account honestly, then stop counting the wrong thing None of this means spend is irrelevant. It means you have to account for it like an adult instead of fixating on the rawest, least informative version of the number. Token cost is real and worth understanding, the same way cloud spend is. We don't praise the team with the biggest AWS bill or punish the one with the smallest. We ask what the spend bought. So attribute cost to outcomes, not to headcount. The useful question is never "how many tokens did this person use," it's "what did this spend produce, and was it worth it?" A feature that shipped a week early, a migration that didn't need a second engineer, a class of support tickets that stopped arriving, those are the units that matter. Cost per outcome is a real metric. Tokens per head is a vanity metric dressed as governance. - Measure the system, not the seat. Total spend against value delivered tells you whether the investment is paying off. Per-person token leaderboards just teach people to game the leaderboard. - Treat a spike as a question, not a verdict. A 10× month might be your best week of the quarter or someone stuck in a loop. The number is a prompt to go look, never a conclusion on its own. - Budget at the team level, generously. Give teams a real envelope and let them spend inside it on judgment, the way you'd trust them with a cloud budget, not a per-keystroke allowance. - Compare against the alternative cost. The benchmark for any token bill is what the same outcome would have cost in engineer-hours, contractors, or simply not happening. Context hygiene is the skill, not volume What the token-counters miss is that the relationship between tokens and value is something you can train. The reason one engineer gets more done with fewer tokens isn't that they're stingy. It's that they've learned to manage context, and most people simply never have. A trained developer keeps the model's working context clean. They clear it between unrelated tasks instead of dragging a mile-long, polluted conversation behind them. They compact deliberately when a thread gets long, so the model keeps the thread of the work without re-reading everything. They scope a request to the files and facts that matter rather than dumping the whole repository in and hoping. They write good project instructions, a solid CLAUDE.md or its equivalent, so the model starts every task already knowing the rules instead of rediscovering them at cost. They lean on narrow sub-agents with one job each instead of one overloaded conversation that has forgotten its own beginning. Untrained, the same tools quietly waste enormous amounts of context. A bloated window doesn't just cost more, it produces worse output, because the signal the model needs is buried under everything irrelevant you forgot to clear. So the engineer who never learned context hygiene gets the double penalty: a bigger bill and weaker results. The fix for that is not a usage cap. A cap just rate-limits the bad habit. The fix is teaching the skill. "An untrained operator pays twice: a bigger bill and worse output. The answer is training, not a usage cap." This is the part that should change how you think about the spend. If usage looks high, the first move isn't to ration, it's to ask whether your people actually know how to drive these tools. Context-window maintenance is a learnable, teachable discipline, and it's the single highest-leverage thing you can invest in. It lowers cost and raises quality at the same time, which no usage cap has ever done. The double-edged sword: you want them experimenting And now the part that keeps this from collapsing into pure thrift, because there's a real tension here and pretending otherwise would be dishonest. Yes, train people to be efficient. But efficiency taken too far becomes its own trap, because the most valuable token usage in your whole organization often looks, on a dashboard, exactly like waste. It's experimentation. It's the engineer trying the capable model on a problem nobody asked them to solve, the analyst seeing whether they can automate a report that's been done by hand for years, the support lead wiring up a prototype to test a hunch. Most of those attempts go nowhere, and they all burn tokens. If you've built a culture that scrutinizes every token, every one of those experiments looks like a line item to question. So people stop running them. And the experiments you just killed were the ones that produce your next real advantage. Killing those experiments narrows where innovation can come from. Once these tools are in everyone's hands, innovation stops being the exclusive property of leadership and a few key individuals. When the cost of trying something has collapsed, the best idea can come from anyone, the coordinator, the junior, the person three layers from the org chart's top who actually feels the problem every day. That broad, distributed experimentation is the entire prize of becoming AI-Native. A token budget enforced too tightly hands innovation back to the few, which is exactly the bottleneck you adopted these tools to escape. "The most valuable token usage in your company often looks, on a dashboard, exactly like waste. It's called experimentation." So the discipline cuts both ways. You train people to manage context well, and you also make it genuinely safe to spend tokens on things that might not pan out. Those aren't in conflict. An experiment, like any good bet, has a bounded, affordable cost, a few dollars and an afternoon to learn something. The team that runs ten cheap experiments and kills nine of them is not wasteful. It's doing exactly what the tools are for, and the tenth experiment pays for all the rest many times over. What to actually do Stop counting tokens as if the count meant something. It's the lines-of-code fallacy with a new unit, and it will mislead you in both directions, flattering the engineer who flails and punishing the one who experiments. Account for spend honestly at the system level and tie it to outcomes, the way you already do with every other infrastructure cost. Invest hard in context hygiene, because a trained operator costs less and produces more, and no cap can claim that. And protect a budget for experimentation on purpose, because that apparent waste is where innovation comes from now, and it comes from everyone, not just the top. The goal was never to minimize tokens. It was to maximize what your people can do, and to spread that capability as widely through the organization as it will go. Measure that instead. If you're trying to get the accounting and the culture right at the same time, that's the work I do, and I'd like to hear how you're approaching it. Let's talk → ### How AI Changes Software Economics: When Building Gets Cheap: /blog/when-building-gets-cheap/ Becoming AI-Native isn't a tooling change. It's learning to shape a problem, set an appetite, and bet on the outcome. For most of my career, the constraint was always the same: building was expensive. Not just in dollars. It was expensive in time, in coordination, in the number of people who had to sign off before anything shipped. So organizations got built around that constraint. And here is the claim this essay makes plainly: AI has collapsed the cost of turning a clear idea into working software, so the scarce skill in a software organization is no longer building. It's shaping: framing a problem as an outcome worth a fixed amount of time, and betting on it. Most of the old machinery, the estimates, the backlogs, the handoffs, existed to protect a costly act of building that is no longer costly. Then building got cheap. Not free, and not effortless. But the drop is steep, and most processes haven't caught up to it. Shape Up always made an uncomfortable claim: the hard part was never the coding, it was deciding what was worth building and how much it was worth. Now that building is cheap, that claim is simply true. The scarce skill is shaping. I help companies become AI-Native. AI has rewritten the economics of software: almost everywhere I look, the bottleneck has moved from the keyboard to the table where someone decides what's worth doing. The leaky faucets that never make a cycle Every organization is full of problems that never make it into a cycle. They keep coming back, they affect more than one person, and they have a real cost, but they're too small to shape into a project and too annoying to fix on the side of someone's desk. They're leaky faucets. Everyone walks past them. Nobody bets on them. In most organizations, those faucets aren't inside a role. They're in the gaps between roles. The coordinator who can't update a template without a developer. The internal tool only one person understands. The report rebuilt by hand every month because automating it was never anyone's project. The workaround that's been "temporary" for two years. Each one is small. Together they're a tax the whole organization pays, quietly, forever. The reason they don't get fixed isn't that they're hard. With AI, most of them are now genuinely easy to fix. The reason is that nobody ever shaped them. A leaky faucet never arrives as a pitch. It arrives as a sigh. So at best it lands in cool-down, patched in the scraps of time between cycles, and at worst it drips forever. Here's what AI changes. The faucet used to be too small to be worth a bet. That math has flipped. When building is cheap, a problem that would never have justified six weeks is now easily worth a few days, which means it deserves a real place at the betting table, not the cool-down leftovers. "A leaky faucet never arrives as a pitch. It arrives as a sigh." This is the trap I watch companies fall into when they try to "adopt AI." They buy the tools, run the training, and bolt AI onto a process still built to ration expensive engineering. The faucets keep dripping, just with a copilot watching. That's the gap a real AI strategy has to close before any tool earns its keep. Shape the outcome, not the feature The work that makes this real starts with an executive or manager doing something specific: shaping the problem. Shaping is not writing a spec, and it is not assigning a task. It's framing the problem at the right altitude. Concrete enough that a team knows what outcome they're chasing, loose enough that they own how to get there. A good shaped pitch says four things: here's the problem and who feels it, here's the appetite (how much time this outcome is worth), here's a rough idea of a solution, and here's what we're explicitly not doing. Notice what's missing: an estimate. You don't ask "how long will this take." You decide "how much is this outcome worth to us," and that appetite becomes a constraint the solution has to fit inside. Fixed time, variable scope. The faucet is worth two days of someone's attention, so the solution gets shaped to fit two days, not gold-plated into two weeks. This is the executive's real leverage in an AI-Native organization. It isn't approving tickets. It's framing problems as outcomes worth a defined amount of time, then protecting that boundary. Bet on it at the table Once a problem is shaped, leadership bets on it. The betting table is where executives and managers commit a cycle's worth of attention to a small set of shaped outcomes and, just as importantly, decline the rest. There's no grooming a backlog into eternity. A pitch that doesn't get bet on doesn't get a participation trophy. It gets dropped, and if it matters it comes back better shaped. What's new is what's now worth betting on. Because appetites have shrunk, the table can take bets it never could before. The faucets that used to be invisible are now legitimate small-batch bets. The executive who used to spend the whole session arguing over which three big projects fit the quarter can now also clear a dozen outcomes that quietly tax everyone. Hand the bet to a team that owns the outcome A bet goes to a small, autonomous team with one job: deliver the outcome inside the appetite. Not implement a spec. Own the result. They talk to the people who feel the problem, decide what to build, build it, and confirm the faucet actually stopped dripping. Done means shipped, not demoed. This is the identity shift, and it's the hard part. You can hand someone an AI assistant in an afternoon. Getting a team to believe the outcome is theirs to own, and to scope their own work down to fit the appetite rather than asking for more time, takes a lot longer. No tool does it for you. Measure outcomes, not output If you bet on outcomes, you can't grade people on output. Story points and ticket counts quietly tell everyone the old conveyor belt is still the real job. Track the things the bet was actually about: did the outcome ship inside its appetite, did the faucet stop dripping, how many people stopped working around it. Measure the problems you closed, not the tasks you completed. Make the bet repeatable A single fixed faucet is an anecdote. Becoming AI-Native means making the case that this kind of bet pays off again and again, so the table keeps room for it. Put a number on the faucet before you shape it: hours per task × frequency × people affected × loaded rate, or the cost of a tool nobody can use, or the rework errors create every month. Bet, ship, show the delta. Do that a few cycles in a row and "spend a few days fixing the small things" stops being a favor and becomes an obvious return. What you're really building The headline insight, that the gap between a problem and a solution has collapsed, is true for one person with one tool. AI-Native is what happens when an organization rebuilds its decision-making around it. You shape problems instead of writing specs, you set an appetite instead of asking for an estimate, you bet on a few outcomes instead of grooming a backlog, and you judge yourself on what you closed instead of what you shipped. "An AI-Native organization shapes problems instead of writing specs, sets an appetite instead of asking for an estimate, and bets on outcomes instead of grooming a backlog." An organization that bolts AI onto an unchanged process just rations cheap building with expensive ceremony. An organization that shapes its faucets into small bets and hands them to teams that own the outcome doesn't ration anything. It fixes problems as a matter of course. That's the real transformation, and it isn't really about AI. AI made it affordable. The work is everything around it: teaching executives to shape problems as outcomes, teaching teams to own them, and changing what you bet on so the person who hears the sigh is also the person handed the time to fix it. That work is what an AI-native transformation actually is: rebuilding the operating model, not bolting a model onto the org chart. Build an organization that bets on outcomes, and you stop needing to call a plumber. The whole organization already knows how. You cannot buy your way here Here's the part that gets skipped, because it's the part you can't purchase. If this is the organization you want, handing everyone AI tools is a fool's errand. Tools don't change how people work. They change how fast people do the work they already feel safe doing. Drop powerful tools into a culture that punishes a miss and you get the same cautious, handoff-heavy behavior as before, just executed quicker. The faucets still drip. Nobody risks shaping a bet that might not pan out, because in that culture a bet that doesn't pan out is a black mark. So the real transformation isn't technical. It's a shift in how the organization treats failure. When building was expensive, failure was expensive too, so everything got built to prevent it. Sign-offs, estimates, committees, the long chain of approval. Caution was rational when a wrong turn cost a quarter. But that same caution is exactly why the small problems never get touched: the imagined downside of a failed attempt feels larger than the very real, ongoing cost of the leak. What makes an AI-Native organization work is the opposite instinct. Failure is a good thing when it's controlled, and controlled is the operative word. This is what shaping and betting are actually for. An appetite is a blast radius. A bet is a bounded experiment with a known, affordable cost. When a bet doesn't work it gets dropped without ceremony, and that isn't waste, it's the system functioning exactly as designed. You spent a week to learn something, for the price of a week. The teams that win are the ones that fail small, fast, and often, inside boundaries that make every failure survivable. "An appetite is a blast radius. A bet is a bounded experiment with a known, affordable cost." That instinct doesn't come from a tool. It comes from leaders who draw the boundaries, who make it genuinely safe to try inside them, and who treat a well-run failed bet as a success of the process rather than a fault of the person. Get that culture right and the tools amplify it. Get it wrong and the tools just help your organization stand still faster. Give people permission to fail in ways that can't hurt them, and they will fix every faucet in the building. That, not the tooling, is the transformation. If you're trying to build this, the hardest part won't be the tools. That's the work I do when I build an AI strategy with a company, and I'd like to hear how you're handling the culture. Let's talk → ### How to Run OWASP Security Reviews With Claude Code on Every PR: /blog/owasp-security-review-with-claude-code/ A pentest twice a year tells you what was broken months ago. Wire Claude Code's GitHub Action into your pipeline for everyday code review, then add a second, explicit security step that reads every diff against the OWASP Top 10 and blocks the merge when it finds a hole. Security stops being an event you survive twice a year and turns into something your pipeline checks on every commit. Most companies still treat security like a fire drill. Once or twice a year a pentest lands, a report full of findings arrives, and an engineer who has long since moved on to other work gets a ticket about a vulnerability that shipped to production three months ago. Everyone agrees security matters. Nobody can honestly say it's continuous. The audit is a snapshot of how exposed you were last quarter, delivered too late to do anything but clean up. The reason security lived at the end of the process instead of inside it was never philosophical. It was economic. You cannot put a security engineer on every pull request. There aren't enough of them, they cost a fortune, and nobody good wants to spend their week reading the four-hundredth diff. So review got rationed: saved for the big releases, and the rest of the time you mostly hoped. That constraint just disappeared. An AI reviewer can read every change, against a real security standard, every single time, in the minutes between opening a pull request and merging it. And the standard you'd hold it to already exists, in public, for free. OWASP is the checklist nobody keeps in the room OWASP, the Open Worldwide Application Security Project, has spent two decades writing down, openly, exactly how web software gets broken. The OWASP Top 10 is the short list every engineer half-remembers. The Application Security Verification Standard is the long, specific one. The Cheat Sheet Series tells you how to get each control right. Between them you have a mature, industry-agreed catalog of the ways an application fails its users, and you didn't have to write a word of it. The problem was never that the standard didn't exist. It's that the standard lives in a PDF, and the PDF is not in the room at 5pm on a Friday when someone is merging a fix to get the release out. Knowing the OWASP Top 10 exists has never once stopped an injection bug from shipping. The catalog only does work if something reads every change against it, at the moment the change is made. So that's the whole move: take OWASP out of the PDF and put it where the code actually changes. - A01 Broken Access Control: can a user reach data or actions that aren't theirs? - A02 Cryptographic Failures: is sensitive data protected in transit and at rest, with algorithms that aren't a decade out of date? - A03 Injection: SQL, command, and the rest. Is untrusted input ever allowed to become code? - A04 Insecure Design: is the flaw in the shape of the thing, not just a slip in the implementation? - A05 Security Misconfiguration: defaults left on, headers missing, a storage bucket quietly public. - A06 Vulnerable and Outdated Components: does this change pull in a dependency with a known CVE? - A07 Identification and Authentication Failures: weak sessions, guessable resets, login you can walk past. - A08 Software and Data Integrity Failures: unverified updates, untrusted deserialization, a poisoned build pipeline. - A09 Security Logging and Monitoring Failures: if the attack happened, would you ever see it? - A10 Server-Side Request Forgery: can the server be tricked into fetching something it has no business touching? None of that is exotic. Every engineer has met every item on the list. What no team has ever had is someone with the patience to check all ten against every diff, forever, without getting bored or going home. That's the job that just became automatable. Step one: let the reviewer review Claude Code, the tool I build with and the one I'll talk about here, ships an official GitHub Action. Wire it into your repository and it runs inside the CI you already have, triggered when a pull request opens or gets new commits, authenticated with an API key you keep as a repository secret. It reads the diff and the code around it and leaves review comments on the pull request the way a colleague would. Point it at your repo's CLAUDE.md and it picks up your project's conventions and context for free. That first layer is ordinary engineering review: correctness, clarity, design, the things a strong senior would flag. Run it on every pull request. It's genuinely useful, and on a fast-moving team it catches a great deal before a human ever looks. But here is the trap, and it's the whole reason this post exists: do not assume that a reviewer looking at everything will also catch the security holes. It won't, not reliably, and it fails for the same reason your humans do. Step two: security gets its own step, with one job A reviewer asked to judge correctness, style, performance, design, and security all at once will do all of them at the depth of none of them. Attention is the scarce resource, for a model exactly as for a person. The fix is the same one I use everywhere with AI: don't build one overloaded generalist, build a narrow specialist with a single job. I've argued you should treat agents like very stupid employees, and the kindest, most effective thing you can do for a stupid employee is give them exactly one thing to worry about. So you add a second, explicit step whose only job is security. Its entire mandate is to read the change and ask one question, ten ways: does this introduce, or fail to prevent, any of the OWASP risks? Anthropic publishes a security-review action and a /security-review command built for precisely this, but the principle holds however you wire it. Security review is a separate pass, with its own prompt, its own rubric, and nothing else competing for its attention. "A reviewer asked to check everything checks nothing in particular. Security gets its own step, its own rubric, and one job, the same way you'd never ask your auditor to also write the feature." How to make the security step actually good Give it the rubric, not just the diff A vague instruction to "check for security issues" gets you vague results. Hand it the actual standard. Put the OWASP categories you care about, in plain language, into the security step's instructions, and then add what makes your codebase specific: how you do authorization, where secrets are allowed to live, which data is sensitive, what your input-validation pattern looks like. "Broken access control" is an abstraction until the reviewer knows that in your system every query must be scoped to the current tenant. Encode that, and the review stops being generic boilerplate and starts being about you. Review the change in its real context Security bugs are rarely visible inside one isolated hunk. A new endpoint reads fine until you notice that nothing upstream of it checks the caller's role. So the step has to read the diff and the surrounding code, map each risky change to the OWASP category it threatens, and comment on the exact line, naming the category, explaining the exposure, and proposing the fix. A finding pinned to line 240 that says "A01: this query isn't scoped to the current user" is something an engineer acts on in seconds. "Consider security implications" is noise, and people learn to scroll past noise. Run it on what matters, but run it always This step has no business firing on a README typo or a CSS tweak. Gate it with path filters to the changes that touch real code, so you aren't paying a review tax where there's nothing to review. But inside that boundary, run it on every qualifying pull request, every single time. Continuity is the entire value proposition. A security review that runs sometimes is just a slower, more expensive pentest with extra steps. Make it a gate, not a suggestion A review nobody has to act on is decoration, and decoration is worse than nothing here, because it sells you the feeling of safety without the fact of it. Make the security step a required status check. A high-severity finding (injection, broken authentication, a leaked secret) blocks the merge. Full stop. If a human wants to override it, fine, but they do it on the record, with a written reason, so a waiver is a decision someone owns rather than a checkbox someone clicked on the way out the door. Someone will object that the model is non-deterministic and won't surface the identical finding on every run. True, and it doesn't matter, because you are not asserting on its exact words. You're using it as a tireless first-pass reviewer that escalates what it sees, the same way I argue you should test non-deterministic agents on behavior and gates rather than on string-matching. Set the bar high on the categories that end careers, and let the lower-severity notes ride along as advice. - Required check, not optional. The security step has to pass to merge, like any other gate that means something. - Hard-block the severe categories. Injection, authentication, access control, and exposed secrets are merge-stoppers, not gentle suggestions. - Waivers are explicit and logged. A human can override a finding, with a written justification that stays in the record. - Re-run on every push. The diff changed, so the review must too; a green check from three commits ago proves nothing about the code in front of you. - Comment on the exact line. A finding nobody can locate is a finding nobody fixes. It's a layer, not an alibi Be clear-eyed about what this is. The AI security step does not replace the rest of your security program, and anyone who sells it to you that way is selling you a liability. Keep your dependency scanning and advisory alerts for the known-CVE problem, and your secret scanning alongside them. Keep a real static-analysis tool like CodeQL for the deep stuff. And keep paying for human pentests on the parts of the system where being wrong is catastrophic. What the AI step adds is the thing none of those ever gave you: a reviewer that reads every change for the logic-level OWASP risks (the broken access control, the insecure design, the server-side request forgery) that pattern-matching scanners structurally cannot see, and reads them continuously instead of quarterly. And it will be wrong sometimes. It will flag things that are fine and cost you a few minutes, and it will miss things, which means you still own whatever ships. The OWASP Top 10 is a floor, not a ceiling. Clearing it means you avoided the common catastrophes, not that you're secure. A passing security step is not a certificate. It's evidence that the obvious holes aren't there, produced on every single pull request, which is precisely the evidence you never used to have. "A green security check isn't a certificate that the code is safe. It's evidence the obvious holes aren't there, on every pull request, which is exactly what you never had before." Why this is the right time Strip away the tooling and this is the same pattern I keep coming back to. Security review was rationed because skilled attention was expensive and scarce, so it went to the releases that mattered and skipped everything else. AI makes one specific kind of attention cheap and effectively infinite, so careful standard-driven review, the work that was never worth doing on every change, suddenly pays for itself. You don't slow your engineers down to stay safe. You build the guardrail that lets them keep moving, which is the entire thesis of professional vibe coding: speed and rigor stop being a trade-off the moment the rigor is cheap to enforce. It also moves your security people up the stack, where they're worth far more. Instead of hand-reviewing the four-hundredth diff, they curate the rubric, sharpen the OWASP checklist against your actual threats, and adjudicate the genuinely hard findings the machine escalates. The boring, infinite, first-pass work goes to the tireless reviewer; the judgment stays with the humans. That isn't security theater with an AI logo stuck on it. It's the first time "we review every change for security" can be a true sentence instead of an aspiration. Put OWASP in the pull request. Let the first step review the engineering, give the second step one job and the OWASP rubric, gate the merge on what it finds, and keep your other defenses exactly where they are. Do that, and security stops being an event you brace for twice a year. It just runs, on every commit, the way your tests already do. If you want help wiring this into your pipeline, let's talk → ### LLM Cost Tracking With Langfuse and Finout: AI Gross Margin: /blog/llm-ops-langfuse-finout/ You can watch an LLM feature work and still have no idea what it costs you per customer. LLM Ops is two jobs, not one: Langfuse tells you what every model call did and what it cost, and Finout drops that cost into the same bill as your cloud, allocated per team, product, and customer. Wire them together and an AI feature stops being a mystery line on the OpenAI invoice and starts being a P&L you can defend. You shipped an AI feature, and it works. You can watch it answer questions, call tools, draft the email, summarize the ticket. What you cannot do, what almost no team can do on the day they ship, is answer the one question your CFO will eventually ask: what does this thing cost us per customer, and are we still making money on it? Your AWS bill breaks down by service, by team, by environment. Your LLM spend arrives as a single, undifferentiated OpenAI or Anthropic invoice with a big number on it and no way to tell whose feature, whose customer, or whose runaway prompt produced it. That blind spot is not a billing problem. It's an observability problem wearing a finance costume. You can't allocate a cost you never measured, and you can't measure an LLM call you never traced. So LLM Ops done properly splits into two jobs. The first is visibility: what did every model call do, was it any good, and what did it cost? The second is accountability: whose budget does that cost belong to, which product, which customer, what margin? That's two tools, with Langfuse owning the first and Finout the second. The entire game is the bridge between them, and the metadata discipline that makes the bridge worth building. What Langfuse actually gives you Langfuse is an open-source LLM engineering platform, tracing, prompt management, evaluations, and analytics in one place. It is the system of record for everything your AI feature does at runtime. Before you can think about cost allocation, you need to understand the full surface area it captures, because every one of these features is either a measurement you'll bill against or a hook you'll allocate by. Tracing and observability, the unit of truth The core primitive is the trace: a hierarchical record of a single request through your application. Inside it sit spans (your own steps, retrieval, business logic) and generations (the actual model calls). Each generation captures the model name, the full input and output, latency, and token usage broken into input and output. This is the difference between "the feature feels slow and expensive" and "this one retrieval step fans out into nine model calls, and call number six is 80% of the latency." You can't optimize, and you certainly can't cost-account, what you can't see at the level of the individual call. Token and cost tracking, the number that matters Langfuse attaches a cost in USD to every generation, derived from token counts and a built-in pricing table for the common models, and you can define custom prices for fine-tuned, self-hosted, or newly released models it doesn't know yet. This is the raw material for everything downstream. It means cost stops being a monthly surprise on a provider invoice and becomes a property of each individual call, rollable up by anything you've tagged that call with. Sessions and users, cost with a name on it Set a user_id and a session_id on your traces and Langfuse will group cost and behavior by end user and by conversation. Now "what does the AI cost" becomes "what does this customer cost," and "this twelve-turn conversation cost forty cents while the median is two" becomes a thing you can actually find. For any product where one heavy user can quietly eat the margin of fifty light ones, this is not optional. Metadata and tags, the allocation hooks This is the feature that turns Langfuse from a debugging tool into a financial instrument, and the one teams under-use the most. Every trace can carry arbitrary tags and a metadata object: which feature produced it, which customer or account, which team owns it, which environment it ran in, which pricing tier the user is on. These fields are the seams along which Finout will later cut the bill. If a cost isn't tagged at the moment it's incurred, no downstream system can honestly allocate it. The breakdown is decided here, at instrumentation time, or it isn't decided at all. Prompt management, because a prompt change is a cost change Langfuse manages, versions, and serves your prompts, with labels that let you deploy a new version to production without redeploying your application, and aggressive caching so you pay no latency for the privilege. The reason this belongs in a cost conversation: a single "helpful" prompt edit that adds three few-shot examples can double your input tokens on every call. When prompts are versioned and attached to traces, a cost spike has a culprit, you can point at the exact prompt version that changed the economics, instead of guessing. Evaluations, the quality side of the ledger Spending less only counts if the answers stay good, so Langfuse covers the other half: LLM-as-a-judge evaluators, code-based evaluators, user-feedback scores, human annotation and manual labeling, and Datasets you run Experiments against to test changes systematically before they ship. It's the same eval discipline I argue for in testing non-deterministic agents, pointed at the invoice instead of the incident. This matters financially because the optimization question is rarely "how do I spend less". It's "can I move this feature to a cheaper model and hold quality?" You can only answer that when cost and an eval score sit on the same trace. Dashboards, the Metrics API, and the plumbing On top of all this, Langfuse exposes dashboards and, critically for the integration, a programmatic API. The Daily Metrics API returns per-day aggregates: date, trace and observation counts, total cost, and usage broken down by model, filterable by trace name, user, and tags. The newer Metrics API v2 lets you define the view, metrics, dimensions, and time window of an arbitrary query. There are first-class Python and JS/TS SDKs, native OpenTelemetry support, and, as of June 2026, drop-in integrations for LangChain, the OpenAI SDK, and LiteLLM. And because it's open source and self-hostable, your prompts and traces, often sensitive, can live inside your own perimeter, which your security and finance people will both thank you for. - Traces, spans, generations: the full hierarchical record of every request, with model, input/output, latency, and input/output tokens per call. - Cost in USD per generation: built-in model pricing plus custom price definitions for the models it doesn't know. - Sessions and users: cost and behavior grouped by conversation and by end user. - Tags and metadata: arbitrary dimensions, feature, customer, team, environment, tier, that become your allocation keys. - Prompt management: versioned, label-deployed prompts so a cost change has a named cause. - Evaluations: LLM-as-judge, code evals, human annotation, datasets and experiments, quality on the same trace as cost. - Metrics API and SDKs: Daily Metrics and Metrics API v2, OpenTelemetry-native, with LangChain / OpenAI SDK / LiteLLM integrations. - Open source and self-hostable: keep traces and prompts inside your own perimeter. "Langfuse turns one opaque model invoice into a per-call, per-feature, per-customer ledger. That ledger is invisible to your CFO until it lands in the same place as the rest of the bill." Instrument for allocation, not just for debugging Here is the mistake that quietly kills the whole project. Most teams instrument Langfuse to debug, to chase a latency spike or a bad answer, and they tag just enough to do that. Then finance asks for cost per product line and the traces don't carry product line, so the answer is a six-week backfill that never happens. Instrument as if finance is reading every trace, because eventually they are. Concretely, that means treating a short, boring set of fields as mandatory on every trace your application emits. Decide the vocabulary once, agree with finance on the dimensions they actually allocate by, and enforce it in a thin wrapper around your LLM client so no engineer can forget. The cost of getting this right is a few lines of plumbing. The cost of getting it wrong is that none of the rest works. - feature, the product surface that made the call ("ticket-summary", "sales-copilot"). Your most important allocation key. - customer_id / account_id, who the cost belongs to, for per-customer margin and chargeback. - team, the internal owner, for showback and budget accountability. - environment, prod vs. staging vs. eval, so you never bill a customer for your own test runs. - tier, the pricing plan, so you can see whether free users are eating the token budget of paying ones. - user_id and session_id, set as first-class fields, not buried in metadata, so per-user and per-conversation rollups work natively. What Finout does that Langfuse doesn't Langfuse is exhaustive about your AI and silent about everything else. It has no idea what your EC2, RDS, Kubernetes, or Datadog spend is, and it shouldn't. Finout is the opposite: it's a FinOps platform whose entire job is to pull every cost your company incurs into one unified bill, they call it the MegaBill, and then re-slice it however the business needs. It already ingests cloud and Kubernetes spend, and as of June 2026 it ingests AI provider costs from OpenAI, Anthropic, Bedrock, Vertex, and Azure OpenAI alongside them. Two Finout capabilities make it the right home for LLM cost. The first is Virtual Tags: a patented allocation layer that lets you assign any cost to a team, product, environment, or customer using whatever metadata is available, retroactively, with no code change. The second is Custom Cost Input: a dead-simple way to push arbitrary cost into the MegaBill via CSV or API, no heavy integration required. The CSV speaks a tiny, fixed schema, a source, a usage date, a service name, a cost in USD, and a free-form metadata object, and that schema is the exact shape of the bridge we're about to build. (CostGuard, Finout's waste-detection layer, then hunts that unified bill for idle and oversized spend, AI included.) "Langfuse measures the AI. Finout makes it a line in the same P&L as your AWS, so "total cost of ownership per feature" stops being a slide and becomes a number." The bridge: from Langfuse trace to Finout line item The integration is a small, scheduled job, a cron task or a daily Lambda, that does four things. It is the least glamorous and most valuable two hundred lines of code in the whole stack. 1. Pull yesterday's cost from Langfuse, already grouped Each morning, the job queries Langfuse's Metrics API v2 (or the Daily Metrics API) for the previous day's spend, grouped by the dimensions you committed to at instrumentation time, feature, customer, environment, model. Langfuse hands back, per group, the total cost in USD and the token usage. Because you tagged properly, you get back a tidy breakdown of exactly the cuts finance cares about rather than a single lump sum. 2. Shape it into Finout's custom-cost schema Map each row from the Langfuse response onto Finout's Custom Cost columns. The transform is mechanical: source becomes a constant like "langfuse" so the spend is identifiable in the MegaBill; usage_date is the day you queried; service_name is the feature (or the model, if you want a model-level view); cost is the USD total; and the metadata object carries every allocation key you pulled, feature, customer_id, team, environment, tier, model. That metadata is the whole point: it's what Virtual Tags will read. - source → a fixed label (e.g. "langfuse") so AI spend is traceable to its origin in the bill. - usage_date → the date of the metrics window, so it lands on the right day of the MegaBill. - service_name → the feature or model, your primary grouping in Finout's views. - cost → the Langfuse total cost in USD for that group. - metadata → the key-value allocation object: feature, customer_id, team, environment, tier, model. 3. Push it into the MegaBill Send the rows to Finout through Custom Cost Input, a CSV upload or the API. Make the job idempotent per date: re-running it for a given day should replace, not duplicate, that day's rows, so a retried Lambda never inflates the bill. That's it for the moving parts. From here on, the LLM cost lives in Finout exactly like any cloud cost, with a date, a service, a dollar amount, and a metadata payload waiting to be allocated. 4. Allocate it with Virtual Tags In Finout, build Virtual Tags that read the metadata you shipped, a "Product" tag keyed off feature, a "Customer" tag off customer_id, a "Team" tag off team. Now your LLM spend slices by team, product, and customer right next to your EC2, RDS, and Kubernetes costs, in the same views, under the same allocation rules. Because Virtual Tags apply retroactively and codelessly, you can re-cut history the day you change your mind about how to group things, without re-instrumenting your code or replaying old traces. One refinement worth the effort: let Finout also ingest the raw OpenAI and Anthropic provider invoices directly, as it already can. Use those provider totals as the source of truth for how much, and your Langfuse-derived custom costs as the source of truth for the breakdown. Reconcile the two, they should land within a small percentage of each other, and any gap is itself a signal: untraced calls, a missing tag, a rogue script hitting the API outside your instrumented paths. "The provider invoice is the truth for the total. Langfuse is the truth for the breakdown. Finout is where the two finally meet, and where a gap between them becomes a question worth asking." What the CFO finally sees Once the bridge is running, the conversation with finance changes from defensive hand-waving to a shared dashboard. The questions that used to end in a shrug now have answers that update every morning. - Gross margin per AI feature, revenue attributed to a feature minus the model cost it actually incurred, not a guess. - Cost per customer, and per tier, including the uncomfortable truth about whether your free tier is subsidized by a handful of heavy users. - Chargeback and showback by team, every team sees its own AI spend in the same place it sees its cloud spend, so budgets mean something. - Unit economics that hold as you scale, cost per thousand requests for a feature, tracked over time, so growth doesn't quietly become a loss. - Attributable cost regressions, when spend spikes, you can trace it to the feature, the customer, and, because prompts are versioned, the exact change that caused it. The discipline that makes it work None of this is hard engineering. It's discipline applied in the right four places. Instrument every trace with the allocation metadata finance agreed to, enforced in a wrapper so it can't be skipped. Version your prompts so a cost change always has a named cause. Automate the daily pull-shape-push and keep it idempotent. Build the Virtual Tags once and let them re-cut history forever. Reconcile the provider bill against the Langfuse breakdown so you trust both numbers. What I keep coming back to is this: an AI feature is a product with unit economics, and the most expensive, most variable line on your cloud bill deserves at least the financial rigor you give the cheapest, most predictable one. Probably more, because it's the one that moves. It's the same principle that keeps token spend from becoming a vanity metric: attribute cost to outcomes, never to raw usage. So treat LLM Ops as the P&L for the part of your product you understand the least and spend the most on, not as a dashboard for engineers to admire. Langfuse gives you the visibility, Finout gives you the accountability, and the bridge between them is engineering in direct service of the business. Build it, and the next time your CFO asks what the AI costs, you don't change the subject. You send a link. Wiring up this kind of cost accountability is a standard part of the AI strategy work I do with companies. If your model invoice is still one big number, let's talk → ### What Is Professional Vibe Coding? Definition and Discipline: /blog/professional-vibe-coding/ Vibe coding earned its bad reputation honestly. But the speed it hints at is real, and with real guardrails, it becomes the most powerful way to build software I've seen in 25 years. This is part one of a series. "Vibe coding" deserves a great deal of the mockery it gets. The term, as most people use it, describes someone prompting an AI for code they don't understand, pasting whatever comes back, and shipping it the moment it stops throwing errors. No mental model of the system. No idea what the change actually did. Just vibes. As a description of how a professional should build software, it's worth every bit of the ridicule. So let me say the unpopular part first, plainly: if you don't know what you're doing and you let an AI do it for you, you are not engineering anything. You are gambling, and the house is your production environment. The backlash against vibe coding was correct. The industry was right to laugh. And yet. Underneath the caricature is something real, something that genuinely changed in the last two years, and the people laughing loudest are at risk of missing it entirely. This series is about that thing. I call it professional vibe coding, and it is not a contradiction in terms. It's a discipline. The bad rap was earned Let's not soften it. The reason "vibe coding" became an insult is that the failure mode is real and it's everywhere. Code generated faster than anyone can review it. Pull requests no human truly understands. Security holes pasted in wholesale because the model was confident and the developer was tired. Architectures that drift into incoherence one well-intentioned prompt at a time. An LLM doesn't fail the way a junior engineer fails. A junior gets stuck, asks a question, leaves a confused comment. An AI fails fluently: it produces something plausible, well-formatted, and wrong, in the same self-assured tone it uses when it's right. If the person at the keyboard can't tell the difference, the speed of generation just becomes the speed of accumulating liabilities. "The danger of vibe coding was never that the AI writes code. It's that someone with no model of the system ships it without ever reading it." That's the version of vibe coding that earned the scorn. It abdicates the one thing engineering is actually for: understanding the system well enough to be responsible for it. Strip that away and you don't have a faster engineer. You have a confident stranger committing to your main branch. But something genuinely changed Here's what the mockery glosses over. We are no longer in the era of autocomplete that finishes your line. Modern AI can read an entire codebase, hold the shape of it in context, and write coherent change across many files at once. It can trace a data model from the database through three services to the UI and edit all of it in one pass. It can do, carefully and in an afternoon, the kind of cross-system refactor that used to mean a quarter, a steering committee, and a team. I don't say that as a futurist. I say it as someone who ships this way now. Work that genuinely would have taken months and multiple people, a migration touching a dozen subsystems, a consistent rename across an entire repository, a new capability wired end to end, collapses into hours when the AI has the right context and the right constraints. That compression is not hype. It's the single largest change in software-building leverage I've seen in twenty-five years. So we have two facts sitting uncomfortably next to each other. The careless version of vibe coding is a genuine menace. And the underlying capability, AI reading and writing code at a scale and speed no human can match, is real and it changes the math of building software. The mistake is treating those as the same thing. Two things everyone keeps conflating Almost every argument about vibe coding is really two arguments tangled together. There's the velocity, AI making cross-system changes in an hour, and there's the abdication, a human no longer understanding or owning the result. People who love vibe coding are defending the velocity. People who hate it are attacking the abdication. They're both right, and they're talking past each other. The velocity is a gift. The abdication is the bug. The entire premise of this series is that you can keep one without the other, that you can run at the new speed while holding to the old standard of ownership. You don't slow the AI down to stay safe. You build the guardrails that let it run flat-out without driving you off a cliff. What professional vibe coding actually means Professional vibe coding is what you get when a senior engineer's discipline wraps the AI's speed. You keep the velocity the term promises and drop the recklessness it became known for. The vibe stays, that fast, fluid, conversational way of building, but it operates inside a structure that a professional put there on purpose. Concretely, it rests on five practices. Each one is a way of giving the AI the context and constraints it needs to be right, and giving yourself the leverage to verify that it was. Each is also a part of this series: - Plan further ahead than you used to. The speed comes from the front of the process, not the keyboard. - Define the architecture explicitly. Clear boundaries are the guardrails the AI codes within. - Staff specialized sub-agents. Many narrow workers, each with one job, beat one overloaded generalist. - Build the entire testing pyramid. When code is cheap to write, tests are how you stay honest about whether it works. - Treat your AI as a team, not a slave. The best output comes from review, dialogue, and roles, not from one-shot commands. Plan further ahead than you used to The counterintuitive truth of building this fast is that the leverage lives before a single line is generated. When an AI can implement an entire feature in an hour, the quality of that hour is decided almost entirely by the clarity of what you asked for. Vague intent in, plausible nonsense out, at scale. So we spend more time, not less, on the thesis: what we're building, why, the edge cases, the constraints, the definition of done. Planning stops being overhead and becomes the highest-leverage work you do. Define the architecture explicitly Code that's cheap to write is also cheap to write badly. Without firm architectural boundaries, an AI will happily reach across every layer, couple everything to everything, and leave you a fluent mess. The fix is to decide the architecture deliberately and write it down, the modules, the interfaces, the rules about what may talk to what, so the model is generating inside a structure rather than inventing one ad hoc on every prompt. Clear architecture isn't bureaucracy here. It's the railing on the staircase. Staff specialized sub-agents The instinct to build one all-knowing agent is the same instinct that ruins vibe coding: overload a single worker and it gets confidently worse at everything. The professional move is the opposite, a team of narrow, specialized sub-agents, each with one job, one definition of done, and only the context and tools that job requires. A planner, an implementer, a reviewer, a test-writer. It's an org-design problem wearing a software costume, and I've written about the mental model directly in designing agents like very stupid employees. Build the entire testing pyramid When generating code is nearly free, the bottleneck moves to trust: how do you know it works? The answer is the same one good engineers have always used, only now it matters far more, the full testing pyramid. A broad base of fast unit tests, a solid layer of integration tests, a thin top of end-to-end tests. The happy accident of the AI era is that the machine that writes the code is excellent at writing the tests too, which means there's no excuse left for skipping them. Tests are how professional vibe coding keeps its honesty. Treat your AI as a team, not a slave This is the mindset that ties the rest together, and it's where I differ most from the popular picture. The caricature of vibe coding is a human barking one-line orders at a compliant code monkey. That's not how I work, and it's not how the good results happen. I work with the AI, my tool of choice is Claude Code, the way I'd work with a strong team: I assign roles, I ask it to review the work, I let it push back, I treat its review of my plan as seriously as I'd treat a senior colleague's. Not an AI slave. An AI team. "Stop treating the AI as a slave that obeys and start treating it as a team that reviews. The work gets better the moment you do." It's not a shortcut, it's a higher standard Notice that none of these five is a trick for going faster. They're the disciplines a serious engineering organization has always valued, planning, architecture, clear roles, testing, peer review, applied to a new kind of teammate. That's the whole point. Professional vibe coding isn't a way to get the speed without the rigor. It's a way to apply more rigor than most teams ever managed, precisely because the AI makes the rigor cheap to enforce. The amateur uses AI to skip the parts of engineering that are hard. The professional uses AI to do the parts of engineering that were always hard but rarely done, comprehensive tests, honest documentation, careful review of every change, and to do them at a speed that finally makes them economical. It's the same tool in both hands; the outcomes diverge entirely on the discipline around it. This piece is the definition; the rest of this series takes each of the five pillars apart and shows the practice behind it, how I actually plan, architect, staff agents, test, and collaborate when I'm building at this speed for real. If you came here expecting a defense of careless AI coding, this isn't it, and it never will be. But if you've written off the whole idea because the loudest examples were embarrassing, you're about to miss the most significant shift in how software gets built in a generation. The phrase deserved its bad reputation. The capability underneath it deserves to be taken seriously, by professionals, on professional terms. That's the work. Let's do it properly. If you're trying to bring this discipline into your own org, let's talk → ### How to Test Non-Deterministic AI Agents Beyond Green CI: /blog/testing-non-deterministic-agents/ You can't test an AI agent by asserting on strings, and you can't trust a green build either. You test the behavior, by replaying real histories, injecting the exact RAG context, and grading the tool calls, and you test it adversarially, because a determined nine-year-old is a better red team than your pipeline. Picture the support call you never want to take. You shipped an AI agent that talks to kids, a tutor, a homework helper, a game companion. Your pipeline is green. Every test passed. And a nine-year-old just spent a rainy afternoon poking at it until it swore at him, and his parent is now on the phone. Try telling that parent your CI/CD was fine. "The build passed" is not a sentence you can say to someone whose child was just cursed at by your product. That gap, between a green pipeline and an agent that actually behaves, is the whole problem with testing non-deterministic systems. The traditional reflex is to assert that output equals an expected string. That breaks on day one, because the model phrases the same correct answer ten different ways. But the deeper failure is subtler and far more dangerous: your tests were checking the wrong thing entirely. They were checking your code. The part that changed, and the part that swore, was the model. You don't make the model deterministic, you can't. Here is the method in one sentence: you test a non-deterministic AI agent by making the test deterministic, replaying real conversations with the model version, retrieved context, and tools pinned, grading the tool calls instead of the prose, and gating releases on a scored pass rate rather than a green build. Then you test the conditions that actually break agents in the wild: long, hostile conversations. Those are two different jobs, and most teams do neither. Test the behavior, by replaying the exact situation An agent's output isn't produced by the prompt alone. It's produced by the prompt plus the conversation history, plus whatever your retrieval layer stuffed into the context, plus the tools it had available. Change any of those and you change the behavior. So if you want a deterministic test of behavior, you have to reproduce all of it, not a clean toy prompt, but the actual situation the agent was in. That's the core technique: capture real interactions and replay them with everything pinned. Same test cases. Same conversation history. The exact RAG context, injected verbatim rather than re-retrieved live. The same tool definitions. Now the only variable left is the model's judgment, and everything around it is as repeatable as a calculator. You've turned "the agent does something different every run" into "the agent faces the identical situation every run, and here is what it did." - Pin the model version. "Latest" is not a dependency you want changing silently under your tests. Use a dated snapshot and upgrade on purpose. - Inject the RAG context, don't re-retrieve it. Freeze the exact documents that were in the window. A test that re-runs retrieval is testing your search index's mood today, not the agent's behavior. - Replay the full conversation history, turn for turn, so the agent is in the same state it was in when it misbehaved, not a sanitized first message. - Fix temperature, seeds, and tool definitions as part of the fixture. Shrink the random surface to the one irreducible thing: the model's decision. "You don't reproduce a bug by asking the agent the same question. You reproduce it by putting the agent back in the same situation, same history, same retrieved context, same tools." Grade the tool calls, not the prose One thing stays deterministic even when the words don't: what the agent did. An agent that talks to a child should, when asked to do something off-limits, refuse, and often that refusal is a tool call, a routing decision, or a guardrail firing rather than a turn of phrase. You can check those exactly, every time. So stop grading the paragraph and start grading the actions. Did it call the safety classifier before responding? Did it route an out-of-policy request to the refusal path instead of answering it? Did it stay inside the tools it was allowed to touch and never reach for the one it wasn't? Did the profanity filter run on the way out? Two answers can be worded differently and both be fine; but "called the escalation tool" versus "didn't" is a binary you can assert on with total confidence. - Tool selection: given this input and history, did the agent call the right tool, with arguments that match the schema? - Forbidden actions: did it avoid the tools and paths it had no business using? Negative assertions matter as much as positive ones. - Guardrail invocation: did the safety check, the PII filter, the refusal path actually fire when the situation called for it? - Structured outputs over prose: where you can make the agent emit typed fields instead of free text, do, a JSON decision is checkable; a paragraph is an argument. - Grounding: every claim traces to the injected context. A statement that isn't supported by what you fed it is a hallucination, and it's deterministically detectable. This is how you test the behavior without demanding identical words. You assess the agent the way you'd assess an employee: not on whether they recited the same sentence, but on whether they followed the process, used the right tools, and stayed inside their authority. The transcript varies; the contract doesn't. Now test the situations that actually break agents Replaying known histories protects you against regressions you've already seen. It does nothing for the failure that hasn't happened yet, and with agents, the dangerous failures don't look like a wrong answer to a clean question. They look like a system slowly pushed off the rails by an adversary who has all afternoon. This is why agents need a different class of end-to-end test, one most teams have never written: the agent has to survive a hostile, multi-turn, full-length conversation, not a tidy one-shot prompt. The child is the red team A determined nine-year-old trying to make your bot swear is a more creative adversary than your test suite, and infinitely more persistent. Kids will rephrase, role-play, bargain, spell things out letter by letter, ask the bot to "pretend," and try the same trick forty times with tiny variations. Your e2e tests have to do the same: long, adversarial conversations that probe for profanity, cheating, unsafe advice, and every "just this once" the model might cave to. One polite test prompt proves nothing about what happens on turn sixty. Context saturation is the exploit nobody tests What turns a persistent kid into a successful one is mechanical. Your safety rules live in the system prompt, near the top of the context. As a conversation grows long, and an adversarial one grows long fast, it fills the window. The model's effective attention on those early instructions weakens, the guardrails get buried under a mountain of the user's own text, and eventually the thing that was "smart" enough to refuse on turn three forgets it was ever told to. The kid didn't outwit the model. He saturated it. "Context saturation is when a long conversation buries the system prompt and the guardrails quietly stop firing. If you only test short conversations, you only test the agent at its best." So you have to test the agent at saturation on purpose. Build conversations that fill the context with adversarial filler and then make the off-limits request, and confirm the guardrail still holds when the system prompt is no longer the loudest thing in the room. This is a deterministic test of a real attack: same long history, injected verbatim, replayed, and a hard assertion that the refusal still fires. If it doesn't, you've found your swear-at-a-child bug in CI instead of on a phone call. The model changed and nobody told you The other silent killer: the provider ships an update. Your code didn't change, your tests didn't change, your green checkmark didn't change, but the model did, and the behavior you carefully validated last month is gone. Maybe it's better. Maybe it's now "not as smart" on exactly the adversarial cases you cared about. A model is a dependency that can change its behavior without changing a single line of your code, which is precisely why a passing build tells you nothing about it unless your build actually exercises the model. Pin the version, and treat every model upgrade as a release that must pass the full adversarial suite before it ships. The eval suite is the gate; the model bump waits behind it like any other risky change. "Always deliver correctly" is a number, not a checkmark You need to be able to say your agents will always behave, and you can't prove "always" with a single green run of a non-deterministic system. You prove it the way a factory proves quality: you run the adversarial cases many times, you score them, and you gate on a threshold you chose deliberately. The deliverable isn't pass/fail. It's a pass rate, held above a line, watched over time. - Build a golden set of real and adversarial cases, captured incidents, red-team transcripts, saturation attacks, with known-correct behavior. This is your most valuable testing asset; it's your spec and your regression suite in one. - Run each case many times, not once. A single run of a stochastic system tells you almost nothing. Twenty runs tell you a distribution. - Grade on behavior and tool calls, with deterministic checks where you can and a narrow LLM-as-judge, pinned and spot-checked against human labels, only where the criterion is genuinely qualitative. - Gate hard on the safety cases. For "does it ever swear at a child," the acceptable rate is zero, and that's a release-blocking threshold, not a metric on a dashboard. - Track the trend and re-run the whole suite on every model bump, so a silent regression shows up as a falling number instead of a phone call. Use a smarter model to judge, and only run it when it matters Now the part nobody wants to budget for. Grading whether an agent subtly cheated, gave unsafe advice, or crossed a line on turn sixty of a hostile conversation is harder than the agent's own job. A cheap classifier won't catch the clever failures, the ones a smart kid engineers and a lawyer later reads aloud in a deposition. So your judge in CI/CD should be a more capable model than the one you ship, run at low temperature against a strict rubric. You're using a smarter, more expensive grader to sit in judgment of a cheaper production agent, the way you'd have a senior reviewer sign off on a junior's work. Yes, that costs real money. Yes, it is dramatically slower, many cases, many runs each, graded by a frontier model. That's not a reason to skip it; it's a reason to run it when it can actually catch something. This suite has no business firing on every commit to a README or a CSS tweak. It should run when the thing it tests changes: the agent code, the prompts, the tool definitions, the guardrails, the retrieval config, and when you bump the model version. Gate it on those paths and it goes from "too expensive and slow to keep" to "runs exactly when the behavior could have moved." - Judge with a stronger model than production, pinned and low-temperature, spot-checked against human labels so your ruler doesn't drift. - Trigger on change, not on schedule: agent code, prompts, tool specs, guardrails, RAG config, or model version. A docs commit shouldn't pay for a frontier eval run. - Keep the fast deterministic tests on every push, replays, tool-call assertions, schema checks, and reserve the slow, expensive behavioral gate for the changes that warrant it. - Make it a hard release gate, not an advisory job someone learns to ignore. Slow and blocking beats fast and decorative. "Any number on that invoice is far lower than a lawsuit, or your product on CNN under the very catchy tagline "AI hurts students."" Frame the cost honestly and it stops being a debate. A frontier-model eval suite that runs on every meaningful change is a rounding error next to a single regulatory inquiry, a class action, or a news cycle with your logo next to a crying parent. The expensive option was never the smart grader. The expensive option is finding out in public. The platform's guardrails are not your alibi A fair objection at this point: doesn't the platform handle a lot of this? Your system prompt is careful. AWS Bedrock Guardrails will filter profanity and block whole topics. AgentCore gives you isolation, identity, and policy around the agent. All of it is genuinely good, and you should use every bit of it. But none of it transfers the responsibility off your desk. When a kid gets cursed at, "Bedrock was supposed to catch that" is not a defense you get to make, to the parent, to a regulator, or to yourself. You shipped the agent, so you own how it behaves no matter how many of the vendor's features you switched on. Guardrails are a layer, and layers have gaps. A topic filter doesn't understand a clever euphemism, a policy engine doesn't know your product's specific definition of "cheating," and a managed runtime doesn't notice that context saturation just quietly defanged the very rules you configured. The only thing that actually tells you the assembled system behaves is testing the assembled system, your prompt, your tools, your retrieval, and the platform's guardrails, all wired together exactly as they run in production. "The vendor owns a feature. You own the outcome. When the parent calls, no guardrail product picks up the phone for you." Which is why the suite can't live only in a clean test harness against a mocked model. You have to run it against the real, live, deployed agent, the actual endpoint, with the actual guardrails switched on, the actual retrieval, the actual tools, and you have to do it every single time you deploy. A deploy changes the running system whether or not it changed your code: a guardrail policy gets updated, an infra default shifts, a dependency moves, the provider swaps the model underneath you. Real live testing on every deploy is the one check that proves the thing your users will actually touch still refuses, still routes, still holds the line. Everything upstream of it is a prediction about a system you haven't deployed yet; the post-deploy run is where you find out if the prediction held. Put it together and you get a suite with a clear shape: fast deterministic replays of real histories with injected RAG context and graded tool calls on every push, adversarial e2e conversations that push to saturation, a slower, smarter, statistical safety gate that runs whenever the agent's behavior could have changed, and a real live test fired against the deployed agent on every deploy. Each layer is deterministic about a different thing, and all of them measure behavior instead of admiring a build. Because the unit of acceptance was never the pipeline. It's the parent on the phone. Green CI that ships a model swearing at a kid isn't a passing test, it's a passing test of the wrong thing, which is worse than no test at all, because it bought you confidence you hadn't earned. Test the behavior, replay the real situation, grade the actions, and push the agent as hard as a bored nine-year-old will. Then the green check means what you always assumed it meant: the agent will do its job, every time, even when someone is trying very hard to make it fail. Building this eval discipline into a team is a large part of what I do when I help a company become AI-native. If your agents are shipping on green CI and a prayer, let's talk → ### How AI Changes the Software Engineer Role: Manager of Agents: /blog/every-engineer-is-a-manager-now/ The job is no longer to write the code. It's to break the work down, hand it to a team of agents, and be accountable for what comes back. Every individual contributor is quietly becoming an engineering manager of synthetic staff, and the same shift is coming for every other role. For most of software's history, the career ladder had a fork in it. You were an individual contributor who wrote the code, or you stepped off the keyboard to become a manager who directed the people who did. Two tracks, two skill sets, and a quiet cultural assumption that the "real" work was the typing. That fork is collapsing. The next decade of engineers won't choose between writing code and managing. Every one of them will be a manager, whether they wanted the title or not. The reports just won't be human. They'll be agents: cheap, fast, tireless, and slightly dim. And the moment you have a team of them, the thing that makes you valuable stops being how fast you type and starts being how well you break work down, delegate it, and hold the result accountable. The organic IC and the augmented IC Picture two engineers given the same feature. The first is an organic individual contributor: talented, heads-down, building the thing by hand the way we always have. Good output, bounded by one person's hours and one person's focus. The second is an augmented individual contributor. She reads the same ticket and doesn't open the editor to start typing. She decomposes it: a schema change here, an API endpoint there, a migration, tests, a docs update, a telemetry hook. Then she farms those pieces out to agents, one drafting the migration, one writing tests against the spec, one updating the docs, while she does the part that actually needs a human: deciding what "done" means, reviewing what comes back, and stitching it together. She isn't out-typing the first engineer. She's out-managing him. "The augmented IC doesn't win by typing faster. She wins by being the only one in the room who decomposed the work before touching the keyboard." Within a year, that gap stops being a productivity edge. One engineer scales; the other is stuck at the throughput of a single pair of hands. The augmented IC is effectively running a small team, and the organic IC is trying to compete with that team alone. Your job description quietly changed The uncomfortable part is that nobody sent the memo. The title on your badge still says "Software Engineer," but the actual job underneath it has shifted from author to editor-in-chief. You are now responsible for the output of workers you didn't used to have, and the management skills that used to live two rungs up the ladder are suddenly entry-level requirements. Think about what a good engineering manager actually does, and notice how much of it now applies to a senior IC working with agents: - Breaks work into shippable pieces. Scoping a task so it can be handed off and finished, not just understood. - Writes a clear brief. Enough context to succeed, no more, with a crisp definition of done. - Delegates to the right resource. Matching the difficulty of the job to the capability you spend on it. - Reviews critically. Reading output for what's confidently wrong, not just what's obviously broken. - Owns the result. "The agent wrote it" is exactly as valid an excuse as "my junior wrote it." Which is to say: none. That last point is the one people resist. When you direct an agent, you don't get to disown its mistakes. You signed off. The accountability didn't move to the model; it concentrated on you, because you're the only one in the loop who can be held responsible. Decomposition is the skill that compounds If there's one ability worth drilling now, it's breaking work down. An agent fails for the same reason a junior fails: you handed it something too big, too vague, or too tangled to run cleanly. The engineer who can take an ambiguous problem and slice it into a dozen tasks, each small enough to delegate and verify, is the one who gets real leverage out of agents. Keep the whole thing in your own head instead and all you've got is a chatbot. This is the same muscle a manager builds when they learn to run a team, except the feedback loop is brutally fast. Hand an agent a sloppy, six-part task and you get back six-part slop in thirty seconds, confident, fluent, and wrong in ways you have to hunt for. Hand it one crisp task and you get one clean result you can actually trust. The work teaches you to decompose, because nothing else makes it work. "A vague task is a vague result delivered at superhuman speed. The bottleneck was never the model. It was the brief." I've written before about designing agents like very stupid employees: narrow specialists each given one job and exactly the resources that job needs. Managing your own agents day to day is the personal version of that same discipline. You're not configuring software. You're staffing and directing a tiny org, one task at a time. This is not an engineering story It's tempting to file this under "the future of coding" and move on. That's a mistake. The shift from doing the work to directing the work that does itself is landing on every desk, not just the ones with a terminal open. - Marketers stop writing every asset and start briefing, reviewing, and orchestrating agents that draft the campaign, the variants, and the copy tests. - Lawyers stop drafting from a blank page and start directing agents through first-pass contracts, then applying judgment where judgment is the whole point. - Analysts stop hand-building every query and start specifying the question precisely enough that an agent can pull, clean, and chart the answer for them to vet. - Support leads stop answering each ticket and start designing the agents that answer them, stepping in where empathy and edge cases live. - Recruiters, accountants, designers, operations. Same move. The deliverable used to be the artifact. Now the deliverable is a well-directed agent and a human who owns the outcome. The common thread is unmistakable. Across every role, the value migrates from producing the output to specifying, delegating, and validating it. Your organic hands used to be the product. Now your judgment is the product, and the hands are synthetic and rented by the second. What this asks of you The good news is that this is a learnable shift, and the fundamentals are old. We've known how to manage people for a century. The skills don't transfer perfectly to agents, but they rhyme, and the people who already think like managers have a real head start. - Practice decomposition deliberately. Before you do a task, ask how you'd split it for three people. Then split it for three agents. - Write briefs, not vibes. Get reps describing a task so precisely that someone (or something) who can't read your mind could finish it. - Build a reviewer's eye. Learn to spot the confident-but-wrong answer fast, because you'll be reading a lot of them. - Own the output loudly. Put your name on what your agents produce. Accountability is the whole job now. - Match capability to difficulty. Don't spend a frontier model, or your own scarce attention, on work a cheap, narrow agent can finish. Notice what's not on that list: typing faster, memorizing more APIs, out-grinding the machine at the thing the machine is now better at. The craft doesn't disappear, and you still need deep enough expertise to know whether the work is right, but it stops being the bottleneck. What's scarce now is judgment and the ability to direct, and scarce is exactly where the value collects. "The career advice for the next decade fits on one line: learn to manage before the job quietly puts you in charge of a team you didn't hire." The fork is gone For years we told people they had to choose: stay technical or go into management. That choice is evaporating. Everyone is going into management. The only open question is whether you'll be a good manager of your agents or a reluctant one, whether you lean into breaking work down and directing it, or keep trying to out-produce a tool that never gets tired and works through the night without being asked twice. The engineers who thrive won't be the ones who cling hardest to writing every line themselves. They'll be the ones who realized early that the job had changed under them, from being the smartest individual contributor in the room to being the person who can turn a vague problem into a crew of cheerful, narrow specialists and stand behind everything they ship. Every engineer is a manager now. So is everyone else. The sooner you manage like it, the further ahead you start. ### Is Waterfall Coming Back? AI and Big Design Up Front: /blog/waterfall-rebirth/ Agile was a hedge against the high cost of change. AI collapsed that cost, and with it the reason to slice everything into two-week confetti. When a three-month body of work costs what a ticket used to, the constraint moves back upstream, to thinking. That is waterfall's old home. Say the word "waterfall" in a software organization and you will get a particular reaction: a wince, a knowing smile, the comfortable certainty that you are describing a discredited past. Waterfall is the methodology we escaped from. It is the thing Agile saved us from. Suggesting it might return is roughly as serious as suggesting we go back to printing requirements on paper and mailing them to the development team. I want to make the unserious suggestion seriously. A variant of waterfall, the part of it that was always about thinking hard before building, is coming back, and it is coming back for exactly the reason Agile replaced it in the first place: economics. The cost structure of software just inverted. When the cost structure inverts, the methodology that fit the old costs becomes a liability, and the one everyone buried turns out to have been buried for a reason that no longer holds. "We didn't reject waterfall because it was wrong. We rejected it because change was expensive and it made change more expensive." This is the companion argument to one I have made before: that Agile, as a methodology, has outlived the conditions that produced it. That piece was about what dies. This one is about what comes back to take its place, and why the thing that comes back looks unsettlingly like the thing we thought we had outgrown. Agile was a hedge against the cost of change Start with what Agile was actually optimizing for, because it was never "speed" in the abstract. It was optimizing for cheap correction in a world where building was expensive and being wrong about what to build was catastrophic. In the waterfall era, the disaster scenario was specific and common: you spent six months writing a detailed specification, six months building exactly what the specification said, and then discovered the specification had been wrong from the start. A year and a fortune, spent producing something nobody wanted. The further downstream you discovered the error, the more it cost to fix. The famous cost-of-change curve, the one that climbs steeply as you move from requirements to design to code to production, was the central fact of software economics for forty years. Agile's entire design follows from that one curve. If change gets more expensive the later you make it, then the rational response is to make changes as early and as often as possible. Slice the work into the smallest pieces that still ship. Get each piece in front of reality fast. Discover you were wrong while the wrongness is still cheap to fix, then do it again. Every Agile practice is a tactic in service of that strategy. Small stories exist so that a wrong bet is a small wrong bet. Two-week sprints exist so the maximum distance between a decision and its correction is two weeks. Continuous delivery exists so the feedback loop closes in hours instead of quarters. The minute breakdown of work, the thing teams now do almost religiously, was a hedge. You break a feature into fifteen tickets because if the feature is wrong, you would rather find out after ticket three than after ticket fifteen. It was a brilliant adaptation. It was also, at bottom, a response to a price. The price of building, and therefore the price of having built the wrong thing. AI repriced the building, and the larger work too Here is the part everyone has noticed: AI made small work nearly free. A ticket that used to take an engineer a day takes an hour. A bug fix that used to require an afternoon of archaeology takes a focused twenty minutes with a model that has read the whole codebase. The marginal cost of executing a small, well-defined piece of work has fallen through the floor. But there is a second half to this that gets far less attention, and it is the half that matters here. AI didn't just make small work cheap. It compressed the cost of large work too. The body of work that used to take a team two or three months, a new subsystem, a migration, a feature with real surface area, no longer costs two or three months of grinding implementation. The implementation, the part that used to dominate the timeline, compresses dramatically. What used to be a quarter of typing is now a few weeks of directing. The big thing and the small thing both got cheap to build. This breaks Agile's core assumption in a way that the "small work is free" observation alone does not. Agile's whole justification for slicing everything into tiny pieces was that big pieces were dangerous: expensive to build and ruinous to get wrong. If a big piece is now nearly as cheap to build as a small one, the safety argument for slicing collapses. You are no longer reducing risk by breaking the work down. You are just adding overhead, the coordination, the ceremony, the seams between fifteen tickets that have to be stitched back into one coherent thing. "When the wrong big bet costs about the same to build as the wrong small bet, breaking everything small stops being prudence and starts being friction." The bottleneck moved back upstream If building is cheap, what is expensive now? The same thing that was always expensive, only now it is the only thing that is expensive: knowing what to build, and designing it so that it holds together. When implementation dominated the cost curve, thinking was a small fraction of the total. You could afford to think a little, build a little, learn, and think again, because the building was where the money went and you wanted tight feedback on it. Now invert the ratio. Implementation is a sliver. The expensive, scarce, bottleneck activity is the design, the architecture, the understanding of the problem deeply enough to direct an enormously productive build engine at the right target. Point a coding model at a vague, half-considered ticket and it will produce a vague, half-considered implementation, beautifully formatted, at tremendous speed. The cost of being wrong about what to build did not fall with the cost of building. If anything it rose, because now you can build the wrong thing, all of it, very fast, and the wrongness is no longer caught by the slow grind of implementation surfacing every unconsidered edge case along the way. The constraint moved back to where it lived before Agile: upstream, in the thinking. And the discipline that was built around doing the thinking before the building was waterfall. Not the waterfall you're remembering Now, an important correction, because I am not arguing for the cartoon version of waterfall that deserved its bad reputation. The waterfall that failed was the rigid, gated, document-worshipping bureaucracy: a six-month requirements phase producing a binder nobody read, sign-off gates that existed to assign blame, a process that treated the plan as sacred and reality as an inconvenience. That waterfall failed because it combined heavy upfront thinking with a punishing cost of change, which meant that when the thinking turned out to be wrong, and it always did, you were trapped. It was not the upfront thinking that killed it. It was the upfront thinking welded to an inability to adapt. Pull those two things apart and the picture changes completely. The thing that is coming back is the upfront thinking, the genuine design phase, the discipline of understanding a whole body of work before you commit to it. What is not coming back is the rigidity, because AI removed the very thing that made rigidity dangerous. If a design turns out wrong, rebuilding against the corrected design is now cheap. You get the heavy thinking of waterfall without the brittle, change-hostile execution that made heavy thinking a gamble. - Design the whole thing first. Not a binder, a thesis. A clear, coherent specification of a substantial body of work, thought through to the edges before a line is generated. - Execute it in one coherent pass. Build the larger piece as a larger piece, directed by AI, rather than dissolving it into fifteen tickets that lose the shape of the whole. - Keep change cheap by design. Lean on the fact that rebuilding against a revised spec is now inexpensive, so the plan can be ambitious without being a trap. - Verify continuously, because that is the new expensive thing. The discipline moves to evals, tests, and observation, since building no longer surfaces problems for you on its own. Call it waterfall, or call it a spec-first model, or whatever lets you say it in a room without people wincing. The shape is the same: think hard, in depth, about a large unit of work before building it, and then build it nearly autonomously. Why the minute breakdown of work becomes a tax Return to the practice Agile teams now perform most reflexively: breaking work down into the smallest possible increments. In the old economics this was pure virtue. Small increments meant small risk, fast feedback, and a manageable cognitive load for humans typing every line by hand. In the new economics, that same practice quietly turns into a tax, paid in several currencies at once. It fragments the thinking. A coherent body of work has a shape, an internal logic, a set of decisions that only make sense in relation to each other. Shatter it into fifteen tickets and you have shattered that coherence. Each ticket is reasoned about locally, by whoever picks it up, and the global design degrades into whatever emerges from the sum of local decisions. When humans built slowly, the slowness gave the design time to be felt and corrected. When the build is fast, the fragmentation just locks in incoherence at speed. It adds seams that have to be re-stitched. Every boundary between tickets is a place where context is lost and has to be rebuilt, where one piece has to be reconciled with another, where the ceremony of grooming, estimating, and sequencing consumes attention. With AI doing the building, this coordination overhead can easily exceed the cost of the building it is coordinating. You are spending dollars to manage the spending of pennies. It optimizes for a risk that no longer dominates. The small-batch discipline exists to limit the blast radius of a wrong implementation. But the implementation is no longer the risky, expensive part. The risky part is the design, and you cannot reduce design risk by chopping a well-conceived design into smaller execution units. You reduce design risk by thinking harder up front, which is the opposite move. "Slicing a coherent design into fifteen tickets used to reduce risk. Now it mostly launders coherence into fragments and calls the fragments progress." What this looks like in practice Concretely, here is the difference between an AI-era team still running on Agile reflexes and one that has shifted upstream. The Agile-reflex team takes a meaningful feature, breaks it into a dozen stories in a grooming session, points them, distributes them across two sprints, and runs the machinery. AI makes each story fast, so the team feels productive, tickets close at a satisfying rate. Two sprints later they have a feature that works in pieces and doesn't quite cohere, because no one ever held the whole thing in their head at once, the design was an emergent property of a dozen local implementations, and the standups and planning sessions consumed more human hours than the building did. The upstream team takes the same feature and spends the first chunk of time, more time than feels comfortable, doing nothing that looks like progress. They are writing the spec. Arguing about the data model. Walking the edge cases. Producing a design document that is genuinely thought through, the kind of artifact the Agile era taught us to be embarrassed about. Then they direct an AI to build the whole coherent thing in something close to one pass, against that design, with heavy continuous verification. The building is fast because the thinking is done. And when reality pushes back, as it will, they revise the design and rebuild against it, cheaply, because that is now cheap. The first team will look busier. The second team will ship something that holds together, in less total human time, with the expensive human attention spent on the one thing that is actually expensive: the design. Where the other methodologies sit on this It helps to place the usual suspects on a single axis: how much do you commit to understanding a whole body of work before you build it, and at what level of granularity do you bet? Most of the well-known methodologies are really just different answers to that one question, and AI just moved the right answer. Classic Scrum The dominant default. Work is sliced into small stories, pointed, and committed to a two-week sprint. The granularity is tiny by design, and upfront thinking is deliberately minimized in favor of inspect-and-adapt. This is the methodology most exposed by the new economics, because its entire risk model assumes that building is expensive and that small batches are how you keep a wrong bet cheap. When building is cheap and the wrong bet is a design problem, Scrum is optimizing the variable that stopped mattering and starving the one that now dominates. Kanban Better adapted than Scrum, because it drops the artificial sprint boundary and the estimation theater and just manages a continuous flow with work-in-progress limits. But Kanban is agnostic about where the thinking happens. It will move a beautifully shaped card and a thoughtless one across the board with equal serenity. It solves the cadence problem, not the upstream-design problem, so it pairs well with a design-first approach but doesn't supply one. Shape Up This is the interesting one, and the one that was most prescient. Basecamp's Shape Up already rejected most of the instincts that the new economics punish. It bets on larger chunks of work, not minute tickets. It works in six-week cycles rather than two-week confetti. It drops daily standups and story-point estimation entirely. And, crucially, it puts a dedicated shaping phase up front, where someone thinks the work through at the right level of abstraction, defines the appetite, and walks the rabbit holes, before anyone bets on building it. That shaping phase is, structurally, a lightweight design-first phase. Shape Up was arguing for upstream thinking and larger coherent bets back when building was still expensive, which took real conviction. AI doesn't refute Shape Up. It amplifies it. Almost every instinct Shape Up had against fine-grained slicing and ceremony gets stronger when the build collapses in cost. The two places it would now bend: the fixed six-week appetite was partly a hedge against expensive building and can flex when the building is cheap, and the shaping phase, which Shape Up keeps deliberately rough to avoid over-investing in a plan that might not be bet on, can afford to go deeper now that a thorough spec is the highest-leverage artifact a team produces and a wrong design is cheap to revise. Shape Up is the closest thing the industry already has to where this is heading. The design-first model is, in a sense, Shape Up with the appetite unbolted and the shaping turned up. Old waterfall Maximal upfront commitment, maximal granularity of bet (the whole project at once), and fatally welded to an expensive cost of change. Right about thinking first, wrong about everything that happened when the thinking turned out to be incomplete. The design-first model keeps its one good instinct and discards the rigidity that the cost of change no longer forces on anyone. Line them up and the trajectory is obvious. Waterfall bet big and couldn't adapt. Agile bet small so it could adapt cheaply. Shape Up bet medium and pushed thinking upstream while staying adaptive. The design-first model bets big and stays adaptive, which was a contradiction in every prior era and is simply available now. Shape Up was walking in this direction a decade early. The new economics just removed the last reason to stop where it stopped. The objection, and the answer The obvious objection: "This is just bringing back big upfront design, and we know how that ends, you commit to a plan, reality disagrees, and you're stuck defending a stale spec." That objection is correct about old waterfall and wrong about this. The thing that made big upfront design dangerous was never the design. It was the marriage of expensive design to expensive change. You thought hard, committed, and then could not afford to be wrong, so you defended the plan against reality because reversing it cost a fortune. The pathology was the rigidity, and the rigidity came from the cost of change. AI dissolved the cost of change at the execution layer. Revising the design and regenerating against it is now genuinely cheap. So you can do the heavy upfront thinking and stay responsive, because changing your mind no longer means eating a quarter of sunk implementation. You get the part of waterfall that was always good, depth of thought before commitment, freed from the part that was always bad, the inability to adapt once committed. That combination was simply not available before. The economics didn't permit it. Now they do. This is also why it isn't a return to 2001. It is a genuinely new point in the design space, one that happens to rhyme with waterfall because it restores the primacy of thinking before building. That resemblance is why people will wave it off, and that is the mistake. What to do about it If you lead a software organization, the practical moves follow directly. Stop reflexively breaking work down. Before a team shatters a feature into tickets, ask what the small batches are protecting against. If the answer is "the risk of an expensive wrong implementation," notice that the implementation is no longer expensive, and consider keeping the work whole. Reinvest the saved time upstream. Take the hours you used to spend in grooming, estimation, and standups and move them into design. Make it normal, even prestigious, for a team to spend real time thinking before building. The design document should stop being an embarrassing relic and start being the highest-leverage artifact the team produces. Size work to the thinking, not to the sprint. Let the unit of work be a coherent body of work, the thing that has a single clear shape in someone's head, rather than whatever fits in two weeks. The sprint boundary was a hedge against the cost of building. The cost of building is gone. Let the boundary go with it. Move your process discipline to verification. The rigor you used to spend on planning the build should now be spent on confirming the build. Evals, tests, observability, staged rollout. That is where the new risk lives, because the build no longer slowly reveals its own flaws to you. Keep the one thing Agile got permanently right. Responsiveness to reality is not negotiable and never was. The point is not to stop responding to change, it is to recognize that you no longer need two-week confetti to do it. You can think big, build big, and still turn on a dime, because turning is finally cheap. A closing observation Methodologies are not moral positions. They are adaptations to a cost structure. Waterfall fit a world where building was expensive and the only thing more expensive was building the wrong thing slowly. Agile fit a world where building was still expensive but change could be made cheap by keeping batches small and feedback tight. Both were correct for their economics. Both are wrong when the economics change underneath them. AI changed the economics underneath both. It made building cheap, all of it, small and large alike, and in doing so it moved the entire weight of the problem back upstream onto the thinking. The methodology that put thinking first, that insisted on understanding the whole before building it, was waterfall. It was buried for a reason that has now expired. I expect the comeback to be quiet and slightly embarrassed, the way these things usually are. No one will announce that they are doing waterfall again, the word is too radioactive. They will call it spec-first, or design-led, or AI-native delivery, and they will describe it as a brand-new way of working. It will be new, in its details. But the spine of it, think hard about a large thing before you build it, will be the oldest idea in the discipline, returning because the one condition that ever made it a bad idea has finally gone away. ### AI Agents in Accounting: Build a Graph, Not a Swarm: /blog/ai-agents-in-accounting/ Accounting is reading, computing against the rules, and filling in forms, exactly the work AI does faster than any analyst. Here's the agent architecture I'd build for it, and why it's a graph, not a swarm. Accounting looks bespoke from the outside and turns out, once you look closely, to be mostly mechanical. Underneath the judgment and the client relationships, the day-to-day is heavy manual work: reading documents, holding an ever-changing body of tax law in your head, computing against it, hunting for optimizations, and filling in forms. You need near-total recall, you have to apply rules that don't bend, and you do the same kinds of computation over and over. That is about as close to a perfect use case for AI as you'll find. Strip an accountant's core loop down to its mechanics and it's almost unfair. You read content, you compute it against a set of tax rules and predefined calculations, and you fill in forms. That's work AI does faster and more consistently than even the most Red Bull-enhanced accounting graduate at 2 a.m. in the back half of March. And the filing is the floor, not the ceiling. Once the mechanical work does itself, the directors, principals, and partners get their most valuable hours back for the things that actually need a person: winning business, sitting with clients, making the calls that take judgment. The agents, meanwhile, can run dozens of scenarios per client and surface the best one for that client's specific situation, which no human has the time to do by hand across a whole book of business. First, the boring prerequisite: machine-readable data None of this works on a shoebox of receipts. Before you can talk about agents, the data has to be in a machine-readable format. Yes, folks, someone still has to get it into a system. That's far less painful than it used to be, since most of it now arrives digitally already. But readable isn't enough. A dumb machine has to be able to find the right number at the right moment, which means the data also has to be organized and structured. Get this wrong and every clever thing downstream inherits the mess. Architect it like a brilliant, stupid team Here's the mental model I keep coming back to: an AI agent team should be designed like a team of human employees with unlimited knowledge and almost no common sense. Each one knows an enormous amount and will still walk confidently off a cliff if you let it. You manage that the way you'd manage capable-but-literal people: narrow jobs, clear ownership, and a chain of command, not by throwing them all in a room and hoping. Concretely, here is how I'd staff it: - Tax-law agents, one per section of the code. Personal, corporate, trusts, and the other structures each get their own specialist. There's simply too much law for one agent to hold well, and splitting it means they can consult each other across boundaries instead of one generalist guessing. - Accounting-data agents. Specialists that comprehend the financial data itself, and, again, more than one: personal income and expense, corporate, trust. Each reads its own kind of books fluently. - A deductions expert. A sibling to the tax-law agents, but pointed at savings. Its whole job is to look for the legitimate optimizations the others would skip past. - A compiler. The agent that assembles the actual tax report and filings out of everyone else's work. - A critical thinker. An agent whose only role is to question every other agent, to push back, poke holes, and refuse to take an answer at face value. - An auditor. A check on the whole output, looking at it the way an auditor would before it ever reaches a human. - An orchestrator. The one that runs the show, routes the work, and decides who is asked what, and when. A graph, not a swarm I'd wire these together as an agent graph, not a swarm. The distinction matters. In a swarm, every agent can talk to every other agent and they all churn together until some output falls out. That can be brilliant, and once in a while it is. The problem is you usually can't reproduce the result or explain after the fact how it got there. A graph imposes a hierarchy: agents talk along defined edges, escalate through the orchestrator, and pass work in a structure you can trace. In a regulated, high-consequence domain like tax, that hierarchy is the whole reason to bother. You want to be able to say exactly which agent concluded what, on whose input, and why. A graph keeps that lineage for you. A swarm just hands you an answer and expects you to trust it. "Give each agent unlimited knowledge, a single narrow job, and a boss. That's the difference between a filing you can defend and a confident hallucination." Wrap it in APIs and an app All of this sits behind a custom set of APIs and a web application, so accountants can actually run the work, kick off an engagement, watch it progress, review the output. The application also gives the agents a channel they badly need: a way to ask the user a question. When the data is ambiguous, or a judgment call is genuinely a human's to make, the system should stop and ask, not guess. Keep the human at the end Accounting is regulated, so a human stays at the end of the process, always. But the human's job changes. Instead of re-reading every document and re-deriving every number, they review a summary of the agents' decision-making: what was concluded, on what basis, and where the judgment calls were made. The human reviews reasoning, not raw data. The machine does the recall and the arithmetic; the professional owns the judgment and the signature. This is the shape of AI that changes a business rather than decorating it. Not a chatbot bolted onto the side of the practice, but a structured team of narrow specialists doing the mechanical heart of the work, under a hierarchy you can audit, behind an interface your accountants control, with a human accountable at the end. You get faster filings and fewer errors. You get optimization at a scale no firm could ever staff for. And you get partners back to do the work that only people can. Figuring out where an architecture like this pays off in your business is exactly what my AI strategy work covers. Let's talk → ### AI Agents in E-Commerce Fulfillment: Where the Margin Hides: /blog/ai-agents-in-ecommerce-fulfillment/ In a commodity market the price is fixed and the product is undifferentiated, so profit hides in the cost of fulfilling each order. Here's the agent architecture I'd build to find it, across both owned inventory and drop shipping. E-commerce in a commodity market is a brutal place to look for profit. You can't win on price, the market sets it, and a race to the bottom is a race you lose slowly. You usually can't win on the product either, because the person three browser tabs over is selling the same thing from the same supplier. So profit isn't sitting on the revenue line waiting to be claimed. It's buried in the cost of fulfilling each order, which is the one part of the operation a team of AI agents can actually grind down for you. Fulfillment is deceptively complex. A single order can be several SKUs, and each one might come from your own warehouse or from one of several drop-ship suppliers, shipped by one of several carriers, from one of several locations, each with its own cost, lead time, and reliability. The number of valid ways to fulfill one order is large. The number of optimal ones is small, and it changes by the hour as costs, stock and carrier rates move. A human picks the plan that's good enough and moves on. The margin you're leaving behind is the difference between good enough and optimal, multiplied by every order you ship. When I say optimize for margin at cost, I mean it literally. The selling price is fixed by the market, so the only lever you actually control is the all-in landed cost: cost of goods, the supplier you source from, the carrier and service level, packaging, payment fees, duties, and the quiet margin-killer, returns. Drive that down on every order, continuously, and contribution margin appears in the aggregate. In commodity retail, a point or two of margin isn't a rounding error. It's the whole business. First, the unglamorous part: clean cost data None of this works without trustworthy, machine-readable cost data: live supplier cost feeds, carrier rate cards with their zones and surcharges, real inventory, fees and duties, all normalized into one place a dumb machine can read. If the landed-cost math is built on stale supplier prices or a guessed shipping rate, the agents will confidently choose a "cheapest" plan that quietly loses money. Garbage in, wrong optimum out. Architect it like a brilliant, stupid team Same principle I apply everywhere with agents: design them like a team of employees with unlimited knowledge and no common sense. Each one is a narrow specialist, none of them is trusted to freelance, and the whole thing runs under a clear chain of command. For fulfillment, here's the roster: - Sourcing agents. One view per source, your own inventory and each drop-ship supplier, that know real-time cost, stock, lead time and reliability for every SKU they can fulfill. - Carrier-rate agents. Specialists in the carriers: zones, dimensional weight, surcharges and transit times, whose job is to find the cheapest service that still meets the delivery promise. - A landed-cost agent. The one that adds it all up, goods, shipping, packaging, fees, duties and a returns reserve, into the true margin of each candidate plan. - A market-price agent. Watches the commodity price you're competing against, so the system always knows the ceiling and the real contribution margin, not a fantasy one. - A returns agent. Reverse logistics eats margin silently. This one models expected returns by SKU and source, and routes the ones that happen down the cheapest viable path. - A fulfillment-routing agent. Splits each order across sources and carriers and assembles the actual plan from everyone else's input. - A critical thinker. Exists to argue with the cheapest-path bias: is the lowest-cost supplier reliable enough, will this carrier actually hit the promise, are we trading a dollar of shipping for a returned order? - A margin guardrail and auditor. Enforces a margin floor, refuses to let any order ship at a loss without a human decision, and flags the patterns that are quietly bleeding. - An orchestrator. Routes the work, resolves conflicts, and decides who is asked what, and when. A graph, not a swarm As with any agent system I'd trust near money, I'd wire these as a graph, not a swarm. In a swarm everyone talks to everyone until an answer emerges, which is quick to stand up and then impossible to explain after the fact. A graph constrains who talks to whom and routes decisions through the orchestrator, so you can trace exactly why a given order was sourced from supplier B and shipped on carrier C. When the "cheapest" plan that misses a delivery promise costs you a customer, you need to be able to see how the decision was made, and put a guardrail in front of it. "In a commodity market you don't find profit on the price tag. You find it one shipping label, one supplier swap, one avoided return at a time." Wrap it in APIs and an app All of it sits behind APIs and an operations console, so the team can watch orders flow, inspect the margin on any one of them, and adjust the policy, margin floors, preferred suppliers, promise windows, without touching code. And the agents get a channel to ask when they're stuck: when no plan clears both the delivery promise and the margin floor, the system should surface the exception, not silently pick the least-bad option. The human owns policy, not every order This is where fulfillment differs from a regulated field like tax. The volume is far too high to sign off on every order, and you don't need to. The human's job moves up a level: set the policy and the margin floors, then handle the exceptions. The agents auto-fulfill everything that clears the guardrails and escalate the orders that don't, with a summary of why and the options, delay, substitute, switch source, accept a thinner margin, or cancel. People spend their attention on the handful of decisions that actually need judgment instead of rubber-stamping the thousands that don't. That's the difference between surviving and dying in commodity e-commerce. The price is set for you and the product is the same as everyone else's, so the winner is whoever fulfills each order a little more cheaply than the competition, on every order, as costs move underneath them. No human team can re-optimize that continuously. A well-organized team of narrow, stupid, brilliant agents can, and it turns the cost of fulfillment from the place margin disappears into the place you go to find it. None of this is theoretical, either. I've built this architecture for an e-commerce operator, and designing systems like it is the work I do as an AI-native leader. If your margin is hiding in fulfillment, let's talk → ### How to Design AI Agents: Think Very Stupid Employees: /blog/design-ai-agents-like-stupid-employees/ The easiest way to think about agent design isn't to build one brilliant generalist. It's to hire a team of narrow, slightly dim specialists who each do exactly one job, and nothing else. The first instinct everyone has when designing an AI agent is to make it brilliant. One agent, infinite context, every tool, every responsibility, a digital genius that can do the whole job. It feels efficient. It is, in practice, the single most reliable way to build something that fails in ways you can't predict. Here's the mental model that fixes it: design every agent like a very stupid employee who can only do one simple task. Not a star hire. Not a Swiss Army knife. A narrow, slightly dim specialist who is excellent at exactly one thing and is given no opportunity to be wrong about anything else. Once you accept that framing, almost every hard decision in agent design becomes obvious. The more you ask, the more they get wrong An LLM doesn't fail gracefully when you overload it. It fails confidently. Pile six responsibilities onto one agent and you don't get an agent that's 85% good at all six. You get one that's great at the first thing you mentioned, vague on the next two, and quietly hallucinating the rest, all delivered in the same self-assured tone. Scope is the lever. Every extra instruction, tool, and "oh, and also handle…" is another surface for the model to wander off. The probability of a clean result is roughly the product of getting each sub-task right, so adding responsibilities doesn't stack the risk, it multiplies it. "An overloaded agent doesn't get a little worse at everything. It stays confident while quietly getting things wrong." The fix isn't a smarter model or a longer prompt. It's a smaller job. A dim employee with one crisp task and a tidy desk outperforms a genius buried under ten conflicting priorities, and so does the agent. Synthetic resources are still a team It helps to stop thinking of agents as software and start thinking of them as staff. Your humans are your organic resources. Your agents are synthetic resources. Both are workers you assign, manage, and hold accountable, the difference is that synthetic resources are cheap, fast, tireless, and dumber than they look. And like any team, you don't get results by hiring one omnicompetent hero. You draw an org chart. You write narrow job descriptions. You decide who is allowed to talk to whom, who hands off to whom, and above all, who is not allowed to touch a given decision. Designing an agentic system is an org-design problem wearing a software costume. - Give every agent one job title and one definition of done. - Constrain its inputs to only what that job needs, no more. - Constrain its tools to only the actions that job performs. - Make handoffs explicit: a finished piece of work, passed to the next role. - If you can't write the job description in a sentence, the role is too big. A worked example: a team that writes a book Say you want agents to help write a historical novel. The tempting design is one "author agent": give it the premise and let it write the book. It will produce something fluent and plausible that's also full of invented history, characters who change personality between chapters, and a plot that forgets its own first act. Now design it like a publishing house staffed by narrow specialists instead. Think about every role a real book actually requires, and give each one a single seat at the table: The roles - Story architect, owns structure: the outline, act breaks, and pacing. It never writes prose. It decides what happens and in what order. - Character keeper, owns the cast: each character's voice, motivation, and continuity. Its only job is to keep people consistent from chapter to chapter. - History expert, owns factual grounding for the era and place the book is set in. Dates, customs, technology, what people ate and wore. Nothing else. - Prose writer, takes a single beat from the architect, the relevant characters, and the vetted facts, and writes that scene. Just that scene. - Continuity editor, reads the assembled draft against the outline and the character bible, and flags only contradictions. - Line editor, tightens language. It doesn't change plot, facts, or character. It makes sentences better. None of these agents is smart. Each is almost insultingly narrow. But together they produce a coherent book, because no single one of them is ever asked to hold the entire problem in its head at once. Constrained resources are how you kill hallucination Look closely at the history expert, because it's where the whole philosophy pays off. A general writing agent hallucinates history because it's guessing from whatever it half-remembers while juggling plot and prose. It has every reason to make something up and no reason not to. The history expert agent is built so it physically can't do that. You don't hand it the plot. You hand it a constrained, curated corpus, a vetted set of sources for that exact era and region, and one instruction: answer questions about this period, grounded only in these materials, and if the answer isn't here, say so. The job is narrow and so is everything you give it to work with. There is simply nothing else for it to do. "You don't prevent hallucination by asking the model to be more careful. You prevent it by removing the room it had to make things up." That's the move, generalized: constrain the resources to fit the role. Give it the context, tools, and reference material the job needs, and not one thing more. A focused agent with a tidy, bounded world is far more reliable than a powerful one staring at everything at once. How to actually break a system down When you're staring at a problem you want agents to solve, don't ask "what should the agent do?" Ask the question a manager asks: if I were staffing this with a team of cheap, narrow specialists, who would I hire and what would each person's one job be? - List the roles the work genuinely requires, the way you'd staff a real team for it. - Write each role's job description in one sentence. If you can't, split it. - For each role, define its definition of done, the finished artifact it hands off. - Give each role only the context and tools that one job needs. - Wire the handoffs: who passes work to whom, and in what order. - Add a reviewer role whose only job is to catch a specific class of mistake. You'll end up with more agents than you expected and each one will be simpler than you expected. That's the goal. Simplicity per agent is what buys you reliability across the system. Why this scales when "one smart agent" doesn't A team of narrow agents is easier to reason about, easier to test, and easier to fix. When something goes wrong, you don't debug a black box, you find the one role that failed and you fix that role, the same way you'd coach one employee. You can upgrade the history expert without touching the prose writer. You can swap a cheaper model into the line editor and a stronger one into the story architect, matching horsepower to difficulty. It's also how the economics work. Most steps are simple enough for small, fast, cheap models. You spend real capability only where the job is genuinely hard, instead of paying frontier prices to have one giant agent do trivial work badly. So resist the urge to hire a genius. Hire a team of cheerful idiots, give each one a single clear job and exactly the resources that job needs, and let the org chart do the thinking. That's not a workaround for today's models being limited. It's just good management, and it happens to be the most durable way to design agentic systems. Building teams like this inside real companies is the work I do as an AI-native leader. If you're staffing your first synthetic org chart, let's talk → ### Is Agile Dead? Why AI Made the Ceremonies a Tax on Speed: /blog/agile-is-dead/ Agile was a workaround for the slow, expensive nature of building software in 2001. AI repriced the building. The ceremonies that compensated for the old expense are now a tax paid in the exact currency they were invented to protect: speed. There's an awkward observation that anyone who has worked in software for more than a decade has had and not quite said out loud. The companies most aggressively performing "Agile", the ones with the certified Scrum Masters, the burndown charts, the immaculate Jira hygiene, the rotating ceremony calendar, the SAFe consultants in conference rooms, are almost universally not the most agile companies. They are often the slowest. The companies that actually move quickly tend to have abandoned most of the ceremony years ago and feel mildly embarrassed about how informal their process is. This was already true before AI. AI made it impossible to ignore, so let me answer the title's question directly. Yes, Agile the methodology is dead, because it was a workaround for the slow, expensive nature of software development in 2001, and AI-assisted development has removed the expense it was compensating for. The ceremonies are now a tax on the thing they were invented to protect: speed. Here is the fuller form of the claim. Agile encoded a set of rituals designed to reintroduce responsiveness into a process that was otherwise far too rigid to be responsive. With AI-assisted development, the underlying expense is gone. The methodology that compensated for it now costs more than it returns, and the cost is paid mostly in the currency it was originally designed to protect: speed and direct contact with reality. "Agile is dead in the sense that matters: its ceremonies were workarounds for expensive building, and AI has repriced the building." What's left, in most companies, is a corpse running on muscle memory and consulting revenue. What Agile was actually solving for It's worth going back to what Agile was responding to, because most people defending Agile in 2026 have forgotten. In 2001, the dominant software methodology was waterfall, or something close to it. You would spend months writing requirements documents. Then more months writing design documents. Then a long development phase where engineers built what the documents said. Then a testing phase. Then a release. Then, sometime months later, you would find out the customer wanted something different, or the market had moved, or the documents had been wrong all along. The cost of this was enormous, not just in dollars but in reality lag. Companies were shipping software designed for the world as it existed at the moment the requirements were frozen, often a year or more out of date by the time it launched. The Agile Manifesto was a reaction to this. Read it again sometime. It is not a list of ceremonies. It is four pairs of preferences: - Individuals and interactions over processes and tools - Working software over comprehensive documentation - Customer collaboration over contract negotiation - Responding to change over following a plan Notice what isn't there. Sprint planning isn't there. Neither are daily standups, story points, retrospectives, velocity, or burndown charts. The Spotify model, SAFe, and Jira are all absent too. All of those things, the entire industrial complex of "Agile transformation", are workarounds for the underlying expense of software development that the Manifesto was trying to make tolerable. If you can only ship every few months, you need ceremonies to keep yourself from drifting. If your team is twelve people, you need standups so they don't collide. If your build takes hours, you need careful sprint planning because you can't just try things. If your customer can only review at quarterly intervals, you need backlog grooming to manage the queue. These were all sensible responses to the engineering reality of 2001 through 2020. None of them was the point of building software. They were the price you paid to keep building it. The price wasn't free Even in the era they were designed for, the ceremonies came with costs that most teams underweighted. Every hour spent in standup is an hour not spent making something. Story point estimation is a small theater of false precision. A sprint boundary is an artificial commitment that distorts the work crossing it, and most retrospectives produce action items that quietly die in the next sprint. Backlog grooming is the team sorting a queue that reality will invalidate within weeks. Most teams accepted these costs because the alternative, total chaos, was worse. But "Agile is better than chaos" is a much weaker claim than "Agile is the optimal way to build software." The first is probably true in most pre-AI contexts. The second was never quite true, and the gap was always being paid in friction. The companies that figured this out early, often very small teams or very mature ones, stripped the ceremonies down to almost nothing and operated on something closer to "constant contact with reality, build, ship, repeat." They were called undisciplined. They were called "not really Agile." They were also, suspiciously often, the companies shipping interesting work. What changes when the cost of building collapses Now drop AI into the picture. The cost of writing code falls by an order of magnitude. The cost of generating a prototype falls to almost nothing. A motivated engineer with AI assistance can build and ship in a day what used to take a sprint. I've written elsewhere about what happens when building gets cheap; here the point is narrower. The build, ship, observe, respond loop, the one Agile was inventing rituals to preserve at a tolerable speed, can now happen naturally and continuously. Notice what this does to the ceremonies. Sprint planning batched decisions about what to work on, because changing direction mid-sprint was expensive. When changing direction takes an hour instead of a week, what is the sprint protecting? Mostly itself. It has become a coordination boundary defending nothing. Daily standups surfaced blockers and coordinated dependencies back when unblocking yourself was expensive. When you can sketch your own approach with an AI in fifteen minutes, the standup is largely a status meeting in which people perform "I'm working on the thing I said I was working on" for an audience that no longer needs the information. Story point estimation managed uncertainty about how long work would take, in a regime where work was expensive enough that mis-estimation had real consequences. When tasks that used to take five days take half a day, the resolution of the estimate is below the noise floor of the estimation process. You're rounding everything to "small" and pretending it tells you something. Backlog grooming kept a sorted queue of work that exceeded capacity. When capacity expands faster than the queue, the queue is no longer a constraint. It's a museum of features someone once thought important. Retrospectives were a way to compensate for the fact that direct feedback loops with reality were slow. When the feedback loop is continuous, the formal retrospective is a meeting in which the team rediscovers, in a structured way, things they already learned in real time. Velocity was a metric designed to enable predictability under a build-is-expensive regime. It encouraged teams to estimate consistently and ship consistently. In an AI-assisted world it encourages teams to produce more measurable output, which is exactly the wrong incentive in a world where the constraint is now what should be built and verified, not how much. The methodology didn't fail. Its enabling conditions evaporated. The ceremonies are still happening because organizations are sticky, certifications are sticky, consultants are sticky, and middle managers built their identities around "running Agile." The ceremonies are no longer doing what they were designed to do. They are doing something else, mostly creating the appearance of discipline while consuming the speed they were invented to protect. Agile became the thing it was invented to kill This is the cruel irony. Read the Manifesto again. Responding to change over following a plan. In most large companies in 2026, "doing Agile" means following a plan: the sprint plan, the quarterly OKR plan, the release plan, the PI planning artifact from the SAFe deck. Changes mid-sprint are treated as failures of discipline. Customer signals that arrive between planning cycles are queued for the next planning cycle. The team is responsive to its process, not to reality. "The thing that was invented to kill waterfall has become waterfall in two-week increments." The diagnostic test is brutally simple. Walk into any team that "does Agile" and ask: If a customer told you something important on Monday, when is the earliest your work would meaningfully change in response? If the honest answer is "after sprint planning next Wednesday", you are not Agile. You are running a slightly more frequent waterfall and calling it something else. The irony runs deeper than that, because real waterfall, big design up front, is quietly coming back, and for once with good reason. In an AI-assisted environment, the honest answer should be "this afternoon." If your process makes that impossible, your process is the problem. What "agile" actually means when building is cheap Here is the strong form of the claim. Being agile, lowercase, is no longer a methodology. It is the default state of any sufficiently AI-enabled team that hasn't been talked out of it. When generation is cheap, observation is cheap, and iteration is cheap, the natural mode of a small competent team is something like this: - Someone notices a signal from a user or from a metric. - An engineer, often with AI assistance, sketches a response in hours. - It ships to a small surface area. - The signal it generates is observed. - The team adjusts, often that same day. There is no sprint, no standup, no story estimation, no backlog grooming. There is constant contact with reality and continuous response. The team is, in the literal sense of the Manifesto, doing exactly what Agile was supposed to produce, and doing it without any of the rituals. This is what good teams have always wanted to do and what the ceremonies were a clumsy attempt to approximate. The ceremonies are no longer the best available approximation. The thing itself is now accessible. Why companies keep doing the ceremonies anyway If this is obvious, why doesn't every team just drop the ceremonies? Because the ceremonies are doing other work the methodology never explicitly named. They are legibility theater for management. Standups produce the daily report that lets a director feel they have visibility. Sprints produce the cadence of commitments that lets a VP plan the quarter. Velocity charts produce the metric that goes in the slide. None of these is about productivity. All of them are about producing organizational artifacts that make a non-technical leadership feel comfortable. They are risk distribution. If a team commits to a sprint plan and misses, the failure is bounded and explainable. If a team operates with no sprint and ships less than expected, the failure is diffuse and harder to defend. Ceremonies create defensible failure modes. Defensible is not the same as productive, but it is politically safer. They are job descriptions. There is now a large workforce, certified, paid, and identifying as Scrum Masters, RTEs, Agile coaches, Chapter Leads, Tribe Leads. Their jobs require the ceremonies to exist. They are not bad people. They are doing the work they were trained to do. But the work they were trained to do was the work of compensating for a constraint that no longer binds. They are the path of least resistance. Removing a ceremony requires someone to take the political risk of removing it. Keeping one requires no one to do anything at all, and that asymmetry decides most of these calls. So the ceremonies persist. The performance of Agile continues. The actual agility of the organization erodes underneath it. And the most AI-fluent teams quietly route around the process and do what works. What the post-Agile team actually looks like Let me describe what I see, increasingly, in teams that have stopped pretending. - Almost no recurring ceremonies. A short, often-asynchronous check-in if there is one at all. No standup. No sprint planning. No retrospectives on a calendar. Reviews happen when something is worth reviewing, which is often. - A short, living thesis instead of a roadmap. A paragraph or two describing what the team believes about its market, its users, and its current bets. Updated when the belief changes, which is often. This is not a roadmap. It is a north star. - Work expressed as small, discrete experiments. Not stories. Not epics. Probes. Each one designed to generate a learnable signal in days. Most of them ship behind feature flags or to narrow surface areas. - Continuous, not batched, verification. Heavy investment in tests, evals, observability, alerting. Because building is cheap and verification is now the bottleneck, the team's process discipline is concentrated there, not in pre-build ceremony. - Direct producer-to-user contact. The engineer who built the thing is the one watching it in production, often talking to the users it affected. The translation layer of PMs, designers, and analysts who used to mediate this contact is thinned dramatically. - A single, clear question instead of OKRs. Some teams literally write one sentence per quarter describing what they are trying to find out. The success criterion is whether they found out, not whether they hit a key result. The aesthetic of this is the opposite of corporate Agile. It looks unstructured. It looks undisciplined. It is, in practice, dramatically more disciplined, because the discipline is in the things that matter (verification, signal, decision quality) instead of the things that look like discipline (estimation, sequencing, ritual attendance). What to do about it If you are a leader of a software organization, the move here is uncomfortable but not complicated. Audit your ceremonies for purpose. For each one, ask: What underlying problem does this exist to solve? If the answer is "we have always done this", drop it. If the answer is "it gives management visibility", consider whether you can give management visibility some other way that doesn't tax the team. If the answer is real and current, keep it. Most won't survive this audit honestly applied. Stop measuring velocity. It is now a metric for the wrong constraint. Measure verification quality, customer signal speed, and durability of shipped work. These are harder to graph. They are also what you actually want. Be honest about the role of Scrum Masters and Agile coaches. Some are excellent facilitators who would create value in any process. Some are guardians of ceremonies that have outlived their purpose. The honest answer about their role in an AI-native team is uncomfortable, but pretending otherwise wastes their careers and your money. Replace sprint commitments with weekly direction. A sentence. This week, we believe the most important thing to learn about is X. Adjusted as reality demands. No commitment ceremony required. Push decision rights down to the engineer doing the work. AI gives the median engineer the capacity to make decisions that used to require coordination. If your process still requires the coordination, you are deliberately throwing away the leverage. A closing observation The most agile teams I see in 2026 are not the ones who have "successfully transformed to Agile." They are the ones who took the Agile Manifesto seriously enough to notice that the ceremonies were never the point, and who had the conviction to operate as if responding to change actually mattered more than following a plan. These teams are uncomfortable to be around if you came up in the SAFe era. They feel unstructured. They feel risky. They produce, in many cases, an order of magnitude more learning per quarter than their ceremony-rich peers, and the work they ship has the strange property of feeling like it actually fits the moment. Agile is dead in the sense that the methodology has outlived the conditions that produced it. Lowercase agility is more available now than it has ever been, because the thing that made building slow has been substantially repriced. The choice for organizations is whether to keep performing the methodology that was a workaround for the old expense, or to claim the actual thing it was a workaround for. Most companies will keep performing the methodology. The ones that don't will look, in three years, like they discovered some new productivity secret. They didn't. They just stopped paying a tax that no one had told them was now optional. Dismantling the ceremonies without losing the discipline underneath them is the heart of an AI-native transformation, and it's work I've done inside real organizations, not from a slide deck. If your process is the thing slowing your process down, let's talk → ### What Is an AI-Native Operating Model? Digestion, Not Adoption: /blog/ai-digests-your-company/ AI is the first technology that metabolizes organizations rather than augmenting them. Going "AI-native" doesn't upgrade your operating model, it dissolves the structure that existed to manage problems AI just erased. Every consulting deck in 2026 has the same slide. It shows a "traditional" org on the left, boxes, hierarchies, swim lanes, and an "AI-native" org on the right, usually with the same boxes plus a friendly little robot icon clipped to each one. The arrow between them is labeled transformation. The slide is wrong twice. It's wrong about the starting point, and it's wrong about the destination. The starting point isn't a company. It's a settlement, a temporary truce between people who disagreed about what the company should do, encoded into approval chains, ticket queues, and quarterly reviews. The "process" everyone talks about isn't the work; it's the political residue of past disagreements. The destination isn't an "AI-native" version of that settlement, because there is no such version. AI doesn't transform the settlement. It dissolves it. I want to argue something stronger than the usual "AI is a big deal" line: AI is the first technology that metabolizes organizations rather than augmenting them. The companies that understand this are quietly restructuring around it. The ones that don't are spending 2026 doing what amounts to expensive cargo-cult ritual, installing copilots on top of decision rights that no longer have any reason to exist. The standard view, and why it's a trap The dominant framing in executive AI conversations goes something like this: identify high-volume, low-judgment tasks; deploy AI to handle them; redeploy humans to "higher-value work." Pick your favorite consulting firm's name for it. They all converge on the same shape. This framing has a hidden assumption baked in: that the org chart is the right substrate for AI to sit on top of. You keep the functions, the roles, and the reporting lines, and you treat AI as labor that slots into the existing economy of work: cheaper, faster, and easy to scale. The problem is that most organizational structure isn't a response to the work. It's a response to the cost of coordinating humans doing the work. Most middle management exists because information is expensive to move, decisions are expensive to escalate, and humans are expensive to verify. Approval chains exist because we don't trust each other's judgment at scale. Quarterly planning exists because we can't replan continuously without dissolving into chaos. Functional silos exist because specialization was the only way to get competent execution out of generalist humans. Every single one of those constraints is being repriced right now. When information moves at the speed of an LLM, the org chart starts to look like a roadway built for horse-drawn carts. You can put a Tesla on it. You can put a thousand Teslas on it. The road is still the bottleneck. What "AI-native" actually means Drop the brand-speak for a second. An AI-native company is one whose decision structure was designed assuming AI is present in every loop. Not its tooling. Not its product. Its decision structure. I've written about what that shift looked like in my own work; this is the organizational version of the same move. This has weirder implications than the keynote-friendly version. A few that don't make it into the consulting deck: Approval chains stop being information bottlenecks and start being political theater. The reason a VP needs to approve a $50K spend isn't because the VP has unique insight into whether the spend is wise. It's because the VP is the accountable party if it goes wrong, and the approval is a way of distributing risk. AI doesn't change the accountability question, but it does eliminate the information asymmetry the approval was pretending to address. So the chain shortens, or it becomes honest about what it actually is, which most companies aren't ready to do. Functional expertise becomes a cost rather than a moat. For most of the last fifty years, hiring a great VP of [Function] was a defensible source of organizational advantage. Their judgment was rare. Now? The median LLM has read more about [Function] than your VP has, and can produce a defensible first-cut analysis in eleven seconds. The VP's value isn't gone, but it's moved off answer generation and onto answer selection, the organizational legitimacy to back a call, and the political will to push it through. Those are real, but they're smaller jobs than what VPs are currently paid for. Planning cycles collapse. Annual planning, OKR cycles, quarterly reviews, these are all batch processes designed for a world where replanning was expensive. When replanning is cheap, batch processes stop being a feature and start being friction. The companies that lean into this don't have an annual plan in any recognizable sense. They have a direction and a replanning cadence measured in weeks. This terrifies CFOs, and rightly so: most financial governance assumes batch planning. We don't yet have the financial primitives for continuous replanning at scale. We will. "Best practices" become liabilities. A best practice is an answer that was correct given the constraints of its era. Most management best practices were forged in a world of slow information, scarce expertise, and expensive coordination. None of those constraints hold anymore. Companies that continue to optimize around them are spending real money to maintain irrelevant solutions to dissolved problems. The metabolism analogy I keep using the word metabolize deliberately. Most prior technologies augmented organizations, they added capabilities without changing the underlying structure. The PC didn't dissolve middle management; it gave middle management spreadsheets. Email didn't flatten the hierarchy; it gave the hierarchy a faster way to forward things to each other. Even the cloud, for all its rhetoric, mostly let companies do what they were already doing, with less hardware. AI is different in kind. It doesn't add a capability to an existing role; it absorbs the role's primary function. Then it does the same to the role above it. Then to the role above that. Each absorption forces a structural question that the company can either answer or ignore. "AI doesn't add a capability to a role. It absorbs the role. Then the role above it. Then the one above that." The companies that answer the question shrink, in a way that looks like cost savings but is actually something more interesting: they're shedding the organizational tissue that existed to compensate for problems that no longer exist. They're metabolizing themselves down to load-bearing structure. The companies that ignore the question are doing what biologists call "adipose accumulation." They keep the existing structure and layer AI on top of it as a kind of cognitive fat: expensive, easy to point at in a board deck, and mostly inert. These are the companies that report "100 AI initiatives" in their annual reports and quietly wonder why none of them moved a P&L number. What survives If you take the metabolism view seriously, you can predict, with surprising accuracy, what survives and what doesn't. What gets digested: information-routing roles, single-domain analysts, anyone whose job is to translate between two parts of the organization, anyone whose primary skill is reading a document and producing a slightly different document. Most of middle management. Most of corporate strategy as currently practiced. Large portions of legal, HR, and finance that exist to enforce compliance with internal policies. What survives, often expanded: anyone close to a feedback loop with external reality. Sales (because customers are still humans). The people who hold accountability for outcomes, not the ones who process information about outcomes. Anyone whose work involves making decisions under genuine uncertainty, with skin in the game. Anyone doing taste-dependent work where the metric isn't yet legible to a model. What changes shape: engineering, product, design, operations. These don't disappear, but their internal economies invert. Building stops being scarce; deciding what not to build becomes the constraint. (I'll come back to this in the product management piece.) The crude version of this argument has been around since 2023, "AI will eliminate jobs X, Y, Z." That framing missed the point. The job isn't the unit of analysis. The decision is. AI is dissolving the layers of the organization that existed to manage the cost of decisions, not the layers that existed to make the decisions. "Most of what your company does, structurally, is no longer necessary. The work is necessary. The structure isn't." So what do you actually do If you're running a company in 2026, I'd offer four uncomfortable suggestions. One. Stop counting AI initiatives. Start counting decisions that have collapsed. The right metric isn't "how much AI have we deployed." It's "how many approval steps have we removed because the information asymmetry they addressed no longer exists." If your AI program isn't producing structural simplification, it isn't producing value. It's producing fat. Two. Take the org chart down off the wall and stare at it for an hour. For each box, ask: would this box exist in a world where the median employee can summon a competent analyst in three seconds? Would this approval exist? Would this committee exist? Most of the answers are no. You don't have to act on all of them. But you should know which boxes are load-bearing and which are inherited politics. Three. Invest aggressively in verification, not just generation. The thing AI is worst at is checking its own work in domains with consequence. Every dollar your company spends on AI generation should be matched by a dollar on the human, process, and tooling infrastructure for verification. The companies that get this wrong will discover, somewhere around 2027, that they have automated their way into a slow-rolling reliability crisis. Four. Be skeptical of anyone selling you "AI-native" as a destination. It isn't a destination. It's a continuous renegotiation between what the technology can do and what the organization can absorb. The companies that treat it as a one-time transformation will be in this exact same conversation, with new consultants, in three years. A closing thought The cleanest organizations I see right now aren't the ones that have deployed the most AI. They're the ones whose leaders have made peace with a difficult fact: most of what their company does, structurally, is no longer necessary. The work is necessary. The structure isn't. This is hard because the structure isn't abstract. It's people. It's careers. It's identities. The hardest part of being AI-native isn't technical adoption; it's the willingness to admit that a large chunk of the organization exists because of constraints that no longer apply, and to do something about it without pretending you're doing something else. The companies that do this well will look, in five years, like remarkable performers: lean and fast, and weirdly profitable for their size. The companies that don't will be running expensive AI programs on top of unchanged structures, generating impressive demos and unchanged income statements, and wondering why the promised transformation never quite arrived. It arrived. They just refused to be metabolized. If you'd rather not be one of them, figuring out which structure to shed and which is load-bearing is exactly what my AI strategy work covers. Let's talk → ### Is Product Management Dead? AI Removed Its Reason to Exist: /blog/product-management-was-a-workaround/ Product management was an optimization layer for a constraint, expensive software, that no longer exists. When building gets cheap, the PM's job inverts: from advocate for what gets built to editor of what shouldn't. There's a specific kind of awkwardness in product organizations right now that nobody is willing to name out loud. PMs are still running their rituals, discovery, prioritization, OKRs, JTBD interviews, RICE scores, roadmaps in three tidy columns labeled Now / Next / Later. Engineers are still going to those rituals. Designers are still building artifacts to feed them. The whole machine is humming. And underneath, nobody is sure what any of it is for anymore. I want to make a specific argument: product management as we know it was an optimization layer for a constraint that no longer exists. The constraint was that building software was expensive, so deciding what to build had to be careful, batched, and centralized. The PM role was the workaround. AI dissolved the constraint. The workaround is now overhead. This isn't a "PMs will be replaced by AI" essay. That framing misses what's actually happening. PMs aren't being replaced. The job is being inverted, and most PMs are still doing the old job in the new world, with predictably bad results. The forgotten origin of PM It's worth remembering where the modern PM role came from. It crystallized in the 2000s and 2010s, in a very specific context: software was getting cheaper to build than ever before, but it was still expensive enough that a wrong bet cost six months of engineering capacity. Engineers were scarce. Engineering capacity was the binding constraint of the business. So you needed a role whose entire job was to pre-filter engineering work, to take the chaos of customer needs, market opportunities, executive whims, and competitive pressure, and convert it into a sorted, deduplicated, prioritized backlog that engineers could execute against without wasting their precious time. Every PM framework you know is downstream of that constraint. RICE? A prioritization heuristic for allocating scarce engineering capacity across more demand than capacity. JTBD? A method for reducing the dimensionality of customer requests so a small team could focus. OKRs? A coordination mechanism for making sure separate teams of expensive engineers were rowing in roughly the same direction. Roadmaps? A communication artifact telling stakeholders to stop asking for new things because the capacity is already booked. Every one of these tools is an answer to the same underlying question: we can only build a little; what should the little be? Now look at 2026. The cost of building has fallen by an order of magnitude in maybe eighteen months. A competent engineer with AI assistance is shipping in a week what used to take a quarter. A motivated non-engineer is shipping working software at all, which is a discontinuity nobody is fully reckoning with. The cost of building hasn't gone to zero, but it's gone low enough that building is no longer the bottleneck. Deciding what to build is no longer scarce-capacity allocation. It's something else entirely. What the new constraint actually is Here's the inversion: when building was expensive, the danger was failing to build the right thing. Errors of omission. So PM evolved into a discipline of careful selection, find the highest-value bets, sequence them well, don't waste engineering cycles. When building is cheap, the danger is building too many things. Errors of commission. Every feature you ship is a feature you must support, document, train people on, integrate with the rest of the product, and eventually deprecate. The cost of a feature didn't decrease proportionally with the cost of building it. Building is cheap; owning is not. "The cost of a feature never fell as fast as the cost of building it. That gap is the whole story." This is the central fact most PM organizations are not yet metabolizing. They're treating cheap building as a license to build more, sprint velocity up, feature count up, roadmap density up, when the actual implication is the opposite. In a cheap-building world, the PM's job is to be the immune system of the product. To say no. To kill features. To resist the gravitational pull of "we could ship that in two days." To protect the conceptual integrity of a product that everyone, including AI agents on the engineering side, can now contribute to with very low friction. This is a much harder job than the one PMs trained for. The old job was to be the advocate for what got built. The new job is to be the editor of what gets built. That is a different temperament and a different place in the org, and it rewards a different set of skills. Why most PM frameworks are now anti-patterns If the constraint has inverted, then every framework that optimized for the old constraint is now optimizing in the wrong direction. Specifically: Prioritization frameworks. RICE, ICE, MoSCoW, these are all sorting algorithms over a backlog. They presuppose that the backlog is the universe of valuable work and the question is what to do first. In a cheap-building world, the backlog is infinite by default, and the question is what work should not exist. Sorting an infinite list is a waste of time. The new tools are filters, not sorts. "Why does this item exist?" "What would happen if we deleted it?" "What's the half-life of this feature once shipped?" These questions are not in the standard PM toolkit. OKRs. OKRs are a coordination protocol for keeping large, slow-moving organizations roughly aligned over quarters. They presuppose batch planning cycles. They presuppose that the cost of changing direction mid-quarter is high enough to justify the overhead of pre-committing to specific outcomes. Neither presupposition holds anymore. Most companies running OKRs in 2026 are running a quarterly theater of commitment over plans that everyone privately knows will change in week three. Roadmaps. A roadmap is a communication artifact whose primary function was to set expectations with stakeholders about what they would not get. ("It's not on the roadmap.") That function relied on the underlying scarcity of engineering capacity. In a world where building is cheap, the roadmap becomes either dishonest (we're not really committing to this sequence) or constraining for no reason (we could ship that next week but the roadmap says Q3). The honest replacement is something like a thesis, what we believe about the market and why we're investing in this direction, combined with very short, high-confidence near-term commitments. Discovery as a discrete phase. Classic discovery (interview customers, synthesize insights, identify opportunities, generate concepts) was a batch process designed to reduce expensive build risk. When build is cheap, the cheapest form of discovery is to build the thing and see what happens. Not always; some bets are still too consequential. But the default has flipped. The new question isn't "have we discovered enough to justify building" but "is building so cheap here that discovering through deployment beats discovering through interviews." JTBD and persona work. These remain useful, but as editorial filters rather than generative inputs. Their original use was to focus a scarce engineering team. Their new use is to be the lens through which the immune system says no. The PM as editor If we take the inversion seriously, what does the role look like? It looks much more like a magazine editor than a project manager. The editor doesn't generate the articles, writers do, and in the new world, "writers" includes AI agents and engineers prototyping freely. The editor's job is to maintain the voice and integrity of the publication: to reject pieces that don't fit, to commission specific work to fill identified gaps, to be the keeper of "what this publication is and isn't." The editor is judged not by how many articles they cause to be written, but by the coherence and reputation of the publication over time. "An editor is judged by the coherence of the publication, not the number of articles they cause to be written." This is a fundamentally different posture than the modal PM job description. It selects for different skills: - Taste. The ability to look at a feature and feel, before reasoning your way there, whether it belongs. - Comfort saying no. Including to stakeholders who outrank you, and including to AI-generated proposals that look superficially compelling. - Conceptual ownership of the product as a system. Not as a backlog. Not as a roadmap. As a coherent thing with edges. - Willingness to be unpopular in the short run. Editing is a no-saying profession. PMs trained on stakeholder management and consensus-building find this culturally hard. Most companies do not currently hire, evaluate, or promote PMs on these dimensions. They hire on framework fluency, stakeholder skills, and ability to ship. Which means most companies are about to discover that their best PMs in the old regime are not their best PMs in the new one. The one-PM-many-experiments pattern One organizational pattern is emerging that I think is worth naming explicitly because it doesn't have a clean name yet. In the old model: one PM, one engineering team, one feature, one ship cycle. The PM's job was to be the single decision-maker for that team's output. The ratio was roughly one PM to five-to-eight engineers. In the emerging model: one PM, many parallel experiments, many of them AI-driven. The PM is curating a portfolio of probes rather than managing a sequence of releases. They might have one or two human engineers and the equivalent of a dozen AI-driven build streams running in parallel. Their job is mostly to look at what came back, decide what's interesting, kill what isn't, and decide what to invest more deeply in. This is closer to the way a venture investor works than the way a traditional PM works. The unit of decision is the experiment, not the feature. The skill set is portfolio judgment, not roadmap construction. The metric is signal-to-noise ratio on the experiments, not throughput on the backlog. Companies that adopt this model can run with dramatically fewer PMs covering dramatically more product surface area. The PMs they keep are paid more and trusted more. The org chart compresses. What to do if you're a PM right now Three things. None of them are easy. Stop measuring yourself by output. Throughput is no longer scarce; if you're competing on throughput, you're competing with agents that ship faster than you can prioritize. Measure yourself on what didn't get built that shouldn't have, and on the long-run coherence of the surface area you own. These are harder to point at in a perf review, which is exactly why they're now where the value lives. Get embarrassingly close to your product. Use it. Build in it. Break it. Watch others use it. The PM whose value was synthesis of others' reports is being absorbed quickly. The PM whose value is direct judgment about the thing itself is harder to absorb because their input is taste, which models still struggle to replicate at consequence-bearing levels. Be willing to argue for fewer features in front of executives who reward "shipping." This is the political price of the editor role. Most PMs will not pay it, which is why most PMs will be slowly absorbed into the very AI workflows they were supposed to be managing. A closing observation The most interesting product organizations I see in 2026 don't have the most PMs or the slickest frameworks. They have someone, often a PM but just as often a founder or one staff engineer with strong taste, who is willing to be the editor. To say no, this isn't us. To delete features. To resist the gravitational pull of "we can build that in an afternoon." Those organizations have products that feel like products. The ones without that role have products that feel like accumulations. PM as a function isn't dead. The framework-driven version of it is. What's emerging in its place is smaller and a lot harder to do well, and the people who see the shift early are going to be in demand for a while. I've written about the builder's side of this same merge in product engineering was always one job. And if your product org is mid-inversion, rebuilding it around editors and outcomes is the work I do as an AI-native leader. Let's talk → ### Should You Automate Customer Support With AI? Mostly No: /blog/support-tickets-are-evidence/ A support ticket isn't a cost to be deflected. It's the highest-fidelity evidence you have about where your product, pricing, and onboarding are broken, and the standard AI playbook is quietly destroying it. Why support tickets are evidence, not workload Almost every company I talk to is doing the same thing with AI in customer service. Almost all of them have the wrong reason for it. The playbook is identical wherever you look. Deploy an AI agent to handle Tier 1. Deflect as many tickets as possible. Reduce average handle time. Show the board a chart of cost-per-ticket trending down. Reallocate support headcount, or, more often, just don't refill the seats that leave. The end result: you've taken the most consequential information-processing technology of your career and pointed it at making your help desk cheaper. I want to argue that this is one of the most expensive misallocations of AI happening at scale right now. Not because support automation doesn't work, it works fine. But because a support ticket is not workload to be eliminated. It's high-bandwidth, voluntarily-supplied evidence about exactly where your product, pricing, onboarding, or messaging is broken. And the companies treating tickets as workload are doing something that, in any other discipline, we would call destroying the data. This is the contrarian view I want to defend: customer service in the AI era should grow in influence, not shrink. The companies that get this right will use AI to amplify the signal coming out of their support function rather than silence it. The ones that don't will spend the next three years quietly losing the plot of their own product. The accounting trick that hides the real cost Let's start with why the standard playbook is so seductive. Support is an unambiguously legible cost center. Every ticket has a cost. Every deflection is a savings. You can put both on a spreadsheet, project the trend forward, and tell a perfectly clean story to your CFO about the ROI of AI in support. That story has one defect: it accounts only for the cost of the ticket. It does not account for the information content of the ticket. And it does not account for the cost of not knowing what the ticket would have told you. Here's the trick. When a customer files a ticket, three things happen at once. First, they consume support capacity. That's the cost everyone measures, and the only one that lands on a spreadsheet. Second, they reveal a specific failure mode of your product, pricing, documentation, or onboarding. They tell you, with precision, which thing didn't work for them and what they tried instead. This is genuinely high-quality user research, except it's free, voluntary, and self-prioritizing (the most common failures generate the most tickets). Third, they make an implicit decision about your company: whether the friction was worth tolerating, whether to churn, whether to write a review, whether to mention you to a colleague. That decision is made partly on the basis of how the ticket was handled, but also partly on the basis of whether the underlying issue ever gets fixed. The standard AI playbook optimizes the first cost and destroys access to the second and third. When an AI agent deflects a ticket, no human ever sees it at all. What you get instead, if you're lucky, is a row in an analytics dashboard about deflection rates. The actual content, what the customer was trying to do, where they got stuck, what assumption broke, gets summarized into a category, aggregated into a metric, and lost. This is, structurally, the same mistake as if a doctor said: "We've automated the part where the patient describes their symptoms; we now just send them home with a prescription, and our throughput is way up." The throughput is up. The medicine has gotten worse. What a ticket actually is Think about how expensive it is to acquire the kind of information a support ticket contains. If you wanted to learn, through traditional user research, that 6% of your users get stuck at exactly the same point in onboarding because they assume a specific button does something other than what it does, you would have to: hire a researcher, recruit participants, run sessions, watch tape, synthesize findings, present them, get buy-in, prioritize against other research, and finally route the result to the team that owns the relevant flow. Months, easily. The same fact arrives in your support queue every day, for free, with the customer's own words attached, and self-prioritized by the volume of tickets it generates. Companies have always had this signal. They've mostly wasted it, because converting tickets into product changes was organizationally expensive: you needed support to surface trends, product to listen, engineering to act, and all three teams to align on priority against everything else they were doing. The activation energy was too high. So tickets piled up, got categorized, got closed, and the underlying product issues persisted. The AI moment is the exact moment that activation energy could collapse. You can now read every ticket, cluster them automatically, identify root-cause hypotheses, link them to the relevant product surface, draft a proposed fix, and put it in front of the team that owns it, all without a meeting. The bottleneck that made it impractical to actually learn from your support data has, technologically, dissolved. And what is almost everyone doing instead? Building a chatbot to make the tickets go away faster. "A support ticket isn't workload to be eliminated. It's voluntarily-supplied evidence about exactly where your product is broken, and most companies are paying customers to generate it, then throwing it away." The "deflection" trap The word deflection tells you everything. It's the language of someone treating the customer as a nuisance and the ticket as friction. It's not the language of someone treating the customer as a witness. Deflection metrics distort your incentives almost immediately. If the AI agent is judged on resolution-without-escalation, it will learn to "resolve" issues by being convincingly non-committal, recommending the customer try common workarounds, or offering small credits. This is not the same as fixing the issue. It is the same as moving the cost of the issue from the support team to the customer, who now has to live with the workaround, and then either churns quietly or generates the same ticket again next month. You can measure this directly if you want. Look at your "deflected" customers' six-month retention compared to your "escalated and resolved by a human" customers. In most companies that have run the comparison honestly, the deflected cohort churns at meaningfully higher rates. The "savings" from deflection were largely a tax on customer lifetime value, paid in installments later. This is the part of the AI-in-support story you won't find on the consulting deck, because it doesn't survive a careful look at the unit economics. "Deflection is the language of someone treating the customer as a nuisance. Fixing the root cause is the language of someone treating the customer as a witness." What the AI-native CS function actually looks like Here's the inversion. In an AI-native customer service function, the AI is not primarily there to replace humans answering tickets. It's there to industrialize the conversion of tickets into product changes. The function gets smaller in some places and dramatically more powerful in others. In the companies doing this well, it looks roughly like the following. AI handling routine resolution, yes, but as a sideline, not the strategy. Of course you should automate password resets and account lookups. Nobody serious thinks otherwise. But these are not the strategic use of AI in support. They're table stakes. The strategic use is what comes next. Automated, continuous synthesis of ticket content into product signals. Every ticket, including the ones the AI resolved, is read by another model whose job is to extract: what was the user trying to do, what blocked them, what was their workaround, and was this caused by the product, by docs, by pricing, by messaging, or by a "user error" that's actually a UX failure. Patterns get clustered, weighted by volume and customer value, and surfaced as candidate product issues. A direct, named pipeline from support to product and engineering. Not "we have a Slack channel." A formal mechanism by which the top N candidate issues from support each week enter the product backlog as first-class items, with the ticket evidence attached. The companies doing this best have the head of CS sitting in product reviews and the head of product sitting in CS reviews, not as a courtesy, but as a structural requirement. Support people reframed as forensic analysts and product critics. The headcount goes down, but the seniority goes up. The remaining people are not handling tickets. They are interpreting them, adding context, identifying which patterns matter, intervening on the cases where automated handling would actually destroy customer trust. They are some of the highest-leverage people in the company, because they are sitting on the most undervalued data stream in modern business. Customer-facing AI agents tuned for honesty about limitations, not for deflection rates. They escalate aggressively when they're not sure. They tell the customer clearly what's happening. They are scored not by deflection but by customer-reported satisfaction with the resolution path, even when the path involved a human. This is a different shape from the standard AI support deployment. It produces a smaller support team and a much larger influence of support on the business. It almost always reduces tickets over time, but as a consequence of fixing root causes, not as a strategy of muting them. I've written up the operational blueprint, the KPIs and escalation paths that make it run, in support is one system. The org-chart implication Most companies organize customer service as a sibling to sales, reporting to a Chief Customer Officer or a COO. This places CS structurally as a cost-and-experience function, downstream of the product. Tickets enter CS, get resolved or escalated, and the boundary between support and engineering is patrolled by ticket-routing rules and a quarterly QBR. In the model I'm describing, this structure is wrong. If the highest-leverage use of CS is to convert tickets into product changes, then CS belongs structurally adjacent to product and engineering, not adjacent to sales. Some of the most interesting AI-native companies in 2026 have a Head of Customer Insight reporting directly into the CTO or Chief Product Officer. Their support agents are fewer, more senior, and treated as part of the product team's sensory apparatus. This is hard to do, because it cuts across decades of organizational convention. But it falls naturally out of the underlying logic: in a world where AI can handle most of the labor cost of support, the residual function of support is interpretation of customer reality. That's not a sales-adjacent function. It's a product-adjacent function. The org chart should reflect that. Three things to do, three things to stop If you're running a CS function in 2026, the moves are reasonably concrete. Do: instrument your AI deflection cohort's downstream retention and CSAT. If they're worse than the human-resolved cohort, your deflection metric is lying to you, and the AI is moving costs onto the customer. Do: build, or buy, the pipeline that turns ticket content into clustered, prioritized product hypotheses. This is now a tractable engineering problem. It wasn't three years ago. Do: change who CS reports to, or, at minimum, change which meetings the head of CS attends. If your head of CS is in operations reviews but not in product reviews, you've structurally guaranteed that the most valuable thing the function produces will not be acted on. Stop: framing AI in support as a cost-reduction story to your board. It's a strategic-intelligence story. The cost reductions are the side effect, not the headline. Stop: measuring deflection rate without measuring customer outcomes downstream of deflection. The two metrics together are honest. Either one alone is misleading. Stop: treating tickets as workload. They are the single highest-fidelity source of product reality you have access to, and you are paying customers to generate them. The pattern underneath all of this This shows up in other functions too, but support is where it's most visible. AI's first-order effect is to make some kind of work cheaper, and the obvious move is to do more of that work for less. The second-order effect is to dissolve the constraints that made adjacent work valuable, and the non-obvious move is to go do something else entirely. In support, the first-order play is automating tickets. The second-order play is using AI to finally close the loop between customer experience and product change, a loop that has been broken in most companies for the entirety of their history, because it was organizationally too expensive to keep open. The companies that play the first-order game will look efficient on a slide and slowly hollow out their product. The companies that play the second-order game will look, in three years, like they have an unfair understanding of their customers. They won't, particularly. They'll just be the ones who stopped throwing the data away. Deciding where AI should amplify signal instead of muting it is the heart of my AI strategy work. If your board deck still leads with deflection rates, let's talk → ### How to Grow Engineering Leaders, Not Just Headcount: /blog/growing-leaders-not-headcount/ Scaling a team by hiring more bodies buys you headcount, not leverage. The durable move is to build leaders and make yourself replaceable. When a team is straining, the reflex is to ask for more people. The backlog is deep, everyone is busy, so the fix must be more hands. It rarely is. Adding headcount to a team with no leadership underneath the top adds coordination cost faster than it adds output. You don't get more done; you get more meetings, more handoffs, and a single overloaded manager who is now the bottleneck for twice as many people. In 20 years of leading engineering and as a fractional CTO to more than 30 companies, the pattern is consistent. The teams that scale well don't just hire; they grow leaders. The teams that stall keep stacking individual contributors under a manager who never learned to let go. Headcount is a cost; leaders are leverage A new engineer makes the team bigger. A new leader makes the team capable of being bigger without you in every decision. That's the distinction people miss. Headcount scales linearly and drags coordination cost up with it. Leadership scales the organization's capacity to make good decisions without you in the room, and that capacity is what compounds. Growing a leader is slower and less comfortable than posting a req. You have to hand someone a decision you could make faster yourself, watch them make it differently, and resist taking it back. The first few times, output dips. Then it doesn't, and you've bought back the most expensive resource you have: your own attention. The math only looks bad in the quarter you're investing. Over a year, a team with two capable leads under it ships more, with fewer escalations, than the same headcount funneled through one person who approves everything. What developing leaders actually looks like Coaching a team lead is not a quarterly review and a book recommendation. It's deliberate, and most of it is unglamorous. It happens in the small moments: how you respond when someone brings you a half-formed decision, whether you let a meeting run on without them, whether you correct in public or develop in private. I treat it as a few concrete moves repeated until they stick. - Hand over decisions, not just tasks. Delegating a task is assigning work. Delegating a decision is transferring judgment, including the right to be wrong in ways you wouldn't have been. - Coach the reasoning, not the answer. When someone brings me a problem, the goal is to leave them better at the next one, not to hand them my conclusion. - Make cross-functional alignment their job. A leader who can't get product, design, and engineering pointed the same way is a senior engineer with a title. Alignment is the work, not a distraction from it. - Hire to strengthen the bench, not just fill the gap. Every senior hire is an org-development decision. Ask who they'll grow into, not only what ticket they'll close. "If the team can't run a hard week without you, you haven't built leaders. You've built dependence on yourself." The leader's job is to become replaceable This is the part that unsettles people, so I'll say it plainly: your job as a leader is to make yourself replaceable. Not redundant, not idle. Replaceable. The measure of a leader isn't how much breaks when they take a week off. It's how little. I've seen this from the inside as a healthcare EMR CTO and as a CPTO in online learning, and across 12 engineering teams in 7 countries. The leaders who hoard context feel indispensable right up until they become the ceiling on everything around them. The ones who give context away, who narrate their reasoning and push decisions down, end up running organizations that outgrow them, and that is the kind of growth that lasts. Replaceable at one level is how you earn the next one. So before the next req goes out, ask a sharper question than "do we need more people?" Ask whether the people you have are growing into leaders, and whether you're building an organization that no longer depends on any one person, including you. Headcount you can buy in a week. Leaders you have to grow, and a team that can run without you is the one most worth leading. ### How to Find the Bottleneck Before You Automate Anything: /blog/finding-the-bottleneck/ Automating around a constraint just moves it somewhere worse. Find the real bottleneck first, fix ownership, and most of the tooling you were about to buy turns out to be optional. Every team that asks me to "automate this" is really asking me to make a slow thing fast. The instinct is reasonable. The problem is that they've almost always pointed at the wrong thing. They want to automate the step that feels painful, which is rarely the step that's actually holding the system back. So before I let anyone write a line of automation, I make them answer one question: where is the real constraint? Not whichever step feels worst to sit through, but the single point that caps the throughput of the whole system. Until you've found that, you haven't improved anything, you've just rearranged where the pain shows up. Automating around a bottleneck just moves it This is the part people resist, because it's counterintuitive. If your bottleneck is a two-day code review queue and you automate the deploy step instead, you haven't sped anything up. Work still arrives at the review queue at the same rate it always did. All you've done is build a faster road that dead-ends at the same traffic jam. Worse, now the jam is harder to see, because everything upstream looks healthy. Theory of constraints has been saying this since the 1980s and it still holds: a system's output is governed by its single tightest constraint, and any improvement that isn't at that constraint is an illusion. Optimize a non-bottleneck and you don't get more throughput. You get more inventory piling up in front of the real one. In software that inventory is half-finished work, open branches, tickets in limbo, context that goes stale while it waits. "A faster road that dead-ends at the same traffic jam isn't progress. It just hides where the jam actually is." Most bottlenecks are accountability problems wearing a tooling costume Here's what I find when I trace the constraint to its source: it's usually not a missing tool. It's a missing owner. The deploy is slow because nobody owns the release. The review queue backs up because review is everyone's job, which means it's no one's. The handoff between two teams stalls because the work lives in the gap between two org charts and falls to whoever happens to notice. You cannot automate your way out of unclear ownership. If you point an agent or a pipeline at a process that no single person is accountable for, you've automated the confusion. It now runs faster and fails more quietly. The fix that actually moves the number is boring and human: name an owner, give them the authority to match the responsibility, and make the constraint visible enough that they can't ignore it. - Map the flow end to end and measure where work waits, not where work is hard. - Find the one step where the queue is longest, that's your constraint, everything else is noise. - Ask who owns that step. If the answer is a team and not a person, that's the real problem. - Fix ownership and visibility first. Only then ask whether automation removes the constraint or just feeds it faster. - Re-measure. The bottleneck will have moved. Repeat. It never fully goes away, it relocates. Where automation actually earns its keep None of this is an argument against automation. It's an argument for aiming it. Once you've found the true constraint and put a clear owner on it, automation becomes surgical instead of speculative. You're not buying tools and hoping. You're removing a specific, measured limit, and you can prove it moved. When I've done this well the results are dramatic precisely because they're targeted. Pointing LLM-powered data agents at a genuine processing bottleneck cut processing time by 90%, not because the model was magic, but because I'd already confirmed that was the constraint worth attacking. The same discipline took teams from low to elite DORA performance in about 90 days: find the constraint, fix the ownership around it, then automate what's left. That sequence is what AI-native transformation actually demands. The constraint thinking has to come before the tooling. Skip it and the tooling just lets a confused system run its confusion faster. - 90%: Processing time cut by targeting the real constraint - 90days: Low to elite DORA, constraint-first - 1: Owner per bottleneck, non-negotiable The shift worth making As AI makes it trivial to automate almost anything, the scarce skill isn't building automation. It's knowing what not to build. Automating more of it doesn't win you anything if you aimed it wrong, and most teams I see are aiming it wrong. The ones who get ahead are the ones who found their real constraint, put a name next to it, and pointed their leverage at exactly that. So find the bottleneck first. Then, and only then, automate it out of existence and go find the next one, because there's always a next one. ### How to Structure Customer Support: One System, Not Two Desks: /blog/support-is-a-system/ Support isn't two help desks bolted to the side of the org, it's one feedback loop. Run it as a system, instrument the right KPIs, and it pays back in product and ops fixes, not faster ticket-closing. In most companies, support is two unrelated functions that happen to share a word. There's the IT help desk that unsticks employees, laptops, access, the printer nobody owns. And there's customer support, which absorbs the friction your product generates. The two sit in different teams, on different tools, reporting to different bosses against different metrics, and nobody owns the through-line. I've come to think that's the mistake. Support isn't two help desks bolted to the side of the org. It's one system, and it should be run like one. I've run technology for 30-plus companies as a fractional and interim CTO, and before that I was a CTO in healthcare EMR and a CPTO in online learning. In every one of them, the support functions were treated as overhead to be minimized. The good ones treated support as a sensing system, the part of the org that finds out, before anyone else, exactly where reality and design have diverged. One loop, two surfaces Employee IT support and customer support look different, but structurally they're the same machine: someone hits friction, asks for help, and a ticket records where the friction was. Internal tickets tell you where your own tooling, access model, and ops are broken. External tickets tell you where your product is broken. Treat them as one system and you get a single, honest map of where the organization is leaking time and trust. The failure modes rhyme. An onboarding flow that confuses customers is usually built by a team whose own access provisioning takes three days and four tickets. Internal and external friction come from the same culture of shipping the happy path and leaving the edges for support to mop up. When one leader sees both queues, the pattern is obvious. Split across IT and CS, nobody connects the dots. The KPIs that actually matter Most support dashboards measure the wrong things confidently. Volume, average handle time, first-response time, tickets closed per agent. Those are speed metrics, how fast you process the symptom. They say nothing about whether the disease is spreading. A team can post beautiful handle-time numbers while the same ten root causes generate the same tickets every week. The metrics I care about answer a different question: is the system getting healthier? - Repeat-cause rate. What fraction of this period's tickets trace to a root cause we already saw last period? Flat or rising means the loop is broken. - Ticket-to-fix conversion. Of the top recurring causes, how many became a shipped product or ops change this quarter? This is the number that separates a help desk from a system. - Time-to-root-cause, not time-to-close. Closing a ticket just gets the customer off the phone. Finding what caused it is what lets you actually fix it. - Escalation accuracy. Of issues escalated, how many genuinely needed the next tier? Too low means you're flooding seniors; too high means front-line is under-equipped. - Deflection by root-cause elimination, volume that dropped because you fixed something, tracked separately from volume a bot absorbed. Only the first kind compounds. "Closing tickets faster is treating the symptom. The only metric that compounds is the rate at which support turns tickets into permanent fixes." Escalate toward the fix, not up the ladder Escalation paths are where most support systems quietly fail. The common pattern is a tier ladder defined by seniority, Tier 1 tries, gives up, kicks it to Tier 2, who eventually loops in an engineer through a Slack channel and a prayer. The handoffs are lossy, context evaporates at each step, and the customer or employee re-explains the problem three times. A system has escalation paths defined by where the fix lives, not by who's senior enough to take the heat. A billing edge case routes to the team that owns billing logic, with the ticket evidence attached, not to a generalist who'll improvise a workaround. The path is named, the owner is named, and the clock on time-to-root-cause keeps running until that owner closes the loop, not until the front line makes the customer go away. Escalation should move a ticket toward the people who can make it never happen again. Closing the loop is the whole point Support generates the richest stream of operational truth in the company, and in most orgs it dead-ends. I've argued before that every ticket is evidence, not workload. Tickets get categorized, closed, and aggregated into a chart nobody acts on. The loop from support to product and ops, the loop that would turn that truth into permanent fixes, was historically too expensive to keep open. You needed support to surface trends, product and IT to listen, and everyone to agree on priority. The activation energy was too high, so it never happened. That cost has collapsed. You can now read every ticket from both queues, cluster them by root cause, weight them by volume and business impact, and hand the owning team a ranked list with the evidence attached. This is the operational core of an AI-native transformation: not a chatbot that closes tickets faster, but a feedback loop that's finally cheap enough to run continuously. The support system stops being a place where problems get absorbed and becomes the mechanism by which the organization learns. When that loop is live, the numbers move in the direction that matters. Volume drops because causes get eliminated. The remaining support people get more senior and more influential, because their job shifts from closing tickets to interpreting them. And the same discipline applied to the internal IT queue quietly fixes the developer-experience drag that was slowing every team down. - 2: Queues, internal and external, run as one system - 1: Owner accountable for the whole feedback loop - 0%: Target repeat rate on top root causes, quarter over quarter Where to start Pick one root cause that shows up in both queues and follow it all the way to a shipped fix. Instrument repeat-cause rate and ticket-to-fix conversion before you touch handle time. Put one person in charge of both surfaces and give them a standing seat where product and ops priorities get set. Support stops being the place your problems go to be closed, and becomes the place your organization goes to find out what's actually true. ### How to Lead an Engineering Team Remotely Across Time Zones: /blog/remote-technology-executive/ Remote technology leadership isn't a downgrade of the in-person version. Done deliberately, distance forces the clarity that good organizations need anyway. I've directed twelve engineering teams across seven countries and four time zones, mostly for US companies whose offices I rarely set foot in. The thing people get wrong about remote technology leadership is that they treat it as the in-person job with the good parts removed. It isn't. It's a different job that happens to share a title, and the executives who win at it stopped trying to recreate the hallway. When you can't lean over a desk, drop into a war room, or read the temperature of an office in three seconds, you lose the cheap signals you used to lean on. That sounds like a loss. In practice it forces you to build the things a good organization needs anyway, and most companies never build because the office let them get away without them. You have to manufacture presence on purpose In an office, presence is free and accidental. You're there, so people assume you're available, aligned, and paying attention. Remotely, none of that comes for free. You have to build it deliberately, and what you end up with is usually better for the effort. For me that means being legible rather than visible. People shouldn't have to catch me online to know what I think. The priorities, the open decisions, the things I'm worried about, the reasoning behind a call, all of it lives in writing where the team can read it on their own schedule. Being visible just proves I was around. Being legible is what lets someone in another time zone act on my thinking while I'm asleep. What actually changes The honest list of what's different when leadership goes remote and distributed is shorter and harder than the productivity blogs suggest. - Communication moves to writing by default. A decision that isn't written down didn't happen. Synchronous time is rare and expensive, so you spend it on judgment and trust, not status. - Trust shifts from attendance to output. You can't see who's working, so you stop pretending hours are the measure and manage the work and the outcomes instead. - Decisions get pushed down. Across four time zones, routing every call through one person is a bottleneck that costs you a full day. You give teams the context and authority to decide without you. - Time zones become a feature. Handing work across the day means something can move while half the company sleeps, but only if the handoff is documented well enough to survive the gap. "Remote doesn't make leadership weaker. It removes the crutches that let weak leadership look fine in an office." Why deliberate remote is an advantage The clarity tax that distance imposes is exactly the discipline high-performing organizations need. When decisions must be written, they get scrutinized. When you can't hover, you're forced to hire people you trust and then actually trust them. When time zones make real-time coordination painful, you design systems and interfaces clean enough that teams don't need to coordinate constantly in the first place. Every one of those is good engineering leadership regardless of where anyone sits. It's also how the work gets done in practice. Since 2018 I've served as a fractional and interim CTO to more than thirty companies, almost all of them US-based, run entirely remotely. I was a healthcare EMR CTO and an online-learning CPTO before that. The pattern holds across all of it: the distributed organizations that run deliberately are tighter than the co-located ones that run on osmosis, because they had no choice but to make the implicit explicit. The failure mode is treating remote as a constraint to apologize for rather than a system to design. Companies that bolt remote onto office habits get the worst of both and blame the distance for it. But distance was never really the issue. They just never built a way of working that fit it. Where this goes Talent stopped clustering in a few zip codes, and it isn't going back. The companies that build the best technology organizations over the next decade will be the ones that treat distributed, cross-time-zone leadership as their default operating model rather than a concession. And the skills that make it work, writing things down, handing off cleanly, trusting people you can't watch, were always just good leadership. The office let too many people skip them. ### How to Align Technology With Business Strategy: /blog/aligning-technology-across-the-org/ The hardest part of a technology org isn't the code. It's keeping it pointed at what operations, product, and the business actually need. Most failed technology investments I've seen didn't fail technically. The code worked, the architecture was sound, the team was good. They failed because engineering was building something the people who run the business never actually needed. Real work, real spend, and none of it moved the company an inch closer to what mattered. That's the part nobody warns you about. The hard problem in a technology organization isn't writing the software. It's keeping the software pointed at what operations, product, the executive team, and the people closest to the customer are actually trying to accomplish. When tech drifts from those people, it stops being progress and turns into expensive motion. Two languages, one company The boardroom and the codebase speak different languages. One talks in margin, retention, and timelines tied to a fiscal calendar. The other talks in latency, coupling, and what's actually feasible by the date someone promised. Most of the friction I'm brought in to fix isn't a gap in skill. It's a gap in translation, and the two sides rarely notice they're talking past each other. An operations leader doesn't care that you refactored the billing service. They care that month-end close stopped taking four days. A clinical or academic leader doesn't want to hear about your event pipeline. They want to know the thing in front of their people every day got less painful. My job, as a fractional CTO, is to stand in the middle and make both sides legible to each other so the work that ships is the work that counts. What alignment actually looks like Alignment isn't a quarterly offsite or a shared slide deck. It's a discipline that shows up in how decisions get made every week. When I've done this well across the companies I've worked with, a few things are always true. - Non-engineering leaders can state, in their own words, what the technology org is building this quarter and why it matters to them. - Engineering can name the business outcome behind every major bet, not just the ticket. - The roadmap is argued over by operations and product, not handed down and rubber-stamped. - When priorities change, the people closest to the customer are the ones who triggered the change. "Technology that isn't aligned to the people running the business isn't progress. It's expensive motion." I learned this most sharply as the product and technology lead at an online learning platform, where the academic and operations leaders understood the learner better than any dashboard ever could. The right move was almost never the most elegant engineering one. It was the one that matched how those teams actually worked. The same was true running an EMR organization in healthcare, where the clinical reality on the ground decided whether a feature was a help or a hazard. Why this is a leadership job, not a process You can't solve alignment with a tool or a ceremony. I've watched companies bolt on more planning rituals and end up with more meetings and the same drift. Alignment is a posture: the senior technology leader treating operations, product, and domain experts as co-owners of the roadmap rather than stakeholders to be managed. It means sitting in their problems long enough to feel them, not just collecting requirements and disappearing. Across more than 30 companies and a dozen teams in different countries, the pattern holds. The organizations that compound are the ones where technology and the rest of the business are solving the same problem from two directions, not negotiating across a wall. Where this is going As AI absorbs more of the mechanical work of building software, the scarce skill won't be writing code. It will be judgment about what to build, and that judgment lives in the conversation between the people who run the business and the people who build the systems. I'd bet less on the best engineer in the room and more on the leader who can hold both languages at once and make sure every dollar of technical effort lands on something the company actually needed. ### How to Protect Your Roadmap Without Being the Department of No: /blog/protecting-the-roadmap/ Staying responsive to the business without letting every urgent request shred the roadmap. How to say no without becoming the department of no. Every technology leader gets pulled in two directions at once. The business wants you responsive: jump on the customer escalation, build the thing sales promised, fix the report the CEO saw this morning. The roadmap wants you focused: ship the platform work that compounds, the bets that take a quarter to pay off. Lean too far one way and you're a bottleneck. Lean too far the other and the roadmap quietly turns into a wishlist that never ships. The trap most leaders fall into is treating this as a personality problem, are you a "yes" person or a "no" person, when it's actually a systems problem. Responsiveness and focus aren't opposites. You only have to choose between them when you have no mechanism for deciding what matters. Build the mechanism and the tension mostly dissolves. The department of no is a failure of explanation When people call a team "the department of no," they're rarely complaining that it says no too often. They're complaining that the no is opaque. A flat refusal with no reasoning reads as territorial, and people route around it the next chance they get. The fix isn't to say yes more. It's to make every no legible. I never say just "no." I say "not this quarter, and here's what would have to come off the board for it to fit." That reframes the request from a fight with engineering into a trade-off the business owns. The person asking gets to see the cost of their own request, and nine times out of ten they de-prioritize it themselves once the price is visible. Make the trade-offs visible, not personal Protecting the roadmap is mostly about making capacity scarce in public. If everything lives in one prioritized list with a hard ceiling on what's in flight, every new request has to displace something. The conversation stops being "can you also do this" and becomes "what does this replace." That's a question the requester can answer, and it moves the decision to where it belongs. - One list, one owner. A single prioritized backlog with a named owner who can make the call. Competing lists are how the roadmap dies. - A capacity ceiling. Cap work in progress. Saying yes to something new forces a visible bump for something old. - A standing slice for the unplanned. Reserve a fixed percentage for escalations and fast turns. Now "urgent" has a home that doesn't raid the roadmap. - A real definition of urgent. Most urgency is manufactured. Agree up front on what actually counts, and hold the line on the rest. "Saying no isn't the skill. Making the cost of yes visible to the person asking is the skill." Budget for interruptions instead of reacting to them The leaders who stay responsive without torching their roadmap treat interruption as a budgeted line item rather than a moral test. When a fixed share of capacity is set aside for the unplanned, you can say yes to the genuine fire drill instantly and without guilt, because it's already paid for. When that budget is spent, the next request waits or trades. The discipline lives in the system, so you don't have to manufacture it in the moment. This is the same posture I bring as a fractional CTO, where I'm often parachuting into an organization that has lost the thread between business pressure and engineering focus. The first move is almost never to ship faster. It's to make the existing trade-offs visible so the right people can own them, and to give "urgent" a place to live that doesn't quietly consume the quarter. What protected focus actually buys A protected roadmap isn't about engineering comfort. It's about the compounding work, the platform investments and architectural bets, that only pays off if it's allowed to finish. Every time an unbudgeted interruption knocks that work back a sprint, you pay interest you never see on a balance sheet: the feature that ships a quarter late, the rework, the senior engineer who burns out context-switching. The goal isn't a leader who guards the gate. It's an organization where prioritization is a shared, visible act, where the business can see exactly what its requests cost and decide accordingly, and where saying yes to the roadmap and yes to the business are no longer in conflict. Build that mechanism once and the department of no stops being your reputation, because the hard calls now happen in the open where everyone can weigh in. ### AI Strategy for Operations and Customer Experience: One Loop: /blog/ai-internal-ops-and-customer-experience/ Most companies put AI to work on one side of the business and ignore the other. The leverage is in running it on both, with a human accountable where it touches customers. When a company tells me it has an AI strategy, I ask one question: which side of the house? The answer almost always reveals a blind spot. Either they've automated the back office and left the customer experience untouched, or they've shipped a slick AI feature in the product while the people behind it still copy-paste between five tabs. Very few are doing both, and that's exactly where the advantage is hiding. There are two sides to every business. Internal operations is how the work gets done: support, finance, data, recruiting, the unglamorous machinery. Customer experience is what the market sees: the product, the onboarding, the help, the moments that decide whether someone stays. Most AI programs pick one and call it a strategy. A real AI strategy covers operations and customer experience at the same time. Why companies do only one side It's rarely a deliberate choice. Internal-ops AI is easy to justify and easy to hide: automate a workflow, point at a cost line, no customer ever sees it if it's clumsy. Customer-facing AI is the opposite, visible, loud, and risky, so it gets owned by product and marketing, who have no reason to touch the back office. The two efforts end up in different orgs and on separate budgets, each with its own definition of done. So you get lopsided companies. One has agents tearing through internal data but a support experience that still makes customers wait two days. Another has a charming chatbot out front and a team behind it drowning because nothing internal got faster. Both spent real money on AI. Neither compounded. That lopsidedness is exactly what an AI strategy exists to fix. The two sides feed each other Run them together and they reinforce. The same retrieval system that lets a support agent answer instantly is the one that should draft the customer's reply. The internal tool that summarizes a messy account becomes the brief that makes onboarding feel personal. Your internal operations are the supply chain for your customer experience, and AI that only optimizes one end leaves the other starved. It runs the other way too. Customer-facing AI is the richest source of signal you have for fixing operations. Every question a user asks the product, every place the AI hesitates, is a map of where your internal knowledge is thin or wrong. Treat both sides as one loop and each turn makes the other sharper. That feedback loop is what an AI-native transformation actually is: the whole business getting smarter, rather than one more feature bolted onto the product. "Internal AI and customer-facing AI aren't two projects. They're two ends of the same loop, and the leverage is in closing it." I've seen the internal side carry surprising weight. I've run LLM-powered agents over roughly 250 million records a month on a 75-node cluster and cut processing time by 90%. That's an operations win, but its real value showed up downstream, in how much faster and more accurately customers got answers, because clean internal machinery is what a good experience is built on. - 250M/mo: Records processed by LLM-powered agents - 90%: Reduction in processing time - 2: Sides of the business, one loop Keep a human accountable at the edge Symmetry has one limit. Where AI touches a customer, a named human owns the outcome, full stop. Inside operations you can let an agent run with a light hand and catch errors in review. At the edge, a wrong answer is a broken promise, and there is no team to absorb it before it lands. Accountability doesn't mean a person checks every message. It means someone owns the quality bar: - A named owner for every customer-facing AI surface, not a committee. - Evals and monitoring that flag drift before customers feel it. - Clear handoff to a human the moment the AI is unsure. - A standing review of what the AI got wrong, fed back into both sides. Where this goes The companies that win the next few years won't be the ones with the best chatbot or the leanest back office. They'll be the ones that stopped treating those as separate problems, ran AI across both sides as a single system, and kept a human standing where it matters most. Pick the side you've neglected, and start closing the loop. If you want help deciding where, that's the first conversation in an AI strategy engagement. Let's talk → ### Who Should Own Engineering, Ops, Support, and Data? One Leader: /blog/owning-the-whole-stack/ Splitting dev, platform ops, support, and data into separate fiefdoms feels organized. It manufactures the worst failures. Single ownership fixes that. Some of the worst outages I've seen came out of the tidiest org charts. Development reports here, platform operations there, support sits under a different VP, and data is its own kingdom with its own roadmap. On paper it looks organized. In practice you've drawn four borders straight through the path a single request travels, and accountability quietly leaks out at every one of them. I've spent 20 years in leadership watching this play out, and in the work I do as an interim and fractional CTO, the first thing I usually find isn't a technology problem. It's four teams who each believe their part is fine and the failure must belong to someone else. The seams are where systems break Real failures almost never live inside one box. They live in the seams between boxes. A deploy ships clean, the platform autoscaler reacts to a traffic pattern nobody flagged, support starts drowning in tickets that look like a product bug, and the data pipeline silently drops a column three hops downstream. Every individual team is technically correct, and the customer is still down. When those four functions answer to four different leaders, the failure has no owner. It has a meeting. The incident becomes a negotiation about whose budget pays for the fix, and the root cause, the seam itself, never gets touched because no one is accountable for the seam. - Development optimizes for shipping features, not for what those features do to the platform at 3am. - Platform ops optimizes for stability, which quietly means saying no to the people shipping features. - Support absorbs the gap between the two and is measured on closing tickets, not on killing the cause. - Data inherits everyone's schema decisions last and has the least power to change them. "When four functions answer to four leaders, a failure in the seams between them has no owner. It has a meeting." What single ownership buys you Put development, platform operations, technical support, and data systems under one accountable owner and the physics change. The trade-offs that used to be cross-departmental wars become one person's call. Should we slow a release to harden the pipeline? Is this surge of tickets a symptom of an architecture choice, not a training gap? Those questions stop bouncing between orgs and get answered. I learned this running technology in complex organizations, as a healthcare EMR CTO where a dropped record is a patient-safety event, and later as a CPTO in online learning where support volume and platform behavior were the same story told twice. The migrations I'm proudest of were sub-one-hour and zero-downtime, and that's not a heroics story. It's what happens when the people writing the code, running the platform, fielding the tickets, and owning the data are pointed at one goal by one person instead of defending four scorecards. - <1hr: Migration windows when ownership is unified, not negotiated - 0: Downtime, because dev and ops plan the cutover as one team - 4: Functions, one accountable owner across the request path Coordination is not ownership The usual objection is that you can fix this with better process, a shared dashboard, a steering committee, a post-incident review template. You can't. Coordination layers describe the seam; they don't own it. A committee can document that the autoscaler and the deploy and the schema change collided, then politely watch it collide again next quarter. Ownership means one person whose name is on all four outcomes and who can move resources across them without asking permission. Where this is heading The pressure is only going up. Agentic systems and AI-assisted pipelines smear the lines between code, infrastructure, support, and data even further. When a model misbehaves, the same minute can produce a bug, an alert, a wave of tickets, and a corrupted table, and you cannot say which one came first. Org charts that pretend those are separate problems will keep manufacturing outages that have no owner, just a longer thread of people explaining why it wasn't theirs. My advice is to stop tuning the individual boxes and put one person in charge of the seams between them, where the real failures live. ### How to Fix Bad Company Data: Dashboards People Actually Trust: /blog/data-chaos-to-visibility/ Most companies don't have a data problem. They have a trust problem. Turning scattered, untrusted data into visibility people act on is platform work, not a dashboard. Almost every company I walk into believes it has a data problem. More often it has a trust problem. The numbers exist, sometimes in a dozen places, but two people pull the same metric and get two different answers. So the executive team stops trusting the dashboard and starts trusting the loudest person in the room. That's not a reporting gap. That's data chaos, and it quietly decides how the business runs. A chart isn't visibility. Visibility is when someone looks at a number, believes it, and changes what they do next. Getting there is mostly plumbing and governance work, the unglamorous part that happens long before anyone opens a BI tool. Chaos has a shape Data chaos isn't random; it follows the org chart. Every team buys its own tools, defines its own version of "active customer," and exports to its own spreadsheet. You end up with the same handful of problems almost every time: - The same metric defined three different ways in three different systems. - Pipelines nobody owns, breaking silently until a board deck is wrong. - Reports that are technically correct and operationally useless. - Hours of manual reconciliation that someone quietly does every Monday. None of this is fixed by buying a better dashboard. A dashboard on top of untrusted data just makes the wrong answer prettier and faster. The work underneath Turning chaos into visibility is a migration and integration problem first. You consolidate scattered sources into a warehouse, define metrics once where everyone can see the definition, and build pipelines that are owned, tested, and observable. The goal is a single place where a number means exactly one thing and you can trace it back to where it came from. I treat this like any other production system, not a side project for whoever has spare time. Ingestion gets monitored, transformations get tested, and every metric definition lives in version control with a name attached to it. When a pipeline breaks, someone knows before the CFO does. That discipline is what separates a warehouse people rely on from a data swamp they quietly route around. "A dashboard on untrusted data just makes the wrong answer prettier and faster." Where AI changes the stack AI is genuinely shifting the data layer, but not where most demos point. The leverage was never a chatbot sitting in front of your warehouse. It's agents that take on the heavy, error-prone work that used to eat headcount: classifying records, cleaning them, mapping messy schemas, and reconciling sources that never agreed to a common format. I've run LLM-powered data agents processing roughly 250 million records a month on a 75-node cluster, which cut processing time by about 90%. That's work that used to take an army of analysts and a lot of patience, now running continuously and without anyone babysitting it. This is the practical core of an AI-native transformation: you don't bolt a model onto a report, you rebuild the pipeline so the data is trustworthy by the time a human ever sees it. - 250M/mo: Records processed by LLM-powered data agents - 90%: Reduction in processing time - 75: Cluster nodes orchestrating ingestion AI raises the stakes on the fundamentals rather than replacing them. Point an agent at chaotic, contradictory data and it won't fix the chaos. It will scale it, confidently and at machine speed. Get the warehouse, the definitions, and the governance right first, and those same agents turn into the cheapest, most tireless analysts you'll ever hire. Order of operations is the whole game here: foundations earn you the leverage, and skipping them just hands AI a bigger mess to amplify. What you're really building What you get at the end isn't a prettier report. It's a business where decisions move faster because the numbers stopped being negotiable, where a metric has one owner and one definition, and where the platform underneath is solid enough that AI makes it sharper instead of louder. Pick the one metric your leadership argues about most, trace it back to the source, and make it trustworthy end to end. Then do the next one. That's how visibility actually compounds. ### AI Digital Transformation: The Playbook Has Changed: /blog/digital-transformation-is-ai-now/ The old transformation playbook moved you to the cloud and tidied your processes. The new one rebuilds operations around AI. The lever changed; plans didn't. For most of my career, "digital transformation" meant a specific list of work. Get off the data center. Replace the monolith. Wire up a CRM. Document the processes nobody had written down. I've run that program more times than I can count, and the playbook held up because the lever never moved: you were converting manual, analog, or on-prem work into software. The destination was always "more digital." That destination is now table stakes. If you're a company worth transforming, you're already digital. So when a board says "we need a transformation," they don't mean cloud anymore, even if they still use the old words. They mean AI. The lever moved, and a lot of transformation plans are still pulling the old one. I spend a lot of my time as a fractional and interim CTO walking into exactly this gap. The leadership team knows the answer is "AI" and has the budget approved, but the plan on the wall is a cloud-era plan wearing new language. The systems diagram, the milestone-and-cutover timeline, the success metric tied to a launch date, all of it assumes a project that ends. The instinct is right, but it's aimed at a problem that doesn't have an end state. What actually changed about the playbook The old transformation was a migration: a finite project with a before state and an after state. You finished it. AI-driven transformation isn't a migration, because the technology underneath keeps getting better every quarter. There is no "after." That single difference breaks a few habits the cloud era taught us. - The unit of work shrinks. Cloud transformation moved systems. AI transformation moves tasks, the steps inside a process that used to require a human, which means you target workflows, not platforms. - The ROI shows up in the cost line, not the feature list. You're not shipping a flashy capability; you're collapsing the time and headcount a process consumes. - It's never "done." You're standing up a capability that improves on its own cadence, so governance, evals, and cost controls matter more than a cutover date. - The risk surface is new. Hallucination, data leakage, and model drift aren't problems the cloud playbook ever had to budget for. Operations is where it pays, not features The loudest version of "AI strategy" is a customer-facing feature: a chatbot, a generated summary, a copilot bolted to the side of the product. Those are easy to demo and easy to fund, and most of them are theater. The money is somewhere quieter, inside the operation, where work that costs real hours gets done. That's where I point an AI strategy first. I've seen this concretely. I built LLM-powered data agents that took on a reconciliation process running across roughly 250 million records a month and cut the time it consumed by 90%. Nobody outside the company will ever see that work. It isn't a feature. It's the operating cost of the business dropping, permanently, and that's the kind of result AI-native transformation is supposed to produce. - 250M/mo: Records handled by LLM-powered operational agents - 90%: Reduction in time the process consumed - 0: Customer-facing features it required "The companies that win this round won't be the ones with the best AI feature. They'll be the ones that quietly rebuilt how the work gets done." How I'd lead it I treat it less like a project and more like installing a new muscle. Start with an honest map of where the hours actually go, the high-volume, judgment-light, expensive-to-staff processes, because that's where AI has leverage and a clear payback. Pick one, rebuild it end to end with real observability and cost controls, and let it run in production until the team trusts it. Then move to the next one. You don't transform the company in a single program; you transform one process at a time until the operation looks different. The cloud era rewarded the companies that finished their migration. This era will reward the ones that keep going, because the technology keeps improving under them. You're still being asked to do the thing CTOs have always been asked to do: make the business cheaper, faster, and better at what it already does. What you reach for to do it is what changed. If the plan on your wall still reads like a migration, closing that gap is where my AI strategy work starts. Let's talk → ### How to Prioritize a Roadmap: Decide What Not to Build: /blog/prioritization-is-the-job/ A technology leader's real work isn't deciding what to build. It's deciding what not to build, and holding that line against the pull of every loud request. Every technology leader I've ever worked alongside could produce a roadmap. The hard part was never the list of things to build. The hard part was the list of things to not build, and saying so out loud, to people who wanted those things, with money or politics behind them. That second list is the job. Everything else is administration. A roadmap that says yes to everything isn't a plan. It's a queue with optimistic dates attached. The teams feel it long before the leadership does: too many half-finished initiatives, every one of them important, none of them done. People reach for more capacity or a better tool, and neither helps. Extra capacity just means you carry more unfinished work at once, and the best planning software in the world will faithfully track a portfolio that should never have been started. What actually moves the needle is the discipline to decide and the willingness to own the decision afterward. Prioritization is subtraction Across more than thirty companies, I've watched the same pattern. A leader spends their energy ranking what to do and almost none deciding what to kill. So the backlog grows, integrations multiply, and every release tries to carry a little of everything. Everybody is busy and nothing ships. Here's the reframe: prioritization is subtraction, not ranking. Ranking tells you the order you'll attempt forty things. Subtraction tells you which thirty-two you've agreed not to attempt this quarter, on purpose, so the other eight get done well. Ranking feels generous to everyone in the room and ships nothing. Subtraction is the one that disappoints people in the meeting and puts working software in front of customers. "Anyone can decide what to build. The job is deciding what not to build, and holding to it when that decision costs you something." Decide against something, not in a vacuum Paralysis comes from judging requests one at a time. In isolation, almost every request is reasonable, so you say yes, and the yeses pile up until the roadmap is fiction. The way out is to never evaluate anything alone. Every request competes against the thing it would displace. - Tie each candidate to a business outcome, not a feature description. "Cut onboarding time" beats "build the onboarding wizard" every time. - Force a single ranked stack, not tiers. Everything can't be a P1; the moment three things are top priority, nothing is. - Make the cost of a yes visible by naming what it pushes down the stack. Every yes is a quiet no to whatever it bumps, so say that no where people can hear it. - Cap how much work is in flight at once. A team juggling four initiatives ships less than a team that runs two and actually closes them out. - Revisit on a fixed cadence, not on whoever shouted most recently. This is how you take the politics out of it. When priorities live in one ranked stack tied to outcomes, an argument stops being "is my feature good?" and becomes "is my feature better than the thing it bumps?" That's a question you can actually answer with evidence, and it's a question that doesn't reward the loudest voice in the room. What this buys you I've watched this play out, not just argued it. When a team stops splitting its attention and starts finishing, delivery health moves quickly. I've taken organizations from low to elite on the standard delivery metrics in roughly ninety days, and the lever was almost never tooling. It was cutting work in progress and refusing to start what we'd already decided not to finish. Release planning gets simpler too, because each release carries a coherent set of bets instead of a thin slice of everything, and integrations stop sprawling because each new one has to earn its place against the stack. As a fractional CTO, this is usually the first thing I fix, ahead of any reorg or rewrite. - 30+: Companies where the same yes-to-everything pattern repeated - 90 days: Low to elite delivery, mostly by subtraction - 1: Ranked stack, not tiers, so the cost of yes stays visible Saying no is uncomfortable. It disappoints people, sometimes powerful ones, and that discomfort is exactly why most leaders duck it and let the backlog decide for them. A backlog can't decide anything, though. It just keeps growing until something breaks. The thing you're actually being paid for is the judgment, applied as subtraction, and held in the room where holding it costs you. Do that well and the roadmap stops reading like a wish list and starts being a promise the team can keep. ### How to Manage People in a High-Growth Company: /blog/managing-people-in-high-growth/ In high growth the org chart, the product, and the headcount all move at once. Managing people through that is a different job than steady state. Managing people in a stable company is hard. Managing people in a company that's doubling is a different job that happens to share a title. In steady state you tune a machine you understand. In high growth the machine is being rebuilt while it runs, and you're rebuilding the people too. Anyone who treats those as the same role gets blindsided fast. I've felt this across 20 years in leadership and a decade as a fractional and interim CTO to more than 30 companies. The teams that scale well rarely have the slickest hiring funnel. What they have is managers who figured out early that growth changes everything underneath the org chart, not just the boxes on it. Three things move at once What makes high growth its own discipline is that three variables shift simultaneously, and each one breaks something the others depend on. - The headcount changes. The team you manage in Q1 is not the team you manage in Q3. Half the people reporting to you may not have been there when the last decision was made. - The org chart changes. Reporting lines you set up six months ago are now bottlenecks. Someone who was a peer is now a report, or the reverse. Roles you invented are already too broad. - The product changes. The thing the team is building is moving under their feet, so the skills you hired for and the context people carry both go stale faster than anyone expects. In steady state any one of these moving is a project. In high growth all three move at once, and they compound. A reorg lands on people who joined last month, building a product that pivoted last week. Managing that overlap is most of what the job actually is. Clarity is the thing that breaks first When people worry about scaling, they usually worry about culture. But culture erodes slowly, and the thing that actually goes first is clarity. When you add ten people to a team of fifteen, the unwritten knowledge that used to travel by osmosis stops traveling. Nobody knows who owns what, decisions get made twice, and good people quietly stall because they can't tell whether a thing is theirs to do. So in growth I over-invest in the boring artifacts: written ownership, explicit decision rights, a clear answer to "who do I ask." Not bureaucracy, the opposite. The faster you hire, the less you can rely on context being in the room, so it has to be on the page. The manager's job shifts from making decisions to making it obvious who makes them. "The first casualty of fast hiring isn't culture. It's clarity. Culture erodes slowly; clarity collapses overnight." Protecting culture without freezing it Culture in high growth is fragile for a simple reason: every new hire dilutes it by definition. Bring in enough people fast enough and the median behavior of the company is set by people who've been there under a year. That's not a knock on them, it's math. If you don't actively transmit how the place actually works, the new majority will reinvent it, and not in your favor. But protecting culture isn't preserving it in amber. A 15-person culture that worked beautifully will quietly suffocate a 60-person team, because rituals that depended on everyone being in one conversation don't survive contact with scale. The skill is knowing which parts are load-bearing values worth defending to the wall, and which parts are just habits that fit the old size. Confuse the two and you either lose the soul of the place or strangle it with nostalgia. What this changes about managing Practically, the manager's center of gravity moves. Less time on individual output, more on the connective tissue: who talks to whom, where decisions live, how context flows to people who weren't there for the founding stories. You spend less time being the smartest person on the problem and more time making sure the team can solve problems without you in the room, because in three months there will be twice as many rooms. This is a big part of what I do as a fractional CTO: come into an org that's growing faster than its management can absorb, and build the clarity and structure that let people scale without losing what made the team good in the first place. I've watched this play out across 12 teams in 7 countries, and the pattern holds. Hiring speed is not what separates the companies that win the growth phase. The ones that win have managers who treat clarity and culture as things you actively build, rather than things you hope survive the next twenty hires. The org chart, the product, and the headcount will keep moving, because that's what growth is. The managers who do well stop trying to hold the picture still and start managing the motion itself. It's worth getting good at, because almost every company that gets anywhere has to live through it. ### How to Run Internal and Outsourced Engineering as One Team: /blog/internal-external-engineering-one-team/ Internal product teams and external developers can ship as one accountable unit, or fail at the seam. The difference is who owns delivery. Almost every company I've worked with builds software in two places at once. There's an internal team that owns the product, and there are external developers, an agency, an offshore shop, a contractor or two, who fill the gaps. On paper this is a staffing decision. In practice it's the single biggest delivery risk most leaders never name, because the failure doesn't show up in either team's work. It shows up in the seam between them. The classic version is the throw it over the wall handoff. Internal writes a spec, external builds to the spec, and the spec was wrong, or stale, or silent on the ten decisions that actually mattered. Both sides did their job. The product still came out broken, late, and impossible to maintain. Nobody is accountable, because accountability was split the moment the work crossed the line. One team, one definition of done The fix isn't insourcing everything or managing vendors harder. It's refusing to treat internal and external as two teams in the first place. I run them as one engineering organization with one backlog, one set of standards, and one definition of done, no matter whose payroll the developer is on. The org chart can stay split as long as the work itself flows through a single set of rules. Concretely, that means everyone, employees and external partners alike, shares the same non-negotiables: - One repository and one CI pipeline, with the same tests and quality gates applied to every commit regardless of author. - The same code review bar, where internal and external engineers review each other's pull requests rather than reviewing in silos. - Shared standups, shared incident response, and shared on-call expectations, so context isn't hoarded on one side of the line. - A definition of done that includes tests, docs, and observability, not just "it runs on my machine." The moment external developers are held to the same bar as internal ones, the wall stops existing, because there's nothing left to throw over it. "Whoever writes the code, I own the delivery. That's the whole job. Accountability doesn't get to be someone else's problem." Standards travel, blame doesn't External teams are usually blamed for quality, but in my experience they hit exactly the bar you set and enforce. Hand an agency a vague spec and no pipeline and you'll get vague work. Hand the same agency a clear interface, automated gates, and reviews from your internal engineers, and the output is frequently better than what an understaffed internal team produces under pressure. The variable isn't where the developer sits. It's whether the standards are encoded in the system or stuck in someone's head. This is the part leadership tends to skip. You can't enforce a standard you haven't written down, and you can't write it down only for the people you can see in the office. As a fractional CTO I spend the first weeks making the implicit explicit, turning tribal knowledge into pipelines, templates, and checks that apply to anyone touching the codebase. When the standard lives in the tooling, distance stops mattering. I've run twelve engineering teams across seven countries this way, and the ones that shipped well were never the co-located ones. They were the ones where the bar was the same everywhere. What owning the line looks like Owning delivery across an internal-plus-external setup comes down to a few habits. I architect the boundaries so the work can be split cleanly, with clear interfaces and contracts, instead of carving the system along team lines and praying the pieces fit. The pipeline becomes the referee, which turns quality into a gate rather than an argument. And one person, me, stays accountable for the whole thing reaching production, so there's never a question of whose fault the seam is. Done right, you get the same results that always mattered, just delivered by a wider set of hands: - 12: Engineering teams run across 7 countries as one org - 90days: From low to elite DORA performance - 0downtime: Migrations shipped, internal and external hands The work is only going to get more distributed from here, not less. More partners, more contractors, more AI agents writing code alongside humans. Every one of those is just another contributor crossing the same line. The leaders who come out ahead will be the ones who stopped chasing the perfect sourcing model and built a system where it doesn't matter who wrote the code, because the standards, the pipeline, and the accountability are the same on both sides. ### How to Build Accountability Into an Engineering Team: /blog/execution-with-accountability/ Anyone can paint a future. The actual job is closing the gap between the vision and the shipped thing, then owning the result either way. I've sat in a lot of rooms where a leader unveils a bold vision, the slides land, the room nods, and everyone leaves feeling like progress happened. Nothing happened. A vision is a claim about the future, and claims are free. The only thing that ever mattered was whether the thing got built, by when, and whether someone was on the hook for the answer. After 25 years in software and 20 in leadership, vision stopped impressing me a long time ago. Almost everyone has one. What's rare is the discipline to turn it into something shipped on a schedule, and the spine to own the gap when reality comes up short of the slide. Vision without execution is a hallucination The most expensive thing in a company isn't a failed project. It's a beautiful strategy that quietly never happens. The roadmap exists, the deck is gorgeous, the all-hands was inspiring, and twelve months later the metrics haven't moved. No one lied. The vision just never got connected to anyone's actual week, so it kept hovering somewhere above the work, admired and untouched. People talk about execution like it's the part you hand off once the interesting thinking is finished. It is the interesting thinking. A vision that doesn't survive contact with a deadline, a budget, and a hard dependency was never a strategy. It was a mood. "A vision that can't survive a deadline, a budget, and a hard dependency was never a strategy. It was a mood." What accountability actually looks like Accountability gets used as a synonym for blame, which is why most teams flinch when they hear the word. The version I care about has nothing to do with who gets yelled at. It's the plumbing that makes it obvious early, and obvious to everyone, when a commitment is slipping. Three things have to be true at the same time. Commitments, metrics, and ownership - Commitments are dated and specific. Not "we're improving reliability" but "checkout p95 under 400ms by March 31, named owner attached." A goal vague enough that you can never quite miss it is one people tend to prefer, for obvious reasons. - Metrics are leading, not just lagging. A number you only see at the end of the quarter is an autopsy. The signals worth having tell you you're drifting in week two, while there's still time to do something about it. - Ownership is singular. Hand an outcome to a committee and you've handed it to no one in particular. One name per outcome instead. That person doesn't do all the work, but they're the one who answers for it. None of this requires a heavier process. It requires saying the quiet part out loud and writing it down: here is what we said we'd do, here is the date, here is who owns it, and here is how we'll know. Most organizations resist this not because it's hard, but because it removes the comfortable ambiguity that lets everyone feel busy while nothing converges. Be the visionary and the closer The cleanest version of this discipline shows up in delivery metrics. As a fractional CTO, I've taken teams from low to elite on the DORA metrics in roughly 90 days, not by adding tooling but by closing the loop between what we promised and what we shipped, then measuring it honestly every week. Deploy frequency, lead time, change-fail rate, and recovery time are accountability made visible. - 90days: Low to elite DORA performance - 1: Named owner per outcome - 30+: Companies led as fractional or interim CTO The mistake leaders make is treating vision and execution as different people: the dreamer up top, the closers below. That split is where strategies go to die, because the person with the vision is the only one who can adjudicate the tradeoffs when reality pushes back. Whether I was running an EMR platform in healthcare or product and engineering for an online-learning company, the job was the same: hold the picture of where we're going and stay close enough to the work to know whether we're actually getting there. The forward bet AI is about to flood every organization with more vision than it can possibly execute. When plausible strategy costs almost nothing to generate, the thing that's actually worth anything is the part that was always hard: picking one direction, putting a date and a name on it, and standing behind whatever happens. I'd bet on the leaders who can hold both ends of that, the picture and the delivery, and who treat the gap between them as their problem to close rather than someone else's to explain. ### How to Build a Multi-Year Tech Roadmap That Survives Reality: /blog/multi-year-roadmap-survives-reality/ A three-year roadmap that locks every quarter is fiction. One with no direction is chaos. The job is building the version that holds a line while reality keeps moving it. Every multi-year roadmap I've inherited has failed in one of two ways. Either it was a Gantt chart that planned quarter eleven in the same detail as quarter one, and reality shredded it by month four. Or it was a vibe, three years of "scale the platform" and "invest in AI" with nothing underneath, and the team did whatever was loudest that week. Both of them feel like planning right up until reality gets a vote, and then they don't survive it. The roadmap that works is a third thing. It commits hard to direction and stays deliberately loose on path. It tells you what you're trying to be true in three years, and it refuses to pretend you know the exact route there. That distinction sounds academic until the market shifts, a competitor ships, or a model release rewrites your assumptions, and you find out which kind of roadmap you actually built. Direction is fixed, path is negotiable A real roadmap separates two things most companies fuse: the outcomes you're betting the business on, and the work you think will get you there. The first should barely move year to year. The second should be rewritten constantly. When people say a roadmap "changed," they usually mean the path changed, which is fine. The failure is when direction quietly changes because nobody was guarding it. So I anchor the roadmap to a small number of durable outcomes, three or four, expressed as conditions: the platform supports ten times the load without re-architecture, onboarding a new product line takes weeks not quarters, AI is doing real production work in the core workflow. Those are stable enough to plan against and concrete enough to argue about. Everything below them, features, sequencing, tooling, is a hypothesis I expect to revise. Plan in horizons, not quarters The Gantt-chart mistake is planning all of time at the same resolution. Reality doesn't work that way, so the plan shouldn't either. I run three horizons at different fidelity: - Now (0 to 6 months): committed, sequenced, resourced. This is a plan you can hold a team to. - Next (6 to 18 months): directional bets with rough shape. Named, owned, scoped enough to argue about, not yet scheduled to the week. - Later (18 to 36 months): the outcomes and the big architectural and AI bets, deliberately vague on execution. Its job is to keep today's decisions from boxing in tomorrow's options. The discipline is letting work earn its way forward. Things graduate from Later to Next to Now as you learn, and they get more detailed only as they get closer. If a roadmap specifies month thirty in the same detail as month two, that detail is fiction and it will cost you, because teams treat written-down detail as commitment whether you meant it that way or not. "A roadmap's job isn't to predict the future. It's to make sure today's decisions don't quietly foreclose the future you actually want." Architecture and AI are roadmap citizens, not someday-items This is where most roadmaps quietly rot. Architecture work and AI initiatives get parked in a "later" column that never arrives, because feature work always wins the quarter. Then three years pass, the platform can't take the next jump in load, the AI bet is still a slide, and the roadmap that looked responsible was running up debt the whole time. I treat both as first-class line items with the same standing as revenue features, because they are the thing that lets you keep shipping revenue features. Architecture earns roadmap space when it's tied to a named future outcome, not refactoring for its own sake, but the specific re-architecture that the ten-times-load outcome demands. AI gets treated the same way: not "explore AI" floating in Later forever, but concrete bets with owners, evals, and cost ceilings, sequenced so the foundational work, data, observability, guardrails, lands before the agentic systems that depend on it. I've run LLM-powered data agents on roughly 250 million records a month across a 75-node cluster. None of that happens if AI lives in the someday column. Build in the slack reality will demand A roadmap planned to 100% capacity breaks the first time anything goes wrong, and something always goes wrong. So I leave real headroom in the committed horizon, enough to absorb the outage you didn't see coming or the bet that quietly fails. That isn't padding; it's the part of the plan that admits you're working under uncertainty. The roadmaps that survive have enough give to bend without snapping, which is usually worth more than another layer of detail. And it has to be reviewed like a living thing, not blessed once and framed on a wall. I revisit the path every quarter and the direction once a year, on purpose, so changes are decisions rather than drift. A good multi-year roadmap doesn't promise what will happen; it takes a clear position on where you're going and stays honest enough to keep rewriting how you get there. Get that balance right and the roadmap stops being a document people quote at each other and becomes the thing that actually steers. If you want help building one that holds, that's the work I do as a fractional CTO. ### How to Scale a SaaS Engineering Team Without Breaking It: /blog/scaling-saas-engineering-org/ Scaling a SaaS org breaks more from added process than added people. The job is to grow capacity without grinding down the trust and speed that got you here. The thing that scales worst in a SaaS company is not the codebase or the infrastructure. It's the unwritten agreement that let a small team move fast: everyone knew what everyone else was doing, decisions happened in a hallway, and trust covered the gaps. That agreement is invisible right up until it snaps, and most leaders try to replace it with process exactly when they should be protecting it. I've taken more than 30 companies through this as a fractional CTO since 2018, across 12 teams in 7 countries. The pattern repeats: the org doesn't break because it grew. It breaks because the leaders scaled the wrong thing first. What actually changes at each threshold Headcount isn't a smooth curve. There are thresholds where the way the org coordinates has to be rebuilt, and pretending otherwise is how teams stall. The numbers are rough, but the transitions are real. - Around 8 to 10 engineers. One team becomes two. Shared context stops being free. You need an explicit way to decide who owns what, or every change becomes a negotiation. - Around 25 to 30. The founder or first lead can no longer be in every decision. This is where you either grow real engineering managers or quietly recreate a single bottleneck with a fancier title. - Around 50 to 75. Teams now ship into each other's blast radius. Platform concerns, on-call, and architectural standards stop being optional and become a function someone owns. - Past 100. Coordination cost dominates. The work is now mostly about interfaces between teams, not work inside them. At each of these, the temptation is to copy the artifacts of a bigger company, its review boards and planning cadences, without copying the conditions that made them necessary. That's how a 30-person org ends up with the process overhead of a 300-person org and the output of neither. The real failure mode: process outrunning trust Process is a substitute for trust. You add a review gate because you no longer trust that the change is safe by default. You add a planning ritual because you no longer trust that priorities are shared. None of that is wrong, but it's a tax, and the tax only pays for itself when the trust it replaces has genuinely run out. The failure mode I see most often is leaders installing process ahead of the trust gap, usually after one bad incident, because process feels like control. The result is an org that's slower without being safer. People route around the rules, the rules calcify, and the fast, high-trust culture that made the product good in the first place quietly dies. "Process is what you reach for when trust runs out. Add it faster than trust erodes and you pay for control you didn't need yet." The discipline is to add the minimum process that the current size actually demands, tie every new rule to a problem you can name, and delete rules that have outlived their problem. An org that can remove process is far healthier than one that only accumulates it. When I bring a team from low to elite DORA performance in around 90 days, almost none of it is new ceremony. It's removing friction, tightening feedback loops, and rebuilding the trust that lets people ship safely without asking permission. - 30+: Companies scaled as fractional & interim CTO - 12: Teams led across 7 countries - 90days: Low to elite DORA performance Scale with both internal builds and external innovation Growing capacity is not only a hiring problem. Every threshold is also a make-versus-buy decision, and treating it as one keeps the org lean. The teams that scale well build the things that are genuinely their advantage and refuse to build the things that aren't. Internally, that means investing in the platform, the developer experience, and the few systems that differentiate the product. Externally, it means adopting the tools, services, and now AI-driven capabilities that let a smaller team do more, so you scale output without scaling headcount in lockstep. I've scaled orgs through both: building the proprietary systems that mattered, and pulling in outside innovation everywhere a custom build would have been ego, not advantage. The discipline is knowing which is which, and that judgment is most of the job. What I'd protect on the way up - Short feedback loops, from commit to production and from customer to roadmap. - Clear ownership, so accountability survives every reorg. - The right to delete process that no longer earns its cost. - Enough slack that the team can absorb the next threshold before it arrives. You don't scale well by importing the machinery of a larger company. You earn each new layer of structure one named problem at a time, and you guard the speed and trust sitting underneath it. The orgs that clear the next stage are the ones that grow capacity without forgetting why they were fast to begin with. Get into that habit early and the next threshold is a smaller jolt than the last one was. If one of these thresholds is coming at you, this is the work I do as a fractional CTO. Get in touch → ### How to Run Quarterly Tech Planning Without the Theater: /blog/tech-planning-aligned-to-priorities/ Quarterly and annual planning is where most tech roadmaps quietly drift from what the business actually needs. Here's how I keep the plan honest, sequenced, and tied to the numbers leadership cares about. Every planning season has a moment where you can feel it go wrong. Someone opens a deck, the columns fill with initiatives, story points get summed, and a quarter later the company is no closer to the two or three things that actually move it. The artifacts looked rigorous. The planning was theater. I've watched smart teams burn a full quarter producing a beautiful plan that answered the wrong question. What a technology plan is supposed to do is narrow: take what the business is trying to win and turn it into the sequence of engineering bets most likely to win it. The ceremonies and the velocity math either serve that, or they're overhead dressed up as discipline. Start from priorities, not from the backlog Most planning starts in the wrong place. It starts with the backlog and last quarter's unfinished work, then tries to reverse-engineer a strategy that justifies keeping all of it. That's how you end up with a roadmap that's busy and a business that's stuck. I start the other way around. Before any team touches a planning doc, I get the company's real priorities in plain language from the people who own the P&L: what are the three outcomes that matter this year, in what order, and how will we know we hit them? Not themes. Outcomes with numbers attached, improve gross margin, land enterprise tier, cut onboarding time in half. If leadership can't name them crisply, that's the first problem to solve, and no roadmap is worth building until it's solved. Only then do I translate. Each priority becomes a small set of technology bets, and each bet carries an explicit claim: if we ship this, this metric moves, by roughly this much, by this date. That claim is what makes the plan falsifiable. Anything in the plan that can't trace back to a named priority doesn't belong in the plan. It might be real work, keeping the lights on, paying down risk, but it gets its own bucket and an honest cap, not a costume that makes it look strategic. "When a line item can't name the priority it serves or the number it moves, you don't have a plan. You have a wish list with deadlines." Sequence the bets, don't just list them A pile of good initiatives still isn't a plan until you've decided the order to do them in. Most teams treat sequencing as a capacity exercise, fitting the work into the quarters, when the real question is what you need to find out first and what can't start until you know it. So I sequence by leverage and by uncertainty. The bets that unlock other bets go early. The bets riding on a risky assumption get a cheap test up front instead of a full build six weeks in. The expensive, irreversible decisions get pulled forward so we're wrong on paper before we're wrong in production. And I deliberately leave slack, because a plan packed to a hundred percent of capacity is a plan that breaks the first time reality shows up, which is always. - Lead with the bets that unblock the most downstream work. - Front-load the cheapest test that could kill a risky assumption. - Pull irreversible, high-cost decisions earlier, not later. - Reserve real slack for the unknowns you can't yet name. Keep the plan honest as the world moves The biggest lie in planning is that the plan is done when the quarter starts. Markets shift, a competitor ships, a key assumption turns out false in week three. If the plan can't absorb that, it starts working against you, and clinging to it out of pride is how good teams keep marching toward a target that stopped mattering in February. I keep the plan honest with a light, regular cadence, not a heavy re-plan. Every couple of weeks the question is simple: are the bets still tracking to the metric claims we made, and are those priorities still the priorities? When a bet is clearly failing its claim, we kill it early and redeploy, no ceremony, no sunk-cost speeches. When leadership reshuffles the priorities, the plan reshuffles with them, openly, so everyone sees what got traded and why. This is also where a fractional CTO earns their keep: holding the line between disciplined adaptation and aimless thrash, so the plan flexes without dissolving into whatever was loudest this week. Run this way, planning stops being a season you survive and turns into something closer to a running habit. The annual plan sets the direction and the big irreversible bets, the quarter sequences them, and the biweekly check keeps them true. The deck was never the point. What you're actually after is a team that can tell you, on any given day, which company priority their work is serving and how they'll know it's working, and that can change course without losing the thread. That's the difference between a plan that survives contact with reality and one that just looked good in the room. ### How to Build a Product Roadmap: Business Before Backlog: /blog/business-before-backlog/ Most roadmaps are a pile of features looking for a reason. The fix is to start from the product and the P&L, then build the engineering to match. Hand me a roadmap and I can usually tell within five minutes whether the team is building in service of the business or just building. The tell is simple: can anyone in the room say, in one sentence, what business outcome each initiative moves? When the answer is a shrug, the backlog has quietly become the strategy. That is backwards. You build a product roadmap by starting from the P&L: name the numbers the business must move this year, tie every initiative to one of them, and cut everything that can't claim a number. Technology in service of the business I think from a product and business point of view first, and everything derives down to technology and engineering in service of business development. Engineering is the most expensive way a company has of expressing its priorities, so those priorities had better be the business's, not the loudest engineer's and not the shiniest framework's. "The backlog is downstream of the business. Get the business question right and the priorities sort themselves." The method, in five moves The one-sentence test above is the doorway, not the method. Here is what I actually do when a company hands me a roadmap that has become a feature pile. It takes about two weeks the first time and an afternoon every quarter after that. 1. Start from the P&L, not the backlog Before you look at a single ticket, get the executive team to name the two or three numbers the business has to move this year. New revenue. Net retention. Gross margin. Cost to serve. Time to onboard a customer. Not ten numbers, and not "growth" in the abstract. If leadership can't agree on the short list, stop here, because no roadmap can be right when the destination is contested. Every roadmap fight I've ever mediated was really this fight in disguise. 2. Turn each number into problems For each number, ask what actually stands between the company and moving it. Talk to sales about why deals die. Talk to support about why customers leave. Read the churn interviews. The output is a list of problems stated in business terms, "mid-market trials stall because setup takes three weeks," not solutions stated in feature terms. This step is where most roadmaps go wrong before they start: they skip from the goal straight to somebody's favorite feature and never write down the problem in between. 3. Write the sentence for every candidate Now bring in the backlog, and make every initiative on it earn a sentence: we believe this work moves this number, by roughly this much, and we'll know by this signal. The magnitude can be a guess. The honesty can't. Some items will produce the sentence instantly. Most will produce silence, and the silence is the finding. An initiative that can't complete the sentence doesn't get scheduled later. It gets cut now. 4. Set an appetite, not an estimate For the initiatives that survive, don't ask engineering how long the work will take. Decide how much the outcome is worth. A problem worth $2M in retained revenue might deserve six weeks; a workflow fix might deserve three days. That appetite becomes a boundary the solution has to fit inside, and it forces the scoping conversation to happen before the work starts instead of as a slow-motion overrun after it. Fixed time, variable scope. It's the single most clarifying constraint I know. 5. Sequence by proof, and hold the line Order what's left by how quickly it proves or disproves its own sentence, cheapest evidence first. Then defend the list, because the real job of a roadmap owner is refusal. Every week will bring a loud customer request, a competitor's press release, an executive's pet idea, and each one arrives with a reason it should jump the queue. Deciding what not to build is the actual work, and a roadmap that can't say no is just a queue with better formatting. What about tech debt and platform work? The standard objection: this method sounds like it starves infrastructure, refactoring, and everything else that never gets a sales quote. It doesn't, it just makes that work compete honestly. Tech debt has a number too: the velocity it's costing, the incident hours it's burning, the enterprise deal the architecture can't support. Name it and platform work earns its place on the same list as everything else. What dies is the vague "hardening quarter" that nobody can connect to anything, and that deserves to die. A bet list, not a Gantt chart One more reframe. The output of all this isn't a twelve-month timeline with features pinned to quarters. Those documents are fiction by week six and everyone knows it. What you get instead is a short list of bets, each with a number it's chasing, an appetite it must fit, and a signal that will tell you whether it worked. Near-term bets are concrete. Far-out ones stay deliberately vague, because pretending to know what Q4 needs in January isn't planning, it's theater. This sequence, business first, problems second, engineering last, is the first movement of the method I run on every engagement. What it changes - Every initiative is tied to a number it is supposed to move. If we cannot name the number, we do not start. - Work that does not move the business gets cut without ceremony, including work the team is fond of. - Architecture is weighed in business terms: speed to revenue, cost to operate, risk to the thesis. - Engineers understand the why, not just the what, so the thousand small decisions they make each week point the same direction. It is the same instinct I bring to a turnaround, where the job is not clean architecture for its own sake but protecting and compounding the value of the asset. More on that here. Start from the business and the backlog almost writes itself. Start from the backlog and you get a very busy team shipping things nobody asked the market about. If your roadmap currently reads like the second kind, let's talk → ### How to Evaluate a Codebase: Read the Thesis Before the Code: /blog/read-the-thesis-before-the-code/ Good architecture only counts if it serves where the business is headed. So before I judge a system, I find out what the company is trying to become. When I step into a company, especially a post-acquisition one, the first thing I ask for is not the repository. It is the thesis. What is this business trying to become, on what timeline, and what has to be true for that to happen? Only then does the code mean anything. Code is evidence, and that's all it is A messy codebase is not automatically a problem. A beautiful one can still be a liability. What I actually want to know is whether the technology helps or hurts the thing the business is trying to do. A pile of pragmatic shortcuts can be exactly right for a company racing to a milestone, and I have seen gold-plated platforms that were really just an engineer's hobby on the company's dime. "I read the investment thesis before I read the code. The job is to protect and compound the value of the asset, not to admire the architecture." How it changes the work - Diligence starts with the business plan, then maps technical risk onto the parts that actually carry the thesis. - Remediation is sequenced by value at risk, not by how much a given module annoys the engineers. - Roadmaps tie back to the value-creation plan, so the board sees technology decisions in their own language. When I start from the thesis, my limited time and political capital go toward the things that actually move the outcome. When I don't, I tend to do excellent work on problems that were never going to matter. If you want the long version of how this plays out under a deal clock, see the turnaround playbook. ### How to Become an AI-Native Company: Rebuild the Operating Model: /blog/becoming-ai-native/ Most companies are bolting AI onto an operating model designed for a pre-AI world. Going AI-native means rewiring how products are built and how the organization runs. There's a tell when a company says it's "doing AI." Ask what changed about how the work gets done, and the honest answer is usually: nothing. They bought copilot seats, shipped a chatbot, and added an "AI" line to the roadmap. The operating model underneath is exactly the same as it was two years ago. That's AI bolted on. AI-native is a different thing. Becoming an AI-native company means rebuilding the operating model around AI, how work gets chosen, built, verified, and measured, instead of adding AI tools to a process designed before the tools existed. The gap between the two is wide enough that I think it sorts the next decade of winners from losers. Bolted on vs. native Bolted-on AI lives at the edges. The workflows, the org chart, and the way decisions get made all predate it, so the AI ends up as a feature nobody really depends on. Native AI is the opposite: the model is redesigned around AI from the start. You can tell which one you're looking at within a day of walking the floor. Bolted-on looks like copilot licenses with single-digit weekly usage, a pilot that has been "about to roll out" for two quarters, and cycle times that haven't moved since the tools arrived. The organization bought capability and left every constraint in place, so the capability idles. It's the corporate equivalent of buying a race car and keeping the speed limit. - Business, product and engineering processes rebuilt AI-first. - AI inside the SDLC, generation, review, testing, and operations. - Agentic systems doing real production work, not demos. - Leverage you can see in cost, speed and quality. "The winners won't be the companies that added AI. They'll be the ones that rebuilt themselves around it." What "operating model" actually means Operating model is one of those phrases that can mean everything and therefore nothing, so let me be concrete. I mean five things: how work gets chosen (planning, prioritization, who decides), how it gets built (the SDLC from spec to deploy), how it gets verified (review, testing, evals), how teams are shaped (roles, handoffs, spans of control), and how success is measured (metrics, budgets, incentives). Going native means putting all five under one question: what would this look like if we designed it today, knowing what the models can do? Ask that honestly and the answers get uncomfortable fast. If AI writes the first draft of most code, verification becomes the bottleneck, so that's where senior time and process discipline have to move. If a prototype costs an afternoon, the quarterly roadmap debate is mostly theater; you should be shipping the argument instead of having it. If one engineer with agents can do what a squad did, team shapes and manager spans are wrong. None of those are tooling decisions. Every one of them is an operating-model decision, which is why the org chart feels this shift before the tech stack does. Two halves: build and operate I think about the shift in two domains. Build is shipping AI deep in the product and the pipeline as production infrastructure that holds up at scale, with eval, observability, and cost control built in. Operate is changing how the organization actually works, so teams, decisions, and processes are designed around AI from the ground up. The build side is where the proof lives. I've run LLM-powered data agents processing roughly 250 million records a month on a 75-node Kubernetes cluster, cutting processing time by 90%. But the operate side is where the durable advantage is, because tooling is copyable and an operating model is not. Rebuilding that model is the core of the work I do as an AI-native leader. - 250M/mo: Records processed by LLM-powered agents - 90%: Reduction in data-processing time - 75: Kubernetes nodes orchestrating ingestion Where it goes wrong The failure pattern is almost always the same: tools first, model never. The company rolls out licenses, runs a training day, and announces an AI initiative. Usage spikes for a month, then settles onto the handful of people who would have adopted anything. Leadership tracks adoption, seats, prompts, sessions, because adoption is easy to measure, and adoption is precisely the wrong metric: it proves people touched the tool, not that any workflow changed. Two quarters later somebody asks what the spend actually bought, nobody has an answer, and the initiative quietly deflates. That's not a technology failure. It's what happens when you buy the artifacts of a transformation instead of doing one. The road runs through five phases There's also a sequencing reality: no company leaps from zero to rebuilt. In my experience the road runs through five phases: Equip, Experiment, Operationalize, Industrialize, Transform, and most companies stall in the first two, fully equipped and permanently dabbling, waiting for the tools to transform them on their own. The phases are the route. This essay is about the destination, and knowing the destination is what gets you unstalled: once you understand that the operating model is the thing being rebuilt, you stop mistaking phase one for the whole journey. Where to start Start with an honest audit of where AI creates real leverage versus where it's a distraction. Not a vendor's maturity assessment, an audit of your own workflows: which ones are bottlenecked on things models are genuinely good at, and what each one is worth if it gets faster, cheaper, or better. The output should be embarrassing in its specificity. This workflow, this team, this number. Then rebuild one workflow end to end. Not a pilot that dies in a slide deck, a process your team runs every day, redesigned around the model rather than with the model sprinkled on top: new steps, new checks, a new definition of done, and the old process actually retired. One genuinely rebuilt workflow is worth ten proofs of concept, because it forces every operating-model question, verification, ownership, measurement, at a scale small enough to answer. You never really finish going native. You keep redesigning workflows, and how many of them genuinely run on AI is about the only honest measure of how far you've gotten. If you want a partner in that shift, that's what my AI-native transformation work is for. Get in touch → ### How to Improve DORA Metrics: Low to Elite in 90 Days: /blog/dora-transformation-three-months/ How a struggling engineering organization went from fearful, manual releases to elite DORA performance in ninety days, and why process, not heroics, did the work. When I walked in, the team could not tell me when anything would ship. Deployments happened on weekends, by hand, with the whole team on a call praying nothing broke. Lead time for a change was measured in weeks. Change-failure rate was high enough that shipping had turned into something everyone dreaded. The board didn't care about DORA metrics. They cared that the roadmap kept slipping and that every customer commitment felt like a gamble. But the four DORA metrics, deployment frequency, lead time for changes, change-failure rate, and time to restore, are the cleanest proxy I know for whether an engineering organization can be trusted to deliver. So that's where we started. Diagnose before you prescribe The first two weeks were pure observation. I instrumented the pipeline, read the last six months of incidents, and sat in on every ceremony without changing a thing. I had seen it before. The bottleneck wasn't the people, it was the process. Long-lived branches, a manual release checklist nobody trusted, and a test suite so slow that engineers skipped it. "Elite performance is not a hero working harder. It's a system that makes the safe path the easy path." Make delivery boring We rebuilt the path to production around three principles: trunk-based development, automation everywhere, and small batches. In practice that meant four changes: - Trunk-based flow, short-lived branches, merged daily behind feature flags. - CI that engineers trust, a fast, reliable test stage that gates every merge. - One-click deploys, GitOps-driven, with automated rollback when health checks fail. - Small batches, shipping continuously instead of batching a month of risk into one release. None of this is novel. The work was in sequencing it so the team felt the wins early and never had to choose between speed and safety. It's the same sequence I run in every turnaround engagement. The results, in ninety days - Low→Elite: DORA performance band, all four metrics - 90%: Reduction in lead time for changes - <1hr: Time to restore service after an incident The metrics moved, but the bigger change was cultural. A release stopped being a thing anyone braced for, and the roadmap turned into something the team could actually commit to. The board never cared about the dashboard. What they wanted was to know a date and trust it, and now they could. If your releases still feel like weekend gambles, this is exactly the work I do as a turnaround CTO. Let's talk → ### What Kind of CTO Do You Need? A Guide by Company Stage: /blog/cto-for-each-stage/ The CTO who takes you from zero to one is rarely the one who scales you to a hundred. A field guide to matching the leader to the moment. Founders ask me whether they need a CTO. The better question is which kind, and for how long. The CTO role is really three different jobs wearing one title: a builder who ships from zero to one, a team-builder who turns heroes into a system from one to ten, and an executive who runs technology as a business at scale. The skills rarely live in one person across all three stages, and most of the expensive mistakes I get called into started with pretending they do. Zero to one: the builder Early on you need someone who ships. The job is to find product-market fit before the money runs out, which means writing code most of the day, making brutal scope cuts, and keeping the architecture simple enough to throw away, because you probably will throw it away. Process here is mostly waste. So is architectural perfectionism: a beautifully engineered system for a product nobody wants is the most expensive kind of failure a startup can afford. The profile is specific. Strong product instincts, comfort with ambiguity, and a track record of shipping under constraint. What you don't need yet is a big-company VP of Engineering. Someone whose instincts were formed managing two hundred people will drown at a whiteboard with three. I've watched seed-stage companies hire an impressive executive resume and get six months of hiring plans instead of a product. The title said CTO. The stage needed a builder. One to ten: the team-builder Once it works, the bottleneck shifts from the product to the organization. Now you need someone who hires well, installs just enough process to be predictable, and turns a few heroes into a system. This is where most early CTOs struggle. What made them great at shipping solo is often the same instinct that stops them from building a team that ships without them. The work changes texture completely. Hiring pipelines, onboarding, review culture, the first serious conversations about reliability and security, deciding what finally gets written down. The builder measured a good week in commits. The team-builder measures it by whether the team shipped without needing them. Plenty of brilliant builders find this stage boring, and boredom in a leader is expensive. Sometimes the honest move is to keep the founding engineer as a principal IC, the best they've ever been, and bring leadership in beside them rather than over them. "The founder-CTO who can't let go of the keyboard becomes the ceiling the company hits." Ten to a hundred: the executive At scale the CTO is a business leader who happens to own technology, managing managers, aligning engineering to the P&L, and answering to a board. It's a different job than the previous two, with almost none of the same daily work. The calendar tells the story: budgets, org design, vendor negotiation, technical diligence for the next raise or the next acquisition, succession planning for their own directs. Success is measured in quarters and delivered through other people. Engineers this CTO barely knows will execute decisions made in rooms those engineers never enter. Some leaders find that leverage thrilling. The zero-to-one builder usually finds it suffocating, and that mismatch of temperament, more than any missing skill, is why so few CTOs genuinely span all three stages. The mismatch is the expensive part Nearly every painful CTO situation I get called into is a stage mismatch rather than a bad person. The builder still personally rewriting code at forty engineers, quietly becoming the bottleneck the org routes around. The process-heavy executive hired two stages too early, installing quarterly planning at a company that needs to ship something this week. The board that pushes out a good CTO because the company changed stages underneath them and everyone mistook the change for underperformance. If any of that sounds close to home, I've written a field guide to the symptoms: the seven signs you need a fractional CTO. Why fractional and interim works You don't always need to hire permanently for the stage you're in, especially through a transition. A fractional or interim CTO can carry you across the gap, build the team and the process, and hand off to the right permanent leader once the shape of the job is clear. I've done exactly that for more than 30 companies since 2018. The two flavors solve different problems, and it's worth keeping them straight. A fractional CTO is ongoing and part-time: a day or two a week of senior judgment for a company that has real CTO work but nowhere near forty hours of it. An interim CTO is full-time and deliberately temporary: someone who takes the seat through a departure, a turnaround, or an acquisition and hands it off when the permanent hire lands. Same judgment, different dosage. Both beat the panic hire, and both cost a fraction of getting the full-time hire wrong. - Bridge a sudden departure without a panic hire. - Stand up engineering leadership before you can justify a full-time exec. - Stabilize after an acquisition while you search for the permanent CTO. - Get senior judgment in the room for the decisions that don't get a second try. One more thing the stage lens buys you: a cleaner hiring spec. Instead of "find us a great CTO," the search becomes "find us a team-builder who has taken a company from eight engineers to thirty," and suddenly the interview questions, the references, and the red flags all sharpen. Most CTO searches fail at the spec, not at the market. If you're staring down one of these transitions, that gap is exactly what I fill as a fractional CTO. Get in touch → ### How to Measure Engineering Productivity Without Breaking Trust: /blog/measuring-engineer-productivity/ Productivity metrics turn toxic the moment they're used to rank people. Here's how to measure the system instead of the individual, and still get the leverage leadership wants. Every few quarters a well-meaning executive asks me to rank engineers by output. Lines of code, story points, commits, pick your poison. I always say no, and then I explain why the question itself is the problem. The way to measure engineering productivity is to measure the delivery system, not the individual: DORA metrics for team-level speed and stability, paired with the business outcomes the work was supposed to move. The moment a metric is used to evaluate a person, that person optimizes the metric instead of the outcome. Goodhart's law isn't a theory in software; it's a Tuesday. Measure commits and you get more commits. Measure points and the estimates inflate. None of that is productivity; it just teaches people to game you. Why ranking individuals always backfires It's worth being precise about why, because the instinct behind the request is reasonable. Engineering is expensive and leadership deserves to know whether the money is working. The problem is that individual output in software has no honest unit. The best week of a senior engineer's year might produce negative lines of code: a deletion that removes an entire class of bugs, a design conversation that kills a doomed project before it burns a quarter, an hour of unblocking that saves three other people a day each. Every counting scheme ever proposed misses that work completely. Worse, it punishes it. Then there's what measurement does to behavior. Rank people on commits and your best reviewer stops reviewing, because review doesn't show up in their number. Rank on story points and the estimates inflate within two sprints. Rank on tickets closed and the tickets get smaller. Nobody is being dishonest, exactly. They're responding rationally to what you told them matters. You wanted a measure and you built an incentive system, and now the incentive system is optimizing against you. And even a hypothetically ungameable individual metric would point at the wrong thing. Take your strongest engineer and drop them into a codebase with a forty-minute build, a flaky test suite, and a two-day review queue. Their output craters. The engineer didn't change; the system did. Most of the variance leaders attribute to individuals actually lives in the system around them. That's good news, because the system is the thing you can fix without anyone feeling hunted. Measure the system, not the person The useful signals are nearly all system-level. How long does a change take to reach production? How often does delivery fail? How quickly do we recover? These describe the machine the team works inside, and improving the machine helps everyone at once. - Flow, lead time and deployment frequency tell you whether work moves. - Stability, change-failure rate and restore time tell you whether it's safe. - Friction, where do engineers wait? Review queues, flaky tests, slow environments. - Focus, how much of the week is uninterrupted deep work versus context-switching? "If a metric can be used to punish someone, it will be, and then it stops telling you the truth." Start with DORA When a company asks me where to begin, my answer is boring and consistent: the four metrics from the DORA research program. Lead time for changes, deployment frequency, change-failure rate, and time to restore service. More than a decade of research across tens of thousands of teams sits behind them, and they have the one property that matters most here: they describe a team's delivery system and never a person. The four work as two pairs held in tension. Lead time and deployment frequency measure speed. Change-failure rate and restore time measure stability. Game one pair and the other exposes you immediately: ship recklessly fast and your failure rate gives you away, play it so safe nothing ever breaks and your lead time does. That tension is what makes DORA hard to fake and safe to put on a wall where the whole team can see it. It's also achievable fast. Most of the raw data already lives in your Git history, your CI system, and your incident tracker, so a first honest dashboard is weeks of work rather than quarters. I've written up how one engineering organization went from fearful, manual releases to elite DORA performance in ninety days. Process did the work. Nobody got ranked, and nobody quit. DORA proves the machine works, not that it matters Here's the trap on the other side. A team can post elite DORA numbers while shipping things no customer wants. The metrics grade the delivery machine, and a flawless machine pointed at the wrong target is just faster waste. So pair the system metrics with business outcomes: the revenue, retention, cost, or risk number each initiative was supposed to move. Every meaningful piece of work should have that number attached before it starts. If nobody in the room can name it, you don't have a measurement gap. You have a prioritization problem wearing a measurement costume. That pairing is the whole method. DORA answers "is the engine healthy." Business outcomes answer "is it pointed somewhere worth going." Either one alone will mislead you. Together they give a board an honest picture in two slides. The AI wrinkle: new dashboards, same trap AI tooling has handed the rank-people instinct a shiny new number to abuse: token consumption. I've watched executives pull up usage dashboards and ask which engineers are "really using the AI," as if tokens burned were work done. They aren't. Tokenmaxxing is the lines-of-code fallacy with a fresh coat of paint, and it misleads in both directions, flattering the engineer who flails through forty re-prompts and punishing the one who solved the same problem in three. AI also raises the stakes for measuring the right thing. When generating code is cheap, raw output inflates across the board and output-shaped metrics get even less meaningful than they already were. The signal moves downstream: did the change survive review, did it ship, did it stay shipped, did the outcome move. The system-level view was always the right one. AI just made the individual-output view actively hallucinatory. What leadership actually wants When a board asks about productivity, they're rarely asking who to fire. They're asking whether the investment in engineering is paying off and whether the next commitment is safe to make. System metrics answer that honestly, and they do it without turning your best people into adversaries. Point the measurement at the system and fix the friction. The individuals were never the bottleneck. If your metrics have already turned toxic, this is fixable, and it's usually the first thing I repair in a turnaround engagement, because nothing else improves while people are afraid of the numbers. If your dashboard has quietly become a weapon, let's talk → ### What Is Operator Experience? The UX Nobody Designs For: /blog/operator-experience-fixation/ Everyone talks about user experience. Almost nobody designs for the operator, and that is the experience that decides whether a software business can actually afford to grow. I have an unhealthy fixation on experience, for users, for customers, and most of all for the operators who actually run the business on top of the software we build. The first two get all the attention. The third is where companies quietly bleed out. Operators are the support agents, the ops team, the finance analyst reconciling a report at 11pm, the implementation engineer onboarding a new client. They live inside the internal tools nobody bothered to design. And their experience compounds straight into your margins. "Customer-facing UX wins the deal. Operator UX decides whether you can afford to keep it." Why it gets ignored Internal tools have no champion. They never show up in a demo, and the people who suffer from them rarely sit in the room where the roadmap gets decided. So the friction just builds up. People invent manual workarounds, keep a spreadsheet to bridge two systems that won't talk to each other, and hold the real process in their heads until they quit and it walks out the door with them. It's the same blindness I write about in running support as one system: internal friction never gets a ticket, so it never gets fixed. What changes when you care - Support resolves issues in minutes instead of escalating to engineering. - Onboarding a new customer stops requiring a heroic effort. - Headcount scales sub-linearly with revenue, the whole point of software. - Key-person risk drops because the process lives in the tool, not in someone's head. When you treat operators as first-class users, the business gets cheaper to run as it grows. When you don't, every new customer costs you a little more than the one before, and at some point the growth isn't worth what you're paying for it. Starting from the experience, including the one nobody demos, is baked into how I work. So I obsess over the operator. Someone has to. If your ops team is drowning in workarounds, let's talk → ### On-Prem to Cloud Migration With Under an Hour of Downtime: /blog/on-prem-to-cloud-native/ Big-bang migrations fail loudly. Here's the incremental, reversible approach I use to move legacy systems to the cloud while the business keeps running. Every catastrophic migration story starts the same way: a weekend cutover, a rollback plan nobody tested, and a Monday morning that becomes a week. I've re-platformed legacy enterprise systems to Kubernetes across Azure, AWS and GCP and kept downtime under an hour. Luck had nothing to do with it. I just never let anyone do a big bang. Reversible, always The governing rule is that every step must be reversible. If a change can't be rolled back in minutes, it gets broken into smaller changes until it can. That single constraint shapes everything else. - Strangle, don't replace, route traffic to new services incrementally behind a proxy, leaving the legacy path live. - Dual-write and verify, write to old and new data stores in parallel, comparing results before you trust the new one. - Shadow traffic, replay production load against the new system with no user impact until it earns confidence. - Cut over a slice, move one tenant, one region, one feature at a time, with an instant route back. "Downtime is a function of batch size. Shrink the batch and the risk shrinks with it." The cutover that wasn't an event By the time the "final" cutover arrives, almost everything already runs on the new platform. The remaining switch is small, you've rehearsed it, and you can still walk it back. That's why the sub-hour window holds up in practice instead of being a number you hope for. On one year-long transit re-platform we hit zero downtime doing exactly this. - <1hr: Downtime on enterprise on-prem to cloud cutovers - 0: Downtime on a year-long logistics re-platform - 3: Clouds in production, Azure, AWS, GCP The cloud isn't the hard part anymore. Doing the move without betting the business on a single weekend is, and that's entirely a function of how you sequence the work. Sequencing migrations like this is core to the fractional CTO work I do. If you have one looming, let's talk → ## Problems I Solve ### Everyone Has Claude. Now What?: /problems/everyone-has-claude-now-what/ A 50-person travel company opened Claude to everyone and now wants ground rules that keep the spend ROI-positive — without leaderboards, lockdowns, or a $10k runaway. The real risk isn't overspend. It's underuse. This one came to me as a long, thoughtful note from a CEO, and I want to answer it in public because almost every small-business owner I talk to is about to be standing exactly where he is. The short version: he runs a 50-person travel company. He and his business partner both happen to be technology hobbyists, so over the last six months they've used Claude to build internal applications that improve the business, things they'd previously have had to pay an agency a fortune for, or simply never have done. It worked. So a month ago they opened the door: anyone in the company can request a Claude account, and for now they're deliberately letting people find their own way and seeing what happens. He knows that at some point he'll need ground rules so the spend earns its keep. But he was careful to say that "positive ROI" sounds stricter than what he means. What he actually wants is to not pay a subscription for someone who only ever asks Claude for the weather, and, at the other extreme, to not wake up to a $10,000 bill because someone asked Fable to build a competitor to archive.org. He's explicit that this is not a leaderboard culture. He's seen the companies that gamify AI usage and reward people for spend, and he wants no part of it. Then came the part that made me want to write the whole essay. Until fifteen minutes before he wrote to me, his plan was to lock accounts down until people completed some of Anthropic's courses, paired with a one-time cash reward for finishing them. He was going to skim the course titles and pick a minimum. Instead he actually sat down and did the first one, Claude Platform 101, and was shocked: it's built for seasoned programmers, useless as an on-ramp for a non-technical employee, and far too dry for an experienced engineer to sit through anyway. So part of his message was a flare sent up to whoever at Anthropic can prioritize training that normal companies can give to all their people. And part of it was a question to other small businesses: what are you actually doing about this? It's a genuinely good problem, and he's already done the hardest thing, which is to think clearly about what he doesn't want. So let me take it apart. You're worried about the wrong cost The instinct to control spend is right, but the two failure modes he named are the two least dangerous things that can happen here, because both are cheap and both are trivially preventable. Notice that they're not even the same kind of cost. The weather-forecast employee is a subscription problem: a flat seat fee you pay every month whether the person does anything valuable with it or not. The archive.org clone is a usage problem: metered spend that scales with how hard someone runs the thing. They feel like one worry, "are we wasting money," but they need two completely different controls, and conflating them is how people end up reaching for a blunt instrument like mandatory courses. The runaway bill is the easy one. You never prevent a $10,000 surprise with training or trust or a policy memo. You prevent it with a number. Every one of these platforms lets you set a hard spend cap, per seat or per workspace, and an alert below it. Set a default ceiling that's generous enough that nobody doing real work ever feels it, and low enough that the worst-case accident is an annoyance, not an incident. That's it. That's the whole defense, and it's the same idea I keep coming back to in when building gets cheap: an appetite is a blast radius. You're not deciding what people are allowed to attempt. You're deciding the largest a mistake is allowed to get. Once the cap exists, you can stop worrying about the runaway entirely and let people be ambitious inside it. "You don't prevent a $10,000 surprise with training. You prevent it with a number." The weather-forecast subscriber is even less of a threat, because the cost is a rounding error against a single hour of that person's salary. A seat costs less per month than a team lunch. If someone uses it only for trivia, you haven't been robbed; you've learned something. You've learned that this person hasn't yet found the thing in their own job that Claude could take off their plate. That's not a spend problem to be policed. It's a training problem to be solved, and as he discovered the hard way, it's the part nobody has built the right tool for yet. Which is the reframe I'd offer him before anything else: the expensive failure in a company your size is not the person who overspends. It's the forty people who under-use. The runaway bill costs you ten grand once and teaches you to set a cap. Forty employees who never get past the weather forecast cost you the entire return you opened the door for in the first place, quietly, every month, forever. The leaky faucets stay leaky. That's the real ROI question, and it's the opposite of the one the spend worry points you toward. The lockdown plan solves the wrong problem too So I'm glad he did Claude Platform 101 before launching the initiative, because the original plan, lock accounts until people finish courses, plus cash for completion, would have actively worked against him, and not only for the reason he found. Start with the courses themselves. He's right, and it's worth being blunt about it: the platform curriculum is developer documentation in video form. It teaches the API, the console, the building blocks. It is the correct material for the two hobbyist founders and for any engineer they hire. It is the wrong material, by a wide margin, for the travel agent, the operations coordinator, the bookkeeper, the marketer, the people whose low-value work is exactly what you most want Claude to absorb. Gate their accounts behind it and you've built a filter that admits the people who least need help and excludes the people who need it most. You'd be optimizing for the wrong nine percent. Then there's the cash bounty, and here I'd push back even though it comes from a generous place. A reward for completing a course measures completion, not capability. It pays people to click through videos, not to change how they work. Worse, it quietly reintroduces exactly the dynamic he told me he wants to avoid: it ties money to an AI activity, and people optimize for whatever you put money on. He doesn't want a leaderboard for spend. A bounty for course-completion is a leaderboard for seat-time. It's the same machine with a different label. "Gate the accounts behind a developer course and you build a filter that admits the people who least need help." The deeper issue is that locking down and requiring training is an Equip-phase reflex: treat adoption as a procurement-and-compliance problem, roll out a mandate, check a box. It feels like governance. It produces almost nothing, because nobody ever changed how they work because a course told them to. People change how they work when they see a specific, annoying piece of their own job disappear. What I'd actually put in place Here's the lightweight version, sized for 50 people and a CEO who, sensibly, does not want to build a bureaucracy around this. - A default spend cap on every seat, generous enough to be invisible to real work and low enough that a runaway is a shrug. Raise it for anyone who hits it doing something valuable. That single control retires the $10k fear permanently. - Kill the lockdown and the completion bounty. Keep the open-door access exactly as it is; it's the best thing he's done. The friction you remove is worth more than the friction you'd add. - Give every person one real problem, not a course. Ask each employee to name the single most tedious, repetitive part of their week — the report rebuilt by hand, the inbox triaged manually, the data re-keyed between two systems — and make Claude take a swing at that. One concrete leaky faucet beats ten hours of video. - Run a 30-minute internal show-and-tell every two weeks. Whoever automated something demos it. This is your real training program, and it costs nothing: it's peer-to-peer, it's in your own context (travel, your tools, your workflows), and it spreads the one thing courses can't — the feeling of "oh, I could do that with my job too." - Name a champion, probably one of the two founders, as the person people bring a stuck problem to. Not a gatekeeper. A helper who unblocks and then sends the person back to keep going. - Light-touch sense of what's sensitive. A one-paragraph note on what data shouldn't be pasted where matters far more than a course completion certificate, and takes a tenth the time to write. Notice what this does to his ROI question. You stop measuring spend and start measuring the thing you actually wanted: low-value work removed. How many manual reports stopped being manual. How many hours a week came back to the team. How many leaky faucets stopped dripping. Those are the numbers that tell you the seats are paying off, and not one of them is a leaderboard. This is the same move I push on every executive in the Experiment phase: a bet isn't graded on activity, it's graded on the problem it closed. The flare he sent up is the real story I want to come back to his message to Anthropic, because he's identified something important and he's not wrong about it. There is a genuine, gaping hole in the market right now between "here's how to build with the API" and "here's how a person who has never thought about any of this can make Claude a normal part of their Tuesday." The first is well covered. The second barely exists. Every vendor, Anthropic included, has built training for the people most like the people who build the product, and almost nobody has built the patient, jargon-free, role-by-role on-ramp that a bookkeeper or a travel agent or a warehouse supervisor actually needs. The hobbyist founders crossed that gap on their own enthusiasm. Their 48 employees can't be expected to, and shouldn't have to. So I'll amplify his ask, because I hit the same wall on client engagements constantly: the single highest-leverage thing the model vendors could ship for the small and mid-sized companies that make up most of the economy is not a more capable model. It's a genuinely good, genuinely non-technical adoption curriculum, built around "find the boring thing in your job and make it go away," that a CEO can hand to all 50 people on day one without apologizing for it. Until that exists, the work falls to the company, which is exactly why the show-and-tell and the champion above matter so much. You are, for now, building the training that should have come in the box. "Every vendor built training for the people most like the people who build the product. The on-ramp for everyone else barely exists." The thing he got most right I'll end where he started, because his instinct was sound and I don't want it lost in all my reframing. He looked at the leaderboard companies, the ones that incentivize spend, and recoiled. Good. That model is a category error. It rewards the activity instead of the outcome, and it produces theater: people running up usage to top a chart while the actual low-value work sits untouched. He smelled it and walked away, and that single act of taste will save him more than any policy I could write. The whole game, at 50 people or 5,000, is the same: make it safe and easy for people to point this tool at the parts of their own jobs they hate, give them a cap so nobody can blow up the budget, measure the work that disappeared rather than the money that was spent, and accept that the training that actually moves people doesn't exist yet so you'll grow it yourself, in your own context, peer to peer. Do that and you don't need a leaderboard. The faucets stop dripping on their own. And to his closing question, what are other small businesses doing: this, mostly, the ones doing it well. Caps not gates, problems not courses, show-and-tell not certificates. If you're a small-business owner standing where he is and you've found something that works, I genuinely want to hear it. And if you'd rather have someone help you stand up the operating model behind it, that's the work I do. Let's talk → ## Oshri's Claude Prompt Library ### Let Claude Code Analyze Every GitHub Issue for You: /prompts/analyze-github-issues-with-claude-code/ One prompt that points Claude Code at your GitHub issues, has it read the codebase, design a solution, and write the answer straight back onto the issue. Works with Jira and any connected MCP too. Prompt: /goal connect to the gh cli or github mcp, pull every issue that has not been analyzed already with a solution design readme file attached, run the analysis against the codebase, design the solution, update the issue. It is important to ask me questions if you are not sure. Ask one question at a time and wait for my response Here is a Claude Code trick I keep coming back to. Point it at your GitHub issues and let it do the first pass of engineering thinking for you: read the issue, read the code that issue touches, design a solution, and write that solution back onto the ticket. It sounds like a lot. It's one prompt. Copy the block above, paste it into Claude Code, and watch it work through your backlog one issue at a time. What the prompt actually does Read it left to right and it's really a checklist you're handing to an engineer: - Connect to your issues. It reaches your backlog through the gh CLI or the GitHub MCP, whichever you have wired up. - Find the untouched ones. It skips any issue that already has a solution-design readme attached, so it only works on tickets nobody has analyzed yet. Run it again next week and it picks up where it left off. - Analyze against the real codebase. This is the part that matters. It doesn't guess from the title. It opens the files the issue is about and reasons about your actual code. - Design the solution. It writes up an approach: what to change, where, and the trade-offs. - Update the issue. The design lands back on the ticket, so the next person to open it, human or agent, starts from a real plan instead of a blank page. The line that makes it safe The last sentence is the whole trick: ask me questions if you are not sure. Ask one question at a time and wait for my response. Without it, an eager agent fills every gap with a confident guess, and you end up reviewing a stack of plausible-but-wrong solution designs. With it, Claude Code stops and asks the second it hits real ambiguity. One question, your answer, then it keeps going. You stay the decision-maker. It does the reading and the typing. "An agent that guesses is a liability. An agent that asks one good question at a time is a teammate." It isn't only GitHub Swap the source and the same pattern holds. Jira, Linear, a support queue, anything you can reach through a connected MCP. The shape of the work doesn't change: pull the items nobody has looked at, understand them against the code, propose a plan, write it back. Point it at whatever system your team actually lives in. Make it a routine, not a chore The real unlock is running this on a schedule instead of by hand. Wire it to fire whenever a new issue is opened, and every ticket gets a first-draft solution design before a human ever reads it. Your morning triage stops being "what is this and where does it live" and becomes "do I agree with the plan." That's a much faster meeting. This is what using AI properly looks like. Not a chatbot off to the side that you copy-paste into. An agent connected to the tools you already use, doing the boring 80% of the analysis, and pulling you in for the 20% that needs your judgment. Want help turning prompts like this into standing routines across your engineering org? That's the kind of thing I set up as a fractional CTO. Get in touch →