For quite some time I have been having the same conversation in private. Sometimes with peers who run engineering and product organisations. Sometimes with people who pay for an hour of my time to ask it. It goes the same way almost every time. They have the licences. They have the policy. They have a number for results sitting in somebody’s goals. And when I ask what it has actually produced, there is a pause, every time, and then a long answer about what is about to happen.
I have wanted to write this down for a while and kept finding reasons not to. The honest version starts with me getting it wrong myself, so it starts there.
When an organisation decides it is serious about AI, it is usually the same sequence. A team or a committee evaluates the tools. Security and legal write the policy. Procurement negotiates. Licences are bought and handed out. And finally a number lands in somebody’s goals for the year.
Then comes the part that was never written on the plan. The waiting. The gains are supposed to arrive faster and bigger than the old way of working could deliver. Mostly they arrive smaller and later than the case promised, and often they cannot be found at all.
If you are somewhere in that sequence, or about to begin it, this is written for you. And it is not because your organisation is doing it badly. Most of the ones I talk to are doing exactly what is being asked of them, competently and in a sensible order. The gap opens anyway, for a reason that has very little to do with effort or with the quality of the tools.
My own version of it was small and ordinary. By the second half of 2025, when the pressure to show something for AI had arrived in most executive goals, we had done everything the process asks for. The tool was evaluated. The reviews were done. The licences were bought and handed out. In fact one engineer started using it on his pull requests and told me it had cut the back and forth down considerably. He was right, he was pleased, and I took it as the start of something that would land the way it had been promised.
Then, for a while, very little else happened. There was not much resistance in any meeting. What followed was quieter, and very different from what I had imagined. Engineers would pick the tool up, use it for a week or so, and go back to working the way they always had.
Out of frustration, I asked one of them why. He said that on the changes he was making, it took him longer to describe what he wanted and check what came back than to make the change himself. He was not being difficult and he was not being nostalgic. He had done the sum on his own work, and it came out against the tool. I put the same question to a few others and got a version of the same answer.
I heard all of it as reluctance. That was my mistake, and I spent longer than I should have waiting for enthusiasm to do the spreading for me.
He was right. That is the part I sat with afterwards. The tool was good. The case for it was sound. And the engineer who had been told to use it was still right not to.
The arithmetic
His sum is worth writing out, because it is most of the story.
On a small, well understood task, an experienced engineer is faster alone. Describing the work to a machine and checking what comes back costs more than doing it.
The tool wins in two situations. When the work repeats often enough that the engineer and the tool learn each other. And when the job is big enough that the overhead disappears against its size.
Everywhere else, asking people to use AI is asking them to be slower on purpose. They feel it in their own working day whether they say so or not, and no amount of encouragement from me was going to change the sum.
Which is why the instruction most organisations settle on, use AI more, does so little. The useful instruction names the work.
Why this moves slowly, and why that deserves some respect
There is a second reason large companies move slowly here, and it gets less credit than it should.
Enterprise software is not judged the way a side project is judged. It carries service commitments, or a client’s experience, or a few thousand colleagues who need it to work on a Tuesday morning. Reliability outranks efficiency there, and it should. So when something genuinely capable arrives, the responsible response is rarely speed. It is a run of dull, necessary questions about failure modes and accountability, and then a decision.
Most of the organisations I work with manage the questions well. Far fewer reach the decision.
There is a harder version of this that many leaders are living through, and it deserves saying plainly. When a company moves through waves of restructuring, its capacity to execute falls. At the same time, expectations for revenue and for the quality of the work do not fall with it, and they should not, because those expectations carry the future of the people who stay. That is what makes those periods hard to lead through. It is also when the temptation is strongest to announce a tool as the answer. The people who remain are being asked to meet an expectation that has not moved, with a capacity that has. Anything that genuinely helps them is a way out. Anything that reads as a plan to need fewer of them is a threat, whatever the slide says.
I have stood in front of a room in that condition. You can feel which of the two things people have decided you are saying, usually inside the first minute, long before anybody asks a question.
Two reactions, and what they have in common
Instead of that decision, one of two things tends to happen.
The first is the stampede. Go all in, buy widely, run a dozen pilots, put up a slide about being AI-first, and treat velocity as evidence of seriousness. From the outside it looks decisive. Inside, it usually means the hard question has been skipped: where does this genuinely improve the work, where does it not, and who will stand behind that answer?
The second is the freeze. The ground is moving too quickly, the winners are not obvious, so the sensible course is to wait and enter later when the picture has cleared. It comes with a paper attached, usually a careful one. The trouble is what that paper looks like a year on, which is the same paper with the dates changed, because nothing happened in between that would have taught anybody anything new. The organisations I have watched wait did not arrive later with better judgment. They arrived with less. Judgment about a technology is built by using it on something that matters.
These two are usually described as opposites, yet they produce the same result. The decision about where the technology belongs has been handed outside the organisation, to the market in one and to time in the other, and no one inside owns having chosen.
There is a great deal of fear in this market, and a number of people busy making a living from it. What a leader can usefully offer in the middle of that is a decision. This is where we move. This is where we deliberately do not. And I will answer for having said so.
You have been here before
Most people reading this have already been in a version of this situation.
Not this technology, but this shape of problem. Something arrives that is genuinely capable and genuinely unfinished. The organisation is unevenly excited about it. The budget wants a line item. And you have to work out what it changes and what it leaves alone while the room watches you decide. A platform migration, a cloud move, an outsourcing wave, a data consolidation: all of them put a leader in the same position.
I have been in it four times. In two the technology was AI. In the other two it was not, and those two carry more weight here, because they are what tells you the pattern belongs to how work is led rather than to any particular technology. A method that only holds for the current generation of tools is closer to a fashion than a method, and it will need replacing as soon as the generation turns.
Here is what stayed constant across all four, and what each one cost me.
What stays constant
Four things. The first is the decision. The other three are what makes it hold. None is complicated, and the difficulty is doing them while the room wants a faster answer.
Decide where it earns a place. This is the whole job, and most organisations move past it too quickly.
Qualify the work first, the way you would qualify anything else asking for budget. The tests below decide whether to hand it to a machine, not whether it was worth doing, and plenty of AI work passes every one of them while being worth nothing to anybody. The letters AI on a proposal are not a reason to skip the question.
What remains is three questions that decide whether, a verdict, and one deliberate exception.
One. Does the value clear the overhead? Does it repeat often enough to compound, or is it one job big enough that the setup cost vanishes against it? Describing work to a machine and checking it is not free, and on small one-off tasks it costs more than doing the work.
Two. Is the shape stable? Can you say what good looks like before you see the output? Novelty defeats delegation.
Three. Can it be checked cheaply? Can a person tell the output is right faster than producing it, or sample a batch and be confident? Otherwise the cost has moved rather than gone.
Most stalled programmes I have seen were funding work that fails the third one. The shape is usually the same. Something is being summarised or drafted, the output looks excellent, and the person receiving it still has to read the source to know whether it is right. Verification cost rarely appears in an AI business case. When a person cannot check the output faster than doing the work, the cost has moved rather than gone, usually somewhere harder to see.
The test bites hardest on work the team could not have done itself. Hiring the machine to do what nobody in the room can do sounds like the obvious use, and it is the most dangerous one, because a team that cannot produce the work often cannot check it either. But that is the line that decides, not the capability gap. Where the output is measurable without the missing skill, a test suite, an experiment, a number that can be looked up, work beyond the team’s reach is one of the best places a machine earns its keep. A cheap check is still not a promise that the machine reaches the bar, only the way to find out quickly. When what comes back keeps failing the check, that is the tests answering no in production, and that work gets parked, with a date.
The verdict comes in three forms, not one. Fund it. Refuse it, and say why out loud, because a refusal with no reason gets quietly reversed by whoever is keenest. Or park it with a date to look again, because work that fails today passes two quarters later, when checking has got cheaper or the models have moved. A parked item with no date is a no that nobody had the nerve to say.
Then one deliberate exception, small and named. Some work cannot answer the first question, because nobody can say in advance what it will find or what the finding will be worth. Some of it is still worth funding, because what you are buying is judgment rather than output. Hold a capped line for it. Say it is there. Keep it capped. Uncapped, that line is the stampede. Missing, it is the freeze.
Give people a bounded space to try. The question that governs this one decides how rather than whether: if it goes wrong, how far does the damage travel? A large blast radius is not a reason to stop. It sets how much freedom you can give.
Most organisations tell their teams to experiment safely. Far fewer write down what that permission covers, so teams guess, and guessing produces paralysis in the cautious and recklessness in the confident.
The version that works is unglamorous and specific. Tier what people may do by the damage a mistake would cause. Small blast radius: let the machine run and sample the output. Medium: let the machine draft and have a person accept. Large: have a person produce and let the machine check. On a single page, circulated, that is what a safe space looks like in practice. A permission slip with a boundary on it, which lets a cautious team move faster and keeps an eager one out of trouble.
Give away authorship. Very little scales out of one team on merit alone. It scales when other leaders have their fingerprints on it.
I learned this later than I should have. Early on I would design something sound, present it well, and watch it stop at a boundary that appears on no organisation chart, which is the edge of my own authority. The approach that works is harder on the ego and better for the outcome. Bring in the leaders whose work the thing touches and have them write their own part of it. What comes back is messier than what I would have written, and it comes back as theirs, which is the only version that survives contact with an organisation you do not control.
Keep a gate, and measure it. Anywhere a person accepts or rejects what a machine produced is a gate. A gate with no name against it becomes a queue. A standard with no gate behind it becomes advice.
So put a name on each one. Measure how often that person accepts. And put the budget for the people at the gate on the same page as the budget for the tools feeding them. Movement in an acceptance rate will tell you more about whether this is working than most of the dashboards you will be shown.
None of that is unfamiliar. It is roughly what leading through a technology change has always required, and it holds when the technology underneath changes, which is the main reason to trust it.
Four times, and what I refused each time
The same shape each time: what I was asked for, what the real choice was, what I refused, and what came of it. Two involved AI and two did not. Every figure comes from that organisation’s own measurement, none are independently audited, and companies are not named.
One. Predictable delivery while the organisation was getting smaller
A global consumer media company. A hundred-plus person product, engineering, design and QA organisation across three countries.
I was asked for delivery predictability while the organisation was also reducing capacity. So the real question was not how to make delivery predictable. It was how to hold it steady while the shape of the team kept changing. A harder problem, and a common one.
Two moves answered it, and only one of them involved AI.
The first was a standard. I partnered with the leaders whose work the delivery lifecycle crossed, engineering, QA, design, research, data science and customer support, bringing product from my side, and the customer’s expectations with it, to define a repeatable path from idea to production: the process, the artifacts, and a gate at every handoff with a named owner. The choice there was genuinely close. I could write it myself in about three weeks, and it would have been the better document. Or I could hand the pen and have each function write its own part, and take a quarter to get something looser than I wanted.
I took the second. It cost a quarter of calendar time and a standard lighter than the one in my head, because a document written by seven people converges on what all seven will actually do. I still wonder what the missing rigour cost us later. The rest was my own time; there was no squad and no budget line behind it. What I bought was adoption: the standard was approved and became the de facto format across the company, in teams I had no authority over, because their own leaders had written it.
The second move was AI, and no one had asked me for it. AI-assisted development was the only answer I could find that closed the capacity gap at the ratio we were working at, and it brought a second problem: a mandate would have been heard as an announcement about people’s own replacement. So what I refused was to run it as an efficiency programme. Not by hiding the numbers; they were the same either way. An efficiency programme is defined by where the gains go. Ours went into output, with the team we had, and none were booked as a reason to need fewer people.
We started with code migration, where the arithmetic worked: a body of work big enough that the tool’s overhead vanished against it, and work that eats engineer bandwidth while adding nothing the business is asked to prioritise. Compressing it with AI freed our most expensive constraint, the time of a shrinking team, for the work that carried revenue. It also gave a large number of engineers a low-stakes place to learn where these tools are strong, where they are unreliable, and what they cannot do at all. By the time anyone was asked to trust AI output on work that mattered, they had formed their own view of its limits rather than inheriting mine.
Then the two moves clicked together. A standard with a gate at every handoff does not care whether a person or a machine produced the artifact in front of it. Ours was written so that a person and a model would read it identically, which is what made AI output reliable enough to ship. Six months after rollout we were running at close to 80 percent of projects on time, through the restructuring, with cycle time down sharply and regression cycles closing in days rather than weeks.
If your organisation is reducing headcount while expecting the same output, the instinct is to announce AI as the answer. That announcement is usually what stops it working.
Two. The payment system I refused to replace
A private-equity-backed B2B sales and marketing data company. Six acquired products sitting over three shared databases.
The brief was speed to market. We were rebuilding the largest revenue product end to end, and the question underneath every decision was what to build ourselves and what to take from somewhere else.
The principle I work to is that you build what the customer will judge you on and you take everything else. Here that meant search and retrieval, the quality of the data coming back, and the interfaces around both.
Order management and payments were where the principle stopped being comfortable. We had an in-house system, built a couple of years earlier, that worked. It was slower and plainer than what we would have got by putting Stripe in.
Two options, both defensible. Integrate Stripe: several weeks of work to wire it in and harden it to a standard we would be comfortable taking real money through, and at the end of that a payment experience no customer would ever remark on. Or keep what we had, ship on the date, and go to market with the least modern part of the product still in place.
Stripe was the better system. That was never in question. I had to decide whether it earned a place in this release, and at that point it did not. The MVP existed to tell us whether customers could find the right information, search it well and get accurate data back. Payment had to complete. It did not have to be elegant.
It cost me weeks of defending a worse answer in front of people who could see the better one sitting right there, inexpensive and available. That is more tiring than it sounds. It also meant that if payments had failed in front of a customer, the reason would have been a call I made, and it could be named.
The product reached market on the schedule that mattered, and the payments decision was taken later against real usage instead of a guess. There is no AI anywhere in that story. It is the same judgment I now apply to AI tools every time. Not whether the new thing is better, but whether it earns a place in the release in front of me.Three. The time AI told us what not to build
The same media company as case one. Advertising, the largest of three digital revenue lines.
I was asked to improve the reliability of a platform on which reliability had never been defined. Everyone could tell you the experience was wrong. Nobody could tell you by how much. No number meant no budget anyone could argue for.
So we built the measure first. One that separated reloads the platform caused from refreshes users had chosen. We measured the estate as it stood and put a rate on it: 20 to 30 percent of sessions, by content type. With that number in hand the work became fundable. Worth noticing on its own: the instrument was the leadership act, and it needed no technology at all.
A page like ours was never just our code. Content, markup, styles, custom pieces, a stack of third-party scripts, all of it running on a browser whose memory management no public document explains. Nobody outside Apple can tell you when Safari decides a page has had enough; you learn it by watching your pages crash. The rabbit hole was genuine. You could optimise layer after layer and never know whether an exit existed.
Generative AI is what made that search finite: it read and mapped the whole monolith and every place a reload could fire, in one sweep. It produced a finding, not a fix. The dominant driver of the reloads sat outside our code: it was how much the pages asked the browser to hold, against device-memory limits that were not ours to move. A cleaner codebase would have bought margin at the edges without moving the number. So the expensive rewrite we had been circling was the wrong answer, and the rest of the hole was not worth the scope we had.
Nobody could have said in advance what that read would find, or what a finding would be worth, so it would never have cleared the first test. It was still the most valuable thing AI did on the whole programme, because what it produced was a reason not to build. That is what the capped line is for, and it is a use I have yet to see on a vendor slide.
What I refused was the obvious lever, which was to pull ads out until the pages behaved. Fewer ads would have improved the reliability numbers inside a quarter, and broken the delivery we had promised advertisers, which is where the revenue lived. I was unpopular for a while for saying so.
Once we knew the biggest lever was not in our code, the job became living within browser and memory limits we could not move. A dedicated squad experimented its way to where those thresholds sat, and the answer came out as a working rule: an ad every third slide, with a hard ceiling on the total in a long gallery. Enough headroom to keep the session alive, enough ads to make good on every campaign. That took eighteen months, and the work was tuning rather than rewriting.
Involuntary session failure fell to a small fraction of what we first measured, viewability rose on inventory we had already sold, and the metric outlived me and remains the company standard.
Four. Seven people, and a date the business believed
A broadcaster and streaming group in India. The data function behind a large advertising business.
I was asked for dashboards that each function could treat as its source of truth: sales, finance, streaming engagement, and more, across streaming and linear television, with delivery dates people could rely on. The team behind that ask was seven people working ad hoc: four data engineers, a BI engineer, a data manager, and one project manager standing between the contractors and every stakeholder with no authority to prioritise anything.
What I refused was everything we could not actually deliver. We productised the data layer itself: one architecture underneath, the tables and pipelines fixed once, so the same data could serve every function’s cut without rework, with the customisation living at the output layer where it belongs. On top of that we built a capacity-based intake, so no date was committed before demand had been reconciled against what seven people could absorb, and the sequence of dashboards was agreed with the leaders whose businesses they served. Some requests waited and some were declined, by me, with a reason attached. And we enforced the measurement licence windows inside the platform rather than in policy, which is the difference between a rule and a rule that gets followed.
It cost me bespoke responsiveness and about two quarters of goodwill. Anyone doing this should know that the unpopularity arrives well before the credibility does.
We finished at 90 percent of data product projects on time, with the remainder delayed against a stated business case rather than simply missed, and platform run cost cut by more than half, partly by finding the licences nobody was using. There is no AI in that story, and it might be the most AI-relevant one here. AI is a layer you add on top of a data foundation, and an organisation that has productised that foundation, and learned to decide what it will not promise, has already done the hardest part of adopting agents.
The strongest argument against all of this
The serious objection is that in a field moving this quickly, deciding early is how an organisation commits to the wrong thing. Optionality has real value. A leader who fixes an allocation of work in March may look foolish by September, when the tools have changed underneath the decision, and the cost of having committed is paid in rework and credibility.
I take that seriously. I also think it argues for a narrower conclusion than it appears to reach. Deciding where something earns a place today is not the same as committing to it permanently. The tests get re-run as verification gets cheaper and the models improve, and work that failed last quarter passes this one. What does not survive the objection is declining to decide at all. Waiting produces neither optionality nor learning. It produces delay, and a year later the organisation has the same decision in front of it with less experience to bring to it. And none of this refuses optionality. It prices it. The capped line is optionality with a number on it, bought on purpose, instead of optionality as the reason nothing gets decided.
There is a related objection to the arithmetic itself, which is that it is already dated. The overhead has fallen since 2025 and will keep falling, so the threshold where these tools start to win keeps moving down, and any rule of thumb built on today’s economics will be wrong soon enough. That is correct, and it changes the answer rather than the method. The threshold is a number to re-check each quarter, not a principle to adopt once.
What I have not seen
I should be straight about the limits of this, since I would not trust anybody who was not.
In my own record, the four have never all shown up in the same place at the same time. Each of the cases above does some of this well, and none of them carried all four across a whole estate, including anything I ran myself. Organisations that carry all four at once no doubt exist. Expect them to be rare, because this work is expensive in a way that does not show up in a technology budget, which is most of why it goes undone.
That does not make the method theoretical. It makes the complete version something to aim at rather than something to assume. If somebody tells you they have transformed an entire operating model in two quarters, ask which gates have names against them, and watch what happens next.
If you are starting on Monday
One exercise, and it takes an afternoon.
Take three places where AI is already touching real work in your organisation, and take them from different parts of it. A team coding with it. A workflow outside engineering that runs on it, in finance, in support, in legal. And somewhere a report or an email being drafted with it. For each one, find the moment where a person accepts what the machine produced or sends it back.
Most leaders can name that person in a minute, and the name is not the point. The point is to go and watch the moment. Somewhere in your three, a person is accepting quickly and shipping more because of it. Somewhere else, a person is spending longer checking the output than the work used to take, and saying nothing. Both look like adoption on a dashboard. They are opposite answers, and the difference between them is the whole subject of this piece: the first is work where the value clears the overhead, the second is work that should never have been handed over, or was handed over with no boundary and no owner.
That afternoon gives you your Tuesday. Run the struggling ones through the three questions, and be willing to park them with a date. Ask the excelling ones what boundary they would want written down, and write it. You will learn more from those two conversations than from your next vendor evaluation.
Then decide one thing you are not going to do this quarter, and tell people why. That is the harder half of the job, and it is the half that makes a leader worth following through a period like this one.
A note on how this was written
I used AI to write this, and I would rather say so than have you wonder.
For years I wanted to write and could not find the hours. Typing was never where my thinking happened. So I talk instead. I ramble through an idea out loud, argue with myself, go over the same point four times, and use AI to structure what comes out, tidy it and catch the grammar I get wrong. I also use it to check myself, to find whether a number I half remember actually holds and where it came from. That is how one claim got cut from this piece.
If you were hoping for something with no AI in it, this is not that, and the next one will not be either.
Notice where the time went instead. Writing this by hand would have taken a few hours across a few sittings, most of it typing and fixing commas. Those hours went somewhere better. They went into being asked whether a claim holds. Into going back through my own record to check that a case happened the way I remember it. Into cutting arguments that did not survive a single follow-up question. The brainstorming, the back and forth, the challenging of my own statements. That is the work. The machine took the typing.
Which is the same argument as the rest of the piece. Used with judgment, on the work where it earns a place, this frees you to think harder rather than produce faster. Use it that way, and give the hours back to the thinking you have been putting off. This piece is my own small example of it.
What this publication is
I am writing in public because I am still working it out, and because the people I most want to hear from are the ones going through it now, not the ones selling a finished answer.
The belief underneath it is unexciting on purpose. AI is a tool. A serious one, and one in a long line of technologies that have changed how enterprises reach their outcomes, which means it deserves neither worship nor cluelessness. It deserves judgment. So I write about where I am finding it adds value, where I am finding myself stuck, and what I would decide differently now. If you are seeing the same problems, or different ones, or solving them another way, say so. That is the kind of learning I cannot do alone, and it is most of why this is public.
It will not only be about AI. Leading teams is my trade: people whose growth I am responsible for, stakeholders who depend on what gets delivered, most of it done from India, inside the operating model of a global capability centre, which has nuances of its own that rarely get written about honestly. So the leadership journey will show up here too. How I have changed over time, what the day-to-day teaches that the long view misses, what building something of my own is teaching me now, and how to carry the uncertain stretches without drama. Not a diary, and not everything that crosses my mind. The thoughts I have spent real time on, worked into something you can use, and only when they are worth yours.
Each of the four cases above will get its own fuller anonymised write-up, with the numbers kept and the identifying details out, and the reasoning I have left out of the short versions. I would rather publish something worth your time than keep to a calendar, so there is no fixed schedule here and no dates on those. Subscribing is how you will know when one lands.
If you are somewhere between the stampede and the freeze and looking for a way through that is neither, this is written for you. If you try any of it and it does not work the way I have described, that reply is more useful to me than agreement, and it usually makes the next piece better.
Evidence notes, dated
The argument stands on the four cases rather than on these figures. Each carries its date so you can judge how stale it has become.
2026. McKinsey’s global State of AI survey: roughly nine in ten organisations report regular use of AI in at least one business function, and the share scaling it enterprise-wide rose to 44 percent from 38 percent the year before. Eighty percent of respondents say AI has improved their individual productivity, while 37 percent attribute any EBIT impact to it, a figure essentially flat year on year, and only about 6 percent report a material one. This is the gap between individual gain and organisational return, measured by somebody other than me. McKinsey, The State of AI
2025. The share of firms abandoning most of their AI projects before full deployment rose to 42 percent, from 17 percent a year earlier. Reported via CIO Dive, 2025. The year the pressure arrived is also the year the abandonment rate more than doubled.
April 2026. WRITER’s survey of 2,400 C-suite leaders and employees: 97 percent call AI beneficial while 29 percent report significant return from generative AI, and 48 percent call adoption at their own company a disappointment. 75 percent say their AI strategy is more for show than for internal guidance. That distance between enthusiasm and result is the stampede, measured. The same release leads on a separate finding, that 60 percent of companies plan to lay off employees who will not adopt AI, which is a fair sample of the fear in this market. This piece takes no position on it. BusinessWire, 7 April 2026
December 2025. Deloitte’s chief technology officer, on the firm’s 2026 Tech Trends work: companies are putting roughly 93 percent of their AI budget into technology and roughly 7 percent into the people expected to use it. Which is what skipping the decision looks like on a budget line. Fortune, 15 December 2025
Later evidence gets appended here with its date. If it contradicts the argument, this piece will say so.
About the author
I have spent close to two decades building products and technology from India for global markets: the United States, Europe, and Asia. Distance from the user raises the bar rather than lowering it. A customer in New York or London does not care where a platform was built; they hold it to the standard of anything built next door, and my teams and I were measured against exactly that. There is no pass for the distance. Understanding users and markets you do not live in is part of the discipline, and it is the harder part.
The route ran through very different rooms: enterprise systems at a digital industrial company, a global media network, a century-old publishing house, a private-equity-backed data SaaS business, and now a venture of my own in the agentic AI space, with consulting alongside. I am an engineer by training and an MBA by choice, so I argue for technology in the language of return, and question business cases in the language of engineering. Watching which lessons transfer between those rooms, and which each industry has to relearn its own way, is where most of my judgment comes from.
The technology has turned over more than once in that time: consumer platforms serving hundreds of millions of monthly users, and holding around a million of them at once on their peak nights, the streaming and data build-outs behind them, a generative AI product shipped to revenue while the category was still forming, and lately coding agents and an AI-native development standard across a hundred-plus person engineering organisation. The lesson that survived every wave is the one this publication is built on: the technology changes, and the job of deciding where it belongs does not. I write about the operating side of that job: how work is allocated between people and machines, how standards and gates get adopted rather than announced, and how organisations decide what to fund and what to refuse.Sources and disclosure
The four cases are from my own operating record at three employers between 2020 and 2026. Companies are not named. Fuller anonymised versions will follow here, each with its own disclosure boundary, and the private named record is available on request. Every internal figure comes from that organisation’s own measurement and is not independently audited, and each case carries a single figure by design. No profit and loss ownership is claimed. The tests and the blast-radius tiering are my own working frameworks. The engineering of agent loops, evaluations and verification is a separate subject and not one I am claiming here.


