Retcon Reckoning: the Great AI Replacement pt 2
"Why did you bite me?" Zarathustra asked the snake đ
In Part One, we established that AI isnât replacing developers; itâs revealing which organizations built their engineering substrate with clear, navigable boundaries and which ones duct-taped microservices together and called it a day. The 5% who succeed by measuring what matters and the 95% who fail by measuring whatâs easy.
What we didnât say, because it hadnât yet become undeniable, is that the 95% failing to generate productivity gains doesnât mean the great AI replacement isnât happening. It means the great AI replacement was never about productivity to begin with.
The Cobra and the Amazon Warehouse
Horse and Buggy Software
Developers Developers Developers
Doing Less with More
Spending Other Peopleâs Money
Everybody Has the Wrong Impression
The Cobra and the Amazon Warehouse
The fact that the production of cobras had become decoupled from the elimination of cobras was, from the perspective of the metric, largely incidental.1
The British Raj introduced a bounty on dangerous snakes around 1875, reasoning, not unreasonably, that if you paid people to bring in dead snakes there would eventually be fewer snakes, which would eventually mean fewer of the approximately nineteen thousand annual deaths that the snake population was extracting from the subcontinent with a consistency that suggested something closer to policy than accident.2 The logic was the kind that looks airtight in a planning document and considerably less airtight in contact with the people it is intended to incentivize, who are, as they have always been, rational actors operating within the incentive structure they have actually been given rather than the one the planning doc assumed.
The story you have heard, if youâve heard it, goes something like this: the locals bred cobras for the bounty, the British figured this out and cancelled the program, the breeders released their now-worthless inventory into the streets, and the city ended up with considerably more cobras than it started with, in the OG demonstration of the cobra effect, which is what happens when your solution becomes the problem, cited approvingly in economics textbooks and airport business books and at least one podcast whose website has become the primary source for a peer-reviewed academic paper, which is a level of epistemic collapse that would have interested the Raj considerably.3
The cobra effect isnât that the bounty created more jobs for cobra hunters.
Itâs that it created more cobras.
Is your AI policy breeding Cobras?
Amazon warehouse workers, it emerged recently, have been using AI tools to perform tasks they donât actually need performed, in order to inflate the usage scores their performance reviews now depend on. When the early reports broke, they focused entirely on the fulfillment center floor: the headlines claimed that Amazon warehouse workers were running fake AI tasks on their terminals to dodge performance metrics.4 It was a perfect modern fable, conjuring images of blue-collar line workers outsmarting an intrusive algorithmic boss.
But the truth of what emerged was both weirder and more indicative of corporate panic: the phenomenon wasnât happening among the conveyor belts at all. It was playing out three layers up, among Amazonâs software developers and corporate engineers, who promptly raced to be the first to call it tokenmaxxing on social media.5
The mandate, by this point, exists everywhere simultaneously and nowhere in particular. Executives announce that the organization must become AI-native to justify billions in capital infrastructure spending. Middle management translates this into utilization targets. Someone, somewhere in executive leadership, decides that if AI adoption is the metric, the mandate should be eighty percent of developers using these tools weekly, and begins tracking token consumption, literally as the raw volume of data processed by a large language model, on internal team dashboards and leaderboards.6 It was as if a taxicab service started giving bonuses for amount of gas used, but it was a decision that made complete sense in a planning doc and somewhat less sense in contact with the engineers it was designed to incentivize, who are, as previously noted, rational actors operating within the incentive structure they have actually been given and somewhat clever about automating things.
Somewhere three layers down, an employee spins up an internal agentic platform like MeshClaw, an agent built to automate email triaging or Slack interactions, and instructs it to run extraneous, low-value loops purely to burn through data.7 Because agentic workflows iterate and loop autonomously, they are spectacularly efficient at generating the precise cryptographic exhaust the company is monitoring. The interaction itself has become economically legible in a way the underlying work is not.
Which is to say: the cobra farm has entered the enterprise.
The difficulty, for leadership, is that the metrics available to measure AI transformation are necessarily proxies. Token consumption. Copilot engagement. Prompt volume. AI-assisted commits. The dashboards are indeed a sight to behold, though the actual underlying productivity gains remain, in many cases, strangely difficult to locate. Amazon officially maintains that token metrics do not directly factor into performance reviews, but when the numbers are visible to managers on a ranked leaderboard, employees recognize the implicit threat. They optimize for the metric to shield themselves from an assumption of underperformance.8
And because organizations optimize around what can be measured rather than what was intended, employees respond accordingly.
A sales organization mandates minimum weekly AI usage targets tied to quarterly reviews. Within two months, representatives begin routing customer call transcripts through summarization agents regardless of whether the summaries are ever consulted again, because unused summaries still count toward utilization. The mandate succeeds. Token consumption triples. Leadership presents the adoption curve at the next board meeting.
An engineering organization introduces âAI-assisted development KPIsâ after reading a consulting report suggesting elite teams will soon generate 40% of production code through copilots.[10] Developers rapidly discover that asking the model to scaffold boilerplate they immediately rewrite still satisfies the reporting layer, while the genuinely difficult architectural work continues happening exactly as before, only now interrupted by the requirement to periodically generate machine-produced code artifacts so the dashboard remains healthy.
A customer support department deploys internal agents intended to accelerate ticket resolution. Workers, recognizing that AI interaction frequency has quietly become a managerial signal associated with adaptability and future promotion, begin invoking the agent for tasks already well within their competence, producing longer handling times alongside dramatically improved AI engagement metrics. Leadership concludes adoption is proceeding successfully.
The employee is not resisting the mandate. (Weâll get to that later.) The employee is complying with the mandate as the organization has operationally defined it, which is always the dangerous moment in any sufficiently abstract transformation initiative. Because âuse AIâ sounds, at the executive layer, like a strategic imperative, but arrives at the operational layer as a measurement system, and measurement systems have a long and distinguished history of producing behavior orthogonal to the outcome they were originally introduced to encourage.
The particularly modern wrinkle is that AI adoption generates unusually rich exhaust. Every prompt countable. Every interaction measurable. Every token billable. Meaning organizations suddenly have the managerial narcotic theyâve always wanted: the ability to quantify something adjacent to cognition itself.
Somewhere, inevitably, there is already an engineer with a private script named goodhart.py quietly generating synthetic copilot interactions to satisfy an internal token quota imposed by someone three reporting layers above them who has never once opened an IDE themselves. Every sufficiently mature enterprise initiative eventually acquires at least one employee who understands the metric better than the people administering it. The tokenmaxxing cobra farmers across Amazon, Meta, and Microsoft are simply producing exactly what the system requested; had they worked in SaaS, they would have called the repository something like ai-adoption-helper and received a discretionary performance bonus for âdriving organizational transformation.â
The fact that the production of tokens has become decoupled from the generation of value is, from the perspective of the metric, largely incidental. And what is being measured is ofc not what is actually happening. Even if it does make for a good story.
Or at least something adjacent enough to place on a quarterly slide deck.
Retcons All The Way Down
The cobra effect, as universally understood, is itself a retcon
The cobra effect story is tidy. It has its own Wikipedia page. Itâs also historically wrong. Or rather, itâs a compression that lost the most important information in the encoding. The bounty regarding dangerous snakes started around 1875 and underwent modification around 1895. The goal was to reduce deaths from snakebite, which hovered consistently around 19,000 per year. In 1892, the Raj paid bounties on 84,789 reptiles. In 1893, they paid on 117,120. It was suspected but never proven that people were breeding snakes for profit. More to the point, the death rate barely moved.9
The death rate hovered, with the kind of statistical indifference that suggests the bounty had not so much failed to reduce the snake population as failed to meaningfully interact with it at all. In 1891, 21,389 people died from snake bites. In 1892, with 84,789 reptiles paid for under the bounty, the number was 19,025. In 1893, with 117,120 reptiles paid for, it was 21,213. The suspicion that people were farming snakes for the bounty was never proven, and the death rateâs failure to decline said less about cobra breeding than about the fundamental intractability of nineteen thousand annual deaths in a country of that size and that relationship to its landscape. The bounty was reduced, eventually (not cancelled, reduced) and the death rate did not noticeably respond to that either, because the death rate had its own agenda, which the bounty had never successfully engaged.10
The myth that the locals caused a cobra explosion is a retrospective compression; a way for administrative systems to blame the targets of an intervention for the baseline failure of the intervention itself. It turns out to be retcons all the way down. The internet desperately wanted the âAmazon warehouseâ narrative to be true because it fits a comfortable, algorithmic-resistance trope: low-wage workers sticking it to the machine by gaming the system. But that narrative serves as a smoke screen for the systemic collapse happening in the corporate tier. The warehouse floor wasnât tokenmaxxing; corporate developers were, because upper management bought into an incomplete, top-down model of âAI transformationâ and needed to justify billions in capital expenditure. When the proxy metrics failed to yield real productivity gains, the system didnât question the metrics, it just monitored the boards harder, forcing its highest-paid engineering talent to spend their days behaving like snake farmers.
What actually had worked in 1890s India wasnât the bounty, and it wasnât a crackdown on mythical cobra breeders. It was âboots on the ground,â if youâll pardon the epxression. Farmers in the fields were dying at rates that declined, measurably and specifically, once they started wearing footwear thick enough to stop a fang.
But nobody talks about the boots, just as nobody wants to talk about the actual code being written at Amazon. The underbrush removal that caused the precise category of adverse effect that good-faith interventions with incomplete models reliably produce, driving rats, and the snakes trailing them, right into the village houses. Nobody talks about that either, because the tidy version, where the subjects are devious, the metrics are absolute, and the failure can be blamed on localized fraud, the story that that incentivisation produces perverse behavior, and perverse behavior destroys the outcome the incentive was designed to produce, sounds like a better story. And better stories circulate faster than the internal Slack logs of a corporate engineering team burning cash to keep a dashboard green, circulate faster than primary sources from the Chambersâs Journal of 1895, which is not available as a podcast.11
Horse and Buggy Software
Technology does not eliminate jobs, the received wisdom goes (and what is received wisdom but a retcon), it relocates them; the received wisdom is not wrong, exactly, in the way that a map is not wrong when it omits the elevation, in describing the territory with enough accuracy to be useful and enough omission to be dangerous, depending on where you are going and how much the climb matters.12
The horse didnât survive the automobile as an industry, but the people responsible for the care and feeding of horses became saddled with the job of maintaining the machines needed for the automobile; they became mechanics, and machinists, and assembly line workers, and the vast ecosystem of human labor required to extract oil from the ground and refine it and move it through a distribution network and deliver it to the stations where the cars that needed it could find it, which over the following century produced employment at a scale that the horse economy could not have imagined, and which is the version of the story that gets told, and which despite also being a retcon, is true, as far as it goes.
What the story tends to gloss over is what happened to the mechanic when the car got computerised. The computer didnât eliminate the mechanic. It shifted the mechanicâs work from the physical diagnosis of mechanical failureâa task that rewarded accumulated craft knowledge, that got better with decades of practice, that was legible to the person doing it in ways that could be taught and refined and passed onâto the act of connecting the car to a device and reading what the device said, and then ordering the part the device specified and installing it, which is a different category of work, requiring less judgment, commanding less pay, and carrying less of what the previous version of the job was made of. Rather than disappearing, mechanic got cheaper (on the supply side that is, on the demand side the price of car servicing has increased substantially, as anyone who has received, following a four-minute consultation between a mechanic and a screen, a bill of $1,753.41 before parts and tax, well knows) and the craft that made the mechanic irreplaceable got thinner, and the knowledge that accumulated over a career got shallower, and none of this showed up as unemployment because the person was still employed, doing something still called by the same name, though his coveralls were noticeably cleaner.
What also got thinner was the mechanicâs understanding of what the device was reading. The device diagnosed. The mechanic knew how to read the device. These are not the same kind of knowing, and the difference between them becomes visible when the device is wrong; when the fault code points to a sensor and the sensor is fine and the actual problem involves something the device doesnât surface, even though the sensor could detect it, something that would have been obvious to the mechanic who could hear it and smell it and feel it through thirty years of accumulated attention to how things fail. That mechanic exists. That mechanic is even more expensive and that mechanic is not the mechanic the diagnostic device made economically rational to employ.
the condition under which âmoving up the stackâ is a description of progress rather than a description of a person standing on a ladder whose lower rungs are being removed.
But someone still had to design the diagnostic device, and the software that ran on the diagnostic device, and the chips the software ran on, and the compilers that translated the code into instructions the chips could execute, and the fabs where the chips were made, and the lithography systems inside the fabs, and the ultra-pure water systems the lithography required, and the supply chains that delivered everything to everything else, and the grid that powered the supply chains, layer beneath layer in a dependency stack that expanded horizontally as it deepened, generating employment at every level, most of it further from the physical world than the level below it, most of it more abstract, most of it more dependent on the layers beneath remaining stable and legible and available to the people working above them, which they were, for a long time, which is the condition under which âmoving up the stackâ is a description of progress rather than a description of a person standing on a ladder whose lower rungs are being removed.
Meanwhile the software got buggier. As a direct consequence of the same optimization that made the mechanic cheaper: a shift from trying to prevent failures to trying to recover from them faster, which makes complete sense when the systems have grown too complex to keep stable and which produces, as a direct and measurable consequence, more failures, more frequently, in systems that a decreasing number of people understand well enough to articulate. As Mitchell Hashimoto recently observed,13 the AI infrastructure build-out is repeating the DevOps debate about MTBF (mean time between failures) versus MTTR (mean time to recovery), and the industry appears to be making the same choice DevOps made, which was to optimize for recovery, and not reduced failures.
Defensible as a choice, but also a choice that accepts failure as a chronic condition and defines success as not staying down too long, which describes a different relationship to reliability than the one the previous generation of engineers was hired to maintain. Anthropicâs service status page offers a public record of this relationship in practice. The brownouts and degradations that have become a feature of operating at the frontier are not accidents of scale (though Anthropic seems to struggle with them much more than the other labs.) They are the output of a system optimized for recovery, running exactly as designed, at a reliability level the optimization target was designed to accept. The software is buggy on purpose.
Developers Developers Developers Developers
Software has gotten easier to write, but harder to understand.
There can be no doubt that the office chair has become more comfortable.14 Still eons behind gamer chairs, tho weâre starting to see more of those at startups. And there can be little doubt that during the same period software has gotten easier to write. What was once a an exercise in muscle memory and meticulous attention to syntactic details had come to be turbo autocomplete where tab becomes the most used key on your keyboard, which from a muscle memory point of view is an unambiguous victory, in the same way that not having to remember phone numbers is an unambiguous victory, and which carries, in both cases, the same quiet implication that you would be in very much over your head if the situation ever arose where you needed to remember one.
Regarding what is actually happening no one as yet has a coherent thesis supported by evidence, so everyone constructs one in real time, as the epistemically appropriate response to a situation in which the signal is genuinely contradictory and the stakes are high enough that silence has acquired the specific social weight of holding an antimemetic opinion, an admission the modern online economy has yet to find a way to prosecute but is working on it. So the thought pieces proliferate. The frameworks multiply. The conference panels convene. And over all of it hangs the industryâs dominant retcon: developers werenât being replaced, theyâre being given an opportunity to move up the stack. Being given a better chair, even.
The argument has genuine historical backing. From assembly to C, from C to Python, from on-prem to cloud, from manual memory management to garbage collection, every major transition involved movement to a higher level of abstraction, and in each case the practitioners who adapted found themselves working at a level that was more powerful, if less granular. The stack got taller. (bigger, and more angular15) The abstraction was genuinely upward: you retained what you knew and gained a higher vantage. The retcon frames the current moment as the latest iteration of this pattern, which would be a reasonable frame if the pattern were actually repeating, which it is not, certainly not in the way that matters.
The distinction the retcon requires you not to make involves the difference between a higher level of abstraction and a higher level of ambiguity. When a C++ developer moved to Python, they did not report brain fog. (They did report other kinds of trauma, but thatâs a story for another day.) When a sysadmin moved to AWS, they did not feel their understanding of networking precariously dissolve. Previous transitions preserved the knowledge of the layer below, automating interaction with that layer without making the layer inaccessible, which is why the developer who moved from C to Python could still reason about memory management when something went wrong at the boundary. The abstraction was upward. What gets reported now differs in kind: engineers losing the ability to hold a codebase in their heads, losing the capacity to reason about the system they are nominally responsible for, losing the thread of what the code is actually doing underneath the layer the agent is operating on.
Brain fog, like Gibsonâs famous future, is not evenly distributed. The engineers who report the most acute version of it are not, by and large, the ones who understood the stack most deeply. They are the ones for whom the agentic tools arrived before the underlying knowledge had fully formed, for whom coding had been something the tooling managed, the syntax something the autocomplete supplied, the architecture something the framework handled, and who were, for a remarkable stretch of time, compensated extraordinarily well for their fluency with the interfaces rather than their understanding of what the interfaces were built on. An entire generation for whom software development was, at some meaningful level, a video game16 that paid exceptionally well and which the agentic tools have now revealed to have been, in part, a video game being played on someone elseâs hardware.
Substitution and augmentation are not the same operation, and the difference
shows up in what the person retains when the tool is taken away.
The retcon interprets this as the temporary disorientation of transition, the expected friction of moving to a new level of abstraction, something that will resolve once the new paradigm is fully internalized. The more accurate interpretation is that it is the perception of a capability being abandoned rather than relocated; not moved up the stack but left behind in it, inaccessible not because it has been abstracted but because the tools that were supposed to augment it are instead substituting for it, and substitution and augmentation are not the same operation, and the difference between them only becomes visible at the moment when the layer below needs to be reasoned about and the capacity to do so is no longer there.
The retcon reframes this revelation as elevation. The skill deficit that the tools have exposed becomes the proof that the tools are needed, which is true as far as it goes, and which goes considerably less far than the retcon requires it to. What the tools are actually doing to the population using them has been described with some precision by the researchers studying it: the use of coding agents is actively diminishing the very skills needed to effectively manage the coding agents, which is the cobra effect stated at the level of individual cognition, running in real time, visible in the data, and re-described as progress by an industry that has staked too much on the productivity thesis to afford the alternative interpretation.
Without it, the industry would have to confront what the transition is actually producing, which is not a workforce moving up the stack but a workforce discovering that the stack they thought they were on was shallower than the salary suggested, and that the AI has not so much replaced their skills as revealed the gap between the skills they had and the ones the salary was pricing.
No One Expects The Comfy Chair
The chair got more comfortable.17 This is what happened between 2010 and 2024 for a significant portion of the people now asking why they are being laid off. The IDE got smarter. The framework handled the architecture. The autocomplete supplied the syntax. The linter caught the errors. The CI/CD pipeline managed the deployment. Each abstraction layer added comfort to the chair, removed one more reason to understand the layer below, and was experienced as progress because the output kept coming and the salary kept arriving and the title kept accumulating and the understanding kept thinning and nobody was measuring the thinning because the output was the metric and the output was fine and the chair even had a tilt lever.
PEBKAC (Problem Exists Between Keyboard And Chair) was the help deskâs sardonic shorthand for user error, the human in the loop as the weakest link, the thing that would be fine if only the person would get out of the way. Agentic tools didnât change the acronym. They changed which side of the keyboard the problem was on. The Problem Exists Between Keyboard And Chair, and the chair is a big part of the problem, and the chair is the thing being removed, and the person in the chair spent a decade being told they were moving up the stack while the stack was quietly being built around them in a way that made the chair unnecessary. As comfortable as it was.
Doing Less With More
The gap between what the invoice said and what the board slide promised was not, in the end, a measurement problem. It was the measurement itself.
The corporate math was presented, in every instance, as self-evident. Headcount is a cost. AI reduces the need for headcount. Reducing headcount reduces cost. The freed cost funds AI investment. The AI investment further reduces headcount. The model is recursive. The flywheel spins. The slide is clean. The coffee is ok.
Gartner surveyed 350 large companies actively deploying AI, by any reasonable standard the organizations most positioned to be getting this right, and found that 80 percent had reduced their workforce as a âdirect resultâ of their AI investments. The finding that followed was the one nobody had put in the deck: the layoffs had zero statistical correlation with improved ROI. Companies that cut the most staff showed returns nearly identical to companies that cut the least.
One mechanism the correlation data cannot explain is knowledge debt. Not the kind that appears on a balance sheet, bc it doesnât appear anywhere a dashboard monitors or a quarterly review surfaces but is the accumulated understanding of a specific system, built by specific people over specific years of decisions that were made for reasons that made sense at the time and were never written down because the person who made them was always available to explain them, until they werenât. The lag between cutting the humans and discovering what the humans were doing is long enough that the decision looks rational at execution. The first quarterly review looks fine. The second looks fine. The third surfaces something that cannot be explained without reference to a person who left eight months ago, and by then the documentation they were supposed to produce before leaving has been discovered to be incomplete in a particularly crucial way, bc the knowledge that matters most is hardest to articulate, which is why it was never articulated, and why it left with them. Which is why we aim organisationally to eliminate irreplaceability.
What the Gartner data also cannot capture is the response from the other side of the transaction. One study found that 29 percent of employees admit to actively sabotaging their companyâs AI strategy, with the number climbing to 44 percent among GenZ, whose resistance is deliberate and tactical: feeding corrupted data into models, routing around corporate control planes with unauthorized tools, intentionally generating low-quality output to demonstrate system fragility while 60 percent of executives plan to lay off employees who wonât adopt AI, 77 percent will bar non-adopters from promotion, and 92 percent openly admit they are cultivating an AI elite. Leadership has spent two years insisting this is a productivity initiative. The workforce isnât so sure, and neither is the data.18
Moreover, the data does not and cannot capture what is happening to the work itself. The model does not argue when you accept its suggestion for another layer of indirection. It does not push back when the product requirement is vague and the prompt is lazier still. The model is the world's most expensive yes-man, and organizations have discovered they like being told yes. The result is not acceleration toward excellence but acceleration toward sufficient because sufficient ships, sufficient meets the KPI, and sufficient can be presented to the board as evidence that the AI investment is bearing fruit, even as the long-term carrying cost of the sufficient codebase (quietly) metastasizes in production.
First The Good News
Q1 2026 earnings calls established is that something is happening.
What they did not establish is what.
Zuckerberg described engineers building in a week what once required dozens and months. IBM disclosed 45 percent productivity gains across its developer workforce and $4.5 billion in internal cost savings since 2023. These are quantitative claims in the technical sense, containing numbers, stated with executive confidence on a live call to institutional investors. They are not quantitative claims in the empirical sense, because none were independently verified, and because the people making them had a structurally motivated interest in demonstrating that the AI investment was generating returns.
But Meta did post the most lucrative opening quarter in its corporate history ($56.31 billion in revenue, $26.8 billion in net income) and three weeks later announced 8,000 layoffs. At the internal town hall, Zuckerberg was direct: getting everyone to use AI tools and doing the work more efficiently is not the thing thatâs driving layoffs. The CEO who had just spent an earnings call crediting AI-enabled productivity gains, now telling his own employees that the productivity gains were not why anyone is losing their job. Cisco ran the same play this quarter. Cloudflare followed suit with an apologetic memo described as âthe first true AI-layoff manifesto.â
So, AI is enabling companies to do more with less (according to the earnings calls) but also AI is not the reason for all the pink slips (according to the layoff announcements.) And then thereâs the CEPR survey of 5,000 firms in the same period found that around 80 percent reported no AI-driven productivity gains at all.19 The Atlanta Fed found measurable gains concentrated in high-skilled tasks, with perceived gains exceeding measured revenue gains, a gap the researchers called the productivity paradox, which might less diplomatically be described as the distance between what gets said on earnings calls and what shows up in the books. McKinsey found only 39 percent of companies reporting any current EBIT impact from AI. But 39 percent is not nothing.
The actual data resolves into a measurement fog dense enough that the same quarter can produce IBMâs $4.5 billion in savings and a survey of 5,000 firms reporting nothing measureable. Both can be accurate while neither indicate what is actually happening, because the proxies measure things adjacent to the outcome rather than the outcome itself ; the metric has decoupled from what it was introduced to track, and the decoupling is invisible from inside the metric. When Uber rolled out Claude Code to its 5,000-person engineering organization in December 2025 and by March had 84 percent of engineers using it (with 70 percent of all committed code originating from AI, the highest publicly reported ratio at any major technology company) it announced it had AI opening 11% of new pull requests, which reveals precisely what percentage of committed code came from AI and nothing whatsoever about whether that code was worth what it cost to produce.
CTOs built internal dashboards ranking engineers by AI consumption, which functionally made them cost-acceleration mechanisms. Like Amazon, who in the same period it was announcing 16,000 layoffs, launched a token consumption rankings page, Meta called theirs âClaudeonomicsâ and handed out badges like âToken Legendâ and âSession Immortalâ â the cobra farm stated as corporate policy, formalized into vocabulary, and gamified into a leaderboard. Uber also built internal dashboards ranking engineers by AI consumption and found their top engineers were burning between $500 and $2,000 each per month. Uberâs full-year AI budget was exhausted by April.20 Microsoft, after a quarter of telling investors that Lloyds Bank was saving each employee 46 minutes per day, quietly canceled most of its internal Claude Code licenses in May 2026, effective June 30, directing its own engineers to its own cheaper tool. And Bryan Catanzaro, Nvidiaâs own VP of Applied Deep Learning, told Axios that for his team, the cost of compute had already exceeded the cost of employees. Taken together, these findings that explain why the earnings call slides require such careful construction.
The retcon goes: layoffs are transformation, headcount reductions are investments, undemonstrated gains are early signals, companies without ROI are simply early on the curve, and the curve bends upward: number go up.
Because large language models remain structurally prone to hallucination and architectural regression, cutting the human layer doesnât eliminate the labor cost, it forces the enterprise to pay twice and both bills are on the OpEx side. Once for the token bills of the autonomous agents, and again for the specialized engineers whose job is to untangle what the agents produced. Salesforce will spend nearly $300 million on Anthropic tokens this fiscal year, against a global engineering payroll of roughly $5 billion for 15,000 engineers. That the token bill is no longer a rounding error in an R&D budget but a massive, recurring, variable operating cost that scales with every line of automated output does not make the financial decision any less rational or the amount any less commensurate, and it also does not replace the humans required to verify that output. It joins them on the invoice, at a rate that makes the original headcount reduction look, in retrospect, like a discount. Which it kind of was.
The Gartner data on who succeeds points elsewhere. The companies doing best are not the ones that replaced humans with AI or maxed out their token spends. They invested in people alongside it, building systems where humans supervise and extend what AI produces (which is not to imply Benioff isnât doing that, I donât have that insider info.) Itâs a narrative thatâs easy to nod along to and not so easy to put into practice. And even when it is put into practice, the bill has a way of coming due anyway, as weâre seeing with current obsession over the token cost reduction.
The âTIPâ (Token Improvement Plan)
It starts as a dark joke inside engineering Slack channels, usually right after an unhandled recursive test loop burns through five figures in API calls over a long weekend: âManagement is putting you on a TIPâa Token Improvement Plan.â
But as the inverse economics of the frontier solidify, the reporting layer is already building the infrastructure to make it real.
The conversation happens behind a closed glass door, guided by a Director of Engineering staring at a Datadog billing visualization:
âYour velocity is fine, but your token consumption exceeds twice your base salary. Youâre appending a 100k-token repository to every prompt just to fix minor layout bugs. We need you to drop your marginal token spend by 45% over the next thirty days, or weâre going to have to route your IDE access through a smaller, open-source model running locally on an older Mac Studio.â
This is the ultimate loop of the âDoing Less with Moreâ doctrine. The enterprise pink slips 15% of the human staff to fund the autonomous transition, only to discover that the remaining humans consume tokens like a runaway process, like a horse fitted with a jetpack (which is obviously cool.)
Doing more with less was the promise. Doing less with more is the condition: less verifiable return, more spend; less human understanding of the systems humans are nominally managing, more dashboards measuring the proxies that replaced that understanding; less signal, more noise; less stability, more recovery. Goldman Sachs forecasts a 24-fold increase in enterprise token consumption by 2030. (Let that sink in.) Gartner projects that even as individual token prices fall 90 percent, total enterprise AI costs will increase, because agents consume tokens at rates that make price-per-token beside the point. (Opus 4.7 consumes 4Ă as many tokens for the exact same prompts.) The math does not improve at scale. It compounds. The invoice arrives later and larger, and the people who receive it are not, in many cases, the people who signed the contract.
The story currently circulating is as clean as retcons get: AI is working, the adoption curve is healthy, the returns are coming, the companies that cut too deep are simply early, and the capital is flowing in the right direction. What capital is actually building is a different question. And unlike the productivity claims, unlike the 46 minutes saved at Lloyds and the 45 percent gains at IBM and the budgets blown away by April, and the dashboards and the tokenmaxing and the shadow IT cobras running under everyoneâs desk; the answer is not a matter of measurement lag or methodology.
Spending Other Peopleâs Money
Signals that resist narrative have a different epistemic status than the ones that welcome it.
In a moment defined by universal narrative constructions, retcons layered on retcons, everyone reaching for a story that organizes the contradictions into something livable, there is one signal that is not a story. It does not care about the productivity paradox or the measurement fog or the 80 percent of firms reporting nothing or the 39 percent reporting something or the earnings calls or the employee sabotage or the token bills or the knowledge debt or any of the other genuinely unresolvable questions that we have left exactly as it found them: unresolved. Capital moves. And the movements are physical, visible, and denominated in a currency that does not so much fluctuate with the quarterly narrative, as underpin it.
PJM, the grid operator responsible for electricity across thirteen states and the District of Columbia, serving approximately 65 million Americans, projects a six-gigawatt shortfall by 2027, which is not a forecast about AI adoption rates or enterprise pilot outcomes or whether the productivity studies will eventually vindicate the investment thesis. It is a statement about watts, which are physical, and about the gap between the watts that will be demanded and the watts that will be available, a gap that doesnât listen to earnings calls and cannot be resolved by a better narrative about where the industry is going.
Buy The Numbers
Five technology companies now spend more on capital expenditure than the entire global oil and gas industry spends on upstream production.21
The International Energy Agency published this finding in April 2026, effectively ending arguments about whether the AI build-out is real, whether the productivity claims justify the investment, whether the companies are getting ahead of themselves in a way that will eventually require correction or any of the other tangents thrown out as âexplanations.â The answer to all of those questions is that theyâre the wrong questions.
The right question is: what does it mean that the largest reallocation of productive capital in the history of industrial civilization is currently underway, and the people doing it are not particularly sure itâs working?
Google, Amazon, Microsoft, Meta, and Oracle have committed between $660 billion and $690 billion in capital expenditure for 2026 (a 77 percent increase over 2025âs record of $410 billion, which was itself a 75 percent increase over 2024.) đ Amazon alone projects $200 billion. Alphabet projects $175 to $185 billion, which will reduce its free cash flow by approximately 90 percent, from $73 billion in 2025 to an estimated $8 billion, a number that would represent, for any other company in any other sector, a distress signal. For Alphabet, it represents the cost of not falling behind.22 Microsoftâs calendar-year 2026 capex came in at $190 billion, $38 billion above analyst estimates, with $25 billion of the overage attributable solely to rising memory and chip costs. Alphabetâs cloud contract backlog reached $460 billion in Q1 2026, roughly double the $240 billion reported three months earlier.
The corporate balance sheets are no longer sufficient to fund the build-out at the required pace. AI-related debt has reached $1.4 trillion, making it the largest single segment within US investment-grade credit markets. The companies are issuing bonds. The bonds are being bought. The debt markets have reached their conclusion, with the precision debt markets apply to such conclusions, that the infrastructure being built is worth financing at scale, regardless of whether the productivity paradox ever resolves.
Valve runs the PC gaming infrastructure of the civilized world with approximately 350 people and an estimated $17 billion in annual revenue (roughly $50 million per employee) outpacing Google, Amazon, and Microsoft by orders of magnitude on the only ratio that matters in a leverage economy. By way of contrast, GitLab employs 2,500 people to generate $600 to $700 million in annual revenue. Which doesnât mean GitLab is a poorly run company; itâs a *normally* run company. The delta between Valve and GitLab, and between both of them and the organizations currently deploying thousands of engineers to produce 70 percent AI-generated code at token costs that exhaust annual budgets by April, is not the tooling, which is increasingly the same. The delta is whether individual human judgment interacts directly with leverage, or is separated from it by the layers of translation, from mandate to KPI, KPI to dashboard, dashboard to utilization metric, utilization metric to quarterly slide, converting the decision to do a thing into the administrative management of it being done, which is not the same thing. Capital already has proof of concept for what high-leverage human-AI substrate produces. The question running through all this, âdoes the investment generate returns?â has a demonstrable answer at the substrate level. The answer is mos def (yes) but under conditions that most large organizations have specifically engineered themselves out of being able to achieve.
What capital is now building is not these conditions at scale.
Itâs building the layer these conditions run on.
The bottleneck to this is not the chips. Itâs the power to run them. Forty percent of announced AI data center projects are currently delayed by power infrastructure constraints, not chip supply, not regulatory approval, not financing. (Electrons, the problem is electrons) The IEA projects that data center electricity consumption will double from 485 terawatt hours in 2025 to 950 terawatt hours by 2030, with AI-focused facilities growing three times faster than that. At current trajectory, global data centers would constitute the fifth largest energy-consuming entity on earth, ranked between Japan and Russia. Virginia, where data center density is highest (and AWS the most unreliable) already routes 26 percent of state electricity to computing infrastructure.
Capitalâs response to this constraint is visible from considerably further up the road. Microsoft has signed a 20-year power purchase agreement with Constellation Energy to restart the shuttered Unit 1 reactor at Three Mile Island which is the first time a retired nuclear reactor in the United States has been brought back to life for a single corporate client. The plant will generate 835 megawatts, equivalent to the annual consumption of 800,000 households, 100 percent of which will go to Microsoftâs data centers in Pennsylvania, Illinois, Virginia, and Ohio. The agreement runs to 2044. Amazon has contracted 1.92 gigawatts from the Susquehanna nuclear plant through 2042 and committed $500 million to small modular reactor development. Meta has announced a 6.6-gigawatt nuclear procurement strategy for its Prometheus AI data center project. Google has contracted a fleet of small modular reactors through the tellingly named Kairos Power. In aggregate, the technology sector signed contracts for more than 10 gigawatts of new US nuclear capacity in the past twelve months, enough to reverse the commercial trajectory of an industry that had been in managed decline for forty years.
The flagship Stargate facility in Abilene, Texas (known internally as Project Ludicrous) received its first Nvidia GB200 server racks last summer. The Stargate project overall, joint venture of OpenAI, SoftBank, and Oracle, has committed $500 billion to 10 gigawatts of AI infrastructure across Texas, New Mexico, Ohio, and expanding to Abu Dhabi, Argentina, Norway, and the UK. OpenAIâs CFO Sarah Friar noted, standing in front of buildings still under construction, that the shovels going into the ground were laying foundations for compute that wonât come online until 2026. âNo one in the history of man,â she said, âbuilt data centers this fastâ inviting more than a few questions weâll have to push to the parking lot.
But despite the historical framing and the unanswered questions of just what kind of data centers have been previously built in the long history of man, none of this is a narrative. It is concrete. It is steel. It is electrons. It is cooling infrastructure and fiber runs and grid interconnection agreements and reactor restart schedules. It is the physical expression of a bet that is no longer subject to the measurement fog of narratives, because it has already been placed, in a currency that does not fluctuate with the quarterly slide deck and is not obligated to the debts so-incurred.
Capital isnât building data centers on the assumption that the tech debt wonât come due. Itâs building them on the assumption that when it does, the debt will be denominated in someone elseâs currency. The engineers who lost the craft. The companies that cut too deep. The governments managing the dislocation. The analysts caught up in the productivity paradox. The next generation inheriting the codebase.
As Douglas Adams observed, this is how a spaceship lands on a cricket pitch without being seen.23
Everyone Has the Wrong Impression
itâs not a claim about whether the transition is smooth. Itâs a claim about who captures the value from the transition regardless of whether it is smooth or not.
Even when itâs not somebody elseâs problem, people have a way of missing the point. Or just getting it completely wrong, like hundreds of replies making fun of an allegedly fake Monet that turned out to be a real Monet (aka âthe strange thing that happened last weekâ24) with an AI being the only participant in the exchange to identify it correctly,25 the one without social incentive to perform certainty in the direction the room was moving. Which is not a story about AI being smarter than art critics, but a story about what happens when the performance of expertise becomes more socially rewarded than expertise itself, which is a condition the knowledge economy spent several decades expertly perfecting, and which the current moment is pricing with the blunt instrument that markets use when they have decided something is overvalued. Or pranksters use to expose folly.
@SHLOMSâ art troll stunt was not only brilliant, but single-handedly put what felt like wind back in the sails of the becalmed NFT market: Wonât Get Fooled
The narrative class, in the thought pieces and the conference panels, the framework decks and the podcast citations of podcast citations that became the primary sources of peer-reviewed volumes, has been retconning in real time because that is literally the job, and the job description was written in a world where the narrative was âload-bearingâ (cough), where the story about what the technology meant was part of what the technology actually meant, and where the people who could explain the system were embedded inside the system in a way that gave them a legitimate claim on its returns. That world has not ended, nor is it going away. It is, however, repricing, and the repricing is happening at the layer of capital allocation, which is the layer that does the actual estimation, and which has expressed its current view in watts and square footage and chip delivery schedules that do not require a thought piece to execute nor do they require additional opinions even if you did think you saw, if only in the briefest flash of a spaceship landing in your peripheral vision.
The labor-to-capital thesisâstated plainly in Teamwork Makes the (AI) Dream Work, before the reckoning arrived to retcon itâsurvives contact with the complication, and survives for the same reason the physical infrastructure survives contact with the productivity studies: it is not a claim about whether the transition is smooth. It is a claim about who captures the value from the transition regardless of whether it is smooth, and the evidence that the transition is not smooth does not affect that claim, any more than a delayed train affects the question of who owns the railroad.
The shift does not require the AI to be good.
It requires capital to keep moving in the direction it is moving,
What the reckoning reveals, if you read it carefully rather than for comfort, is not that the thesis was wrong but that the mechanism is more interesting than the summary. The productivity gains are not materializing at the level of the worker, which was always the wrong level to look for them, because the productivity gains were never the point; the labor-to-capital shift was the point, and you can accomplish the shift without the gains, through the replacement cost math that didnât get run before the cuts, through the knowledge debt that is now compounding below the level of any dashboard, through the brain fog that is revealing a skill deficit the salary was confidently mispricing, through the cobra farms that are inflating the usage scores on which the performance reviews now depend. The shift does not require the AI to be good. It requires the capital to keep moving in the direction it is moving, which it will, because the $740 billion and the 6 gigawatt shortfall and the $80 billion in GPUs waiting for power are not contingent on the Gartner numbers improving. They are prior to the Gartner numbers. They are the physical fact that the Gartner numbers are a lagging indicator of.
The boots worked, in the Raj. The one intervention that actually reduced deaths was the unglamorous, non-narrativised, completely obvious one that required no theory of perverse incentives and no understanding of behavioral economics and no podcast to explain it, merely the observation that farmers were dying in fields and the field contained snakes that accessed the farmer through the foot and a thick boot prevented that, which is more like a fact than a story, which is why nobody tells it. The equivalent intervention in the current moment is equally unglamorous and equally hard to package for a conference panel: measure what actually matters, rebuild what was cut, pay for the knowledge before it walks out, find the people who understand the layer below and keep them involved with the layer above. This is not a framework. It doesnât have a trendy name. It is just what actually works, available as always, sitting there, waiting for someone to prioritize the evidence over the story.
Make no mistake though, the story is quite good. The story about the cobra effect is vivid and memorable and useful as a shorthand for a real phenomenon, and it circulates because it is a good story, and it is a good story because it is clean, and omits details like the boots, and the underbrush, and the 19,000 deaths that didnât move, and the primary sources from 1895 that nobody consulted because there was no podcast. The story about AI replacing workers is also vivid and memorable, and it circulates for the same reasons, and it is clean in the same way, and what it omits is structurally similar: the mechanisms that actually matter, the interventions that actually work, the layer below the narrative where the physical facts are doing what physical facts do, which is persist regardless of whether the story accounts for them.
You knew these facts when you accepted the terms. Not AI, specifically, but the deal. The deal the knowledge economy offered everyone at every layer of the stack, from c-suite all the way down the forward deployed engineer26 â your expertise is valuable, your judgment is irreplaceable, the layer you work at is the layer that matters, and the compensation reflects all of this indefinitely. But terms, like all facts, were always finite. The compensation was always a market price for a capability that markets re-price when they find a cheaper source, which they always eventually do, which is not a betrayal of the deal; it *is* the deal, read more carefully than when it was accepted, by someone who wasnât even in a hurry, per se (unless they were an all-in accelerationist) and who, if we are being precise about what was always true, knew exactly what they were the whole time even if they didnât admit to it. Call it corporate mĂŠconnaissance.
The underbrush has been cleared and the snakes are inside, where the music is still going while the chairs are disappearing and the failures keep coming and keep getting recovered from and the data centers are still being built and the boots are still doing what they have always done, what they were made to do, which is the only thing in this entire story that isnât a retcon.27
The boots were made for walking over the snakes.28
Notes
This forms the structural definition of Goodhartâs Law, famously framed by Marilyn Strathern as: âWhen a measure becomes a target, it ceases to be a good measure.â See Strathern, M. (1997). âImproving ratingsâ: audit in the British University system.â European Review, 5(3), 305-321.
The British administrationâs introduction of rewards for killing venomous reptiles in colonial India, alongside detailed mortality tracking, is thoroughly documented in administrative health records. See Fayrer, Joseph. (1872). The Thanatophidia of India: Being a Description of the Venomous Snakes of the Indian Peninsula, with an Account of the Influence of Their Poison on Life. London: J. & A. Churchill.
The anecdotal trajectory of the âreleased cobra population explosionâ represents a classic case of an unverified narrative circulating into institutional economics textbook lore without robust primary sourcing. The foundational modern text popularizing the specific term âCobra Effectâ is Siebert, Horst. (2001). Der Kobra-Effekt: Wie man Fehlsteuerungen in der Wirtschaft vermeidet. Munich: Deutsche Verlags-Anstalt.
Initial viral social media chatter and early aggregated technology newsletters incorrectly localized âtokenmaxxingâ behavior to hourly log-in routines and scanning terminals used by Amazon fulfillment center floor associates. (Ca. new mediaâs early bias problem).
The phrase tokenmaxxing emerged organically across software engineering communities (such as Hacker News and Redditâs r/webdev and r/LLMeng) to describe the calculated inflation of token usage. See âAmazonâs âTokenmaxxingâ Problem,â Financial Times (May 12, 2026).
Amazonâs 80% Developer Adoption Mandate: Executive directives implemented internal team leaderboards tracking LLM utilization data, setting explicit targets requiring 80% or more of engineering personnel to trigger active AI features weekly. FT, ibid.
MeshClaw is an internal Amazon agentic automation infrastructure platform capable of handling environment deployment, email triage, and Slack interactions. Developers configured autonomous loops running trivial scripts to trigger massive, compounding programmatic iterations. Cf. Tomâs Hardware, âBig Tech Has a Tokenmaxxing Habitâ (May 12, 2026).
While Amazon corporate officially messaged that token counts would not act as strict input factors for standard performance loops, workers documented intense peer pressure and a defensive strategy of metric-bloating to avoid appearing as âlow adoptionâ laggards on visible leaderboards. TechRadar Pro (May 13, 2026).
Primary statistical logs show that despite vast payments for destroyed reptiles, regional snakebite fatalities stayed locked in historical ranges. See âReptile Destruction Records,â Chambersâs Journal of Popular Literature, Science, and Art (Edinburgh: W. & R. Chambers, 1895).
Official Raj mortality statistics confirm 21,389 deaths in 1891, fluctuating to 19,025 in 1892, and returning to 21,213 in 1893âcompletely independent of the 117,120 bounties distributed that exact year. Chambersâs Journal, ibid.
A concluding observation on media ecosystems; clean, moralistic narratives of localized exploitation consistently outpace dull systems-level analysis, both in Victorian economic journals and modern digital distribution channels.
The city of Brussels is a wonderful example of this, where navigating your way around the city with Google maps can easily lead you to a spot where you can see your destination right in front of you (below you!) but with no easy way of getting there.
https://hachyderm.io/@mitchellh/116580433508108130
Steve Ballmer, who as Microsoft's CEO famously bellowed "developers" fourteen times at a company event until his shirt was soaked through, was also known to yeet chairs in reaction to engineer departures; less musical chairs and more somebody call HR
As Terry Gilliam brilliantly depicts in Zero Theorem.
I detest having to footnote this, although not quite as much as I detest discovering that a reference once ambient enough to function as atmospheric background radiation now apparently requires attribution; âno one expects the comfy chairâ is from the âSpanish Inquisitionâ sketch in Monty Python's Flying Circus..
WRITER and Workplace Intelligence, "AI Adoption in the Enterprise" April 7, 2026
The same survey found those same firms forecast AI will boost productivity by 1.4 percent over the next three years; literally the productivity paradox in two data points.
âThe budget I thought I would need,â Uber CTO Praveen Neppalli Naga confirmed directly to The Information âis blown away already.â
For a full breakdown of these shifting capital flows, see âBig Tech AI Spending Tops $400B, Now Exceeds Oil And Gas Investment,â Yellow.com (April 2026), https://yellow.com/news/big-tech-ai-spending-tops-oil-gas. But also, in this age of illiteracy Iâm scared people will think âBuy The Numbersâ is the actual idiom.
A cost that, increasingly, seems to weigh on everything.
To give away the joke for those have and have not read the book, those are all somebody elseâs problem. Which is the exact type of field spaceships generate to land unnoticed on Earth, in those final days before its destruction by the Vogons to make way for an Intergalactic Highway.
I donât have the link handy but Claude was the only one (besides @quackimaduc) to not be fooled, and immediately rejected the idea it was fake, with descriptions of the brushwork as supporting evidence. I saw the exchange, which was remarkable, but donât have the link.
a16zâs coinage of forward-deployed engineer as âan engineer who lives inside the customer's problem rather than behind a product roadmapâ is patently a solutions engineer by any other name, but names have expiration dates now.
Not to be overly reductive about it, but the boots are intervention, in the precise sense of being the opposite of narrativization.






