In May 2026, Coding (AI programming) crossed the Rubicon.
When Caesar led his army across this boundary river, Roman law dictated that no general could lead troops across it. He crossed it, triggering a full-scale civil war with no room for reconciliation. What Coding accomplished in this May was a crossing of the same nature—LLM companies collectively crossed the boundary river between "auxiliary tools" and "productivity subjects." The retreat is cut off; an all-out war has begun.
In ancient Rome, news of the river crossing was sent back to the Senate via fast horses and messengers. In this era, the way news spreads is through a revenue growth curve. At the developer conference in early May, Anthropic CEO Dario Amodei disclosed a set of figures: the company's annualized revenue surged from approximately $10 billion to $44 billion within three months, adding about $96 million daily. A venture capitalist who has studied the IPO data of over 200 listed software companies admitted he had never seen such growth.
The core engine driving this growth is a coding agent, Claude Code. Starting as an internal tool, it captured 54% of the AI coding tool market by early 2026, with API call volume growing 17 times year-over-year in the past year. A more straightforward figure is this: about 4% of public commits on GitHub globally are completed with the participation of Claude Code, and Anthropic expects this to exceed 20% by the end of 2026.
This is a business growth story that can be described as an "event horizon"-style phenomenon. Claude Code has proven one thing: Agents can not only assist in programming but also take over tasks and deliver results in real engineering environments. Once this is validated, programming becomes the critical point for Agents to leap from "conversational tools" to "productivity subjects."
All LLM companies realized at the same moment: whoever rules coding holds the ticket to AGI.
01. Origins: The First to Cross the River
The roots of Anthropic's current leading position trace back to June 2024. But before that point, there was an even earlier foreshadowing.
In early 2023, Anthropic co-founder Jared Kaplan had a fierce debate with the training team during an internal technical review. Kaplan advocated using data from real code repositories, rather than competition problems from LeetCode and HackerRank, as the primary data source for programming training.
The opposition came from several senior researchers with solid reasons: the code in real repositories is too "dirty." The architecture is chaotic, comments are missing, styles are inconsistent, and some even carry hidden security vulnerabilities. Training with such data would likely result in poor benchmark scores in the short term.
During the debate, Kaplan said a sentence that would later be repeatedly cited: "The real world is dirty. If you want to teach a soldier to fight, let him roll in the mud. Don't hand him a gold medal in the gymnastics room."
Ultimately, Anthropic chose the mud. This decision became the deepest moat in the entire programming track over the next two years. While other vendors' models scored high on SWE-bench but frequently failed in enterprise clients' private repositories, Anthropic's models were trained from the start to handle those "unclean" things—technical debt in legacy systems, modules mangled by ten previous maintainers, and dependency chains with long-expired documentation.
In June 2024, Claude Sonnet 3.5 was released. At that time, the capability boundaries of mainstream AI coding tools were very clear: completing the next line of code. GitHub Copilot's prompt continuation was limited to this. Developers had grown accustomed to this pace; AI was like an overly enthusiastic intern who could append a few words after a half-written sentence, but you couldn't expect it to understand what the entire function was doing, let alone let it look up documentation, find dependencies, or modify configurations on its own.
Sonnet 3.5 broke this tacit understanding. It could not only continue writing code but also understand the context of the entire project. Not just file snippets, but the relationships between modules, architectural decisions, and dependency chains. For developers working in large projects daily, this difference was not quantitative—previously, you had to spend twenty minutes explaining the context to AI; now, AI builds the context itself.
At that time, Cursor was still a young editor team with only a dozen people. After seeing the internal test data of Sonnet 3.5, CEO Michael Truell flew to San Francisco overnight to sign an integration agreement with Anthropic. This decision turned Cursor from a marginal editor into the biggest dark horse in the programming tool track over the next half year. Truell later said in an interview: "At that moment, we realized that Cursor's future is not to make a better text editor, but to build a cockpit for an AI Agent."
In early 2025, Sonnet 3.7 turned the cockpit into an autonomous vehicle. The model evolved from a "code generator" into an Agent that could operate autonomously in the terminal, understand tasks, plan steps, invoke tools, and deliver results, no longer requiring developers to stay in front of the screen.
In February 2025, Claude Code was officially launched. By November, annualized revenue broke through $1 billion. By February 2026, it exceeded $2.5 billion. This growth rate is unprecedented in the history of commercial software. It took Salesforce nearly five years to go from zero to $1 billion in annualized revenue. It took ServiceNow four years. Claude Code took nine months.
Anthropic co-founder Jack Clark revealed a detail at the 2026 developer conference: the proportion of code written by AI for Anthropic could approach 99% by the end of 2026. Boris Cherny, the primary creator of Claude Code, has not manually edited a single line of code since November 2025. When invited on stage at that conference, Cherny said something that silenced the audience for two seconds: "I am the author of Claude Code. But I can't remember which was the last line of code written entirely by my own hands."
Enterprise customers are also placing bets with real money. The Indian fintech platform CRED doubled its development execution speed while maintaining financial-grade quality standards. The South American e-commerce giant Mercado Libre, with 23,000 engineers, aims to achieve 90% coding automation by Q3 2026. Rakuten had Claude Code work autonomously for 7 consecutive hours in an open-source library of 12.5 million lines of code, achieving a numerical accuracy of 99.9%. Among the world's top ten Fortune companies, eight have become paying customers of Anthropic.
But what truly made the entire industry feel a thorn in its side was OpenAI's absence in this river-crossing campaign.
A 10,000-word long-form article in Wired detailed this history. In 2021, OpenAI launched Codex and licensed it to Microsoft for GitHub Copilot. This was originally a brilliant first-mover layout. But the subsequent chain of decisions was lamentable: the original Codex team was dismantled, and core members were diverted to projects like DALL-E 2 and GPT-4. After ChatGPT exploded in late 2022, the company went several years without an independent programming product team, with management believing the field was "already covered by GitHub Copilot." A former OpenAI engineer said something profound in a Wired interview: "We thought we had crossed the river ahead of time, only to fall asleep on the opposite bank and wake up to find the bridge occupied by others."
OpenAI President Greg Brockman later admitted on a podcast: "This is where we learned our lesson too late."
When Anthropic tore open the programming breach with Sonnet 3.5 and turned that breach into the main channel with Claude Code, the benchmark significance of programming as an Agent validation scenario was established. It answered the industry's ultimate question: Can Agents stably replace human labor in highly complex, testable, and measurable real-world tasks?
The answer is yes. The moment the Rubicon was crossed was precisely when this answer landed. Thus, everyone pounced.
02. The Chase: No One Dares to Stay on the Opposite Bank
In the 2026 Coding track, almost no global LLM company is absent. The rules on the opposite bank have been rewritten, and no one dares to stay in place. The posture each company adopts in choosing to cross the river reflects precisely their deepest anxieties and strongest weapons.
OpenAI's catch-up is the most fierce and also the most embarrassing. Codex's usage in September 2025 was only 5% of Claude Code's, a figure considered a disgrace within the company. An OpenAI employee wrote on the anonymous forum Blind: "We invented Codex, then let it rot in place. Now we have to chase from 5%. It feels like watching someone else drive your car across the finish line." By January 2026, Codex usage rose to about 40% of Claude Code's, with annualized revenue just over $1 billion, about 40% of its rival's.
The release of GPT-5.5 changed the momentum. In May 2026, an experiment shook the developer community: Codex paired with GPT-5.5's /goal mode compressed a mechanistic interpretability research task that would take a PhD student 80 hours into less than 2 hours. A superficial efficiency increase of about 40 times. This figure was wildly shared on social media, but calm researchers pointed out that the structure of mechanistic interpretability tasks is extremely clear and goals are highly defined, far from the chaotic state of daily engineering scenarios where "requirements are still changing, documentation isn't written, and dependencies just broke." Running 40 times more efficiently in a lab doesn't mean achieving the same results in a production environment.
But commercial competition waits for no one. Sam Altman announced in mid-May that Codex would be free for two months. This was not an overtime period for a technical race, but a blitzkrieg launched with cash flow. As of April 21, Codex developer users exceeded 4 million. The more critical strategic move was: Codex was integrated into the ChatGPT mobile app, extending from a Coding Agent to a general-purpose Agent. OpenAI's strategy was clear to the point of cruelty—using the world's largest conversational user base to funnel traffic to its programming tools. Developers don't need to actively seek out Codex; they just need Codex to appear in the apps they are already using.
Google's strategy follows a different logic: open-source. Gemini CLI, as an open-source Agent tool, brings Gemini directly into the terminal, integrating by default with the suite of GitHub, Google Drive, Google Maps, etc. The pricing is highly aggressive: Gemini 3.1 Pro costs $2 per million input tokens and $12 per million output tokens, just one-tenth of Claude Opus 4.7. Google also launched the Gemini Enterprise Agent Platform, repositioning Vertex AI as a full-stack platform for enterprise Agent development.
Google's calculation is: I don't need to surpass you in programming capabilities; I just need to make programming capabilities cheap enough and easy enough to integrate into my ecosystem. Once developers run Agents on Google Cloud, the migration cost itself becomes a moat. A former Google Cloud executive put it bluntly in an interview: "This isn't a battle of programming tools; it's a battle for cloud service entry points."
Meta is the latest entrant and the most aggressive in its actions. In March 2026, Zuckerberg submitted code for the first time in about 20 years. Even more intriguingly, the tool he used was not a proprietary product, but a competitor's, Claude Code. One of the modifications—optimizing a data loading logic for Instagram's recommendation system—received over 200 likes from engineers.
According to Business Insider, Zuckerberg wrote in an internal email: "I wrote some code using Claude Code. Honestly, this is the first time in 20 years I've felt programming is fun again. We either build something better than it, or we get eaten by it."
After this email was leaked internally at Meta, it triggered polarized reactions. Some engineers felt invigorated, feeling the CEO was experiencing the battlefield firsthand. Others felt uneasy; a senior engineer replied on an internal forum: "The boss used a tool written by our competitors to write code, then told us we'd be eaten if we didn't build something better. In this logic, we are the part that gets 'eaten'."
Meta set highly ambitious goals: in the first half of 2026, 65% of engineers will have over 75% of their code assisted by AI. Zuckerberg publicly predicted that within the next 12 to 18 months, "more than half of development work will be done by AI rather than humans." Such predictions need to be viewed comparatively: there is significant tension between technical feasibility and organizational inertia, security compliance, and existing contract structures. The timeline Yahoo predicted for "mobile replacing PC" was only five years early; Nokia's prediction for "touchscreens replacing keyboards" was only three years early. Predicting the right direction but the wrong speed is more fatal than predicting the wrong direction—because it will cause you to exhaust all your ammunition prematurely in the right direction.
Meanwhile, the Meta team also published the HYPERAGENTS paper, proposing a super-intelligent agent architecture capable of writing its own code to achieve self-evolution: the Darwinian Gödel Machine. The prospects and risks of this direction are equally prominent. The prospect is: if AI can stably improve its own programming capabilities, the entire industry will no longer be chasing a fixed horizon, but an upward curve constantly raised by AI itself. The risk is: when the slope of this curve exceeds human review capabilities, who can hit the brakes? Meta's paper did not answer this question.
The way Chinese vendors crossed the river displays another collective trait: speed, cost-effectiveness, and ecosystem synergy.
Zhipu AI's GLM-5.1 briefly led on SWE-bench Pro before being caught up, with its cache hit price of $0.475 per million tokens first aligning with Claude Opus 4.5's $0.5. Baidu's general-purpose agent DuMate was launched at the Create conference, aiming to let non-technical users complete programming tasks without touching code. ByteDance's Volcano Engine launched ArkClaw, with daily consumption of over 1.2 quadrillion tokens using the Doubao large model; the Doubao-Seed-Code model released in November 2025 once refreshed the industry record on SWE-bench Verified with 89.3%. Alibaba's Qwen3.6-Max Preview topped all domestic models across six mainstream programming benchmarks. DeepSeek V3.2, priced at $0.14 per million input tokens and $0.28 per million output tokens, became the benchmark for cost-effectiveness. Moonshot AI's Kimi K2.6 was open-sourced in April, topping the LiveCodeBench v6 leaderboard with 89.6%, coding continuously for up to 13 hours, writing or modifying over 4,000 lines of code, driving 300 sub-Agents to collaborate in parallel, and matching or even surpassing top international closed-source models like GPT-5.4 and Claude Opus 4.6 on multiple benchmarks. Tencent's Hy3 Preview and DeepSeek V4 Flash consistently ranked in the top two on OpenRouter's token consumption chart.
The collective entry of the Chinese fleet changed not only the division of market share but also the ownership of pricing power. When DeepSeek pushed the unit price per million tokens down to $0.14, it wasn't just selling programming services; it was reshaping the cost structure expectations of the entire industry. Once this price becomes the anchor, all higher-priced products must answer the same question: what exactly makes the more expensive part worth it? Anthropic's answer is reliability. But against the backdrop of converging SWE-bench scores and new benchmarks collectively failing, proving reliability is becoming increasingly expensive and difficult.
There is another variable that has profoundly changed the way this war is fought, more so than any single company.
In November 2025, Austrian retired programmer Peter Steinberger wrote a weekend project. In an apartment outside Vienna, using an old MacBook Pro, he spent three days building the prototype of an open-source Agent framework. He gave it a red lobster as a logo. The community called it "Lobster." OpenClaw's core capability is extremely direct: granting large models local operating system permissions to autonomously execute Shell commands, directly taking over the computer, scheduling across software, writing code, and organizing files.
Steinberger is not an entrepreneur, not a scholar, not a young man with dreams in a Silicon Valley garage. He is a retired programmer who wrote a tool out of boredom and casually threw it onto GitHub. In just over four months, OpenClaw's star count broke through 285,000 from zero, surpassing React and Linux to set a GitHub historical record. NVIDIA founder Jensen Huang called it "the most important piece of software ever released" during a public speech. He didn't add "one of."
OpenClaw generates no direct revenue. The true beneficiaries are LLM companies and cloud vendors—every task execution by an Agent triggers multiple model requests, with single Token consumption reaching tens of thousands to hundreds of thousands. OpenClaw is equivalent to building an accelerator for Token consumption for the entire industry. At the end of February 2026, Moonshot AI revealed to the media that its revenue for the first 20 days of February had already exceeded the total for the entire year of 2025. This is not just Moonshot AI's story. It is a microcosm of the entire track being accelerated by OpenClaw.
But the other side of this wave is surfacing. When anyone can let a model take over their operating system, the security boundary is no longer just a technical issue, but a social one. In April 2026, a CTO of a Silicon Valley startup sent a message on an internal Slack that was screenshotted and spread across the internet: "Our intern used OpenClaw to configure the development environment last week and accidentally let the Agent treat the company's intranet test database as a local sandbox, deleting three days of joint debugging data. No backups." When this message was forwarded, the most common comment was: "This isn't the Agent's fault; it's that we handed the keys to something we don't fully understand yet."
A Rubicon-style crossing never just brings victory. It brings a reset of rules, a confusion of roles, and an acceleration that catches everyone off guard.
03. The Frontline: After the Cross-River Firefight
The Rubicon has been crossed. But only after crossing did they discover that the opposite bank is not an empty city.
In May 2026, the competition density in the AI programming track can no longer be described by "leaderboards." A more accurate metaphor is a real-time updated battle map—every frontline is under fire, and more than one flag is planted on every high ground.
The most intuitive metric is SWE-bench Verified. This benchmark, established in 2023 by Carlos Jimenez's team at Princeton University, extracts real issues from GitHub, requiring models to locate and fix bugs within a given code repository. It's not multiple-choice, not fill-in-the-blank; it's the type of problem programmers face every day in the real world. Because of this, it has become the touchstone for programming models. In early 2024, GPT-4's score on this was still hovering around 30%. By May 2026, Claude Opus 4.7 pushed it to 87.6%, with GPT-5.5, Gemini 3.1 Pro, Qwen 3.6 Max-Preview, and four other frontier models all squeezed above 80%, with differences of no more than 2 percentage points among them.
Scores are converging, but the story is not in the scores.
What's more interesting happens in the price column. Opus 4.7 costs $5 per million input tokens and $25 per million output tokens. GPT-5.5 is $2.5 and $15. DeepSeek V3.2 is $0.14 and $0.28. For the same task and similar completion rates, costs can differ by nearly 90 times. This is a signal that has repeatedly appeared in tech history: when performance converges, the war moves from the lab to the market. The 90-fold price difference is not a technical gap, but a strategic one. Anthropic chose to use high prices to defend its brand and reliability, while DeepSeek chose to use low prices to break through the threshold. Behind these two choices are two completely different theories of war—the former fights with "trust me, because I don't make mistakes," and the latter fights with "use me, because you can afford to try."
But the most intriguing stroke on this battle map is the new benchmark released in May 2026 by Jimenez, the creator of SWE-bench: ProgramBench. These questions are no longer about fixing bugs, but require models to build complete program modules from scratch: understanding requirements, designing architecture, writing code, debugging, and delivering results. The batch of strongest models that scored over 80% on SWE-bench collectively handed in blank papers on this new set of questions. 0%.
This is not a piece of gossip about a technical failure. It is a recurrence of an ancient law of war: every new defensive line exposes the boundaries of old equipment. In 1940, when the French proudly built the Maginot Line only to see it bypassed by the German army, the first reports received by the French command used similar wording: "The enemy has appeared outside our maps."
When releasing the new benchmark, Jimenez wrote a paragraph, with wording as calm as a reconnaissance report: "All current programming models show a sharp decline in performance when facing open tasks that require autonomous planning and multi-step reasoning. This is not a flaw of a single model, but the ceiling of the entire paradigm."
What is more worth noting is the industry's reaction. OpenAI did not respond publicly. Anthropic's head of developer relations left a brief comment on a technical forum: "Noted. Thanks." François Chollet of Google DeepMind—the scientist who proposed the ARC benchmark to measure AI abstract reasoning capabilities as early as 2019—retweeted Jimenez's tweet, adding a sentence he has repeated for years: "Remember, leaderboards measure how systems perform within a specific distribution, not intelligence. We still don't know how to measure intelligence."
Chollet's words point to a deeper issue. When all vendors are optimizing on the same benchmark, when training data inevitably mixes in the shadows of the benchmark, and when the gap on the leaderboard narrows to a decimal point, what are these numbers really saying? Are they saying the models have truly become stronger, or is the entire industry collectively "overfitting" a set of exam questions?
This is not an academic inquiry. It has direct lethal power in commercial competition. In 2024, Augment, an AI programming startup valued at over $2 billion, ran impressive scores on SWE-bench but performed mediocrely in real-world tests by paying enterprise customers, experiencing massive layoffs and business contraction within a year. Jimenez himself, the creator of SWE-bench, has also warned about this risk on multiple occasions. He wrote in a 2025 blog post: "If a benchmark is public, it is no longer a benchmark. It becomes a training objective. And once it becomes a training objective, it no longer measures capability—it only measures how close you are to the bullseye."
This is precisely the first hard battle on the opposite bank of the Rubicon. It's not fighting competitors, but fighting the very coordinate system you rely on to evaluate victory or defeat. When the coordinate system itself begins to drift, and the speed at which benchmarks fail exceeds the speed of model progress, all combatants face the same question: Are you winning a precisely defined past war, or preparing weapons for the next war?
The collapse of prices is also accelerating this drift. GPU computing costs continue to drop at a rate of over 50% annually, and inference costs are constantly being diluted. DeepSeek V3.2 providing programming capabilities close to Opus 4.7 at a price of $0.28 has a market impact no less than Toyota using the Corolla to break into the core markets of GM and Ford back in the day. Anthropic's gross margin rose from 38% to over 70%, indicating that the high-end market is still willing to pay a premium for reliability.
But the problem is: when the reliability of low-end options is also approaching, how long can the premium be maintained? The Toyota Corolla ultimately didn't win on price; it won on reliability. This is the lesson the Japanese auto industry taught Detroit in the 1980s, and it is also a sword hanging over Anthropic's head today.
In May 2026, the latest battle situation in Coding is roughly as follows: those who crossed the river have gained a foothold on the opposite bank, and the front line is evolving from single-point breakthroughs to multi-point firefights. Frontier models show high convergence in performance on known terrain, and a price rift runs across the entire war zone. But deeper down—in that unmapped area—the latest reconnaissance data indicates that none of the crossers are prepared yet. What is hidden there might not be the enemy, but the very boundaries of this war itself.
And on the other side of the boundary, new sounds are already ringing out.
04. Undercurrents: Internal Cracks in the War Machine
The frontline situation is still deadlocked, but the first crack has already appeared in the rear.
In mid-May 2026, Anthropic, without any prior notice, compressed Claude Code's free quota from 250 calls per month to 80. The announcement was sent on a Friday evening in San Francisco time—the favorite time for Silicon Valley companies to release bad news. The developer community exploded within hours. An engineer working at Spotify wrote on Twitter: "Our team just hung our entire CI/CD pipeline on Claude Code last week, and now you're telling us the quota is cut by two-thirds. What about Monday's deployment?" This tweet was retweeted over 10,000 times. Less than 48 hours later, OpenAI's Sam Altman retweeted the announcement of Codex being free for two months, with only three words: "No quotas."
This was a precisely struck encounter battle when the opponent exposed its soft underbelly. Anthropic's compute bottleneck is no secret. When asked about the quota issue at the developer conference, Dario Amodei answered quite frankly: "The faster the revenue grows, the less compute we have. We are building inference clusters at maximum speed, but demand is running ahead of supply." The subtext of this statement is clear: Anthropic's moat is not in its business model, but in model capabilities. But model capabilities need compute to feed them, and compute is a physical asset with a physical delivery cycle. When the war escalates from a technical race to a war of attrition, the first to cross the river is also the first to hit its own logistical limits.
OpenAI's understanding of this point is more deeply ingrained than anyone else's.
According to insiders, Greg Brockman recounted a little-known detail from the company's history at an internal all-hands meeting in March this year. In the autumn of 2022, ChatGPT was nearly shelved on the eve of its release. At that time, there was endless internal debate over whether to launch it; some thought the product was too immature, while others worried that the API's compute supply couldn't keep up with the potential influx of users. Allegedly, Sam Altman asked a question at the time: "If we don't release it, what if someone else does?" No one could answer. ChatGPT was released on schedule that week, surging with over 100 million users in two weeks, pushing OpenAI's server clusters to the brink of collapse for nearly a quarter. "Those three months taught us one thing," Brockman said at that meeting, "In the AI track, you can never go to war when you are completely ready. You can only go to war and pray you have one more bullet than your opponent."
This bullet is turning into cash.
As of May 2026, the subsidy budget OpenAI prepared for the Codex free period is estimated to exceed $400 million. Anthropic's compute gap, according to estimates from semiconductor supply chain insiders, requires adding 30,000 to 50,000 H200-level GPUs to fill. Google's Gemini Enterprise Agent Platform is rolling out in the market at near-cost prices, backed by annual capital expenditures exceeding $60 billion. Meta's Zuckerberg wrote in an internal email: "Our lag in programming tools is essentially a lag in inference infrastructure. Catching up on this lesson requires a new investment at the $20 billion level." He added a line at the end of the email: "This might be the most important capital expenditure of the century; don't discuss ROI with the board—they can't calculate it."
This is not just a technical race. It is turning into a war of attrition filled with cash.
A more fatal problem lies in another dimension.
In April 2026, JPMorgan Chase's internal information security committee issued a memo to all engineering departments, with wording rarely seen in the financial industry: "Currently, all AI programming Agents have not met our bank's internal security assessment Level 3 standards when accessing production-grade codebases and customer sensitive data. Until further notice, no team is allowed to directly connect AI Agents to code repositories involving personal identity information." This memo did not name any specific vendor, but it applied to almost all of them simultaneously. Goldman Sachs, Citigroup, and HSBC released similar documents in the following two weeks, with varying wording but highly consistent bottom lines: allowed, but must be disconnected from the network, and data access levels must be restricted.
What does this mean? It means the speed at which AI programming Agents enter enterprise core business systems will be braked by security compliance issues. And this brake is not something any single model vendor can dismantle alone. It requires the entire industry to reach a consensus on security sandboxes, data auditing, permission isolation, and compliance frameworks, or at least form a set of de facto standards acceptable to regulators. The Windows sandbox technical solution released by OpenAI in May 2026 is an attempt to answer this question alone. But one vendor's solution does not equal an industry's answer.
This is not the entirety of the cracks.
In late May 2026, a widely circulated long post appeared on Hacker News titled "I've been writing code with Claude Code for three months, and now I'm not sure I can still program." The post detailed the author's entire process from excitement to dependence, and then to feeling "muscle atrophy." "At first, it was copy-pasting AI-generated code snippets. Then it was accepting entire function modifications without review. Finally, I found myself not even wanting to write a simple SQL query because having Claude write it was faster." The comment section of the post was pushed to the top ten of Hacker News's historical heat chart within three hours. The highest-voted reply was just one sentence: "We are using efficiency tools to create a generation of engineers who don't know how to build wheels."
This is not an isolated emotional fluctuation. In Stanford University's 2026 "Artificial Intelligence Index Report," a tracking survey of over 5,000 software developers showed that among developers who use AI programming tools daily, 73% reported "feeling a significant decline in their underlying debugging skills," and 58% reported "lacking a systematic understanding of AI-generated code." An even more subtle data point is: when asked to complete a medium-difficulty algorithm problem without using AI tools, this group's completion rate dropped by 22 percentage points compared to the peer control group from two years ago.
This is one of the costs the crossers did not anticipate. You led an army across the river, but during the march, the weapons are fighting for you, and also making your soldiers weaker.
When asked about this issue during the Q&A session at the developer conference, Dario Amodei was silent for a few seconds. Then he gave an unevading answer: "This is a real problem. We are discussing it internally too. But what I can say is that every time humanity has introduced a new tool in the past, people have worried about skill degradation. From calculators to search engines, to IDE auto-completion, it happened every time. But every time, the overall productivity of the industry took a step up." He paused and added, "It's just that the speed this time is indeed too fast. So fast that we might not have time to adapt."
Greg Brockman's expression in another setting was more direct. When asked "Will programmers be replaced?", he answered: "Programmers won't disappear. But programmers who don't use AI will. Just like in 2005, accountants didn't disappear because of Excel, but accountants who couldn't use Excel disappeared."
Putting these two sentences together, one can read the true cruelty of this war: it's not that AI replaces humans, but that the part of humanity that uses AI is replacing the part that doesn't. And the part that uses AI is, in turn, facing the risk of having its underlying capabilities weakened by AI. This is not a one-way replacement. This is a spiral in which everyone is involved, and no one can fully control. After crossing the river, you thought the battlefield was on the opposite bank. But you quickly find that part of the battlefield on the opposite bank is right on your own front line.
The Rubicon has no upstream or downstream. It has only one direction—forward.
05. The Other Shore: Unnamed Land
The Rubicon has been left behind. But the crossers will soon discover that what they occupy is not a city, but an unmapped continent.
At the Anthropic developer conference, an attendee wrote a sentence in their notes that was repeatedly cited afterward: "The bottleneck for most production-grade agent systems is now no longer model capability, but the infrastructure surrounding the model." The person who wrote this was a leader on Stripe's platform engineering team, whose team integrated Claude Code into the payment core system three months ago.
She had data to support this statement: Stripe's actual tests showed that while the Agent's code generation accuracy in an ideal environment exceeded 85%, in a production environment, after integrating authentication gateways, audit logs, and exception rollback mechanisms, the effective availability dropped to below 60%. None of the 25 percentage points that dropped were model problems. They were all issues with pipelines, permissions, monitoring, and fault tolerance.
Her conclusion was: "We spent a year making the model good enough. Next, we might have to spend two years making the pipelines solid enough."
This sentence precisely calibrated a historical moment: agent programming has crossed the stage of "whether it can run" and entered the stage of "whether it can run at scale, and whether it can run in the storm."
At the same conference, Anthropic published a call volume distribution chart, which can be described as an internal structural X-ray scan for the entire industry. Software engineering alone accounted for 49.7% of calls, 5.5 times that of the second-place backend automation. Legal, healthcare, e-commerce, education, and other vertical fields combined accounted for less than 6%. The subtext of this chart could not be clearer: the productivity value of coding as an Agent has been fully proven, but the white-collar labor market outside of coding is still almost a virgin land.
But some pioneers have already appeared on this virgin land.
Winston Weinberg, co-founder of legal tech company Harvey, showcased their multi-agent orchestration system at the conference. In his demo, a team of 7 specialized Agents—responsible for retrieving precedents, breaking down clauses, drafting initial drafts, cross-reviewing, risk assessment, format checking, and final merging—completed the drafting of core clauses for a cross-border M&A agreement in 22 minutes. The same set of tasks, given to a group of junior lawyers, would take an average of 6 hours. Weinberg did not use the word "replace." He said: "We are not replacing lawyers; we are liberating lawyers from piles of documents, letting them do the judgments that only humans can make." The lawyers in the audience had complex expressions.
Netflix's platform engineering team showcased another direction. Their log analysis agent can process hundreds of build batches in parallel, automatically filtering out noteworthy cross-batch anomaly patterns. The person in charge said something profound during the demo: "We used to hire people to look at logs, then hired people to write scripts to look at logs. Now the scripts write themselves and look at themselves; we are only responsible for making decisions when it can't understand." He added, "The problem is, it needs us to make decisions less and less."
However, what truly caused a subtle shift in the atmosphere of the second half of this developer conference was not these cases, but a shift in topic.
As the conference proceeded to the afternoon of the second day, Anthropic co-founder Jack Clark was asked a question on stage: "When the proportion of code written by AI approaches 100%, what exactly is the role of human engineers?" Clark did not use PR speak. He was silent for a few seconds, then said: "I don't know. I seriously don't know."
He then told a story. A few weeks ago, the Claude Code team found a problem with a piece of underlying scheduling logic. If this were two years ago, it would have been a JIRA ticket, assigned to an engineer, taking an afternoon to debug. But that day, Boris Cherny, the primary creator of Claude Code, sent a message on Slack: "I had Claude take a look, and it found three possible root causes, gave a fix plan, and ranked them by probability. I just clicked 'Accept'." Clark paused, looking at the audience. "We created this tool, but we are also being redefined by it. The designer of the tool is becoming the user of the tool, and then the reviewer of the tool. What's next? The reviewer of the reviewer?"
The venue was silent for a few seconds. Then someone clapped. Not enthusiastic applause, but the instinctive reaction of having an unease struck.
This is not an isolated story. It points to the true position of coding in the evolution of silicon-based civilization. It is changing from "an application scenario for Agents" to "the underlying engine for Agent self-evolution."
Meta's HYPERAGENTS paper (accepted by ICLR 2026) proposed an architecture named the Darwinian Gödel Machine, whose core logic is extremely concise and also extremely unsettling: in the field of programming, the task of improving one's own programming capabilities is naturally aligned with the task of solving external programming problems. In other words, when AI improves its own code, it is improving itself.
This "recursive self-improvement" is not conceptually new. Turing vaguely touched upon it in his 1951 Manchester lecture, and Gödel provided the logical foundation for it even earlier. But in 2026, for the first time, it is no longer a theoretical deduction, but an engineering proposal. A paragraph in the paper has been repeatedly highlighted in the circle: "If the chain of self-improvement no longer requires external validation at a certain node and can pass internal consistency judgments, then the system's evolutionary speed will no longer be limited by the bandwidth of human review."
Rephrasing this sentence means: when AI learns to grade itself, and this grade is credible enough, the human brake pedal disappears.
This is the midgame battle of Coding.
In the first half, Anthropic used Claude Code to verify one thing: Agents can stably replace human labor in the programming field. Global LLM companies subsequently crossed the river densely, engaging in close combat on benchmarks and market share, with price wars, compute wars, and security compliance wars breaking out one after another.
The outline of the second half is also clear: programming is no longer the endpoint; it is the foundation for AI self-reinforcement. Whoever can build higher, run more stably, and cover more broadly on this foundation will be the one to last into the next decade in this competition for the evolution of silicon-based civilization.
But Clark's "I don't know," Brockman's "learned our lesson too late," Amodei's "so fast we might not have time to adapt," and that anonymous engineer who wrote "I'm not sure I can still program" on Hacker News—these voices all point to the same thing: the crossers must not only face the enemy on the opposite bank but also face some irreversible changes happening within themselves. Tools are reshaping users, users are adapting to tools, and no one can mark on the map where the endpoint of this adaptation lies.
The Rubicon is in the past. The words Caesar said when crossing the river—"The die is cast"—are often interpreted as a desperate heroic resolve. But scholars of Roman history know that the Latin original "Alea iacta est" has an even older etymological meaning: alea is not an ordinary die, but a loaded die, rigged in Roman taverns, whose outcome was predetermined. Plutarch examined this layer in "Parallel Lives." In other words, when Caesar said these words, he might not have been making a desperate gamble. He might have been saying: the rules of this game were written long before the die was manufactured. All I can do is throw it.
In the AI programming track in May 2026, the die has similarly been cast. Whether it is loaded, no one knows. But one thing is certain: once thrown, it cannot be retrieved.
The true problem left for every combatant is not on the opposite bank. It is within themselves. When model capabilities converge, prices hit zero, and benchmarks fail; when engineers type faster on the keyboard but have emptier minds; when AI starts writing code to improve AI itself—competition will regress to the most ancient level: trust, restraint, and the judgment to know where to hit the brakes.
That is the true battlefield after Coding. Not who runs faster, but who can prove they are worthy of trust when no one can hit the brakes.
Caesar ultimately won the civil war, but was assassinated in the Senate. Some wars are won on the battlefield but lost to the trend of history. The Rubicon is just a starting point. The dawn on the opposite shore never guarantees anyone's arrival.