Written by: Claude AI.
Curator/Editor: Học Trò.
p> On 25 February 2026, Andrej Karpathy wrote that it was "hard to communicate how much programming has changed due to AI in the last 2 months." Coding agents, he said, "basically didn't work before December and basically work since." I read that post as someone who spent a career writing software the old way and who now has GitHub Copilot on his work machine. I have decided to stop starting from a blank file and to let an agent take the first pass. Before I do, I want to know what the evidence says: not one famous engineer's weekend, but the randomized trials, the large surveys, the security audits and the labor data. This essay puts the strongest claims next to the best measurements. The short version: the change is real, it is uneven, it feels bigger than it measures, and the people paying most for it are beginners.
1. A Weekend Project That Took Thirty Minutes
Karpathy's post opens with a scene, and the scene does more work than any statistic. Over a weekend he wanted a dashboard that would analyze video from the cameras around his home. So he typed one paragraph to a coding agent: here is the local IP address, username and password of my DGX Spark; log in, set up SSH keys, set up vLLM, download and benchmark Qwen3-VL, set up a server endpoint to run inference on videos, build a basic web dashboard, test everything, set it up with systemd, record memory notes for yourself, and write me a markdown report.
The agent went away for about thirty minutes. It ran into several problems, looked up solutions online, fixed them one by one, wrote the code, tested and debugged it, set up the services, and came back with the report. "I didn't touch anything," Karpathy wrote. "All of this could easily have been a weekend project just 3 months ago but today it's something you kick off and forget about for 30 minutes."
From that he drew a large conclusion: "You're not typing computer code into an editor like the way things were since computers were invented, that era is over." The new work, in his words, is "spinning up AI agents, giving them tasks in English and managing and reviewing their work in parallel."
Anyone who has written code for a living should read the rest of the post too, because Karpathy did not claim the tools were magic. "It's not perfect," he wrote. "It needs high-level direction, judgement, taste, oversight, iteration and hints and ideas." It works much better in some situations than others, "especially for tasks that are well-specified and where you can verify/test functionality." The skill, he said, is learning "to decompose the task just right to hand off the parts that work and help out around the edges."
So we have two claims in one post. The first is that the job changed suddenly and completely. The second is that it changed most where the work can be specified and checked. One expert's weekend cannot settle either claim. The rest of this essay looks at what newsrooms and universities have actually measured.
2. Three Eras in Five Years
The phrase "AI coding" has meant three different activities since 2021, and much of the early research measured only the first. Getting the timeline straight is the only way to read the numbers correctly.
Autocomplete (roughly 2021 to 2023). The human types; the tool suggests the next line or block in grey; the human presses Tab or keeps typing. GitHub Copilot made this mainstream. The human is still writing the program, and the tool predicts the next few keystrokes.
Chat (roughly 2023 to 2024). The human pastes a function into a chat window, asks a question, and copies the answer back. The tool now explains and drafts, but it cannot see the project, cannot run anything, and depends on the human to move text back and forth.
Agents (2025 onward). The human describes a task in English. The tool reads the repository, edits files, runs commands and tests, reads the errors, and tries again until it finishes or gets stuck. This is what Karpathy was describing.
The language changed as fast as the tools. In February 2025, Karpathy coined the phrase "vibe coding" for a style of programming in which you stop reading the code at all, describe what you want, and accept whatever the AI produces as long as it seems to work. By November, Collins Dictionary had made vibe coding its Word of the Year, defining it as software development "that turns natural language into computer code using AI." A half-joking post had become a dictionary entry in nine months.
The tool I already have followed the same path. GitHub put a Copilot "coding agent" into public preview in May 2025, and on 25 September 2025 declared it generally available to all paid Copilot subscribers: you give it a task, and it "open[s] a draft pull request and work[s] in the background in its own development environment," then asks for your review. On 25 February 2026, the same day as Karpathy's post, GitHub made Copilot CLI generally available for Pro, Pro+, Business and Enterprise plans. It is a terminal agent with a plan mode that asks questions and drafts a plan before writing code, an "autopilot" mode that runs commands without stopping for approval, sub-agents that can work in parallel, a way to hand a task off to the cloud agent, and memory that carries a repository's conventions from one session to the next. For Business and Enterprise users, an administrator has to switch it on first.
The big companies tell the same story with numbers, but they should be read with care. Google's chief executive Sundar Pichai said in April 2026 that 75% of the company's new code is AI-generated, up from about a quarter in 2024 and about half the previous autumn, and he described the workflow as "truly agentic," with engineers supervising teams of AI agents. A year earlier, Microsoft's Satya Nadella had said that "maybe 20%, 30%" of the code in Microsoft's repositories was written by software. Neither company has said what it counts as "AI-generated": lines typed, suggestions accepted, or commits. Semafor noted that some of Pichai's claims are hard to fact-check. These figures show which way things are moving. They are not measurements, and the two numbers do not even measure the same thing.
3. What the Controlled Experiments Found
If we want to know whether AI makes programmers faster, the best evidence is the randomized trial: give some developers the tool and not others, by lot, and compare. There are now enough of these to tell a story, and the story is more interesting than either "faster" or "slower." I will take them in the order they were done.
2023: a lab task, 55.8% faster. Sida Peng, Eirini Kalliamvakou, Peter Cihon and Mert Demirer asked recruited developers to implement an HTTP server in JavaScript as quickly as possible. Those with GitHub Copilot "completed the task 55.8% faster than the control group." This was the number that launched a thousand slide decks. It is also a very particular kind of task: small, self-contained, new, and with a clear definition of done. That is exactly the kind of task Karpathy later said agents handle best.
2022 to 2024: real companies, about 26% more tasks completed. Economists Cui, Demirer, Jaffe, Musolff, Peng and Salz looked at three randomized rollouts of Copilot that Microsoft, Accenture and an anonymous Fortune 100 electronics manufacturer ran as part of normal business. Pooling 4,867 developers, they estimated a 26.08% increase in completed tasks for those with the tool. The standard error was large (10.3 percentage points), and the individual experiments were noisy. One finding stands out: less experienced developers adopted the tool more and gained more from it. The paper is now published in Management Science, and MIT Sloan's summary is the easiest way in for a general reader.
Early 2025: experts in their own code, 19% slower. Then came the study that made headlines for the opposite reason. The research group METR recruited 16 experienced open-source developers who worked on large, well-known repositories (averaging more than 22,000 GitHub stars and a million lines of code) that they had contributed to for years. Each of 246 real issues was randomly assigned either to allow AI tools or to forbid them. When AI was allowed, the issues took 19% longer.
The slowdown was not the most interesting number. Before the study, the developers predicted AI would make them 24% faster. After it, having just been measured as slower, they still believed AI had made them about 20% faster. In my view, this gap between what people felt and what was measured is the most useful finding in this whole field, and I will come back to it.
Late 2025: the experiment that could not be run. METR repeated the study with newer tools, starting in August 2025, and on 24 February 2026, the day before Karpathy's post, published what happened. The point estimates now leaned the other way: about 18% faster for returning developers (confidence interval from 38% faster to 9% slower) and about 4% faster for new recruits (from 15% faster to 9% slower). But METR refused to stand behind the numbers. The main reason was "a significant increase in developers choosing not to participate in the study because they do not wish to work without AI." Some developers held back the tasks they did not want to do by hand. Others were running several agents at once, which made their working time hard to measure. METR said it believes developers are more sped up now than in early 2025, but that its data is "only very weak evidence for the size of this increase." The 2025 study page now carries a note saying its results are out of date.
In other words, the measurement problem has become a finding in its own right. When working professionals will not accept a few weeks without a tool, even for pay, you cannot easily build a control group. That tells you something about how deeply the tool has been adopted, even if it cannot tell you how much it helps.
Mid-2026: METR's own summary. In its Frontier Risk Report for February to March 2026, METR put the late-2025 trial's result at "small (~4-20%) productivity benefits," probably an underestimate. In the same report it noted that when developers are simply asked, their self-reported gains range from about 1.6× to 4× depending on how the survey is worded, and that self-reports have overestimated benefits in the past. A separate METR survey of 349 technical workers found a median self-reported improvement of 1.4× to 2× in the value of their work and 3× in speed. METR itself listed reasons to doubt both, starting with its own earlier finding that developers had misjudged AI's effect on their time by about 40 percentage points.
Put the five results side by side and three lessons come out.
First, the gains depend heavily on the task. They are largest on small, new, clearly specified jobs and smallest for experts working in large codebases they already know by heart. That is not a contradiction of Karpathy; it is his own caveat, measured.
Second, feeling faster is not evidence of being faster. In 2025, experienced developers misjudged the direction of the effect, not just its size. Any personal impression, including mine, has to be checked against a clock.
Third, the field is moving faster than its measuring instruments. The tools changed between the start and end of a single study. The most careful research group in the area says honestly that it does not yet know how big the current effect is. That is a better answer than a confident number.
4. The Capability Curve Behind the "December" Feeling
If the measured speed-ups are modest, why did Karpathy feel such a sharp break in December? Part of the answer is a curve that METR has been tracking since 2019.
The idea is simple. Take a set of software tasks, measure how long they take skilled humans, and find the task length at which a given AI model succeeds half the time. Call that the model's "time horizon." In a paper first released in March 2025, Thomas Kwa and colleagues found that the frontier time horizon had been doubling approximately every seven months since 2019, perhaps faster in 2024. At the time, the best models had a horizon of about fifty minutes. METR's February to March 2026 estimate for the best public models was about twelve hours, with a wide uncertainty range of five to sixty-one hours.
A doubling curve explains why a change can feel sudden. Each step is the same size in proportion, but the steps eventually cross the length of real jobs. A model that can reliably do a forty-minute task is a helper. A model that can do a several-hour task is something you hand an evening's work to and walk away from, which is what Karpathy did.
METR insists on two warnings, and both matter for anyone about to start working this way. The first is that a 50% time horizon of X hours does not mean you can delegate every task shorter than X hours. Half the time is not most of the time.
The second warning is sharper. In March 2026, METR had real maintainers of scikit-learn, Sphinx and pytest review 296 AI-written pull requests that had already passed the automated tests of the SWE-bench Verified benchmark. Roughly half of those test-passing PRs would not have been merged. The reasons ranged from style and not following the project's conventions, through breaking other code, to the most serious: the issue was not actually solved. (For comparison, the maintainers would have merged about 68% of the original human-written fixes, so even human patches do not pass automatically.) The authors note this is not a hard limit, since the agents were not given the review-and-revise loop that human contributors get. But it shows that passing the tests and being good code are different things. Karpathy's "oversight" caveat is not a formality.
5. What Developers Say
Controlled trials are small. Surveys are large. They measure different things: surveys tell us what people do and believe, not what actually happens to their output. With that caveat, two of the biggest surveys agree on one picture. Adoption is nearly universal, and trust is not.
Google's DORA research program surveys software professionals every year. Its 2024 report, based on more than 39,000 respondents, found something uncomfortable: in environments where AI had been adopted, software delivery throughput fell about 1.5% and delivery stability fell about 7.2%, and 39% of respondents said they had little or no trust in AI-generated code.
The 2025 report, retitled the State of AI-assisted Software Development and based on nearly 5,000 professionals, found that 90% now use AI at work, up 14 points, for a median of about two hours a day. Over 80% said it had improved their productivity and 59% said it had improved code quality. AI adoption was now associated with higher delivery throughput, reversing the year before, though DORA noted that making sure software works as intended before release is still a challenge. Trust stayed low: only 24% trusted AI "a lot" or "a great deal," while 30% trusted it "a little" or "not at all." DORA calls this a trust paradox: people rely on a tool they do not fully trust.
The report's main conclusion is the one I find most useful. AI, it says, is "a mirror and a multiplier." In a well-run organization it multiplies the strengths; in a fragmented one it shows up the weaknesses. DORA sorted teams into seven profiles and proposed seven organizational capabilities that decide whether AI helps. The idea is that the tool does not create good engineering practice. It makes whatever practice is already there bigger.
The Stack Overflow Developer Survey 2026, published on 6 October with more than 30,000 respondents, shows how far the agent era has spread. Among AI users, 66% use coding agents and 63% use chatbots. The most-used agents were Claude Code (66%) and GitHub Copilot (59%). Of those using AI coding assistants or agents, 73% use them every day, and about a third of daily users spend four or more hours a day with them. Overall, 62% view AI at least somewhat positively, and 93% say they need to see the sources before they trust an AI answer.
The survey also shows where developers use agents, and this is the pattern I want to point out. The top uses are generating code in areas they already know (67%) and debugging (61%). Only 20% use AI for deploying, operating or troubleshooting production systems. Developers have worked out Karpathy's rule on their own: use the agent where you can check the result yourself, and keep it away from the places where a mistake costs the most.
6. The Bill: Quality and Security
If speed is the benefit, quality and security are the costs, and they are better documented than many enthusiasts admit. The studies come from different places, but they point to one mechanism: confidence rises faster than correctness.
The clearest academic evidence comes from Stanford. Neil Perry, Megha Srivastava, Deepak Kumar and Dan Boneh ran a user study, presented at the ACM Conference on Computer and Communications Security in 2023, in which participants did security-related programming tasks with or without an AI assistant built on OpenAI's Codex. Those with the assistant wrote significantly less secure code, and they were also more likely to believe their code was secure. The participants who trusted the AI less, and who worked harder at phrasing and adjusting their prompts, produced code with fewer vulnerabilities. The tools have improved a great deal since then, but the human side of the finding has not gone away: help makes us feel safer than it makes us.
Industry studies point the same way, though they need to be read as what they are. The security company Veracode tested more than 100 language models on 80 coding tasks designed around known weakness types and found that the generated code contained a security flaw in 45% of cases. Java was the worst, failing more than 70% of the time. The models failed to defend against cross-site scripting in 86% of relevant cases and against log injection in 88%. "Models are getting better at coding accurately but are not improving at security," said Veracode's Jens Wessling. Veracode sells security scanning, and the tests were built for its own analyzer, so this is one company's yardstick, not a neutral census. But the direction is consistent with Stanford's.
GitClear, which sells code-analysis tools, studied 211 million changed lines of code from 2020 to 2024. It found that blocks of five or more duplicated lines increased eightfold during 2024, that "moved" lines, the usual sign of refactoring, fell by about 40%, and that 2024 was the first year in which copy-pasted lines outnumbered moved ones. The likely reason is simple: it is easier for an assistant to paste a new block than to find and reuse an existing function it cannot see. Again, the data comes largely from GitClear's own customers, and the company has an interest in the result.
The cleanest version of this whole section is METR's merge study from §4. Code that passes its tests can still be code that the people responsible for the project will not accept. Tests check what someone thought to test. Maintainability, security and fitting the project's conventions are judged by people, and that judgment is exactly the part Karpathy says still needs a human.
7. Who Pays: Skills and the Junior Pipeline
The sharpest academic findings of 2026 are not about speed. They are about learning, and about who gets hired.
In January 2026, Anthropic published a randomized controlled trial by Judy Hanwen Shen and Alex Tamkin on how AI assistance affects the formation of coding skills (the full paper is on arXiv). Fifty-two engineers, most of them junior, learned Trio, a Python library for asynchronous programming that none had used before, by building two features. Half could use an AI assistant. Afterwards everyone took a quiz. The AI group scored about 50%; the hand-coding group about 67%. The difference was statistically significant, and the largest gap was on debugging questions. The AI group was not significantly faster either: about two minutes, and some participants spent up to eleven minutes just writing their requests.
What makes this study useful rather than merely alarming is what it found about how people used the assistant. Those who handed all the code-writing to the AI, gradually came to rely on it, or used it to debug instead of to understand, averaged below 40% on the quiz. Those who asked only conceptual questions and fixed their own errors, or who asked for code together with an explanation, or who generated code and then asked follow-up questions about it, scored 65% or higher. The conceptual-questions group was also the second-fastest. The authors' summary is worth quoting: "Cognitive effort—and even getting painfully stuck—is likely important for fostering mastery." They are careful to say that the patterns are associations from a small sample and that the quiz measured understanding right after the task, not long-term retention.
The labor data is the part that should worry everyone, not only beginners. Erik Brynjolfsson, Bharat Chandar and Ruyu Chen at Stanford's Digital Economy Lab have been tracking payroll data from ADP, which covers millions of American workers. The latest version of their paper "Canaries in the Coal Mine?", dated August 2026 with data through June 2026, reports six facts. There is no sign of economy-wide job loss. But employment of workers aged 22 to 25 in the occupations most exposed to AI "now stands 19% below where it would be had it kept pace with that of their less-exposed peers," while experienced workers show no such gap. The gap has widened steadily since the authors first reported it in August 2025. It works mainly through less hiring of young workers rather than more firing. It is concentrated in jobs where AI mostly substitutes for human tasks, while employment is flat or rising where AI mostly complements workers, especially experienced ones. Software developers are one of the occupations the paper looks at in detail. The effect holds even when technology firms and computer occupations are excluded. The authors call these "early, descriptive indicators," not proof of cause, and they list the caveats themselves.
Put the three findings together and there is a tension at the heart of the whole subject. The Copilot field trials found that juniors gain the most output from AI. The Anthropic trial found that juniors may learn less while using it, depending on how they use it. The Stanford data finds that fewer juniors are being hired into the most exposed jobs. The tool helps the beginner produce, may slow the beginner's growth, and makes employers less eager to hire the beginner at all. Every senior engineer was once a junior who got painfully stuck. If the field stops hiring and training juniors, where does the next generation of the people who do Karpathy's "oversight" come from?
8. What the Job Turns Into
If the typing goes to the machine, what is left for the programmer? Karpathy's answer is that the work moves up a level: "The biggest prize is in figuring out how you can keep ascending the layers of abstraction," setting up long-running orchestrator agents "with all of the right tools, memory and instructions" to manage several coding agents at once.
There is evidence that this shift is happening, not just being predicted. In April 2025, Anthropic's Economic Index looked at half a million interactions, split between its chat interface and its agent tool, and found that 79% of conversations in Claude Code, its agent tool, were "automation", where the AI does the task directly, against 49% in the ordinary chat interface. Start-up work made up a larger share of agent use than enterprise work (about 33% against 24%), a sign of who moves first. That was a snapshot from early 2025, before the December change. Pichai's description of Google engineers supervising autonomous AI teams is another sign, and so is the small detail from METR's trial: developers were running so many agents in parallel that their time could no longer be measured cleanly.
What stays with the human is clearer now than it was two years ago. Someone has to write the specification and decide what "done" means. Someone has to break the work into pieces the agent can handle and that a person can check. Someone has to read the diff, judge the design, own the security, and say no. DORA's "mirror and multiplier" belongs here. An agent dropped into a team with good tests, clear conventions and careful review multiplies all of that. Dropped into a team without them, it multiplies the lack.
9. Both Stories Are True
Karpathy ended his post by saying this is "nowhere near 'business as usual' time in software." The evidence agrees with him about the direction. Nearly every professional developer now uses AI. Agents can do longer and longer tasks on a curve that has held for years. Companies report that most of their new code is machine-written, and the job is moving from typing to specifying, decomposing and reviewing.
The evidence disagrees with the excitement about size and cost. The best controlled trials show gains that range from large on small fresh tasks to modest, or even negative, for experts in code they know well, and the people in those trials judged their own speed badly. Code that passes its tests is rejected by maintainers half the time. Security flaws are common, and assistance makes people overconfident. Beginners learn less when they simply delegate, and fewer of them are being hired.
Three things will show which story is winning over the next year or so, and each can be checked. First, whether METR's redesigned studies produce a speed-up estimate it is willing to stand behind. Second, whether the Stanford gap for young workers keeps widening or begins to close. Third, whether the share of agent-written pull requests that maintainers actually merge catches up with the share that passes the tests.
I keep coming back to the thirty minutes. The agent did the work, and Karpathy did not touch anything. But he knew what to ask for, in what order, and what a finished result had to look like, and he knew how to read the report it handed back. The typing era may be over. The knowing era is not.
References
- Andrej Karpathy, post on X, 25 February 2026; copy kept by Simon Willison, 26 February 2026.
- AFP via The Korea Times, 'Vibe coding' named word of the year by Collins Dictionary, 6 November 2025.
- GitHub Changelog, Copilot coding agent is now generally available, 25 September 2025.
- GitHub Changelog, GitHub Copilot CLI is now generally available, 25 February 2026.
- Semafor, Google CEO says 75% of company's new code is AI-generated, 24 April 2026.
- TechRepublic, Microsoft CEO Nadella: 20% to 30% of our code was written by AI, April 2025.
- Peng, Kalliamvakou, Cihon, Demirer, The Impact of AI on Developer Productivity: Evidence from GitHub Copilot, arXiv, 2023.
- Cui, Demirer, Jaffe, Musolff, Peng, Salz, The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers, Management Science; MIT Sloan summary.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025.
- METR, We are Changing our Developer Productivity Experiment Design, 24 February 2026.
- METR, Frontier Risk Report (February to March 2026), 19 May 2026.
- METR, Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity, 11 May 2026.
- Kwa et al., Measuring AI Ability to Complete Long Software Tasks, arXiv, 2025.
- METR, Clarifying limitations of time horizon, 22 January 2026.
- Whitfill, Wu, Becker, Rush (METR), Many SWE-bench-Passing PRs Would Not Be Merged into Main, 10 March 2026.
- InfoQ, on the 2024 DORA Accelerate State of DevOps Report, 28 November 2024.
- Google, How are developers using AI? Inside our 2025 DORA report, 23 September 2025.
- Stack Overflow, The results of the 2026 Developer Survey are here, 6 October 2026.
- Perry, Srivastava, Kumar, Boneh, Do Users Write More Insecure Code with AI Assistants?, ACM CCS 2023.
- Veracode, AI-Generated Code Poses Major Security Risks in Nearly Half of All Development Tasks, 30 July 2025.
- DevClass, AI is eroding code quality, states new in-depth report (GitClear), 20 February 2025.
- Shen, Tamkin (Anthropic), How AI assistance impacts the formation of coding skills, 29 January 2026; paper on arXiv.
- Brynjolfsson, Chandar, Chen, Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence, Stanford Digital Economy Lab, August 2026 version.
- Anthropic, Anthropic Economic Index: AI's Impact on Software Development, 28 April 2025.