AI agents write more code, but no significant rise in the share of resolved tasks was found: what a study of 700 firms showed and how to measure the benefit in your own team
Topics: AI, Automation, Small business

Short answer: according to a study retold by Ars Technica, after AI agents were introduced for writing code, firms saw lines of code rise by 30%, commits by 20% and pull requests by 23%. At the same time, the authors found no statistically significant change in the share of resolved tasks in the task-tracking system, Jira. The review process, however, became longer: the time from submitting a pull request to merging it rose by 49% on average. The authors write that there is little evidence that firms increase software output or reduce employment. For the owner of a company with a small development team, the practical conclusion is this: judging the benefit of AI by the number of lines is unwise; it is better to look at resolved tasks and at time to merge. Below are the numbers, their limits and three steps.
What was measured
A few words without which the numbers are hard to follow. A commit is a saved batch of changes to code. A pull request is a request to accept changes into the main version of a program: another person reviews it first. An epic is a large task that groups many small ones. Jira is a program where teams keep track of tasks.
The study was carried out by Harvard researchers Fiona Chen and James Stratton. They used aggregated data from Jellyfish, a company that measures the work of development teams. Everything below retells Ars Technica; I did not open the study itself.
- Where the data comes from. About 300 million "work events," such as commits and pull requests, and data from task-tracking systems. It covers more than 700,000 employees at more than 700 firms, from 2021 to March 2026.
- How the moment of adoption was determined. From measured AI usage and from activity on GitHub. The authors distinguish assistants, which suggest continuations of code written mostly by a person, from agents, which mostly write and submit code on their own on request.
- How it was calculated. They compared indicators before and after adoption, which happened at different times at different firms. The method is called "difference-in-differences."
What came out
- lines of code: up 30%;
- commits: up 20%;
- pull requests: up 23%;
- share of resolved tasks and epics in Jira: no significant change;
- time from submitting a pull request to merging it: up 49% on average;
- share of requests where changes are asked for: nearly doubled;
- comments per request: up 35%;
- share of employees doing reviews: up 14%.
Three more points from the source. The authors also found no shift in the size and complexity of the tasks tracked in Jira. They do not attribute significant changes in employment to AI. And on AI review: by March 2026, 80% of firms used some form of it, but agents, according to the study, wrote only 23.3% of all review comments and 10.8% of pull requests. Ars draws the conclusion that people were still doing the overwhelming part of this work.
Ars points out limitations. The data ends in March 2026, and agents have been updated since then. According to Ars, by the time of publication 95% of the firms in the study had adopted agents, and many, in the outlet's view, are still going through a learning period about when and how to use them. The balance between time to write and time to merge may change.
What this means for a small team
These are my conclusions from the numbers above, not the authors' words.
In the study, "lines, commits, requests" rose, while the authors found no statistically significant change in the share of resolved tasks. If you judge the benefit of AI by the first group, they will not show what matters to you: how many finished tasks you close. Three steps to measure the benefit in your own team.
- Choose the indicator that matters to you. Closed tasks in your tracking system per month, not lines of code.
- Who does it: the head of development.
- How to check: take the last three months before AI was introduced and the first three months after. If your tracking system cannot tell a small task from a large one, compare at least the number of tasks of one type, for example bug fixes.
- Measure time to merge. How many hours or days on average pass from submitting a change to merging it. This is calendar time; it includes waiting, not just the reviewer's work.
- Who does it: the head of development; the data comes from GitHub or GitLab.
- How to check: take the 20 most recent merged requests before adoption and 20 after, and compare the averages. If the time has grown, as at the firms in the study, the requests are worth examining more closely to understand the cause.
- Set aside time for review in advance. If the same people review and write, part of their working time goes to review, and without a plan the requests wait.
- Who does it: the head of development.
- How to check: the team's schedule has explicit review hours, and the number of requests waiting for review for more than a day is recorded and not growing.
What it looks like in numbers: an example with illustrative figures
Suppose a change now waits for merging 2 days on average. If after agents are introduced this time grows by 49%, as it did on average at the firms in the study, it will be about 3 days: a day longer. This is calendar waiting time, not the reviewer's working hours, and salary should not be calculated from it. But you can compare it with how much faster you release finished tasks. If there are more requests but the same number of tasks are closed, the benefit remains unproven. The numbers are illustrative: substitute your own.
When this does not concern you
If you do not develop software but use ready-made programs, the study does not directly concern you. If you write code alone, your case may differ: the study covers firms, and from Ars's retelling it is impossible to tell how work was divided in each of them.
Summary
According to the study by Fiona Chen and James Stratton, as retold by Ars Technica, AI agents increased the volume of code firms wrote, but no significant change in the share of resolved tasks was found, and the time from submission to merge rose by 49%. The data ends in March 2026, and the picture may change. The benefit is worth judging by resolved tasks and by time to merge, not by lines.
I work on AI agents and automation. If you want to look at my projects or discuss your own task, visit my portfolio.
Sources
- Ars Technica, 2026-10: AI coding agents generate more code, but not more software (a retelling of the study by Fiona Chen and James Stratton, Harvard University, based on Jellyfish data). https://arstechnica.com/ai/2026/10/ai-coding-agents-generate-more-code-but-not-more-software/