research

Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests

The study asks how often merged AI-agent PRs receive follow-up fixes, who authors those fixes, and how merge-time signals differ between agent PRs with and without a later fix.

Published
Published
Reviewed
Reviewed
Next review due
Review due
Version
Version 2

By

SOFTWARE_ENGINEERINGEMPIRICAL
About this BaitaPhish analysis and its review
Trust and provenance

Editorial record

AI-assistance disclosure

Research Intelligence analysis generated with AI and checked against cited source evidence.

A human review was recorded.

Sources

  • arxiv.org2609.26847v2

    Claims attributed to the linked primary source in this content record.

    Version
    2609.26847v2
    Retrieved
    Reuse
    link-only

TL;DR

  • The study asks how often merged AI-agent PRs receive follow-up fixes, who authors those fixes, and how merge-time signals differ between agent PRs with and without a later fix.

    Source: [2], [15], [22]

  • Candidate fixes had to merge in the same repository strictly after the original PR and within 30 days, carry an AIDev fix-type tag, share at least one edited file with the original PR, and include at least one non-boilerplate file; PRs without a full 30-day observation period were excluded.

    Source: [5], [7]

  • A candidate pair counted as a verified fix only when labeled Direct fix; two annotators independently labeled a 50-pair sample, and the LLM judge labeled the remaining candidate population using the PR information provided to human annotators.

    Source: [16], [23]

  • Within the shared observation window, 3.68% of agent merges and 2.34% of human merges had verified fixes; stratifying over 218 repositories with both cohorts, the Mantel–Haenszel odds ratio was 1.62 (95% CI 1.10–2.39, p=0.015).

    Source: [9]

  • The authors note that their file-co-location and 30-day candidate-linking criteria can miss cross-file fixes, fixes taking longer than 30 days, and fixes in PRs not tagged as fix-type; they therefore treat verified-fix rates as floors. Commit retrieval was capped at 30 per PR, affecting fewer than 2% of PRs.

    Source: [6]

Why This Matters

Source-paper contributions

The authors contribute a large-scale empirical study of follow-up-fix authorship after agent PR merges, human validation and an LLM judge for candidate fixes, an analysis of merge-time signals, and a released replication package.

Source: [11]

The authors recommend monitoring agent PRs beyond merge, noting that aggregate merge-time signals offer little separation between PRs that later receive fixes and those that do not; they also identify future research on monitoring frameworks and on how agent/model ownership affects code after a model is discontinued.

Source: [17], [26]

What the paper contributes

Read the finding above.

Evaluation environment

The agent cohort consists of 6,774 merged PRs from five agents in 891 repositories in AIDev-pop; the human cohort consists of merged PRs from AIDev’s curated human-PR table, restricted to the same repositories and merge-only conditions.

Source: [3]

Key Findings

Paper reports

The authors report that, among 4,505 agent merges with a fully observed 30-day window, 22.9% had a candidate fix and 4.5% had a verified fix.

Source: [20]

The authors report that half of the 30-day verified-fix incidence accumulated within the first week in both cohorts, and that the cumulative verified-fix incidence reached 4.5% for agent PRs and 2.6% for human PRs by day 30.

Source: [1]

Among 263 verified fixes on agent merges, 69.6% were opened by the same agent, 27.4% by a human, and 3.0% by another agent; among 110 verified fixes on human merges, 89.1% were opened by humans.

Source: [24]

For verified-fix PRs opened by agents other than Codex, commits between first push and merge were on average 87.4% agent-authored per PR, and 76.4% of those PRs had all commits classified as agent-authored; the authors exclude Codex from the commit measure because its commits lack agent markers.

Source: [13], [18]

In pooled comparisons, merge-time signal differences between agent PRs with and without a follow-up fix were negligible to small; in within-repository comparisons, a PR with ten times as many commits had 6.1 times the odds of a verified follow-up fix (p<0.001).

Source: [14], [21]

Read the finding above.

Limitations

The authors scope their findings to five AIDev agents in repositories with at least 500 stars; other agents, less-active repositories, and behavior after the data window are outside the stated population.

Source: [8]

The authors state that commit-author classifications measure who committed or marked a change rather than who wrote it; because Codex commits carry no marker, they exclude Codex from commit decomposition and treat reported agent-authored shares as lower bounds.

Source: [6]

Read the limitation above.

Read the limitation above.

How the method works

The study workflow extracts candidate follow-up-fix pairs, validates candidates with human annotation and an LLM judge, classifies fix-PR authorship, classifies commit authorship, and compares merge-time signals for agent PRs with and without verified fixes.

Source: [25]

The authors re-fetched commits, diffs, reviews, and comments for human PRs through the GitHub REST API because the original AIDev human-PR table contained PR-level columns only.

Source: [3]

The within-repository analysis used conditional logistic regression with repository controls and controlled for the number of non-boilerplate files shipped; the reported odds ratio compares merges ten times apart on a signal. The within-repository associations do not establish that a signal causes a later fix.

Source: [10], [12]

Read the finding above.

Read the finding above.

Measurement conditions

For the cross-cohort comparison, the authors use a shared observation-window cutoff of June 28, 2025; the comparison includes 2,012 agent PRs and 4,063 human PRs.

Source: [4], [19]

On binary Direct-fix versus non-fix labels, human annotators agreed on 45 of 50 sample pairs (Cohen’s κ=0.77); the LLM judge achieved κ=0.78 against human relabeling, and its Direct-fix precision was 90% on both cohorts (pooled 54/60, Wilson 95% CI 80–95%).

Source: [16], [23]

Research question and scope

Read the finding above.

Tested scope and boundaries

AIDev-pop applies a popular-repository threshold of more than 500 GitHub stars as of June 22, 2025; the agent PRs span December 24, 2024 to July 30, 2025, and the human PRs span January 1 to June 28, 2025.

Source: [3]

Paper Details

Software Engineering · Empirical

Original research: Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests · 2609.26847v2

Paper authors: Wannita Takerngsaksiri, Nhat Duong, Scott Barnett

Source license: CC BY-SA 4.0. This article summarizes and interprets the source using AI. Attribution does not imply endorsement by the source authors.

This adapted analysis is shared under the same CC BY-SA 4.0 license. This brief uses the sampled human-reviewed reader and evidence-bound editorial corrections. Historical model verdicts are retained separately; they do not evaluate changed wording.

Canonical source identity
arXiv 2609.26847
Analyzed source version
v2
Source retrieved
BaitaPhish analysis published
BaitaPhish analysis reviewed

Evidence & Provenance

Show evidence locators

Evidence labels locate support in the original paper; they do not establish independent replication.

  1. [1] · page 9 — Source passage: Admitted source passage
  2. [2] · page 9 — Source passage: Admitted source passage
  3. [3] · page 6 — Source passage: Admitted source passage
  4. [4] · page 8 — Source passage: Admitted source passage
  5. [5] · page 7 — Source passage: Admitted source passage
  6. [6] · page 16 — Source passage: Admitted source passage
  7. [7] · page 7 — Source passage: Admitted source passage
  8. [8] · page 16 — Source passage: Admitted source passage
  9. [9] · page 9 — Source passage: Admitted source passage
  10. [10] · page 13 — Source passage: Admitted source passage
  11. [11] · page 3 — Source passage: Admitted source passage
  12. [12] · page 13 — Source passage: Admitted source passage
  13. [13] · page 11 — Source passage: Admitted source passage
  14. [14] · page 14 — Source passage: Admitted source passage
  15. [15] · page 12 — Source passage: Admitted source passage
  16. [16] · page 7 — Source passage: Admitted source passage
  17. [17] · page 15 — Source passage: Admitted source passage
  18. [18] · page 12 — Source passage: Admitted source passage
  19. [19] · page 8 — Source passage: Admitted source passage
  20. [20] · page 8 — Source passage: Admitted source passage
  21. [21] · page 14 — Source passage: Admitted source passage
  22. [22] · page 7 — Source passage: Admitted source passage
  23. [23] · page 8 — Source passage: Admitted source passage
  24. [24] · page 10 — Source passage: Admitted source passage
  25. [25] · page 5 — Source passage: Admitted source passage
  26. [26] · page 15 — Source passage: Admitted source passage