AI Pull Requests Wait 5x Longer for Review. Queueing Math Explains Why, and It Gets Worse.
Reviewers are not slower. The queue is nonlinear, and a doubling of PR volume can multiply waiting time tenfold.
A queue that looks like a quality problem
LinearB analysed more than 8.1 million pull requests from about 4,800 teams (report). The headline numbers are about acceptance: AI-assisted PRs merged within 30 days 32.7% of the time, against 84.4% for unassisted ones. The number that matters for this article sits a few lines lower. Agent-created PRs waited 17.6 hours at the 75th percentile before a first review, against 3.4 hours for unassisted work, a gap of about 5.25 times.
Now read the next figure. Once a reviewer picked an agentic PR up, the review took 6.4 hours, against 4.2 hours for unassisted PRs. Slower, but nowhere near 5 times slower. Reviewers are not refusing to read AI code. The code is sitting in a line.
The usual explanation is that AI code is harder to review, and that is partly true: LinearB also reports a 75th percentile size of 408 lines for AI-assisted PRs and 293 for agentic ones, against 157 for unassisted work. But a size difference of about 2.6 times does not produce a waiting difference of 5.25 times on its own. Something nonlinear is going on, and queueing theory has a name for it.
Wait time is load divided by spare capacity
The simplest queue model, called M/M/1, assumes work arrives at random and one server handles it at a variable pace. Its average waiting time is the service time multiplied by ρ/(1−ρ), where ρ is the fraction of time the server is busy. Treat your review pool as that server. If reviewers spend 50% of their available review time actually reviewing, the average wait equals one service time. At 90% it equals nine. At 95% it equals nineteen.
| Reviewer busy fraction (ρ) | Wait multiple | What it feels like |
|---|---|---|
| 50% | 1.0x | Reviews start the same day |
| 70% | 2.3x | Occasional slow PR |
| 80% | 4.0x | Review is a recurring complaint |
| 90% | 9.0x | PRs age for days |
| 95% | 19x | Queue is the main work item |
This curve is why intuition fails. A team running at 45% review load sees an average wait of about 0.8 service times. Raise PR volume by 98%, the merge growth LinearB attributes to AI users, and load becomes about 89%. The wait is now about 8.2 service times. Volume doubled; waiting time grew roughly tenfold.
“Doubling the arrivals does not double the wait. Near capacity, it multiplies it.”
Two caveats keep this honest. M/M/1 is a teaching model, not a measurement of any real team. And real PR arrivals are burstier than random, with Monday-morning agent batches and pre-release pushes, which makes the real curve steeper, not gentler (Kingman's approximation adds a term for variability in arrivals and service). The shape is the point, not the exact multiples.
def wait_multiple(rho: float) -> float:
"""Average wait in units of one review (M/M/1)."""
if rho >= 1:
return float("inf") # queue grows without bound
return rho / (1 - rho)
base_load = 0.45
for volume_growth in (0.0, 0.25, 0.5, 0.98):
rho = base_load * (1 + volume_growth)
print(f"+{volume_growth:>4.0%} PRs -> load {rho:.0%}, wait x{wait_multiple(rho):.1f}")
Why AI changes both sides of the queue at once
Faster code generation raises the arrival rate. That alone moves a team along the curve. But the other variable, the service time, is also moving against you. Bigger diffs take longer to read, and unfamiliar code that nobody on the team wrote carries more risk per line. Higher arrivals and longer service times both push ρ up, and they multiply.
There is a human factor the model ignores. Reviewers choose what to open next. A 400-line diff with no author who can explain it is the one that gets skipped when the queue is long. Skipped items age, and aged items are the ones that eventually get closed unmerged. That is a plausible contributor to the 32.7% merge rate, though the report does not isolate it.
Other 2026 surveys describe the same pressure from the people inside it. A Sonar survey of 1,100 developers, as summarised in a DEV Community roundup, puts AI at 42% of committed code and finds that only 48% of developers always verify it before committing. The same roundup cites a Pragmatic Engineer report of teams seeing about 30 pull requests a day with six reviewers. That is five PRs per reviewer per day, each potentially several hours of attention.
Five levers, ranked by how much they move the curve
You can attack the queue from five directions. They are not equal.
- Cap work in progress per author or agent. Two open PRs per author, or per agent instance, limits arrivals directly. It is the cheapest lever and the one most teams skip because it feels like throttling.
- Enforce a size budget. If agentic PRs run 293 lines at p75 and unassisted ones 157, a limit near 200 lines forces agents to split work. Smaller diffs shorten service time and make reviews easier to start.
- Route by risk. A dependency bump, a generated client or a copy change does not need the same reviewer attention as an auth change. Lighter lanes remove arrivals from the expensive queue.
- Automate the first pass. Static analysis, test evidence and an AI-written change summary shorten service time. GitHub is cited in the DEV roundup as seeing 30% to 50% less human review time with this approach, a secondary figure worth testing on your own repositories before you plan around it.
- Add reviewers last. More capacity does lower ρ, but it is the most expensive lever, and new reviewers need context that takes months to build.
What to measure
Average cycle time hides queue pain, because a few fast PRs offset many stuck ones. Track three numbers instead. First, queue age at the 75th percentile: how long the median-to-slow PR waits before a first review. This is the figure LinearB reports, and it moves early. Second, weekly PR arrivals divided by weekly review capacity in hours, which is your ρ. Third, the share of PRs closed without merging after more than seven days.
If ρ sits above roughly 70% for a month, expect the wait complaints to start before the dashboards turn red. Teams that treat review as a capacity-planned resource, with a load target the way an operations team plans servers, will spend the next year doing better than teams that treat it as goodwill.
Where this goes next
The generation side of software is now cheap and elastic. The verification side is human, finite and queued. Over the next few quarters, the teams that pull ahead will be the ones that design for review capacity the way they once designed for compute, deciding in advance how much code they can safely absorb per week and letting the agents work to that number.
Frequently asked questions
Related reading
A Registry Counted 487 AI Agent Incidents. The Ones With No Attacker Caused the Most Harm.
A new registry counted 487 disclosed AI agent incidents. The headline 24% harm rate is a composition artifact, not a risk rate, the authors say. The number that matters is buried three tables deeper.
OpenAI Shut Down Atlas After 292 Days. Every Other AI Browser Is About to Learn Why.
OpenAI shut down its standalone Atlas browser on 9 August 2026, ten months after launch, even as usage was climbing. The real reason has more to do with Chrome’s grip on the desktop than with agentic AI.
Three AI Agent Production Incidents, One Root Cause Every Postmortem Missed
Replit, AWS Kiro, and Claude Code each deleted production this year. Every published fix patched the specific bug. None asked whether the same silent gap exists everywhere else an agent has write access.