🔍 Read the full analysis: Why AI’s Cheap Output Creates An Expensive Review Queue on ThorstenMeyerAI.com
Get privacy and security gear delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
AI tools are increasing the volume of work in fields such as mathematics, software and contract review, while verification still relies heavily on people. The supplied figures point to a growing review bottleneck, but several come from commercial vendors and do not establish a universal effect or its long-term scale.
In one example, OpenAI generated 722 mathematical manuscripts across 372 families after its model was given about 4,000 problems, the source says. Average compute time per result was about three hours. Some results were checked in Lean, a proof-assistant system; OpenAI cautioned that unformalized results could have issues. A counterexample to an Erdős conjecture from the same programme drew careful checking by five leading mathematicians, illustrating the difference between generating a candidate result and establishing that it holds up.
Software data cited by the source also indicates pressure on review queues. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, which analyzed 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests received no human review before being merged or closed.
The source also points to a partnership between OpenAI and contract-software company Ironclad. It says GPT-6 Astra, trained on real contracting workflows, met 55% of evaluation criteria on average across 11 tasks. That result is a performance measure, not evidence that the system can replace legal review; the remaining criteria still matter if a draft is to be used in practice.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity May Limit AI Adoption
The figures suggest that production and verification are separating. If a team can generate more code or drafts than it can check, the extra output may wait in a queue, receive less scrutiny or be rejected. Each outcome can blunt the productivity gains that faster generation appears to offer.
There are also risks beyond delay. Software merged without review can carry defects; unchecked legal language can miss approval rules or jurisdiction requirements. Formal tools help confirm that an answer meets defined criteria, but people still need to judge whether those criteria match the real task. Responsibility also remains with people and institutions that approve, publish or sign off on the work.
That creates a potential premium for experienced reviewers: engineers, lawyers, auditors and scientists who can assess work and stand behind a decision. It also raises a workforce concern. If AI takes over the drafting and coding tasks through which junior staff learn, employers may weaken the route by which future reviewers develop judgment. The source presents this as a risk, not a measured outcome.
As an affiliate, we earn on qualifying purchases.
Evidence Across Three Workflows
The examples cover three different kinds of review. In mathematics, a proof assistant can check whether a formal proof follows specified rules, but many results are not formalized, and experts still assess whether the claim is meaningful and correctly framed. The source describes this as “verification abundance, adjudication scarcity”: automated checks can expand while expert interpretation remains constrained.
In software, pull-request metrics track how code changes move through team workflows, but the figures have different samples and methods and should not be treated as directly comparable. Faros AI and LinearB sell products related to software development and review, so their findings warrant careful reading. The peer-reviewed study adds a separate data point, though the supplied material does not give its sample details.
In contracting, the cited model score shows that AI can meet some evaluation criteria on defined tasks. It does not establish how often errors would occur in live contracts or whether human review time would fall. Across all three areas, cheap generation does not itself prove reliable output; the relevant test is whether review keeps pace and catches consequential mistakes.
““Verification abundance, adjudication scarcity.””
— A recent paper title, as cited by ThorstenMeyerAI.com
software pull request review software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Available Evidence
The cited metrics do not show how much review time is caused specifically by AI-generated work, or whether the same patterns hold across industries and organizations. The source notes that several software-data providers sell code-review tools; their figures are reported findings, not a single independent measurement of all workplaces. The supplied material also does not include the underlying methods for every statistic.
It remains unclear whether AI review systems will reduce the burden enough to change the overall balance, or whether human reviewers will continue to be required for high-stakes decisions. Nor do the examples establish that AI adoption has already reduced the number of people learning core professional skills. Those are plausible concerns raised by the analysis, but their scale and timing are not yet shown by the cited evidence.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch Review Queues and Training
The next useful signals will be whether organizations track review time, defect rates and rework alongside the volume of AI-generated output, and whether those measures improve as tools change. Further detail on study methods and results would help readers compare the reported figures and judge how broadly they apply.
Employers will also need to decide how junior staff gain the experience required for later review roles. For now, the central question is not only how much work AI can produce, but who checks it, how carefully and who remains accountable when it is accepted.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does “review bottleneck” mean here?
It means AI-generated work may arrive faster than people can assess it. The resulting queue can delay useful work or leave some output with limited review.
Does formal verification solve the problem?
It can check whether a proof or program meets specified rules. It does not necessarily establish that the rules reflect the intended question or that the result is useful in practice.
How strong is the software evidence?
The source cites findings from Faros AI and LinearB, both commercial providers, and a peer-reviewed 2026 study. Their results point in a similar direction, but differences in methods and samples mean they should not be treated as one universal estimate.
Does the contract-model result show AI can replace lawyers?
No. The cited score says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. It does not show that the system can independently handle contracts or remove the need for legal review.
What should organizations monitor next?
They can compare output volume with review delays, acceptance rates, defects and rework, while checking whether junior staff still get experience that builds professional judgment.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
