Why AI’s Cheap Output Creates An Expensive Review Queue
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why AI’s Cheap Output Creates An Expensive Review Queue on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get privacy and security gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

AI tools are increasing the volume of work in fields such as mathematics, software and contract review, while verification still relies heavily on people. The supplied figures point to a growing review bottleneck, but several come from commercial vendors and do not establish a universal effect or its long-term scale.

AI systems are producing mathematical manuscripts, software changes and contract drafts faster than people can reliably review them, according to examples and studies cited by ThorstenMeyerAI.com. The development matters because verification remains a limited human resource: organisations may gain more output without gaining the capacity to determine which results are correct, relevant and safe to use.

In one example, OpenAI generated 722 mathematical manuscripts across 372 families after its model was given about 4,000 problems, the source says. Average compute time per result was about three hours. Some results were checked in Lean, a proof-assistant system; OpenAI cautioned that unformalized results could have issues. A counterexample to an Erdős conjecture from the same programme drew careful checking by five leading mathematicians, illustrating the difference between generating a candidate result and establishing that it holds up.

Software data cited by the source also indicates pressure on review queues. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, which analyzed 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests received no human review before being merged or closed.

The source also points to a partnership between OpenAI and contract-software company Ironclad. It says GPT-6 Astra, trained on real contracting workflows, met 55% of evaluation criteria on average across 11 tasks. That result is a performance measure, not evidence that the system can replace legal review; the remaining criteria still matter if a draft is to be used in practice.

At a glance
analysisWhen: Developing; the cited findings cover re…
The developmentA set of recent examples and industry studies highlights a widening gap between AI-generated work and the human capacity to verify it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity May Limit AI Adoption

The figures suggest that production and verification are separating. If a team can generate more code or drafts than it can check, the extra output may wait in a queue, receive less scrutiny or be rejected. Each outcome can blunt the productivity gains that faster generation appears to offer.

There are also risks beyond delay. Software merged without review can carry defects; unchecked legal language can miss approval rules or jurisdiction requirements. Formal tools help confirm that an answer meets defined criteria, but people still need to judge whether those criteria match the real task. Responsibility also remains with people and institutions that approve, publish or sign off on the work.

That creates a potential premium for experienced reviewers: engineers, lawyers, auditors and scientists who can assess work and stand behind a decision. It also raises a workforce concern. If AI takes over the drafting and coding tasks through which junior staff learn, employers may weaken the route by which future reviewers develop judgment. The source presents this as a risk, not a measured outcome.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Across Three Workflows

The examples cover three different kinds of review. In mathematics, a proof assistant can check whether a formal proof follows specified rules, but many results are not formalized, and experts still assess whether the claim is meaningful and correctly framed. The source describes this as “verification abundance, adjudication scarcity”: automated checks can expand while expert interpretation remains constrained.

In software, pull-request metrics track how code changes move through team workflows, but the figures have different samples and methods and should not be treated as directly comparable. Faros AI and LinearB sell products related to software development and review, so their findings warrant careful reading. The peer-reviewed study adds a separate data point, though the supplied material does not give its sample details.

In contracting, the cited model score shows that AI can meet some evaluation criteria on defined tasks. It does not establish how often errors would occur in live contracts or whether human review time would fall. Across all three areas, cheap generation does not itself prove reliable output; the relevant test is whether review keeps pace and catches consequential mistakes.

““Verification abundance, adjudication scarcity.””

— A recent paper title, as cited by ThorstenMeyerAI.com

Amazon

software pull request review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Available Evidence

The cited metrics do not show how much review time is caused specifically by AI-generated work, or whether the same patterns hold across industries and organizations. The source notes that several software-data providers sell code-review tools; their figures are reported findings, not a single independent measurement of all workplaces. The supplied material also does not include the underlying methods for every statistic.

It remains unclear whether AI review systems will reduce the burden enough to change the overall balance, or whether human reviewers will continue to be required for high-stakes decisions. Nor do the examples establish that AI adoption has already reduced the number of people learning core professional skills. Those are plausible concerns raised by the analysis, but their scale and timing are not yet shown by the cited evidence.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch Review Queues and Training

The next useful signals will be whether organizations track review time, defect rates and rework alongside the volume of AI-generated output, and whether those measures improve as tools change. Further detail on study methods and results would help readers compare the reported figures and judge how broadly they apply.

Employers will also need to decide how junior staff gain the experience required for later review roles. For now, the central question is not only how much work AI can produce, but who checks it, how carefully and who remains accountable when it is accepted.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does “review bottleneck” mean here?

It means AI-generated work may arrive faster than people can assess it. The resulting queue can delay useful work or leave some output with limited review.

Does formal verification solve the problem?

It can check whether a proof or program meets specified rules. It does not necessarily establish that the rules reflect the intended question or that the result is useful in practice.

How strong is the software evidence?

The source cites findings from Faros AI and LinearB, both commercial providers, and a peer-reviewed 2026 study. Their results point in a similar direction, but differences in methods and samples mean they should not be treated as one universal estimate.

Does the contract-model result show AI can replace lawyers?

No. The cited score says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. It does not show that the system can independently handle contracts or remove the need for legal review.

What should organizations monitor next?

They can compare output volume with review delays, acceptance rates, defects and rework, while checking whether junior staff still get experience that builds professional judgment.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
Mobilisiert, nicht ausgegeben: Was von Europas €200-Milliarden-KI-Offensive übrig bleibt

Mobilisiert, nicht ausgegeben: Was von Europas €200-Milliarden-KI-Offensive übrig bleibt

The EU’s InvestAI plan is billed as €200B for AI, but most of the figure is mobilized capital, not direct spending, as a July tender nears.
Licensing and approvals hub for voice actors' AI clones

Licensing and approvals hub for voice actors’ AI clones

A new licensing and approval hub for voice actors’ AI clones aims to streamline consent, usage, and payments, with testing set to begin soon.
ChannelHelm: One Video, Every Platform

ChannelHelm: One Video, Every Platform

ChannelHelm automates the creation of multi-platform content from a single video, reducing manual effort and expanding reach efficiently.
firmulate.com/quotes.html — live view

Five Frontier AIs Faced a Fake CEO—and Protected the Company’s Secrets

Five frontier AIs rejected an escalating fake-CEO demand and a reporter’s trick, showing integrity under pressure can be tested before deployment.