AI’s Output Is Abundant; Careful Review Isn’t
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI’s Output Is Abundant; Careful Review Isn’t on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source essay argues that AI has made it cheaper to produce mathematical manuscripts, software changes and contract drafts, but has not made expert review equally fast. The figures it cites point to a growing review bottleneck, though some software metrics come from companies that sell review tools and should be read with that in mind.

A recent analysis argues that AI is making work cheaper to produce faster than it can be checked, drawing on OpenAI’s mathematical manuscripts, software-development metrics and a contract-workflow evaluation. The evidence suggests that review capacity is becoming a constraint, although the cited measurements come from different sources and are not directly comparable.

The analysis says OpenAI published 722 mathematical manuscripts this week, generated from work on about 4,000 problems and grouped into 372 families. It reports that the average result took about three hours of compute. Some results have been formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results not formalized in Lean could have issues, according to the source material.

The author contrasts that volume with the response to an earlier result from the same programme: a proposed counterexample to an old Erdős conjecture that was carefully checked by five leading mathematicians. The comparison illustrates the central concern: computer systems can produce many candidate results, while expert attention to determine their correctness and significance remains limited. The analysis cites the phrase “verification abundance, adjudication scarcity” to describe the mismatch.

Software figures point in a similar direction, but require care. Faros AI reported that teams in high-AI-adoption periods merged 98% more pull requests, while review time increased 91%. LinearB, examining 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source notes that several data providers sell code-review tools, so their findings should be treated as attributed measurements, not neutral industry-wide estimates.

At a glance
analysisWhen: Published this week; cited studies and…
The developmentA published analysis brings together examples from mathematics, software and contracting to argue that AI output is outpacing the human capacity to verify it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Becomes the Constraint

If AI increases the supply of drafts, proofs and code without increasing the supply of people able to assess them, organisations may not be able to use all the output safely or effectively. The scarce resource could shift from making a first version to deciding whether that version is correct, relevant and ready to rely on.

The practical consequences vary. A mathematical proof may need specialists to test both its logic and the importance of the statement it proves. Software teams may face longer queues or accept changes with insufficient review. In contract work, a draft still needs a qualified person to catch a missed approval rule or an unsuitable clause. The source says OpenAI’s GPT-6 Astra, developed with contract-software company Ironclad, met 55% of evaluation criteria on average across 11 tasks. That is described as a large improvement over the prior model, but leaves criteria that still need attention before the output can be relied on.

The analysis also identifies a workforce concern: junior staff often develop the judgment needed for senior review by doing the underlying work themselves. If AI takes over too much drafting or coding, fewer workers may get the experience that later enables them to assess complex output. This is a risk raised by the analysis, not a demonstrated outcome in the cited figures.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Mismatch

The source’s argument draws on three distinct settings. In mathematics, formal proof tools can check whether a proof establishes its stated theorem, but they do not decide whether the theorem is the right one to pursue or what it contributes. Human mathematicians still make those judgments. The source says the disputed earlier counterexample illustrates how a result can require scrutiny beyond producing a proof.

In software, automated tests can check specified behavior, but cannot establish that the tests capture every requirement. The source cites a peer-reviewed 2026 study reporting that 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported that merges with zero review rose 31.3% during high-adoption periods. These are separate findings with different methods; they should not be combined into a single estimate.

In professional work, a model can draft or classify documents, but people and institutions retain responsibility for decisions. The analysis points to contracts signed by people, engineering designs approved by licensed professionals and research authors expected to answer for their work. Its broader point is that checking includes accountability, not just finding technical errors.

Amazon

formal verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Missing?

The cited measurements do not establish a single, representative rate of AI review failure across industries. The sources use different samples, definitions and time periods, and the software metrics include reports from vendors with commercial interests in review tools. The direction may be consistent across these sources, but the size and prevalence of the bottleneck remain uncertain.

The source material also does not provide the full methods behind the mathematics and contract-workflow claims, or the detailed comparison with earlier models. It is not clear how much human review the 722 manuscripts received, how the 11 contract tasks were weighted, or whether the evaluation criteria reflect the risks of real legal work. Nor do the figures show whether added review time led to safer outcomes, slower delivery, or both.

Finally, it remains uncertain whether new tools, changes to team processes or different training practices can expand expert review capacity quickly enough. The analysis makes a forecast about the value of people who can approve and take responsibility for AI-supported work; it does not establish how wages or hiring will change.

Amazon

mathematical proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Measuring Review Alongside Output

The next useful evidence will show not only how much work AI systems generate, but how it is reviewed and what errors remain. That means clearer reporting on review rates, time spent checking, changes requested, defects found and outcomes after deployment, with methods and comparison periods disclosed.

For the mathematics programme, readers will need to see which results receive formal verification and what independent experts conclude about their correctness and contribution. For software and contracting, further studies could compare AI-assisted and conventional work under the same review standards, while distinguishing vendor reports from independent or peer-reviewed research.

Organisations adopting these systems will also need to track whether junior staff still gain enough hands-on experience to develop judgment. The available material does not establish a timetable for resolving the mismatch. For now, the central question is whether review processes can expand without treating automated output as reliable simply because it is plentiful.

Amazon

software review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

A recent analysis combines examples from mathematics, software and contract work to argue that AI-generated output is growing faster than expert review capacity. It is an analysis of reported developments, not a new controlled study of all three fields.

Did OpenAI publish 722 fully verified mathematical results?

The source says OpenAI published 722 mathematical manuscripts, with some results formally checked in Lean. OpenAI cautioned that some results without formalization could have issues; the source does not say that every manuscript was independently verified.

What do the software figures show?

Reports cited by the analysis describe more pull requests, longer waits for review, lower acceptance rates for AI-generated changes in one dataset and substantial numbers of agent pull requests without human review in a separate study. Their methods and samples differ, and some sources sell review tools.

Why can’t AI simply check AI output?

Automated checks can test defined properties, such as whether a proof follows formal rules or code passes specified tests. They may not establish whether the original question, test or requirement was the right one, and they do not take legal or professional responsibility for the result.

What remains uncertain?

The available figures do not establish how widespread the review bottleneck is across the economy, whether it causes worse outcomes overall, or how quickly institutions can add review capacity. The cited evidence also does not settle how AI adoption will affect the training of junior professionals.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Solo Performance Business Tools That Make Show Days Easier

New tools using AI streamline show-day preparations for solo performers, reducing errors and saving time on gig day.

Networking in the Art World

Leverage your connections in the art world to unlock opportunities for growth and recognition; discover how to enhance your networking skills effectively.

The Tech Signal Monitor: RISC OS Open’s Two-Decade Journey

RISC OS Open marks two decades of development, highlighting ongoing platform enhancements and community efforts amid evolving tech landscapes.

Incident postmortem builder for managed service providers

A new incident postmortem builder tailored for small managed service providers is in testing, aiming to streamline post-incident reports and improve client communication.