AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The New Bottleneck For AI: People Who Can Check Its Work on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source article points to a growing mismatch between AI-generated work and the human capacity to verify it, citing examples in mathematics, software and contract workflows. Its figures come from a mix of company analyses, a peer-reviewed study and vendor claims, so the scale is uncertain; the broader question is how organisations preserve expert review and train future reviewers.

AI-generated work is expanding faster than the capacity to check it, according to an analysis drawing on examples from mathematical research, software development and contract work. Its central concern is that generating drafts and results is becoming cheaper, while deciding whether they are correct, relevant and safe to use still depends heavily on scarce human expertise.

The analysis reports that OpenAI published 722 mathematical manuscripts produced by a model posed about 4,000 problems. It says the manuscripts cover 372 families and that the average result took about three hours of compute. Some results were formally checked using Lean, a proof-assistant system; OpenAI cautioned that unformalized results could have issues. The source contrasts this output with the careful verification by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an old Erdős conjecture.

In software, the analysis cites company datasets that report more code being merged alongside increased review pressure. Faros AI found teams merged 98% more pull requests while review time rose 91% across its comparison of lower- and higher-AI-adoption periods. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. Those are findings reported by the firms, not a single controlled comparison of all software teams.

The analysis also cites a peer-reviewed 2026 study in which 61% of AI-agent pull requests received no human review before being merged or closed. In contract work, it describes OpenAI’s partnership with contract-software company Ironclad and says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. The source characterizes that as an improvement over an earlier model, but the score also indicates that substantial evaluation criteria were not met.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA report argues that the expanding volume of AI-generated research, code and contract work is making qualified human review a constraint.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Human Review Sets the Deployment Limit

The practical consequence is that an organisation may be able to generate far more drafts, code or research than it can confidently put into use. Review capacity can become the constraint: unexamined outputs may be shipped, while careful teams may see work accumulate in queues. The figures cited suggest both risks, including pull requests that reportedly receive no review and AI-generated changes that take longer to reach a reviewer.

The analysis also raises a workforce issue. Experienced reviewers usually develop judgement by doing the underlying work: engineers learn through writing and debugging code, lawyers through drafting and negotiating, and researchers through producing and examining arguments. If AI tools remove too many entry-level tasks, organisations could weaken the training pipeline for future experts even as they rely more on those experts to validate machine output.

This does not establish that every industry will face the same shortage, or that AI output is generally unreliable. It does identify a management challenge: output volume is not the same as usable work. Employers may need to account for review time, assign responsibility for decisions and preserve opportunities for junior staff to build the knowledge that later supports sound adjudication.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Verification Problem

The source frames mathematics, software and professional contracting as examples of a shared gap between producing an answer and deciding whether it answers the right question. In mathematics, formal systems such as Lean can verify that a proof follows specified rules, but human experts still need to judge whether the theorem is meaningful and whether the statement matches the intended problem. The source describes this distinction as “verification abundance, adjudication scarcity.”

Software tests can show that code behaves as specified, but they cannot by themselves establish that the tests captured the real need. Likewise, a technically valid contract draft may still miss an approval requirement or use an unsuitable clause. The analysis argues that review also carries an accountability function: professionals and institutions, rather than AI systems, remain responsible for many consequential decisions.

The cited software statistics should be read with care. The source notes that Faros AI and LinearB sell code-review products, which may shape how their findings are framed. Their reported figures are not directly interchangeable, and the source does not provide enough detail here to independently assess every method or comparison. The peer-reviewed study is a distinct source, but one result does not establish a universal rate across organisations.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Wide Is the Review Gap?

The source material does not establish that its examples represent every mathematics department, software team or legal practice. The software figures come from different datasets and comparison methods, and some are reported by vendors with commercial interests in review tools. Their precise implications depend on how each study defines AI-generated work, review time and adoption levels.

It is also unclear how much review can be automated without creating new errors or shifting responsibility rather than resolving it. The material provides no full evaluation details for the 11 contract tasks, no breakdown of which criteria GPT-6 Astra missed and no evidence that a 55% average translates directly into real-world contract performance. Nor does it quantify whether AI is reducing the number of junior workers who later become senior reviewers.

The broad claim that checking is becoming more expensive is an interpretation, not a universal measured result. Review may be improved by better tools or workflows in some settings. The available examples show a reported mismatch in particular datasets and tasks; they do not establish a single economy-wide rate or prove that the mismatch will persist.

Amazon

contract review AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch Review Queues and Training

The next useful evidence will come from more transparent, independently assessed measurements of review quality and workload. That includes reporting how long AI-generated work waits, how often it is changed or rejected, what errors reviewers find and whether outcomes differ by task complexity. For high-stakes uses, organisations will also need to say who is responsible for approval and what checks are required before work is relied upon.

For the partnerships and systems cited, the source material gives no specific upcoming release, deadline or new evaluation date. Further results from AI-assisted mathematics, software-review studies and contract evaluations may clarify where automation genuinely reduces review effort and where it mainly increases the number of items requiring attention.

Employers and educators will also face decisions about entry-level work. If routine drafting and coding are increasingly automated, training plans may need to give junior staff structured practice in making, testing and explaining work—not only in reviewing machine output. Whether organisations take those steps, and whether review capacity becomes a lasting limit on AI adoption, remains unresolved.

Amazon

proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

The analysis says AI output is growing faster than the human capacity to evaluate it, using examples from mathematical manuscripts, software pull requests and contract tasks. It is an analysis of cited developments, not a report of one newly announced industry-wide finding.

Were all 722 mathematical manuscripts verified?

No. The source says some results were formally checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not state that every manuscript was independently verified.

What did the software studies report?

Faros AI reported more merged pull requests alongside longer review time in its adoption-period comparison. LinearB reported longer waits before review began and lower acceptance for AI-generated changes than for human-written ones in its dataset. The figures use different methods and should not be treated as one combined estimate.

Does this prove AI work is less reliable?

No. The cited results point to review and acceptance challenges in particular datasets and evaluations. They do not establish that all AI-generated work is less reliable or that the same pattern applies to every task or organisation.

Why could this affect junior workers?

Experienced reviewers often gain judgement by performing the work themselves. If AI reduces opportunities to draft code, contracts or research, organisations may need deliberate training pathways so junior staff can develop the expertise needed for future review roles.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stripe’s Vision: Prioritizing AI Development Over Traditional Models

Stripe acquires OpenRouter for an estimated $7.5 billion, signaling a focus on AI token metering rather than routing technology, reshaping AI infrastructure.

Show HN: Bramble – Local-first Password Manager

Bramble, an open source password manager with peer-to-peer sync, has released Android and iOS apps, expanding its cross-device local-first approach.

Apple Silicon’s Quiet Memory Advantage

Apple Silicon’s unified memory architecture offers a significant capacity advantage for large AI models, despite lower bandwidth and speed compared to NVIDIA GPUs.

The Real Cost of a Local-Inference Rig in 2026

Analyzing the hardware costs, VRAM constraints, and strategic choices for running AI models locally in 2026.