🔍 Read the full analysis: The New Bottleneck For AI: People Who Can Check Its Work on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
A source article points to a growing mismatch between AI-generated work and the human capacity to verify it, citing examples in mathematics, software and contract workflows. Its figures come from a mix of company analyses, a peer-reviewed study and vendor claims, so the scale is uncertain; the broader question is how organisations preserve expert review and train future reviewers.
AI-generated work is expanding faster than the capacity to check it, according to an analysis drawing on examples from mathematical research, software development and contract work. Its central concern is that generating drafts and results is becoming cheaper, while deciding whether they are correct, relevant and safe to use still depends heavily on scarce human expertise.
The analysis reports that OpenAI published 722 mathematical manuscripts produced by a model posed about 4,000 problems. It says the manuscripts cover 372 families and that the average result took about three hours of compute. Some results were formally checked using Lean, a proof-assistant system; OpenAI cautioned that unformalized results could have issues. The source contrasts this output with the careful verification by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an old Erdős conjecture.
In software, the analysis cites company datasets that report more code being merged alongside increased review pressure. Faros AI found teams merged 98% more pull requests while review time rose 91% across its comparison of lower- and higher-AI-adoption periods. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. Those are findings reported by the firms, not a single controlled comparison of all software teams.
The analysis also cites a peer-reviewed 2026 study in which 61% of AI-agent pull requests received no human review before being merged or closed. In contract work, it describes OpenAI’s partnership with contract-software company Ironclad and says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. The source characterizes that as an improvement over an earlier model, but the score also indicates that substantial evaluation criteria were not met.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Human Review Sets the Deployment Limit
The practical consequence is that an organisation may be able to generate far more drafts, code or research than it can confidently put into use. Review capacity can become the constraint: unexamined outputs may be shipped, while careful teams may see work accumulate in queues. The figures cited suggest both risks, including pull requests that reportedly receive no review and AI-generated changes that take longer to reach a reviewer.
The analysis also raises a workforce issue. Experienced reviewers usually develop judgement by doing the underlying work: engineers learn through writing and debugging code, lawyers through drafting and negotiating, and researchers through producing and examining arguments. If AI tools remove too many entry-level tasks, organisations could weaken the training pipeline for future experts even as they rely more on those experts to validate machine output.
This does not establish that every industry will face the same shortage, or that AI output is generally unreliable. It does identify a management challenge: output volume is not the same as usable work. Employers may need to account for review time, assign responsibility for decisions and preserve opportunities for junior staff to build the knowledge that later supports sound adjudication.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Verification Problem
The source frames mathematics, software and professional contracting as examples of a shared gap between producing an answer and deciding whether it answers the right question. In mathematics, formal systems such as Lean can verify that a proof follows specified rules, but human experts still need to judge whether the theorem is meaningful and whether the statement matches the intended problem. The source describes this distinction as “verification abundance, adjudication scarcity.”
Software tests can show that code behaves as specified, but they cannot by themselves establish that the tests captured the real need. Likewise, a technically valid contract draft may still miss an approval requirement or use an unsuitable clause. The analysis argues that review also carries an accountability function: professionals and institutions, rather than AI systems, remain responsible for many consequential decisions.
The cited software statistics should be read with care. The source notes that Faros AI and LinearB sell code-review products, which may shape how their findings are framed. Their reported figures are not directly interchangeable, and the source does not provide enough detail here to independently assess every method or comparison. The peer-reviewed study is a distinct source, but one result does not establish a universal rate across organisations.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Wide Is the Review Gap?
The source material does not establish that its examples represent every mathematics department, software team or legal practice. The software figures come from different datasets and comparison methods, and some are reported by vendors with commercial interests in review tools. Their precise implications depend on how each study defines AI-generated work, review time and adoption levels.
It is also unclear how much review can be automated without creating new errors or shifting responsibility rather than resolving it. The material provides no full evaluation details for the 11 contract tasks, no breakdown of which criteria GPT-6 Astra missed and no evidence that a 55% average translates directly into real-world contract performance. Nor does it quantify whether AI is reducing the number of junior workers who later become senior reviewers.
The broad claim that checking is becoming more expensive is an interpretation, not a universal measured result. Review may be improved by better tools or workflows in some settings. The available examples show a reported mismatch in particular datasets and tasks; they do not establish a single economy-wide rate or prove that the mismatch will persist.
As an affiliate, we earn on qualifying purchases.
Watch Review Queues and Training
The next useful evidence will come from more transparent, independently assessed measurements of review quality and workload. That includes reporting how long AI-generated work waits, how often it is changed or rejected, what errors reviewers find and whether outcomes differ by task complexity. For high-stakes uses, organisations will also need to say who is responsible for approval and what checks are required before work is relied upon.
For the partnerships and systems cited, the source material gives no specific upcoming release, deadline or new evaluation date. Further results from AI-assisted mathematics, software-review studies and contract evaluations may clarify where automation genuinely reduces review effort and where it mainly increases the number of items requiring attention.
Employers and educators will also face decisions about entry-level work. If routine drafting and coding are increasingly automated, training plans may need to give junior staff structured practice in making, testing and explaining work—not only in reviewing machine output. Whether organisations take those steps, and whether review capacity becomes a lasting limit on AI adoption, remains unresolved.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The analysis says AI output is growing faster than the human capacity to evaluate it, using examples from mathematical manuscripts, software pull requests and contract tasks. It is an analysis of cited developments, not a report of one newly announced industry-wide finding.
Were all 722 mathematical manuscripts verified?
No. The source says some results were formally checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not state that every manuscript was independently verified.
What did the software studies report?
Faros AI reported more merged pull requests alongside longer review time in its adoption-period comparison. LinearB reported longer waits before review began and lower acceptance for AI-generated changes than for human-written ones in its dataset. The figures use different methods and should not be treated as one combined estimate.
Does this prove AI work is less reliable?
No. The cited results point to review and acceptance challenges in particular datasets and evaluations. They do not establish that all AI-generated work is less reliable or that the same pattern applies to every task or organisation.
Why could this affect junior workers?
Experienced reviewers often gain judgement by performing the work themselves. If AI reduces opportunities to draft code, contracts or research, organisations may need deliberate training pathways so junior staff can develop the expertise needed for future review roles.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
