🔍 Read the full analysis: The Verification Gap: AI Makes, Humans Have To Check on ThorstenMeyerAI.com
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems to a model, while the source account says the work still requires expert scrutiny. Data cited from software teams shows AI-related output can bring longer review waits and, in some cases, less human review. The figures suggest verification capacity may constrain how much AI-generated work organisations can safely use, though some studies come from companies that sell review tools.
OpenAI published 722 mathematical manuscripts this week, after a model was given about 4,000 problems, highlighting a growing challenge as AI produces work faster than people can check it. The source account describes one earlier result from the same programme—a proposed counterexample to an Erdős conjecture—as receiving careful review from five leading mathematicians, illustrating the effort required to establish whether AI-generated results are sound and useful.
According to the source material, the 722 manuscripts cover 372 problem families, and the average result took about three hours of computing time to produce. Some results have been formally checked using Lean, a proof assistant. OpenAI cautioned that some results without formal verification could have issues. The source does not give a complete breakdown of how many manuscripts have been formally checked or independently assessed.
The same verification problem appears in software metrics cited by the source. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are findings attributed to the named firms, not universal measures of software work.
A peer-reviewed 2026 study cited in the source found that 61% of AI-agent pull requests received no human review before being merged or closed. The source also describes OpenAI’s partnership with contract-software company Ironclad: on 11 tasks, a model called GPT-6 Astra met an average of 55% of evaluation criteria. That result indicates improvement over its predecessor, according to the source, but leaves the remaining criteria for people or other systems to assess.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity May Limit AI Use
The practical issue is not just how much work AI can produce, but how much an organisation can check, approve and take responsibility for. If review queues grow faster than output, companies may face delayed releases, errors that pass unchecked, or reviewers who have to prioritise work without enough information about its quality.
The source points to three possible responses already visible in the cited material: work can be merged without review, reviewers can push AI-generated work down the queue, or the producer can decide which outputs deserve attention. Each has risks. Skipping review leaves errors undiscovered; broad suspicion can delay useful work; and relying on a producer’s own selection does not replace independent scrutiny.
This may also affect hiring and training. Experienced reviewers learn through doing the underlying work: writing software, constructing proofs or drafting contracts. If AI absorbs much of that early-career work, organisations could weaken the route by which people develop the judgement needed to assess more advanced AI output. The source argues that skilled reviewers may become a constraint on adoption, but it provides no labour-market data establishing the scale or timing of that effect.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Review Pressures
In mathematics, formal verification can establish that a proof follows from its stated assumptions. It cannot by itself establish that the theorem addresses the right question, that the result is significant, or that its assumptions fit the problem people intended to solve. The source describes this distinction as a gap between checking a proof and adjudicating its meaning.
Software has a related limitation. Tests can show that code passes the tests written for it, but cannot prove the tests cover every requirement or expose every failure. The source says reviewers find AI-generated code taxing to check because it can look clean while offering little indication of where a mistake might be. The cited metrics come from different studies and organisations, so they should not be treated as a single, directly comparable measurement.
In professional workflows, a model can meet some evaluation criteria without satisfying all the requirements for use. Contract review, for example, can depend on jurisdiction, approval rules and the facts of a particular agreement. The source’s Ironclad example is a limited assessment across 11 tasks, not evidence that the model is ready to make contractual decisions without human oversight.
As an affiliate, we earn on qualifying purchases.
How Much Work Has Been Checked?
The source material does not specify how many of the 722 manuscripts have been formally checked, independently reviewed or corrected since publication. It also does not provide enough detail about the five mathematicians’ review of the earlier proposed counterexample to establish the scope or final status of that assessment.
The software figures come from separate datasets and measures, and some cited providers sell code-review products. That does not invalidate their findings, but the figures need to be read with their methods and commercial interests in mind. The source does not give study designs, time windows or comparison baselines for every statistic, so they cannot establish a single industry-wide effect.
It is also unclear whether AI review systems can reduce the human burden at a comparable pace to the growth in generated work. Automated checks can catch some errors, but the source argues that deciding whether a task was framed correctly and assigning responsibility remain human or institutional functions. The size of any effect on jobs, entry-level training and reviewer workloads has not been quantified here.
As an affiliate, we earn on qualifying purchases.
Evidence Needed on Review Outcomes
The immediate test is whether the new mathematical manuscripts receive formal checks and independent expert scrutiny, and whether concerns or corrections are reported. The source material does not give a schedule for that work. Clear reporting on which results were checked, by what methods and with what outcomes would help readers distinguish generated output from established findings.
For software, organisations and researchers will need to track not only how many changes AI helps produce, but also review wait times, defect rates, reversals and the share of work merged without human assessment. Comparisons should state the time period, method and baseline so the results can be interpreted consistently. In professional workflows, evaluations should identify which criteria models miss and who is accountable for checking them.
Longer term, employers and training programmes will have to decide how junior workers can gain practical experience if AI handles more first drafts and routine tasks. The figures cited here describe a possible mismatch between production and review, not a settled forecast. Whether human verification becomes a lasting bottleneck depends on how AI checking tools, workplace practices and training develop.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI publish this week?
According to the source, OpenAI published 722 mathematical manuscripts across 372 problem families after a model was given about 4,000 problems. Some results have been formally checked in Lean; OpenAI cautioned that unformalized results could have issues.
Does formal verification prove an AI result is useful or correct for its intended purpose?
Formal verification can check that a proof or program meets a specified formal standard. It does not by itself establish that the original question, assumptions or tests match the real-world need.
What do the software figures show?
The source cites separate findings from Faros AI, LinearB and a peer-reviewed 2026 study. They report increased output alongside longer review waits or limited human review in some settings. Their methods and populations differ, so the figures are not a single measure of all software teams.
Is human review becoming unnecessary as AI improves?
The cited evidence does not establish that. Automated checks can assess defined requirements, but questions about whether the requirements are right and who accepts responsibility remain unresolved in the examples described.
What remains unknown about the manuscripts?
The source does not state how many of the 722 manuscripts have received formal verification or independent expert review, or provide a complete account of corrections and review outcomes.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
