The Verification Gap: AI Makes, Humans Have To Check
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Verification Gap: AI Makes, Humans Have To Check on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get office and shipping supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems to a model, while the source account says the work still requires expert scrutiny. Data cited from software teams shows AI-related output can bring longer review waits and, in some cases, less human review. The figures suggest verification capacity may constrain how much AI-generated work organisations can safely use, though some studies come from companies that sell review tools.

OpenAI published 722 mathematical manuscripts this week, after a model was given about 4,000 problems, highlighting a growing challenge as AI produces work faster than people can check it. The source account describes one earlier result from the same programme—a proposed counterexample to an Erdős conjecture—as receiving careful review from five leading mathematicians, illustrating the effort required to establish whether AI-generated results are sound and useful.

According to the source material, the 722 manuscripts cover 372 problem families, and the average result took about three hours of computing time to produce. Some results have been formally checked using Lean, a proof assistant. OpenAI cautioned that some results without formal verification could have issues. The source does not give a complete breakdown of how many manuscripts have been formally checked or independently assessed.

The same verification problem appears in software metrics cited by the source. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are findings attributed to the named firms, not universal measures of software work.

A peer-reviewed 2026 study cited in the source found that 61% of AI-agent pull requests received no human review before being merged or closed. The source also describes OpenAI’s partnership with contract-software company Ironclad: on 11 tasks, a model called GPT-6 Astra met an average of 55% of evaluation criteria. That result indicates improvement over its predecessor, according to the source, but leaves the remaining criteria for people or other systems to assess.

At a glance
reportWhen: This week; software and workflow figure…
The developmentOpenAI’s release of 722 mathematical manuscripts has renewed attention to a broader verification gap, with human review struggling to keep pace with AI-generated work.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity May Limit AI Use

The practical issue is not just how much work AI can produce, but how much an organisation can check, approve and take responsibility for. If review queues grow faster than output, companies may face delayed releases, errors that pass unchecked, or reviewers who have to prioritise work without enough information about its quality.

The source points to three possible responses already visible in the cited material: work can be merged without review, reviewers can push AI-generated work down the queue, or the producer can decide which outputs deserve attention. Each has risks. Skipping review leaves errors undiscovered; broad suspicion can delay useful work; and relying on a producer’s own selection does not replace independent scrutiny.

This may also affect hiring and training. Experienced reviewers learn through doing the underlying work: writing software, constructing proofs or drafting contracts. If AI absorbs much of that early-career work, organisations could weaken the route by which people develop the judgement needed to assess more advanced AI output. The source argues that skilled reviewers may become a constraint on adoption, but it provides no labour-market data establishing the scale or timing of that effect.

Amazon

AI verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Review Pressures

In mathematics, formal verification can establish that a proof follows from its stated assumptions. It cannot by itself establish that the theorem addresses the right question, that the result is significant, or that its assumptions fit the problem people intended to solve. The source describes this distinction as a gap between checking a proof and adjudicating its meaning.

Software has a related limitation. Tests can show that code passes the tests written for it, but cannot prove the tests cover every requirement or expose every failure. The source says reviewers find AI-generated code taxing to check because it can look clean while offering little indication of where a mistake might be. The cited metrics come from different studies and organisations, so they should not be treated as a single, directly comparable measurement.

In professional workflows, a model can meet some evaluation criteria without satisfying all the requirements for use. Contract review, for example, can depend on jurisdiction, approval rules and the facts of a particular agreement. The source’s Ironclad example is a limited assessment across 11 tasks, not evidence that the model is ready to make contractual decisions without human oversight.

Amazon

proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Work Has Been Checked?

The source material does not specify how many of the 722 manuscripts have been formally checked, independently reviewed or corrected since publication. It also does not provide enough detail about the five mathematicians’ review of the earlier proposed counterexample to establish the scope or final status of that assessment.

The software figures come from separate datasets and measures, and some cited providers sell code-review products. That does not invalidate their findings, but the figures need to be read with their methods and commercial interests in mind. The source does not give study designs, time windows or comparison baselines for every statistic, so they cannot establish a single industry-wide effect.

It is also unclear whether AI review systems can reduce the human burden at a comparable pace to the growth in generated work. Automated checks can catch some errors, but the source argues that deciding whether a task was framed correctly and assigning responsibility remain human or institutional functions. The size of any effect on jobs, entry-level training and reviewer workloads has not been quantified here.

Amazon

code review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Needed on Review Outcomes

The immediate test is whether the new mathematical manuscripts receive formal checks and independent expert scrutiny, and whether concerns or corrections are reported. The source material does not give a schedule for that work. Clear reporting on which results were checked, by what methods and with what outcomes would help readers distinguish generated output from established findings.

For software, organisations and researchers will need to track not only how many changes AI helps produce, but also review wait times, defect rates, reversals and the share of work merged without human assessment. Comparisons should state the time period, method and baseline so the results can be interpreted consistently. In professional workflows, evaluations should identify which criteria models miss and who is accountable for checking them.

Longer term, employers and training programmes will have to decide how junior workers can gain practical experience if AI handles more first drafts and routine tasks. The figures cited here describe a possible mismatch between production and review, not a settled forecast. Whether human verification becomes a lasting bottleneck depends on how AI checking tools, workplace practices and training develop.

Amazon

formal verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI publish this week?

According to the source, OpenAI published 722 mathematical manuscripts across 372 problem families after a model was given about 4,000 problems. Some results have been formally checked in Lean; OpenAI cautioned that unformalized results could have issues.

Does formal verification prove an AI result is useful or correct for its intended purpose?

Formal verification can check that a proof or program meets a specified formal standard. It does not by itself establish that the original question, assumptions or tests match the real-world need.

What do the software figures show?

The source cites separate findings from Faros AI, LinearB and a peer-reviewed 2026 study. They report increased output alongside longer review waits or limited human review in some settings. Their methods and populations differ, so the figures are not a single measure of all software teams.

Is human review becoming unnecessary as AI improves?

The cited evidence does not establish that. Automated checks can assess defined requirements, but questions about whether the requirements are right and who accepts responsibility remain unresolved in the examples described.

What remains unknown about the manuscripts?

The source does not state how many of the 722 manuscripts have received formal verification or independent expert review, or provide a complete account of corrections and review outcomes.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple earnings: Tim Cook is heading out on top as stock surges to lead the Mag 7 in 2026

Apple’s stock has surged to lead the Mag 7, driven by record earnings under CEO Tim Cook, signaling strong investor confidence amid positive financial results.

The Future Of Microphone Technology: Top AI Picks For 2026

Exploring the latest AI-powered microphone technologies set to transform audio recording and communication in 2026.

Transform Your Viewing Experience With AI Soundbars In 2026

Discover how AI-powered soundbars are revolutionizing home entertainment in 2026, offering immersive sound, smart features, and seamless integration.

DeepSeek-V4-Flash-High And Its Ninth Point: The Future Of Cheap AI Proofs

DeepSeek-V4-Flash-High’s recent update shows significant capability gains through post-training, highlighting a shift in AI development costs and strategies.