The Referee Shortage Behind AI’s Low-Cost Output
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Referee Shortage Behind AI’s Low-Cost Output on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get smart everyday buys delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI reported producing 722 mathematical manuscripts from about 4,000 problems, while checking and judging results still depend heavily on human expertise. Data cited in the source material points to similar review pressures in software and contract work, though some figures come from companies that sell code-review tools. The scale of the gap and its effects on hiring, quality and accountability remain uncertain.

OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems to a model, according to source material from ThorstenMeyerAI.com. The output highlights a growing challenge across mathematics, software and professional services: AI can produce drafts and results quickly, but human review remains slower and limited, leaving questions about quality, accountability and how much of the work organisations can safely use.

The manuscripts came in 372 problem families, and the average result reportedly took about three hours of computing time. Some results have been formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results without formal verification “could have issues.” The source says one earlier result from the programme, described as a counterexample to an Erdős conjecture, received careful verification from five leading mathematicians. It does not provide enough detail here to establish the status of every manuscript or the full verification process.

Software data cited in the source points to a similar tension. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, which analysed 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes took 4.6 times longer to reach the start of review and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are findings reported by the named providers, not universal rates for all software teams.

A peer-reviewed 2026 study cited in the source found that 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported a 31.3% rise in merges with no review during high-adoption periods. The source notes that several organisations behind these figures sell code-review products, a commercial interest that calls for care when interpreting their measurements.

At a glance
reportWhen: OpenAI’s mathematics results were publi…
The developmentA report on OpenAI’s latest mathematics output and software-review data argues that AI is increasing the volume of work faster than human verification capacity can grow.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets AI’s Real Limits

When generation becomes faster and cheaper but review remains a human task, verification can become the constraint on what an organisation can safely publish, merge or sign. More output does not automatically mean more usable work: each result may still need someone with the right expertise to check whether it is valid, relevant and fit for purpose.

The consequences can include work that is accepted without adequate scrutiny, or good work delayed because reviewers lack time or trust its source. The source says LinearB found 38% of reviewers deliberately deprioritise AI-generated changes. That figure describes the company’s reported findings, but it illustrates a possible trade-off: treating machine-written work as especially risky may protect quality while slowing delivery.

The issue also reaches workforce development. Senior reviewers generally gain judgement through years of doing the work they later assess. If AI takes over too much entry-level drafting or coding, organisations may have fewer opportunities to train future reviewers. That is a concern raised in the source material, not a demonstrated outcome across all workplaces. Still, it points to a practical question for employers: how will they build expertise if junior staff mainly supervise drafts rather than learn by producing them?

Amazon

code review tools for software development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Mathematical Proofs to Contracts

The source frames the review gap across three areas. In mathematics, formal tools such as Lean can verify that a proof follows rules for a stated theorem. They do not by themselves decide whether the theorem addresses the right question, whether the result matters, or what it means for the field. The distinction is between checking a formal result and exercising expert judgement about its significance and correctness in context.

In software, automated tests can show whether code passes the tests that have been written. They cannot prove those tests capture every user need or prevent every defect. Reviewers also have to interpret unfamiliar code and assess its effects. In legal work, a contract can meet many evaluation criteria and still contain a consequential omission. The source reports that OpenAI’s partnership with contract-software company Ironclad involved training GPT-6 Astra on real contracting workflows; across 11 tasks, it met an average of 55% of evaluation criteria, an improvement over the prior model. The remaining criteria still require examination before the work can be relied on.

The shared point is not that AI output is necessarily wrong. It is that validation has different requirements from production. A proof assistant, test suite or model evaluation can help with checking, but each operates within defined assumptions. Human expertise and institutional responsibility still matter where decisions have consequences.

Amazon

mathematics proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Work Is Being Checked?

The source material does not state how many of the 722 manuscripts have been independently reviewed, how many were formally verified, or how the overall quality compares with work produced through conventional research. Nor does it establish whether the reported software figures apply across industries, programming languages or different levels of AI adoption.

Several software statistics come from companies that sell tools for code review, and their samples and measurement methods may differ. The source flags this commercial interest but does not provide study links, definitions or enough methodological detail to assess every result independently. The reported contract-model score also covers 11 tasks; it does not show how performance varies across contract types or how often human reviewers catch consequential errors.

It is also unclear whether AI systems will eventually reduce the time needed for review as effectively as they reduce drafting time. Automated verification may help, but the source argues that it cannot settle questions of relevance, intent or accountability. The size of any long-term reviewer shortage remains unquantified.

Amazon

AI-powered pull request review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Review and Training

The next useful evidence will show not only how many drafts, proofs or code changes AI systems produce, but also how they are reviewed and corrected. For mathematics, that means clearer reporting on formal verification and independent expert assessment. For software, comparable data across organisations could clarify review times, defect rates and what “no review” means in practice. In professional services, evaluations need to show performance across a wider range of tasks and the role human sign-off plays.

Employers will also need to decide how to preserve training routes for junior staff while using AI tools. The source does not identify a new policy or scheduled milestone that resolves this issue. For now, the central test is whether productivity gains translate into work that can be checked, defended and responsibly used—not simply whether AI can produce more of it.

Amazon

formal verification software for mathematicians

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI publish?

According to the source material, OpenAI published 722 mathematical manuscripts produced after its model was posed about 4,000 problems. They were grouped into 372 problem families. The source says some results were formally checked in Lean, while others were not.

Does formal verification prove a mathematical result is important?

No. A proof assistant can check that a proof establishes its stated theorem under the system’s rules. It does not decide whether the theorem asks the right question or matters to the field. Those judgements still require mathematical expertise.

What do the software statistics show?

Companies and a peer-reviewed study cited in the source report that AI adoption is associated with more pull requests, longer waits for review and, in some cases, changes merged without human review. The results have different samples and methods and should not be treated as universal rates.

Why might AI increase pressure on reviewers?

AI can produce more drafts or code changes, but reviewers still have to check whether they are correct and fit for purpose. That work may require subject knowledge and responsibility that automated tests or formal tools do not fully supply.

Is there proof that AI is causing a long-term shortage of experts?

No. The source raises a concern that reducing junior staff’s opportunities to practise could weaken future reviewer training, but it does not establish that a shortage has already occurred or quantify its likely scale.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Saturation. The ten-essay framework, closed.

The ten-essay framework on European sovereign AI has reached a saturation point, with no new structural insights expected before August 2026, marking a strategic editorial closure.

Unlocking Humanity’s Potential By Favoring The Best AI Model Over Sovereignty

A detailed analysis of why organizations should favor top AI models over sovereignty concerns, highlighting costs, risks, and strategic implications.

The Truth Behind August 2 And AI’s Future

Key developments on August 2, 2026, reveal deferred AI compliance deadlines and ongoing obligations under the EU AI Act, shaping AI regulation’s future.

Week Three — Foundation model vs Brownian motion. Kronos on five-minute BTC.

Kronos, a foundation model for financial time series, does not outperform the traditional Brownian motion model in short-term BTC predictions, according to recent testing.