🔍 Read the full analysis: The Referee Shortage Behind AI’s Low-Cost Output on ThorstenMeyerAI.com
Get smart everyday buys delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI reported producing 722 mathematical manuscripts from about 4,000 problems, while checking and judging results still depend heavily on human expertise. Data cited in the source material points to similar review pressures in software and contract work, though some figures come from companies that sell code-review tools. The scale of the gap and its effects on hiring, quality and accountability remain uncertain.
OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems to a model, according to source material from ThorstenMeyerAI.com. The output highlights a growing challenge across mathematics, software and professional services: AI can produce drafts and results quickly, but human review remains slower and limited, leaving questions about quality, accountability and how much of the work organisations can safely use.
The manuscripts came in 372 problem families, and the average result reportedly took about three hours of computing time. Some results have been formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results without formal verification “could have issues.” The source says one earlier result from the programme, described as a counterexample to an Erdős conjecture, received careful verification from five leading mathematicians. It does not provide enough detail here to establish the status of every manuscript or the full verification process.
Software data cited in the source points to a similar tension. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, which analysed 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes took 4.6 times longer to reach the start of review and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are findings reported by the named providers, not universal rates for all software teams.
A peer-reviewed 2026 study cited in the source found that 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported a 31.3% rise in merges with no review during high-adoption periods. The source notes that several organisations behind these figures sell code-review products, a commercial interest that calls for care when interpreting their measurements.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets AI’s Real Limits
When generation becomes faster and cheaper but review remains a human task, verification can become the constraint on what an organisation can safely publish, merge or sign. More output does not automatically mean more usable work: each result may still need someone with the right expertise to check whether it is valid, relevant and fit for purpose.
The consequences can include work that is accepted without adequate scrutiny, or good work delayed because reviewers lack time or trust its source. The source says LinearB found 38% of reviewers deliberately deprioritise AI-generated changes. That figure describes the company’s reported findings, but it illustrates a possible trade-off: treating machine-written work as especially risky may protect quality while slowing delivery.
The issue also reaches workforce development. Senior reviewers generally gain judgement through years of doing the work they later assess. If AI takes over too much entry-level drafting or coding, organisations may have fewer opportunities to train future reviewers. That is a concern raised in the source material, not a demonstrated outcome across all workplaces. Still, it points to a practical question for employers: how will they build expertise if junior staff mainly supervise drafts rather than learn by producing them?
code review tools for software development
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Mathematical Proofs to Contracts
The source frames the review gap across three areas. In mathematics, formal tools such as Lean can verify that a proof follows rules for a stated theorem. They do not by themselves decide whether the theorem addresses the right question, whether the result matters, or what it means for the field. The distinction is between checking a formal result and exercising expert judgement about its significance and correctness in context.
In software, automated tests can show whether code passes the tests that have been written. They cannot prove those tests capture every user need or prevent every defect. Reviewers also have to interpret unfamiliar code and assess its effects. In legal work, a contract can meet many evaluation criteria and still contain a consequential omission. The source reports that OpenAI’s partnership with contract-software company Ironclad involved training GPT-6 Astra on real contracting workflows; across 11 tasks, it met an average of 55% of evaluation criteria, an improvement over the prior model. The remaining criteria still require examination before the work can be relied on.
The shared point is not that AI output is necessarily wrong. It is that validation has different requirements from production. A proof assistant, test suite or model evaluation can help with checking, but each operates within defined assumptions. Human expertise and institutional responsibility still matter where decisions have consequences.
mathematics proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Much Work Is Being Checked?
The source material does not state how many of the 722 manuscripts have been independently reviewed, how many were formally verified, or how the overall quality compares with work produced through conventional research. Nor does it establish whether the reported software figures apply across industries, programming languages or different levels of AI adoption.
Several software statistics come from companies that sell tools for code review, and their samples and measurement methods may differ. The source flags this commercial interest but does not provide study links, definitions or enough methodological detail to assess every result independently. The reported contract-model score also covers 11 tasks; it does not show how performance varies across contract types or how often human reviewers catch consequential errors.
It is also unclear whether AI systems will eventually reduce the time needed for review as effectively as they reduce drafting time. Automated verification may help, but the source argues that it cannot settle questions of relevance, intent or accountability. The size of any long-term reviewer shortage remains unquantified.
AI-powered pull request review tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Tracking Review and Training
The next useful evidence will show not only how many drafts, proofs or code changes AI systems produce, but also how they are reviewed and corrected. For mathematics, that means clearer reporting on formal verification and independent expert assessment. For software, comparable data across organisations could clarify review times, defect rates and what “no review” means in practice. In professional services, evaluations need to show performance across a wider range of tasks and the role human sign-off plays.
Employers will also need to decide how to preserve training routes for junior staff while using AI tools. The source does not identify a new policy or scheduled milestone that resolves this issue. For now, the central test is whether productivity gains translate into work that can be checked, defended and responsibly used—not simply whether AI can produce more of it.
formal verification software for mathematicians
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI publish?
According to the source material, OpenAI published 722 mathematical manuscripts produced after its model was posed about 4,000 problems. They were grouped into 372 problem families. The source says some results were formally checked in Lean, while others were not.
Does formal verification prove a mathematical result is important?
No. A proof assistant can check that a proof establishes its stated theorem under the system’s rules. It does not decide whether the theorem asks the right question or matters to the field. Those judgements still require mathematical expertise.
What do the software statistics show?
Companies and a peer-reviewed study cited in the source report that AI adoption is associated with more pull requests, longer waits for review and, in some cases, changes merged without human review. The results have different samples and methods and should not be treated as universal rates.
Why might AI increase pressure on reviewers?
AI can produce more drafts or code changes, but reviewers still have to check whether they are correct and fit for purpose. That work may require subject knowledge and responsibility that automated tests or formal tools do not fully supply.
Is there proof that AI is causing a long-term shortage of experts?
No. The source raises a concern that reducing junior staff’s opportunities to practise could weaken future reviewer training, but it does not establish that a shortage has already occurred or quantify its likely scale.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
