AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI Changed The Economics Of Making And Checking on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the little things that make your day delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

AI systems can generate mathematical manuscripts, software changes and contract work faster and more cheaply, but the supplied report says expert review remains slow and limited. Figures from software studies point to longer waits and less human review, though several sources sell review tools and the data does not establish one universal effect.

A report published this week argues that AI has made it cheaper to produce work in fields including mathematics and software, while human capacity to check that work remains limited. Its examples range from 722 mathematical manuscripts attributed to OpenAI to software-industry data showing increased review delays, though the figures come from separate sources and do not establish a single, economy-wide trend.

The report says OpenAI’s system was given about 4,000 mathematical problems and produced 722 manuscripts across 372 families, with an average result taking about three hours of compute. Some results were checked using the Lean proof assistant. OpenAI cautioned that unformalized results “could have issues,” according to the source. The report contrasts this volume with the careful verification by five leading mathematicians of an earlier result from the same programme: a proposed counterexample to an old Erdős conjecture. It does not provide the names of those mathematicians or details of their review process.

In software, the source cites several studies with different samples and measures. Faros AI reported that teams in high-AI-adoption periods merged 98% more pull requests, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. The measures are not directly interchangeable.

The report also points to OpenAI’s partnership with contract-software company Ironclad, saying GPT-6 Astra was evaluated on 11 contracting tasks and met 55% of evaluation criteria on average. That result is described as an improvement over a previous model, but the source gives no earlier score, task-level breakdown or independent assessment. The remaining criteria indicate potential gaps in performance; they do not, by themselves, show how often errors would occur in actual legal work.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA report argues that AI is widening the gap between the cost of producing work and the capacity to verify it, citing examples from mathematics, software and contracting.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the Pace

If these examples reflect a broader workplace pattern, the constraint on using AI may shift from producing drafts to checking, correcting and taking responsibility for them. More output does not automatically mean more usable output: organisations still need people who can decide whether a result answers the right question, fits the situation and is safe to rely on.

The report describes several possible responses when review cannot keep pace: work may be merged without review, reviewers may delay machine-generated submissions, or producers may decide which results merit attention. Each has costs. Unchecked output can carry errors forward; blanket suspicion can slow useful work; and relying on the producer’s own selection can leave independent scrutiny thin. These are risks identified by the source, not proof that every organisation is experiencing them.

That imbalance could also affect staffing and training. Experienced reviewers typically develop judgment through years of doing the underlying work. If entry-level employees mainly edit AI output rather than learning to write code, draft contracts or prove results themselves, employers could weaken the future supply of people qualified to review complex work. The source frames expert judgment as a possible economic advantage, but offers no wage or hiring data to measure such a “referee premium.”

Amazon

AI review tools for software development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Different Checks

The examples share a broad question—how to verify abundant machine-generated work—but the standards differ by field. In mathematics, formal tools such as Lean can check whether a proof follows from a stated theorem. They cannot decide on their own whether the theorem is the right one to investigate or whether a result is important. OpenAI’s warning about unformalized manuscripts also means the 722 outputs should not be treated as 722 independently verified discoveries.

Software review is measured through pull requests, review starts, acceptance and whether a human examined a change. Those indicators describe workflow, not necessarily the correctness or impact of each change. The report notes that several cited software-data providers sell code-review tools, a commercial interest readers should keep in mind when interpreting their findings. It says the studies point in a similar direction, but their differing methods and samples limit direct comparison.

Contract work adds legal and institutional accountability. A model may draft or analyse language, but a person or organisation still has to decide whether the text meets the client’s requirements and applicable rules. Across these areas, the report’s central distinction is between checking a result against a specified standard and deciding whether that standard captures what people actually need.

““verification abundance, adjudication scarcity.””

— The title of a recent paper cited in the supplied source

Evidence Has Important Limits

The supplied source does not give publication links, dates or full methods for all cited figures, and the studies cover different populations and periods. The software findings therefore cannot establish that AI caused every reported change in review time, acceptance or human oversight. The source itself advises caution because some data providers sell review products.

It is also unclear how many of OpenAI’s 722 mathematical manuscripts were formally verified, how the 4,000 problems were selected, and how the manuscripts were assessed for significance. For the Ironclad evaluation, the source does not identify the criteria, provide a comparison score for the earlier model or say whether the assessment was independent. It does not quantify how much review work AI itself can take on, or whether new processes can expand expert capacity.

Track Review and Training

The next useful evidence will be more detailed, comparable reporting: how AI-generated work performs across matched tasks, what share receives substantive human review, how often reviewers find consequential errors, and how much time correction takes. In software, tracking quality alongside pull-request volume would help distinguish faster production from genuinely faster delivery of reliable changes.

For mathematics and legal work, clearer disclosures about formal verification, evaluation criteria and independent review would make output counts and model scores easier to interpret. Employers will also need to watch whether junior staff continue to gain hands-on experience in the work they may later be asked to supervise. The source offers no specific policy or next milestone; whether AI expands checking capacity or shifts more burden to scarce experts remains an open question.

Key Questions

What is the main development described?

The report argues that AI is lowering the cost of producing work faster than it is lowering the cost of verifying and accepting that work. It supports the argument with examples from mathematics, software and contract workflows.

Were all 722 mathematical manuscripts verified?

No. The source says some results were formally checked in Lean and quotes OpenAI warning that unformalized results “could have issues.” It does not state how many manuscripts were formally verified.

What did the software figures measure?

The cited sources measured different aspects of pull-request workflows, including review time, time before review began, acceptance rates and whether a human reviewed a change. Their findings are not one unified measurement, and several sources sell code-review tools.

Does the report show that AI work is less reliable?

Not conclusively. The cited figures raise questions about review and acceptance, but they do not provide a common, independent measure of correctness across fields or prove that AI caused every reported difference.

Why could this affect junior workers?

The report argues that people often learn to review work by first doing it themselves. If AI replaces too much entry-level drafting or coding, organisations could have fewer experienced people prepared to judge machine-generated work later.

Source: ThorstenMeyerAI.com

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Unveiling AI: A Guide To The 12 Questions On Everyone’s Lips

Explore the 12 essential questions about AI, how it works, its limitations, and what it means for the future, based on insights from Thorsten Meyer AI.

10 Best Computers, Tablets & Components For Flexible Work In 2026

Discover the 10 best computers, tablets, and components for flexible work in 2026, based on expert evaluations of OS, performance, and durability.

Stay Ahead With These 14 AI Note Apps For Students In 2026

Discover the 14 best AI-powered note-taking apps for students in 2026, offering features like transcription, summarization, and seamless organization to enhance learning.

A Skill Is A Folder, Not A Prompt: What Anthropic Learned Running Hundreds Of Them

Anthropic reveals that ‘Skills’ are folders containing instructions, scripts, and knowledge, transforming AI prompt engineering into durable, institutional assets.