🔍 Read the full analysis: AI Changed The Economics Of Making And Checking on ThorstenMeyerAI.com
Get the little things that make your day delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
AI systems can generate mathematical manuscripts, software changes and contract work faster and more cheaply, but the supplied report says expert review remains slow and limited. Figures from software studies point to longer waits and less human review, though several sources sell review tools and the data does not establish one universal effect.
A report published this week argues that AI has made it cheaper to produce work in fields including mathematics and software, while human capacity to check that work remains limited. Its examples range from 722 mathematical manuscripts attributed to OpenAI to software-industry data showing increased review delays, though the figures come from separate sources and do not establish a single, economy-wide trend.
The report says OpenAI’s system was given about 4,000 mathematical problems and produced 722 manuscripts across 372 families, with an average result taking about three hours of compute. Some results were checked using the Lean proof assistant. OpenAI cautioned that unformalized results “could have issues,” according to the source. The report contrasts this volume with the careful verification by five leading mathematicians of an earlier result from the same programme: a proposed counterexample to an old Erdős conjecture. It does not provide the names of those mathematicians or details of their review process.
In software, the source cites several studies with different samples and measures. Faros AI reported that teams in high-AI-adoption periods merged 98% more pull requests, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. The measures are not directly interchangeable.
The report also points to OpenAI’s partnership with contract-software company Ironclad, saying GPT-6 Astra was evaluated on 11 contracting tasks and met 55% of evaluation criteria on average. That result is described as an improvement over a previous model, but the source gives no earlier score, task-level breakdown or independent assessment. The remaining criteria indicate potential gaps in performance; they do not, by themselves, show how often errors would occur in actual legal work.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the Pace
If these examples reflect a broader workplace pattern, the constraint on using AI may shift from producing drafts to checking, correcting and taking responsibility for them. More output does not automatically mean more usable output: organisations still need people who can decide whether a result answers the right question, fits the situation and is safe to rely on.
The report describes several possible responses when review cannot keep pace: work may be merged without review, reviewers may delay machine-generated submissions, or producers may decide which results merit attention. Each has costs. Unchecked output can carry errors forward; blanket suspicion can slow useful work; and relying on the producer’s own selection can leave independent scrutiny thin. These are risks identified by the source, not proof that every organisation is experiencing them.
That imbalance could also affect staffing and training. Experienced reviewers typically develop judgment through years of doing the underlying work. If entry-level employees mainly edit AI output rather than learning to write code, draft contracts or prove results themselves, employers could weaken the future supply of people qualified to review complex work. The source frames expert judgment as a possible economic advantage, but offers no wage or hiring data to measure such a “referee premium.”
AI review tools for software development
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three Fields, Different Checks
The examples share a broad question—how to verify abundant machine-generated work—but the standards differ by field. In mathematics, formal tools such as Lean can check whether a proof follows from a stated theorem. They cannot decide on their own whether the theorem is the right one to investigate or whether a result is important. OpenAI’s warning about unformalized manuscripts also means the 722 outputs should not be treated as 722 independently verified discoveries.
Software review is measured through pull requests, review starts, acceptance and whether a human examined a change. Those indicators describe workflow, not necessarily the correctness or impact of each change. The report notes that several cited software-data providers sell code-review tools, a commercial interest readers should keep in mind when interpreting their findings. It says the studies point in a similar direction, but their differing methods and samples limit direct comparison.
Contract work adds legal and institutional accountability. A model may draft or analyse language, but a person or organisation still has to decide whether the text meets the client’s requirements and applicable rules. Across these areas, the report’s central distinction is between checking a result against a specified standard and deciding whether that standard captures what people actually need.
““verification abundance, adjudication scarcity.””
— The title of a recent paper cited in the supplied source
Evidence Has Important Limits
The supplied source does not give publication links, dates or full methods for all cited figures, and the studies cover different populations and periods. The software findings therefore cannot establish that AI caused every reported change in review time, acceptance or human oversight. The source itself advises caution because some data providers sell review products.
It is also unclear how many of OpenAI’s 722 mathematical manuscripts were formally verified, how the 4,000 problems were selected, and how the manuscripts were assessed for significance. For the Ironclad evaluation, the source does not identify the criteria, provide a comparison score for the earlier model or say whether the assessment was independent. It does not quantify how much review work AI itself can take on, or whether new processes can expand expert capacity.
Track Review and Training
The next useful evidence will be more detailed, comparable reporting: how AI-generated work performs across matched tasks, what share receives substantive human review, how often reviewers find consequential errors, and how much time correction takes. In software, tracking quality alongside pull-request volume would help distinguish faster production from genuinely faster delivery of reliable changes.
For mathematics and legal work, clearer disclosures about formal verification, evaluation criteria and independent review would make output counts and model scores easier to interpret. Employers will also need to watch whether junior staff continue to gain hands-on experience in the work they may later be asked to supervise. The source offers no specific policy or next milestone; whether AI expands checking capacity or shifts more burden to scarce experts remains an open question.
Key Questions
What is the main development described?
The report argues that AI is lowering the cost of producing work faster than it is lowering the cost of verifying and accepting that work. It supports the argument with examples from mathematics, software and contract workflows.
Were all 722 mathematical manuscripts verified?
No. The source says some results were formally checked in Lean and quotes OpenAI warning that unformalized results “could have issues.” It does not state how many manuscripts were formally verified.
What did the software figures measure?
The cited sources measured different aspects of pull-request workflows, including review time, time before review began, acceptance rates and whether a human reviewed a change. Their findings are not one unified measurement, and several sources sell code-review tools.
Does the report show that AI work is less reliable?
Not conclusively. The cited figures raise questions about review and acceptance, but they do not provide a common, independent measure of correctness across fields or prove that AI caused every reported difference.
Why could this affect junior workers?
The report argues that people often learn to review work by first doing it themselves. If AI replaces too much entry-level drafting or coding, organisations could have fewer experienced people prepared to judge machine-generated work later.
Source: ThorstenMeyerAI.com
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
