Who Checks The Work When AI Makes More Of It?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Who Checks The Work When AI Makes More Of It? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source account describes a widening gap between AI-generated work and the human capacity to verify it, citing mathematical manuscripts, software pull requests and contract analysis. The examples point to review as a growing constraint, but some figures come from vendors, and the long-term effects on quality and professional training remain uncertain.

AI systems are producing research manuscripts, software changes and contract work faster than people can reliably review them, according to examples compiled in a report from ThorstenMeyerAI.com. The account says OpenAI published 722 mathematical manuscripts this week after posing roughly 4,000 problems to a model, while verification still depends heavily on scarce human expertise. The gap matters because organizations may be able to generate more work than they can safely approve or use.

The mathematics example illustrates both the scale of output and limits to automated checking. OpenAI’s model produced the manuscripts across 372 problem families, with an average result taking about three hours of compute, according to the source account. Some results were formally checked using Lean, a proof assistant. OpenAI cautioned that some results not formalized in Lean could have issues. The account also contrasts the volume with the response to an earlier result from the same program: a counterexample to an Erdős conjecture received careful scrutiny from five leading mathematicians.

In software, the source cites data from several organizations. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analyzing 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are reported findings, not a direct measure of code quality; the source notes that several providers sell code-review products.

A peer-reviewed 2026 study cited in the account found that 61% of AI-agent pull requests received no human review before being merged or closed. In another example, OpenAI’s partnership with contract-software company Ironclad involved training GPT-6 Astra on contracting workflows. The source says Astra met 55% of evaluation criteria on average across 11 tasks, an improvement over its predecessor. That result also leaves criteria unmet; the account does not specify the tasks’ full scoring method or the consequences of individual errors.

At a glance
reportWhen: The source account describes developmen…
The developmentA report on AI-generated research, software and legal work highlights that producing material is becoming faster than checking it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Is Becoming the Constraint

The immediate issue is not simply whether AI can generate plausible work. It is whether an organization has enough qualified people to determine what is correct, relevant and safe to use. If output grows faster than review capacity, teams may delay good work, approve work without adequate scrutiny, or rely on the producer’s own judgment about what deserves attention.

The source account identifies those patterns in its examples: some pull requests reportedly go without human review, while reviewers may deprioritize machine-generated changes because they expect more problems. In mathematics, the people producing a body of results also select which results to publish. Each response can be understandable, but each shifts risk or responsibility. The data does not establish that every AI-produced item is unreliable; it does show why output volume alone is an incomplete measure of productivity.

There is also a workforce question. Senior reviewers typically build judgment through years of doing the underlying work. If junior staff mainly supervise machine drafts rather than learn to write code, draft contracts or develop proofs themselves, organizations could weaken the future supply of experienced reviewers. That is a plausible risk raised by the source, not a demonstrated economy-wide outcome.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Verification Gap

The examples span fields with different standards of evidence. In mathematics, formal proof systems can check whether a proof follows from stated assumptions, but they do not decide whether the theorem addresses the intended question or whether it is important. In software, tests can check specified behavior but cannot guarantee that the tests cover the real requirements. In contract work, a reviewer may need to identify a missed approval rule or an unsuitable jurisdiction clause, tasks that require understanding the deal and its setting.

The source account argues that automated checks cannot by themselves settle every question of intent or accountability. A proof assistant evaluates formal statements; it does not replace expert judgment about what should be proved. Similarly, a model can draft or analyze a contract, but people and institutions remain responsible for deciding whether it is fit for use. The examples therefore distinguish checking a formal result from adjudicating its meaning and consequences.

The underlying comparison is between the falling cost of generating material and the time required for accountable review. The reported figures come from different datasets and methods, so they should not be treated as one unified measurement. The source also notes that some software figures come from companies selling review tools. Their results offer evidence of a possible bottleneck, but do not by themselves establish its size across all industries.

“Verification abundance, adjudication scarcity.”

— Description of a recent paper cited by ThorstenMeyerAI.com

Amazon

formal verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The available examples do not establish how often AI-generated work contains material errors, or how much human review is needed to catch them. The cited software measures cover different periods and definitions, and some come from vendors with a commercial interest in review tools. The source does not provide enough methodological detail to compare those figures directly or generalize them to all software teams.

It is also unclear how the mathematics manuscripts were selected for publication from the roughly 4,000 problems, how many results were formally checked, and what happened to the rest. The account says the model’s contract evaluation met 55% of criteria on average across 11 tasks, but does not list the criteria or explain whether misses were minor or consequential. Nor do the cited examples show whether review capacity is already reducing quality across organizations or whether new processes will close the gap.

Amazon

contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Review and Training

The next useful evidence will show not only how much AI-generated work organizations produce, but also how it is checked: review rates, time to approval, error detection, and the impact of missed errors. For research, the number of formally verified results and the methods used to select manuscripts would clarify what publication counts represent. For software and contracts, independent evaluations could help distinguish vendor-reported trends from broader patterns.

Organizations also face a practical staffing question: how to use AI while giving junior workers enough direct experience to develop professional judgment. The source account does not identify a specific policy or next milestone. For now, the central uncertainty is whether review methods and training can expand quickly enough to match rising output without weakening accountability.

Amazon

AI research manuscript verification

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

The source account says AI is generating large volumes of mathematical research, code and contract work, while qualified human review remains limited. It cites OpenAI’s 722 mathematical manuscripts and separate software and contracting examples.

Were all 722 mathematical manuscripts formally verified?

No. The account says some results were checked in Lean and quotes OpenAI warning that some unformalized results could have issues. It does not state how many manuscripts received formal verification.

Do the software figures prove AI-generated code is worse?

Not on their own. The cited figures concern review waits, acceptance rates and review coverage, which are not interchangeable measures of code quality. The data comes from different sources, including companies that sell review tools.

Why can’t AI simply check AI’s work?

Automated tests and formal systems can verify defined properties, but may not establish that the right question, requirement or criteria were used. Human reviewers also provide professional judgment and accountability.

What remains unknown?

The evidence cited does not establish the overall rate or severity of errors in AI-generated work, how much review is sufficient, or whether organizations can expand review capacity while preserving training for future experts.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What China’s New AI Compute Satellite Means For Space-Based Computing

A Tech Times headline reports an AI compute node in orbit with SenseTime and Zhipu AI backing, but launch and operating details are unavailable.

Japan suspected of near $30bn interventions in thin Golden Week trading

Japanese authorities likely intervened in currency markets with over $28.8 billion during Golden Week, raising concerns about market stability and policy stance.

Apple releasing 20th anniversary iPhone, AirPods with cameras next year: report

Apple plans to release a special 20th anniversary iPhone and new AirPods featuring cameras next year, according to reports from 9to5Mac.

Can Mukesh Ambani pull off his biggest gamble yet?

Mukesh Ambani is undertaking his most ambitious business move yet, with the outcome uncertain. This development could reshape his empire and the Indian market.