Skip to content
live markets
S&P 5007,785.76▲ 3.21%NASDAQ26,729.16▲ 2.38%DOW53,732.41▲ 2.33%GOLD4,434.90▲ 11.27%WTI82.15▲ 4.05%BRENT88.47▲ 5.03%EUR/USD1.1583▲ 1.21%USD/JPY159.06▼ 2.04%DXY99.55▼ 1.17%BTC$62,736▼ 0.40%ETH$1,872▼ 0.40%SOL$74.43▼ 1.30%TOTAL CRYPTO$2.24T▼ 0.38%
pulseofnations.
UTC --:--NYC --:--LON --:--WAW --:-- telegram ↗ bluesky ↗ Join the wire

AI Agents Re-Ran 2,226 Conference Papers, Contested 496 Claims

A Hugging Face-led hackathon used 1,221 volunteers and coding agents to reproduce ICML 2026 experiments, finding contested or falsified claims in 496 papers.

Partner Surfshark VPN

Hugging Face organized a 19-day hackathon in which 1,221 participants used AI coding agents to re-run experiments from 2,226 papers accepted at ICML 2026, the machine learning conference’s largest-ever batch, and published every reproduction attempt in a public log.

The effort, called Open Reproductions, tasked participants with using tools including Claude Code, Codex, and Cursor to re-run each paper’s core claims. Agents assessed a total of 35,908 individual claims across the paper set. Of the 2,226 papers examined, 496 had at least one claim that was contested or falsified during the reproduction process, according to results published on Hugging Face.

Scale Demands Automation

ICML 2026 accepted 6,352 papers, roughly double the previous year’s count. The volume has outstripped what traditional peer review and manual reproduction efforts can handle, making the hackathon a test case for whether AI agents can serve as a scalable check on published research.

Every reproduction attempt, including failures, was logged in a public Trackio logbook hosted on Hugging Face. Each entry includes the code used, artifacts produced, and optionally the full agent execution trace, creating an auditable record per paper that did not previously exist at this scale.

Falsification Is the Point

The project distinguishes itself from typical reproduction efforts by treating contested and falsified results as valuable outputs rather than embarrassments. Of the 35,908 claims assessed, the vast majority held up under re-examination. The 496 papers with at least one flagged claim represent roughly 22% of the total set, a figure that could reshape how researchers and institutions evaluate the reliability of published ML work.

The hackathon format also exposed practical limitations. AI agents excelled at running straightforward experiments but struggled with papers requiring specialized hardware, proprietary datasets, or undocumented environment configurations. The results provide a realistic picture of what automated scientific verification can and cannot handle today.

Hugging Face, which has positioned itself as the central infrastructure provider for open-source AI, framed the effort as a step toward continuous post-publication review rather than a one-off event. The public logbook remains accessible, allowing other researchers to build on the reproduction attempts or audit the methodology.

The project arrives as the machine learning community grapples with a reproducibility crisis amplified by the sheer volume of accepted papers. Conference organizers have debated stricter acceptance thresholds, but Open Reproductions suggests that automated agent-based checking could complement traditional review without requiring the field to publish fewer papers.

Sources: AI/TLDR daily digest (August 16, 2026); Hugging Face blog; ICML 2026 proceedings

React to this dispatch
Share this dispatch Telegram X WhatsApp Report an error

discussion

Join the discussion

Your email address will not be published. Required fields are marked *

Next dispatch Z.ai Ships GLM-5.3 With Major Coding Gains From Post-Training Only Read →