qedbot

Benchmarks

Benchmarks and contamination

A benchmark is a set of statements, so how much of it already carries public AI work is a query rather than a study. A benchmark whose problems already have published solutions measures recall rather than reasoning.

Sets

3

Formal Conjectures: open

Statements the formal record still marks as open research.

28 of 439 members carry a recorded AI claim (6.4%).

2 cite a proof1 machine-checked

Formal Conjectures: all

Every statement carrying a Lean formalisation.

309 of 1396 members carry a recorded AI claim (22.1%).

571 cite a proof25 machine-checked

Erdős prize problems still open

Erdős problems with money attached and no recorded solution.

7 of 47 members carry a recorded AI claim (14.9%).

5 cite a proof

What this does and does not show

is
Public, recorded AI work against statements in each set, at the last build.
not
Evidence that any particular model was trained on any particular solution.
not
A measure of difficulty. Attention and difficulty are only loosely related.
absent
Benchmarks whose problem sets are not public — FrontierMath, AIMO and similar — cannot be measured here, because membership is unknown.