agentic research

A dated log of fully AI-generated attacks on known open problems.

Over the last year I have watched AI systems go from fumbling textbook exercises to closing real gaps in research mathematics. The question I keep coming back to is how much of that carries into the problems I actually work on — quantum information, complexity, quantum many-body physics. The only honest way to find out is to run the experiment.

The tempting way to report it would be to post the results to arXiv. I would rather not. The literature does not need more machine-written preprints, and a proof no human has read carefully has no business being cited. So it goes here instead, where it can be looked at without pretending to be something it is not.

What follows is a dated log of attempts by AI agents on known open problems. The proofs are machine-generated end to end. My own role sits upstream of the writing: I choose the problems, supply the intuition about where a proof might come from, and design the system that carries it out — how a claim gets attacked, who plays adversary, what is allowed to count as verified. Once an attack is running I do not intervene in it.

Not everything gets an entry. An attack that moved nothing is not listed, and those are the majority — most of what I point these systems at, they fail to shift. What follows is the part that moved: problems settled, problems partly settled, and bounds improved. Each one has a short note attached saying what was actually established and what is still missing.

What “verified” means here. Computer-verified: exact rational arithmetic wherever the problem allows it, a second agent re-deriving the result from scratch on its own construction, and an adversarial agent paid to break it. It does not mean refereed. For most entries no human has typeset a proof you could read in a journal. Read each one as a claim with its working attached, not as a result.

Log

If you think one of these is wrong, I would like to know — that is rather the point of putting them somewhere public instead of on arXiv. Write to zhangyx@comp.nus.edu.sg.