★ Featured Guide · News
A benchmark published this month put frontier models at 3 to 15 percent on generating research hypotheses. Three weeks earlier an internal OpenAI model resolved ten problems open for a decade or more, and published machine-checkable proofs. Both results stand. Reconciling them tells you exactly where AI is useful in research work and where it is not.
Aug 23, 2026
·
9 min read