SUN, OCTOBER 04, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ General

The Model Found Counterexamples. It Did Not Prove the Theorems.

Meta published six mathematics papers with Muse Spark, five answering open questions and two disproving conjectures by counterexample - run through the ordinary meta.ai chat interface with no custom scaffolding, and with each paper marking which passages the AI drafted.

By AIToolsRecap October 4, 2026 6 min read 27 views
Home › Articles › General › Meta's AI Helped Disprove Two Maths Conjectures

What Was Published

Meta released six mathematics research papers produced with Muse Spark 1.1 and 1.2, with human mathematicians, and says five of them answer questions that were genuinely open.

The results:

  • Probability - a sharp threshold for fitting random Gaussian points to ellipsoids in high dimensions.
  • Differential equations - finite-time wave collapse in the mass-critical biharmonic nonlinear Schrodinger equation.
  • Group theory - disproved the conjecture that semiabelian groups must be monomial, by finding a counterexample.
  • Optimization - established when cycle-based relaxations exactly capture binary polynomial optimization.
  • Arithmetic physics - connected p-adic string theory calculations to height functions on curves.
  • Evolution algebras - also disproved a conjecture, again by counterexample.

Named collaborators include Aykut Arslan, Leonard Dinh, Joseph Brennan, Milana Golich, Andres Barei, Anindya Dey, Gabriel Herczeg, An Huang and Nicolas Jaramillo Torres, with peer review from further mathematicians.

The Detail Nearly Everyone Will Skip

The models ran in Thinking Mode through the standard meta.ai chat interface, with no custom scaffolding.

Not a specialised research system. Not an agent harness built for mathematics. The consumer chat product.

That matters more than the results do. A bespoke system solving open problems tells you Meta can build a bespoke system. The same work coming out of the interface anyone can open says the capability is already distributed, and that the limiting factor is who is sitting at the keyboard.

Look at Which Problems It Cracked

Two of the six results are disproofs by counterexample. That is not a coincidence and it is the most useful thing in the announcement.

Disproving a conjecture means finding one object that breaks it. That is a search problem - define the space, generate candidates, test them, keep going. A model that can write and run search programs tirelessly is extremely well suited to it, and Meta confirms the model "generated search programs and identified counterexamples".

Proving a theorem is a different act. It requires an argument that holds for every case, which is not search.

So read the distribution rather than the headline: the model is strongest where the work is search and weakest where it is insight. That is a real and significant capability - mathematicians have hunted counterexamples by hand for centuries - but it is a specific one, and "AI is doing mathematics now" flattens it into something less true and less interesting.

What Each Side Actually Did

The modelThe mathematicians
Generated search programsChose the problems
Identified counterexamplesGuided research direction
Worked through calculationsVerified results
Tested possible argumentsRefined arguments
Reframed problems, drafted sectionsPeer reviewed

Meta's own framing is explicit: "A team of mathematicians guided the research."

The Standard Meta Just Set

Here is the part other labs should be asked about: each paper marks which passages were primarily drafted by researchers and which were drafted by AI.

Per-passage attribution, in the published work. Not a blanket acknowledgement in a footnote, not a disclosure statement - a mark on the actual text.

That is a higher bar than almost any journal currently requires, and it arrives in the same week arXiv capped submissions at two a month per author because AI-assisted volume overwhelmed its reviewers. If AI-assisted papers carried this kind of attribution as standard, the review problem would be considerably easier, because a reviewer would know where to look hardest.

The Context

Meta frames this as the next question after competition performance: its models reached gold-medal level across five competitions in mathematics, physics and chemistry, and competition problems have known solutions. The question here was whether a model contributes when the problem is genuinely open and no solution path exists.

On the evidence of six papers, the answer is a qualified yes - qualified by who chose the problems, who verified the results, and which kind of problem the model happened to be good at.

FAQ

Did AI solve these maths problems on its own?

No. Meta states a team of mathematicians guided the research, chose the problems, verified results and refined arguments. The model generated search programs, found counterexamples, worked through calculations and drafted sections.

Which model was used?

Muse Spark 1.1 and 1.2, running in Thinking Mode through the standard meta.ai chat interface with no custom scaffolding.

What conjectures were disproved?

The conjecture that semiabelian groups must be monomial, and a conjecture in evolution algebras. Both were disproved by counterexample rather than by proof.

How do the papers credit the AI?

Each paper marks which passages were primarily drafted by researchers and which by AI - per-passage attribution within the text itself, which is stricter than most journals require.

Tags
AI NewsGenerative AI2026
⚑

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 AI Tools prompts
Prompt Guide
Best Claude AI Prompts for SEO (2026) — Content, Technical, and Comparison SEO
Claude Sonnet 5 and Opus 5 are strong for SEO work that requires writing quality, structured analysis, and long-form content generation. With 1M context, Claude can analyse an entire site's content structure, compare competing pages, and write complete article drafts in one session. These prompts cover the full SEO workflow: keyword research synthesis, content briefs, on-page optimisation, meta descriptions, technical audit interpretation, and comparison content that ranks above AI Overviews.
Get Prompts →
Prompt Guide
Best ChatGPT Prompts for SEO (2026) — GPT-5.6 and Browse
ChatGPT with GPT-5.6 Sol and Browse enabled is a capable SEO research tool — it can search the live web, analyse SERP results, and synthesise content briefs in a single session. GPT-5.6 Terra at $2.50/M offers a cost-efficient option for high-volume SEO content generation. These prompts are optimised for ChatGPT Plus with Browse, the ChatGPT Work product for larger projects, and the OpenAI API with web_search tool enabled.
Get Prompts →
Prompt Guide
Best Claude Opus 5 and Sonnet 5 Prompts for Writing (2026)
Claude Opus 5 and Sonnet 5 consistently produce the highest-quality long-form writing of any AI model in July 2026 — a lead documented across writing benchmarks and user testing since Claude 3 Opus. With 1M context and 128K output on Opus 5, Claude can write book chapters, complete reports, and long-form content without truncating. Sonnet 5 at $2/$10/M (intro through August 31) is the best value writing model available. These prompts are optimised for claude.ai Pro/Max, Claude Cowork, and the API.
Get Prompts →