Meta's Muse Spark Helped Mathematicians Solve Five Open Research Problems
Meta's Muse Spark worked alongside mathematicians through the standard chat interface to co-author six research papers, five of which answer previously open questions.
- Meta released six math papers co-authored with Muse Spark, five resolving previously open problems.
- Researchers used the Meta.ai chat interface in Thinking Mode, with no custom research scaffold or agent harness.
- Problems span probability, PDEs, group theory, optimization, arithmetic physics, and non-associative algebra.
- Muse Spark generated GAP search code that found a 384-element counterexample to Kida's 2024 semiabelian conjecture.
- Three results were independently proven by other teams in August using different approaches.
- Papers explicitly label which passages were drafted by humans versus AI, and credit prior work.
Meta tests Muse Spark on open math research
Meta says mathematicians used versions 1.1 and 1.2 of its Muse Spark reasoning model to help produce six research papers through the ordinary Meta.ai chat interface. The company released six papers and says five resolve questions that lacked known answers when the collaborations began. Some results overlap with work developed independently during the same period.
Meta previously reported Olympiad results equivalent to gold-medal performance across five high-school competitions in mathematics, physics, and chemistry. Those competitions contain fixed problems with known solutions. The new papers examine active research, where mathematicians must define a path, establish the result, and check that every step holds. The cases offer evidence about a research workflow, although six selected collaborations cannot establish a general success rate.
Research through the chat box
The mathematicians accessed Muse Spark through Meta.ai’s consumer interface and its Thinking Mode. According to Meta, the setup excluded a custom agent harness, a tool-calling pipeline connected to Lean or Mathematica, and a dedicated retrieval system for papers on arXiv.
The collaboration protocol kept responsibilities and provenance explicit. Mathematicians selected the problems and directed the exploration, while a separate group reviewed the model’s output. Each paper labels passages drafted by humans and by Muse Spark, cites prior work, and acknowledges teams that independently reached overlapping results.
Six narrow, technical targets
- Probability: The paper identifies a sharp threshold for fitting random Gaussian points exactly to an ellipsoid in high dimensions. An exact fit exists with high probability below the cutoff and almost never exists above it. Three independent groups produced concurrent proofs using different methods.
- Differential equations: For radially symmetric solutions with negative energy, the paper proves finite-time blow-up in the mass-critical biharmonic nonlinear Schrödinger equation in two or more dimensions. This fourth-order wave equation can develop a singularity where a smooth solution ceases to exist. The result settles a question open since 2015 within that symmetric setting.
- Group theory: The work gives a counterexample to Kida’s 2024 conjecture connecting two technical classes of finite groups, semiabelian and monomial groups. Muse Spark generated a program for GAP, a computer algebra system for group theory, that found a 384-element counterexample. Mathematicians verified the example and completed the proof.
- Optimization: The paper derives a criterion for when a cycle-based relaxation of a binary polynomial problem is exact. Such relaxations replace difficult optimization over zero-or-one variables with constraints that are easier to solve. The result addresses a 2026 question from Del Pia and Khajavirad.
- Arithmetic physics: The authors prove that a string-theory two-point function equals an arithmetic height function on a broad class of p-adic curves. Two-point functions measure correlations, while p-adic geometry uses number systems whose notion of distance is based on divisibility. The paper extends a direction Yuri Manin proposed in the 1980s.
- Non-associative algebra: The paper constructs a three-dimensional counterexample to a proposed test for solvable evolution algebras, where multiplication need not satisfy
(ab)c = a(bc). It also supplies an alternative characterization formulated at the subspace level.
From proof sketches to search code
Across the papers, Muse Spark contributed several kinds of intermediate research work:
- Proof development: It generated candidate arguments and possible proof strategies for experts to inspect and repair.
- Problem reformulation: In the optimization project, it suggested a probabilistic formulation that shaped the eventual solution.
- Computational search: It wrote GAP code that searched finite groups and found the 384-element counterexample.
- Technical drafting: It produced portions of the manuscripts, including three central sections of the arithmetic-physics paper.
- Counterexample generation: It proposed mathematical objects that researchers subsequently checked against the relevant definitions and conjectures.
Final validation remained with the mathematicians. They selected the research agenda, filtered candidate approaches, repaired gaps, checked computations, and approved each claim. In the group-theory project, for example, the model found a candidate through code; the researchers verified its properties and supplied the completed argument.
Evidence from a small sample
FrontierMath and similar benchmarks evaluate models on fixed sets of difficult problems whose answers can be checked, with reported accuracy commonly ranging from 10% to 30% across undergraduate, advanced, and research-oriented tiers. Meta’s papers add case studies from active projects. Reliable generation of useful proof sketches or search programs could reduce time spent exploring dead ends, writing one-off code, and drafting routine technical passages.
Several limits constrain what can be inferred from the release:
- Selection: The announcement presents successful collaborations without a common denominator for abandoned problems, failed approaches, total prompts, or researcher hours.
- Concurrent discovery: Other teams independently proved three of the results around the same period, indicating that at least some questions were already close to resolution.
- Human labor: Experts chose the problems, supplied domain context, steered the conversations, corrected proofs, and verified every result.
- Scope: Six specialized questions from one model family provide too little evidence to estimate performance across mathematical fields.
- Speed: The release provides no controlled comparison of how long the same teams would have taken without Muse Spark.
- Reproducibility: Exact replication depends on prompts, sampling settings, complete transcripts, model versions, and continued access to the same proprietary model snapshot.
A reproducible working pattern
The cases describe an expert-led workflow that researchers can evaluate without building custom agent infrastructure:
- Define a bounded claim, its assumptions, and the prior results on which it depends.
- Ask the model for multiple proof plans, counterexamples, reformulations, or search scripts.
- Run generated code in a controlled environment and inspect its assumptions, outputs, and failure modes.
- Verify each mathematical step manually or with a proof assistant, computer algebra system, or independent calculation where appropriate.
- Record the model version, prompts, outputs, human edits, and failed branches needed to reconstruct the process.
- Attribute model-generated text and ideas according to the publication venue’s authorship and disclosure rules.
Across these projects, Muse Spark generated proof candidates, code, counterexamples, and prose inside a consumer chat product. Mathematicians retained responsibility for problem selection, error correction, and final verification. The papers document useful intermediate contributions under close expert supervision; the general success rate and time advantage remain unmeasured.