Recent progress in AI-generated mathematics has made a question that used to feel hypothetical suddenly feel very concrete.
OpenAI has reported a series of results on long-standing open problems in mathematics and theoretical computer science, and in September announced an AI-generated proposed solution to the Navier–Stokes Millennium Prize Problem.1
The reaction from mathematicians has not been uniformly celebratory. Some are excited. Some are intimidated by the pace. Others are worried about what mathematics becomes if solving famous open problems increasingly turns into an AI benchmark. The Leiden Declaration captures exactly this mixture of enthusiasm and concern.2
At first, the obvious question seems to be:
The same question applies to science more generally.
If AI systems can search hypotheses, run simulations, write code, analyze data, design experiments, and eventually solve scientific problems much faster than humans, what happens to scientists?
I do not think the anxiety behind this question is irrational.
But I increasingly think we may be worried about the wrong thing.
What exactly is being disrupted?
Science has always celebrated people who ask great questions.
The names that remain foundational in a field are often not simply those who solved the largest number of difficult problems. They are the people who defined new objects, connected previously separated areas, proposed questions that occupied generations of researchers, or created entirely new ways of representing a problem.
And yet this is not how most scientific careers are actually evaluated.
For most researchers, the heaviest currency is still much more concrete:
- What experiment did you run?
- What data did you collect?
- What theorem did you prove?
- What empirical result did you discover?
- How many important papers did you publish?
You can write a perspective. You can write a review. You can propose a beautiful question.
But sooner or later, the system asks: what did you actually solve?
For a long time, this made perfect sense.
Problem solving was expensive.
A scientific question could require years of experiments, thousands of failed attempts, new instruments, a large research team, or an extraordinarily difficult proof. The ability to turn a well-defined problem into a verified result was genuinely scarce.
So perhaps what AI is destabilizing first is not science itself.
That distinction matters.
But why was problem solving so valuable in the first place?
This is where I think the story becomes more interesting.
Suppose the only value of solving a scientific problem were the answer itself.
Then an AI that produces the answer faster should simply be better.
There would not be much more to discuss.
But historically, solving a problem almost never produced only an answer.
And sometimes that understanding turned out to be much more important than the original problem.
Consider polynomial equations.
Mathematicians trying to solve cubic equations were not setting out to invent a huge new branch of mathematics. They were trying to find roots.
But the formulas they encountered forced them to deal seriously with quantities involving square roots of negative numbers.
If the only goal had been numerical prediction, one could imagine treating these strange intermediate quantities as an inconvenience and moving on.
Instead, mathematicians tried to understand them.
What kind of object is this?
How should arithmetic work on it?
Why does an apparently “imaginary” intermediate quantity appear when the final solution is real?
That detour helped open the door to complex numbers and the enormous mathematical landscape that followed.
The quintic gives an even cleaner example.
After formulas had been found for quadratic, cubic, and quartic equations, the obvious next problem was:
Mathematicians searched for such a formula and eventually discovered that, in the expected sense, it does not exist.
But the really important step was not simply proving “there is no formula.”
The question changed.
Instead of asking “What is the formula?”, mathematicians began asking something closer to:
Now the problem was no longer merely about computing roots.
It became a problem about permutations, symmetry, and structure.
And from that shift emerged Galois theory and group-theoretic ways of thinking.
The original problem was not merely solved.
Now imagine that AI had already existed
This thought experiment is what made the issue click for me.
Suppose mathematicians asking about a quintic could simply hand the equation to an AI.
The AI gives the roots to arbitrary numerical precision.
Problem solved.
Or perhaps it finds some entirely different computational procedure that works perfectly well in practice.
From the standpoint of task completion, there is no problem anymore.
Why spend years asking whether a beautiful symbolic formula exists?
Why stare at the failed approaches?
Why investigate the strange symmetry of the roots?
And yet those apparently unnecessary questions were exactly where a much larger piece of mathematics came from.
This makes me wonder whether we sometimes confuse two very different things:
They are not the same.
In fact, the most efficient path to an answer may occasionally bypass the most scientifically fertile path.
A failed proof can reveal a new structure.
An awkward intermediate object can become a new concept.
A second-best solution can reveal a more general principle than the optimal one.
A problem that refuses to yield under the current representation can force us to invent a better representation.
So perhaps there is an important distinction between task utility and epistemic fertility.
This is where AI changes the old relationship
Historically, several things were bundled together:
A scientist trying to solve a problem would see the failed attempts, discover the anomaly, develop intuition, change representations, and eventually produce a result.
The answer and the understanding grew together.
AI begins to separate them.
We may increasingly have:
This can happen without the scientific community having gone through the process that produced that discovery.
The theorem may be correct.
The simulation may work.
The molecule may satisfy the objective.
The algorithm may outperform everything before it.
But what have we actually learned?
Why did it work?
What did the failed alternatives reveal?
Which part of the solution is specific to this case, and which part transfers?
Could there be a more explanatory proof?
Does this result suggest that we were using the wrong representation all along?
And, perhaps most importantly:
This is where I think “understanding” becomes much more than a matter of intellectual curiosity.
Understanding is part of the engine of science
I used to think the argument for understanding AI-generated science was mainly about human agency.
If humans do not understand what AI discovers, how can they judge it, trust it, extend it, or remain in control?
I still think that matters.
I explored this gap between fast-moving results and the understanding needed to judge or act on them in The Understanding Bottleneck.
But I now think there is a stronger argument.
A scientific cycle is not simply Q → A. It is closer to:
- Understanding
- Question
- Discovery
- New understanding
- New question
Here U is our current understanding, Q the questions that understanding makes visible, and D the discoveries that come from investigating them.
AI is rapidly getting better at the middle: Qt → Dt.
But the next scientific frontier depends on what happens afterward.
A discovery has to change our model of the world.
That changed model makes different questions salient.
Sometimes it makes entirely new questions expressible.
In that sense, understanding is not something we add after discovery because humans happen to like explanations.
It is one of the mechanisms by which one discovery creates the next generation of problems.
Perhaps this is why the future looks so different depending on how you frame it
If you think science is primarily the production of answers, recent AI progress can look like the beginning of the end.
Machines are becoming extraordinarily good at producing answers.
But if science is also the continual creation of better questions and better representations, the picture looks very different.
AI may remove an enormous amount of technical burden from scientists.
Search that once took years may take hours.
Proof directions that were too expensive to explore may become cheap.
Experiments may be simulated before they are run.
A single researcher may be able to explore a space that previously required an entire lab.
Maybe that does not reduce the room for scientific imagination.
Maybe it increases its leverage.
The scarce question may gradually shift from:
to
and eventually to
I do not think this means humans will have a permanent monopoly on problem formation.
Future AI systems may themselves become excellent at inventing representations, connecting fields, and defining research programs.
If that happens, the picture will change again.
But at least today, most of what we call AI for Science is still much stronger at searching, generating, and testing than at producing the kind of conceptual reorganization that defines a new field.
And even if machines eventually become excellent at that too, there remains a separate question:
So what should AI for Science optimize for?
Certainly not less discovery.
We should let AI search farther and faster than humans ever could.
But perhaps the scientific workflow should not end when the answer arrives.
The arrival of an AI-generated result may increasingly be the beginning of a second task:
Can we reconstruct the key conceptual move?
Can we find a more illuminating proof?
Can we distinguish the accidental trick from the transferable principle?
Can we identify the anomaly that deserves to become a new object of study?
Can we use the result to change how we represent the problem?
And can that new understanding generate questions that were invisible before?
Perhaps this is the real transition ahead.
For a long time, one of the central scarcities in science was our ability to solve hard problems.
AI may dramatically reduce that scarcity.
The next bottleneck may be our ability to turn abundant discoveries into better ways of thinking.
And from those better ways of thinking, better questions.
AI can make answers abundant. The future of science may depend on whether we can turn those answers into understanding—and that understanding into questions that did not exist before.