Back to Blog

AI, science & understanding

The Future of AI for Science: From Problem Solving to Problem Formation

When answers become abundant, what will help science ask questions it could not ask before?

Hanbo Xie

AI、科学与理解

AI for Science 的未来:从解决问题到提出问题

当答案越来越多,科学如何提出过去无法提出的问题?

Hanbo Xie

Recent progress in AI-generated mathematics has made a question that used to feel hypothetical suddenly feel very concrete.

OpenAI has reported a series of results on long-standing open problems in mathematics and theoretical computer science, and in September announced an AI-generated proposed solution to the Navier–Stokes Millennium Prize Problem.1

The reaction from mathematicians has not been uniformly celebratory. Some are excited. Some are intimidated by the pace. Others are worried about what mathematics becomes if solving famous open problems increasingly turns into an AI benchmark. The Leiden Declaration captures exactly this mixture of enthusiasm and concern.2

At first, the obvious question seems to be:

If AI can solve problems that once occupied some of the best mathematicians in the world, what will mathematicians still be needed for?

The same question applies to science more generally.

If AI systems can search hypotheses, run simulations, write code, analyze data, design experiments, and eventually solve scientific problems much faster than humans, what happens to scientists?

I do not think the anxiety behind this question is irrational.

But I increasingly think we may be worried about the wrong thing.

What exactly is being disrupted?

Science has always celebrated people who ask great questions.

The names that remain foundational in a field are often not simply those who solved the largest number of difficult problems. They are the people who defined new objects, connected previously separated areas, proposed questions that occupied generations of researchers, or created entirely new ways of representing a problem.

And yet this is not how most scientific careers are actually evaluated.

For most researchers, the heaviest currency is still much more concrete:

  • What experiment did you run?
  • What data did you collect?
  • What theorem did you prove?
  • What empirical result did you discover?
  • How many important papers did you publish?

You can write a perspective. You can write a review. You can propose a beautiful question.

But sooner or later, the system asks: what did you actually solve?

For a long time, this made perfect sense.

Problem solving was expensive.

A scientific question could require years of experiments, thousands of failed attempts, new instruments, a large research team, or an extraordinarily difficult proof. The ability to turn a well-defined problem into a verified result was genuinely scarce.

So perhaps what AI is destabilizing first is not science itself.

It is the incentive structure built around the scarcity of problem solving.

That distinction matters.

But why was problem solving so valuable in the first place?

This is where I think the story becomes more interesting.

Suppose the only value of solving a scientific problem were the answer itself.

Then an AI that produces the answer faster should simply be better.

There would not be much more to discuss.

But historically, solving a problem almost never produced only an answer.

The process itself generated understanding.

And sometimes that understanding turned out to be much more important than the original problem.

Consider polynomial equations.

Mathematicians trying to solve cubic equations were not setting out to invent a huge new branch of mathematics. They were trying to find roots.

But the formulas they encountered forced them to deal seriously with quantities involving square roots of negative numbers.

If the only goal had been numerical prediction, one could imagine treating these strange intermediate quantities as an inconvenience and moving on.

Instead, mathematicians tried to understand them.

What kind of object is this?

How should arithmetic work on it?

Why does an apparently “imaginary” intermediate quantity appear when the final solution is real?

That detour helped open the door to complex numbers and the enormous mathematical landscape that followed.

The quintic gives an even cleaner example.

After formulas had been found for quadratic, cubic, and quartic equations, the obvious next problem was:

What is the general formula for a fifth-degree equation?

Mathematicians searched for such a formula and eventually discovered that, in the expected sense, it does not exist.

But the really important step was not simply proving “there is no formula.”

The question changed.

Instead of asking “What is the formula?”, mathematicians began asking something closer to:

What structure determines whether an equation is solvable by radicals?

Now the problem was no longer merely about computing roots.

It became a problem about permutations, symmetry, and structure.

And from that shift emerged Galois theory and group-theoretic ways of thinking.

The original problem was not merely solved.

The space of possible mathematical questions changed.

Now imagine that AI had already existed

This thought experiment is what made the issue click for me.

Suppose mathematicians asking about a quintic could simply hand the equation to an AI.

The AI gives the roots to arbitrary numerical precision.

Problem solved.

Or perhaps it finds some entirely different computational procedure that works perfectly well in practice.

From the standpoint of task completion, there is no problem anymore.

Why spend years asking whether a beautiful symbolic formula exists?

Why stare at the failed approaches?

Why investigate the strange symmetry of the roots?

And yet those apparently unnecessary questions were exactly where a much larger piece of mathematics came from.

This makes me wonder whether we sometimes confuse two very different things:

Answering the questionSolving the current problem
Learning from the searchExtracting all of its scientific value

They are not the same.

In fact, the most efficient path to an answer may occasionally bypass the most scientifically fertile path.

A failed proof can reveal a new structure.

An awkward intermediate object can become a new concept.

A second-best solution can reveal a more general principle than the optimal one.

A problem that refuses to yield under the current representation can force us to invent a better representation.

So perhaps there is an important distinction between task utility and epistemic fertility.

The best answer to the problem we currently know how to ask is not necessarily the answer that creates the most science.

This is where AI changes the old relationship

Historically, several things were bundled together:

Problem solving≈Discovery+Understanding+New problem formation

A scientist trying to solve a problem would see the failed attempts, discover the anomaly, develop intuition, change representations, and eventually produce a result.

The answer and the understanding grew together.

AI begins to separate them.

We may increasingly have:

Problem→Machine search→Discovery

This can happen without the scientific community having gone through the process that produced that discovery.

The theorem may be correct.

The simulation may work.

The molecule may satisfy the objective.

The algorithm may outperform everything before it.

But what have we actually learned?

Why did it work?

What did the failed alternatives reveal?

Which part of the solution is specific to this case, and which part transfers?

Could there be a more explanatory proof?

Does this result suggest that we were using the wrong representation all along?

And, perhaps most importantly:

What new questions become possible now?

This is where I think “understanding” becomes much more than a matter of intellectual curiosity.

Understanding is part of the engine of science

I used to think the argument for understanding AI-generated science was mainly about human agency.

If humans do not understand what AI discovers, how can they judge it, trust it, extend it, or remain in control?

I still think that matters.

I explored this gap between fast-moving results and the understanding needed to judge or act on them in The Understanding Bottleneck.

But I now think there is a stronger argument.

Understanding is itself part of how science generates its next questions.

A scientific cycle is not simply Q → A. It is closer to:

Understanding makes the next question possible
  1. Understanding
  2. Question
  3. Discovery
  4. New understanding
  5. New question

Here U is our current understanding, Q the questions that understanding makes visible, and D the discoveries that come from investigating them.

AI is rapidly getting better at the middle: Qt → Dt.

But the next scientific frontier depends on what happens afterward.

A discovery has to change our model of the world.

That changed model makes different questions salient.

Sometimes it makes entirely new questions expressible.

In that sense, understanding is not something we add after discovery because humans happen to like explanations.

It is one of the mechanisms by which one discovery creates the next generation of problems.

Perhaps this is why the future looks so different depending on how you frame it

If you think science is primarily the production of answers, recent AI progress can look like the beginning of the end.

Machines are becoming extraordinarily good at producing answers.

But if science is also the continual creation of better questions and better representations, the picture looks very different.

AI may remove an enormous amount of technical burden from scientists.

Search that once took years may take hours.

Proof directions that were too expensive to explore may become cheap.

Experiments may be simulated before they are run.

A single researcher may be able to explore a space that previously required an entire lab.

Maybe that does not reduce the room for scientific imagination.

Maybe it increases its leverage.

The scarce question may gradually shift from:

Can we solve this?

to

What is worth asking?

and eventually to

What did the last discovery teach us that allows us to ask something we could not ask before?

I do not think this means humans will have a permanent monopoly on problem formation.

Future AI systems may themselves become excellent at inventing representations, connecting fields, and defining research programs.

If that happens, the picture will change again.

But at least today, most of what we call AI for Science is still much stronger at searching, generating, and testing than at producing the kind of conceptual reorganization that defines a new field.

And even if machines eventually become excellent at that too, there remains a separate question:

Will humans understand those conceptual shifts well enough to participate in them?

So what should AI for Science optimize for?

Certainly not less discovery.

We should let AI search farther and faster than humans ever could.

But perhaps the scientific workflow should not end when the answer arrives.

The arrival of an AI-generated result may increasingly be the beginning of a second task:

What is here for us to understand?

Can we reconstruct the key conceptual move?

Can we find a more illuminating proof?

Can we distinguish the accidental trick from the transferable principle?

Can we identify the anomaly that deserves to become a new object of study?

Can we use the result to change how we represent the problem?

And can that new understanding generate questions that were invisible before?

Perhaps this is the real transition ahead.

For a long time, one of the central scarcities in science was our ability to solve hard problems.

AI may dramatically reduce that scarcity.

The next bottleneck may be our ability to turn abundant discoveries into better ways of thinking.

And from those better ways of thinking, better questions.

AI can make answers abundant. The future of science may depend on whether we can turn those answers into understanding—and that understanding into questions that did not exist before.

最近,AI 在数学领域取得的进展,让一个过去听起来像假设的问题,突然变得非常具体。

OpenAI 报告了一系列关于数学和理论计算机科学长期开放问题的成果,并在九月宣布,AI 提出了纳维-斯托克斯千禧年大奖难题的一个解答。1

数学家的反应并非一致的欢庆。有些人兴奋,有些人被变化的速度震住,也有人担心:如果著名的开放问题越来越像测试 AI 的基准,数学将变成什么?《莱顿宣言》恰好呈现了这种兴奋与忧虑交织的心情。2

乍看之下,最直接的问题似乎是:

如果 AI 能解决那些曾让世界上最优秀的数学家花费大量心力的问题,我们还需要数学家做什么?

同样的问题也适用于更广义的科学。

如果 AI 能搜索假设、运行模拟、编写代码、分析数据、设计实验,并最终以远快于人的速度解决科学问题,科学家又会怎样?

我不觉得这种焦虑不合情理。

但我越来越觉得,我们或许担心错了地方。

究竟是什么正在受到冲击?

科学一向赞美那些提出伟大问题的人。

一个领域里留下奠基性影响的,往往不只是解出最多难题的人。他们还定义了新的研究对象,连接起原本分离的领域,提出让几代研究者持续思考的问题,或者创造出理解问题的全新方式。

但大多数科研生涯并不是这样被评价的。

对大多数研究者来说,真正被当作硬通货的,仍然是更具体的东西:

  • 你做了什么实验?
  • 收集了什么数据?
  • 证明了什么定理?
  • 发现了什么实证结果?
  • 发表了多少重要论文?

你可以写观点文章、写综述,也可以提出一个漂亮的问题。

但迟早,评价体系还是会问:你到底解决了什么?

很长一段时间里,这完全说得通。

因为解决问题的成本很高。

一个科学问题可能需要多年的实验、上千次失败、新仪器、一整个团队,或者极其艰难的证明。把一个定义清楚的问题变成经过验证的结果,确实是一种稀缺能力。

所以,也许 AI 首先动摇的不是科学本身。

它动摇的是围绕“解题能力稀缺”建立起来的激励机制。

这个区别很重要。

但解决问题为什么曾经那么有价值?

故事到这里,我觉得才真正有意思起来。

假设解决科学问题的唯一价值就是得到答案。

那么,AI 更快给出答案,自然就是更好的。

似乎也没什么可讨论的了。

但在历史上,解决问题几乎从来不只产出一个答案。

探索的过程本身,也孕育了理解。

有时候,这种理解甚至比最初的问题更重要。

想想多项式方程。

研究三次方程的数学家,起初并不是要开创一个庞大的新数学分支。他们只是想求出方程的根。

但求根公式迫使他们认真面对一些含有负数平方根的量。

如果目标只有数值预测,我们完全可以想象,他们会把这些奇怪的中间量当成麻烦,绕过去就是了。

可他们选择去理解它们。

这到底是什么样的对象?

它应该遵循怎样的运算规则?

为什么最终答案明明是实数,中途却会出现看似“虚构”的量?

这段弯路,帮助人们打开了通往复数以及此后广阔数学天地的大门。

五次方程是一个更鲜明的例子。

二次、三次、四次方程都有求根公式之后,下一个自然的问题是:

五次方程的通用公式是什么?

数学家寻找这个公式,最后发现:至少按照原本期待的根式形式,它并不存在。

但真正关键的,不只是证明“没有这样的公式”。

关键是问题变了。

人们不再只问“公式是什么”,而开始问一个更深的问题:

什么样的结构,决定一个方程能否用根式求解?

这样一来,问题就不只是计算方程的根了。

它变成了关于置换、对称性与结构的问题。

正是这种转变,催生了伽罗瓦理论和群论式的思考。

原来的问题得到的不只是一个结论。

数学能够提出的问题,整个空间都变了。

设想那时已经有了 AI

这个假想实验,让我更真切地意识到问题在哪里。

假设研究五次方程的数学家,可以直接把方程交给 AI。

AI 给出任意精度的数值解。

问题解决了。

或者,它找到另一套实际使用起来毫无问题的计算方法。

从完成任务的角度看,这个问题已经不存在了。

那为什么还要花几年去问,一个漂亮的符号公式到底存不存在?

为什么要盯着那些失败的尝试?

为什么要研究方程根之间奇特的对称关系?

可更大一片数学天地,恰恰是从这些看似不必要的问题中长出来的。

这让我怀疑,我们有时混淆了两件很不同的事:

完成眼前任务解决当前问题
从探索中学习发掘问题蕴含的全部科学价值

它们不是一回事。

事实上,通往答案最高效的路径,有时会绕开最能孕育科学洞见的路径。

一次失败的证明,可能揭示新的结构。

一个别扭的中间产物,可能成为新的概念。

一个次优解,可能比最优解更能揭示普遍原理。

一个在现有表征下怎么也解不开的问题,可能逼着我们发明更好的表征。

所以,也许需要区分任务效用与认知上的生长力。

对眼下这个问题最好的答案,未必是最能催生新科学的答案。

AI 正在改变过去的关系

过去,几件事常常捆在一起:

解决问题≈发现+理解+形成新问题

科学家在解题过程中,会看见失败的尝试,发现异常,积累直觉,改变表征,最终得到结果。

答案与理解,是一起长出来的。

AI 开始让它们分开。

我们可能越来越常见到:

问题→机器搜索→发现

而科学共同体并没有一起经历产出这个发现的过程。

定理可能是对的。

模拟可能有效。

分子可能满足目标。

算法可能超过此前所有方法。

但我们究竟学到了什么?

它为什么有效?

失败的替代方案揭示了什么?

解法的哪一部分只适用于这个案例,哪一部分可以迁移?

有没有更能说明问题的证明?

这个结果是否说明,我们一直用错了表征方式?

也许最重要的是:

现在,有哪些新问题变得可以提出?

正是在这里,我觉得“理解”远远超出了满足求知欲的范畴。

理解是科学运转的动力之一

过去,我以为理解 AI 生成的科学成果,主要是为了维护人的主体性。

如果人不理解 AI 的发现,又该如何判断、信任、拓展这些发现,或者继续掌握方向?

我仍然觉得这很重要。

我在此前的《理解的瓶颈》中讨论过这个问题:成果来得太快,负责判断或接续它的人,未必来得及形成足够的理解。

但现在,我觉得还有一个更强的理由。

理解本身,就是科学生成下一个问题的过程的一部分。

科学循环不只是 Q → A。它更接近于:

Understanding makes the next question possible
  1. Understanding
  2. Question
  3. Discovery
  4. New understanding
  5. New question

这里的 U 是我们当下的理解,Q 是这种理解让我们看见的问题,D 是探索这些问题带来的发现。

AI 正在迅速变得更擅长中间这一步:Qt → Dt。

但科学的下一片前沿,取决于此后会发生什么。

一次发现,必须改变我们对世界的认识。

新的认识,让不同的问题变得值得关注。

有时,它甚至让过去无法表达的问题第一次变得可以提出。

所以,理解不是因为人类碰巧喜欢解释,才在发现之后额外添上的东西。

它正是一次发现得以生成下一代问题的机制之一。

为什么换一种理解科学的方式,未来就会看起来很不一样?

如果你认为科学主要是在生产答案,最近的 AI 进展看起来可能像终局的开始。

机器正变得异常擅长给出答案。

但如果科学也在不断创造更好的问题和更好的表征,画面就完全不同了。

AI 可能替科学家卸下大量技术负担。

过去需要几年的搜索,也许几小时就能完成。

曾经太昂贵、无法尝试的证明方向,也许很快就能探索。

实验可以在实际开展之前先经过模拟。

一个研究者,也许就能探索过去需要整个实验室才能涉足的空间。

这未必会压缩科学想象力的空间。

也许反而会让想象力发挥更大的作用。

真正稀缺的问题,也许会逐渐从:

我们能解出来吗?

变成:

什么值得问?

再进一步变成:

上一次发现究竟教会了我们什么,让我们终于能够提出过去提不出的问题?

这并不意味着人类会永久垄断提出问题的能力。

未来的 AI 也可能很擅长发明表征、连接领域、规划研究方向。

如果那一天到来,整个局面又会改变。

但至少在今天,我们所说的“AI for Science”,在搜索、生成和检验方面,仍比在重组概念、开辟新领域方面更强。

即便未来机器也擅长后者,仍有另一个独立的问题:

人类是否能充分理解这些概念上的转变,从而真正参与其中?

那么,AI for Science 应该追求什么?

当然不是减少发现。

我们应该让 AI 搜索得比人类更远、更快。

但科学的工作流程,也许不该在答案到来时就结束。

AI 给出一项成果,可能越来越像是第二项任务的起点:

这里有什么值得我们理解?

我们能不能还原关键的概念转变?

能不能找到一个更能启发人的证明?

能不能分清偶然奏效的技巧与可以迁移的原理?

能不能找到那个值得成为新研究对象的异常?

能不能借这个结果,改变我们表征问题的方式?

而新的理解,能不能生成过去看不见的问题?

也许,这才是接下来真正的转变。

很长一段时间里,科学的一项核心稀缺能力,是解决难题。

AI 可能大幅降低这种稀缺性。

下一个瓶颈,可能是我们能否把大量涌现的发现,变成更好的思考方式。

再从这些新的思考方式里,长出更好的问题。

AI 可以让答案变得丰沛。科学的未来,也许取决于我们能否把答案变成理解,再把理解变成过去根本不存在的问题。