メインコンテンツに移動

From Navier-Stokes to manuscripts by the kilo: AI’s impact on quant research

As AI produces results in bulk, finance needs trusted ways to assess provenance, reliability and significance

Art sketch image vertical collage of scales weight knowledge comparison book stack artificial intellect head cyber innovation.

At the start of October, we were discussing whether a Navier-Stokes-style controversy was ever likely to arise in quantitative finance.

Our first intuition was that it was much less likely in a field without equally prominent problems or rich bounties attached to their solution. Our second was that we should nonetheless remain on our toes: proprietary ideas, data and partial results could still enter AI systems – and any suspicion of appropriation in finance would probably be much harder to detect.

October 6 changed the picture entirely.

That was the day OpenAI released 722 mathematical manuscripts produced by an internal AI model. The results have not all been independently verified, but their scale is the story. OpenAI says the model was presented with approximately 4,000 problems and produced manuscripts grouped into 372 families.

Only a month earlier, OpenAI’s proposed solution to the Navier-Stokes existence and smoothness problem appeared to be a singular event. Now, it looked more like one result produced by a new species of research-generating machine.

It’s a turn of events that brings to mind the short poem Frimärken by Swedish poet Siv Widerberg:

Jag samlade frimärken.
Pappa gav mig ett halvt kilo.
Jag samlade inte frimärken mer.

Loosely translated:

I used to collect stamps.
One day, my dad gave me half a kilo of stamps.
I don’t collect them anymore.

Delivered by the kilo, the joy of collecting is removed from the equation.

It’s not that research loses its value simply because more of it can be produced. But part of the worth and significance of building a collection lies in the process: searching, selecting, comparing and understanding what each addition contributes.

When results arrive by the kilo, it’s a different proposition.

But let us revisit our Navier-Stokes considerations anyway.

Until now, the most visible demonstrations of artificial intelligence in mathematics have centred on individual achievements. Navier-Stokes is one of the seven Millennium Prize Problems, for which the Clay Mathematics Institute (Oxford) offers a $1 million prize.

According to OpenAI, its final effort involved roughly 10,000 agents working for 88 hours. The announcement offered a simple and globally interpretable signal of AI capability that was guaranteed to turn heads around the globe: a celebrated human intellectual challenge appeared to have fallen.

It was dubbed maths’ Deep Blue vs Kasparov moment.

When answers become abundant, the societal agreement of what is worth our attention persists. It is just as – if not more – valuable than ever before

Quantitative finance has no equivalent list of universally recognised problems. Its important questions are fragmented across asset classes and applications. Even Risk.net readers might disagree over what would constitute the field’s greatest outstanding challenge.

This seemed to make finance less attractive to an AI company seeking spectacular, publicity-generating demonstration of capabilities in the field. We were assessing the odds of this reasoning in our discussions right up until OpenAI’s multiple kilos of manuscripts made the news.

The release of hundreds of mathematical manuscripts forces us to rethink our argument. It now seems that an individual problem no longer needs global prestige to be worth feeding into the solution-generating machine. The ability to attack thousands of questions and produce hundreds of potentially significant results is itself a demonstration of technological capability.

Quant finance offers no shortage of specific research questions that AI could tackle in parallel, across derivatives pricing and hedging, portfolio construction, execution, market-impact estimation, volatility forecasting, stress-testing, limit-order-book modelling, synthetic market generation and many more.

No individual result needs to make global headlines for the collective effort to be potentially valuable. Nor would an AI company need to become a hedge fund to extract value from these results. Financial research could be monetised in many ways – through improved systems, software infrastructure or partnerships with banks and asset managers, for example. Searching across thousands of technical questions might identify a small number of results capable of improving these services.

This week’s developments suggest our original questions – whether AI will solve one or many celebrated problems in quantitative finance or whether the field’s problems are prominent enough to be worth addressing at all for AI companies – are the wrong way to think about what these advances might mean for quant finance. The more urgent question is what happens when candidate models and research ideas can be generated in bulk.

A quieter provenance problem?

Another aspect of the Navier-Stokes comparison from our original discussion remains relevant: could an AI-generated result in quantitative finance provoke a similar public controversy over provenance? We believe the answer to that probably remains negative – for very practical reasons. To make a comparison let’s start with a quick recap of which aspects of the controversy we discussed.

The Navier-Stokes announcement produced a controversy over the possible role of unpublished human research. New York University professor Tristan Buckmaster and Anthropic researcher Levent Alpöge had used several AI systems, including OpenAI’s Codex, while developing blowup results for related fluid equations. This raised the possibility that their intermediate work had passed through tools operated by a company subsequently working on a related problem.

In many practical financial settings, the stakes are monetary, and profiting from an idea may depend on keeping it secret

OpenAI said neither its researchers nor its agents had accessed the pair’s work. Following an investigation, it said Buckmaster’s Codex prompts from the preceding two months could not have influenced its system, including through training. Buckmaster remained unconvinced.

This week’s OpenAI announcement changes the balance of circumstantial evidence on this matter too. But let us pause and rewind to September 10 – the moment of the Navier-Stokes dispute’s zenith – because this is the moment that highlighted that we don’t yet have the necessary infrastructure and frameworks to handle such provenance disputes, as Buckmaster’s note points out.

It helps to distinguish three related issues: whether unpublished material entered the AI system’s weights – and could be accessible in one way or another to competitors; whether it contributed to a particular result; and what this means for attribution. This week’s announcement may soften the worries for the second question here, but leaves open questions about how to handle provenance in future and how researchers should be credited.

In practice, provenance disputes or suspicions in finance would probably take a different form and attract far less public attention – which would not make them any less important.

An excellent trading strategy, data analytics insight or execution method can be directly monetised and a pricing technique can become valuable when incorporated into a company’s financial products. And with that, the financial version of a (suspected) leak controversy would probably be much less visible. 

In academic mathematics, the main reward is often public credit, so disputes tend to produce competing papers, public statements and demands for attribution. In many practical financial settings, the stakes are monetary, and profiting from an idea may depend on keeping it secret. Someone whose work is affected by a leak may struggle to prove what happened – or decide that publicising the loss is not worth the cost. In both settings, ordinary precautions are increasingly important.

An abundance of answers

When answers become abundant, the societal agreement of what is worth our attention persists. It is just as – if not more – valuable than ever before.

In mathematics, a result can at least in principle be checked against a formal statement. Lean software can verify that a proof follows from specified assumptions, provided the formalisation faithfully captures the original argument.

Quant finance doesn’t always offer such a clean certificate.

Some claims can be established mathematically – a pricing construction can be proved arbitrage-free – but many similarly valuable claims are empirical. Does a hedge remain effective during a regime change? Does a market simulator reproduce the dynamics relevant to the decision for which it will be used?

Multiple testing, leakage, overfitting and the winner’s curse do not disappear when the researcher is a machine, and automation can amplify them. A good research record should transparently document the process behind the winning model: which alternatives were tested, which data informed their selection, and how much and what kind of optimisation preceded the reported result.

Independent, leakage-resistant benchmarks consequently become more important.

It is human communities and institutions that determine which problems are worth posing, which evidence is persuasive, which trade-offs are acceptable and which contributions should shape the field

Evaluation should use locked holdout data, multiple market regimes and, where possible, forward testing. Benchmarks must also be renewed: once the same test set has repeatedly been queried by researchers or AI agents, it no longer provides genuinely out-of-sample evidence.

Following the Navier-Stokes announcement, 25 Fields medallists warned against treating a verified answer as a substitute for understanding why a result holds. This distinction is just as important in finance. A model may be mathematically consistent and empirically successful on a particular dataset without revealing what economic mechanism it has captured – or what breaks if the assumptions change.

This creates an essential role for researchers, practitioners and specialist publications. Their task may increasingly become to identify hidden assumptions, distinguish robust findings from statistical coincidence, provide context and decide which contributions deserve attention.

This may sound self-serving in a specialist publication, but this last filtering problem of what merits attention remains a societal consensus. Capturing and reflecting that consensus transparently is ever more important.

For now, this part of the research process – the societal agreement about what is relevant, interesting or fashionable, and what will turn heads – remains unchanged

To this point, the Navier-Stokes problem would likely never have been selected for OpenAI’s late-summer campaign had generations of mathematicians and institutions not already established its exceptional status. AI did not decide that Navier-Stokes mattered – it inherited that judgement from us.

The same will be true in quantitative finance. AI may produce more derivations, write more code and generate more candidate models. But it is human communities and institutions that determine which problems are worth posing, which evidence is persuasive, which trade-offs are acceptable and which contributions should shape the field.

In the coming flood of results, it therefore becomes imperative that we build, recognise and nurture trust in the channels through which we discuss and determine what matters: universities, journals, specialist publications, independent benchmark providers and professional communities.

How these channels should operate, how their independence should be protected and how they can retain legitimacy when no human can read everything is yet to be determined.

AI may increasingly generate the answers. The societal consensus that gives those answers direction and value remains ours.

It just remains to be seen how.

Blanka Horvath is associate professor of mathematics at the University of Oxford

Mauro Cesa is editor, Quant Finance at Risk.net

 

コンテンツを印刷またはコピーできるのは、有料の購読契約を結んでいるユーザー、または法人購読契約の一員であるユーザーのみです。

これらのオプションやその他の購読特典を利用するには、info@risk.net にお問い合わせいただくか、こちらの購読オプションをご覧ください: http://subscriptions.risk.net/subscribe

現在、このコンテンツをコピーすることはできません。詳しくはinfo@risk.netまでお問い合わせください。

Most read articles loading...

You need to sign in to use this feature. If you don’t have a Risk.net account, please register for a trial.

ログイン
You are currently on corporate access.

To use this feature you will need an individual account. If you have one already please sign in.

Sign in.

Alternatively you can request an individual account here