Advancing Physical Understanding with Interpretable Machine Learning
Thanks to the extremely large datasets and computing power that have become available in recent years, a new paradigm in scientific discovery has emerged. This new approach is purely data driven, using large amounts of data to train machine-learning models―typically neural networks―to predict the behavior of the natural world [1]. The most prominent achievement of this new methodology has arguably been the AlphaFold model for predicting protein folding (see Research News: Chemistry Nobel Awarded for an AI System That Predicts Protein Structures) [2]. But despite such successes, these data-driven approaches suffer a major drawback in that they are generally “black boxes” that offer no human-accessible understanding of how they make their predictions. This shortcoming also extends to the models’ inputs: It is often desirable to build known domain knowledge into these models, but the data-driven approach excludes that option. Ziming Liu at MIT and colleagues have now made a notable step toward addressing these challenges by developing a machine-learning method designed to discover simple, interpretable laws from data (Fig. 1) [3]. This method could potentially enable the automated discovery of the physical laws governing a wide range of systems.
Since the Scientific Revolution, scientific progress has mostly been made by discovering simple, human-understandable principles that approximately govern systems in the natural world. Prominent examples include Newton’s laws of motion, the laws of quantum mechanics, and those of special and general relativity. Many other simple principles have also been discovered for more specific systems. These simple principles are used to build mathematical models of the underlying natural processes, which, when solved either exactly or numerically, produce predictions that can be tested against experiments and used in downstream engineering tasks. The new data-driven approaches represent a departure from this principles-driven method of investigation.
Methods for discovering physical laws from data are not new. Non-neural-network-based techniques have previously been used for this task―for example, optimization methods based on sparsity-promoting regularizers from compressive sensing [4–6]. More recently, a neural-network-based approach called physics-informed neural networks (PINNs) has gained widespread success [7]. PINNs can be used both to encode physical constraints in the form of differential equations into a neural-network model and to learn differential-equation models from data. What is new about the recent work of Liu and colleagues is the introduction of a novel network architecture that aims to improve the interpretability of the model.
This architecture, called Kolmogorov-Arnold networks (KANs), is based upon the solution to a famous problem in mathematics, called Hilbert’s 13th problem. Roughly speaking, this problem was concerned with whether an arbitrary function of many variables can be represented using compositions of sums and univariate functions. It was famously shown by Andrei Kolmogorov and Vladimir Arnold that this is indeed always the case [8]. KANs have since been used to parameterize multivariate functions via such a composition of sums and univariate functions and learn the univariate functions in this representation from a set of training data. Since the learned part of a KAN is a collection of univariate functions, one can potentially gain insight into the function represented by the KAN―and thus understand what the KAN has learned―by inspecting these univariate functions after training. This arguably makes KANs more interpretable than standard neural networks.
Liu and colleagues apply the KANs architecture to raw dynamical data from a range of test cases. The network identifies energy and angular momentum as being conserved quantities for a 2D harmonic oscillator, finds the Lagrangians for a simple pendulum and for a relativistic mass propelled by a constant external force, discovers the transformation that reveals a hidden symmetry associated with a nonrotating black hole, and quantifies the stress–strain relationship for a neo-Hookean solid. In all these examples, KANs are used in conjunction with domain knowledge, which suggests the correct form of the physical law, to discover physics from data.
By necessity, these initial tests are toy experiments where the correct answer is already known. It will be exciting to see how KANs perform on problems of real scientific interest, where the correct physical laws are not yet known. Applied to such problems, this approach has the potential to significantly accelerate the scientific process. Indeed, in the next few decades the application of machine-learning methods, including KANs, to the scientific process appears poised to lead to breakthroughs that would not otherwise have been possible. The approach looks particularly promising in areas such as many-body quantum systems, chemistry, or materials science―systems where large amounts of data are available but whose computational complexity effectively prohibits ab initio calculations.
References
- Y. LeCun et al., “Deep learning,” Nature 521, 436 (2015).
- J. Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature 596, 583 (2021).
- Z. Liu et al., “Kolmogorov-Arnold networks meet science,” Phys. Rev. X 15, 041051 (2025).
- D. L. Donoho, “Compressed sensing,” IEEE Trans. Inf. Theory 52, 1289 (2006).
- E. J. Candes et al., “Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency information,” IEEE Trans. Inf. Theory 52, 489 (2006).
- H. Schaeffer, “Learning partial differential equations via data discovery and sparse optimization,” Proc. R. Soc. A. 473, 20160446 (2017).
- M. Raissi et al., “Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,” J. Comput. Phys. 378, 686 (2019).
- A. Kolmogorov, “On the representation of continuous functions of several variables by superpositions of continuous functions of a smaller number of variables,” Proc. USSR Acad. Sci. 108, 179 (1956).




