Large-time behavior of optimal and ensemble Kalman filters in the linear setting
Data Science Seminar
Sangmin Park
Caltech
Abstract
Filtering, or data assimilation, is the problem of estimating the signal process, initial state of which is unknown, given its partial and noisy observations. Applications of filtering arise ubiquitously engineering and the applied sciences.
From a probabilistic perspective, the best estimate is given by the conditional law of the signal at each time given all observations up to that time, which is called the optimal filtering distribution. However, direct approximation of the optimal filter is computationally intractable in high dimensions, as it requires approximating the Bayes update at each step. The ensemble Kalman filter (EnKF), introduced by Evensen, is a filtering algorithm that is computationally efficient in high-dimensional settings; yet, despite its empirical success, theoretical understanding of its properties in relation to the optimal filter is in its infancy.
In this talk, we will characterize the large-time behavior of the optimal and ensemble Kalman filtering distributions in the setting where the signal and observation dynamics are linear and satisfy the detectability condition. In particular, we will deduce that, for almost every observation path, the (mean-field) EnKF distribution, regardless of initialization, converges asymptotically at an exponential rate towards the optimal filter distribution in the Wasserstein distance.
This talk is based on a joint work in preparation with Minh van Hoang Nguyen, and a joint work with Franca Hoffmann and Andrew M. Stuart (all Caltech).
Brief bio: Sangmin Park is a von Karman Instructor in the Department of Computing and Mathematical Sciences at Caltech. He received his PhD at Carnegie Mellon University in 2025.