vix.ing · top · new · best · stats · spec

A Data-Driven Method for Automated Data Superposition with Applications in Soft Matter Science

2022/04/20 by Kyle R. Lennon, Gareth H. McKinley, Lennon, Kyle R. +3 · 1 citation
Biochemistry, Genetics and Molecular Biology · Materials Science · Neuroscience · #Computational Physics (physics.comp-ph) #Data Analysis #FOS: Computer and information sciences #FOS: Physical sciences #Functional Brain Connectivity Studies #Machine Learning (cs.LG) #Machine Learning in Materials Science #Metabolomics and Mass Spectrometry Studies #Soft Condensed Matter (cond-mat.soft) #Statistics and Probability (physics.data-an)

paper · pdf · doi:10.48550/arxiv.2204.09521

openalex publication_date 2022/04/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The superposition of data sets with internal parametric self-similarity is a longstanding and widespread technique for the analysis of many types of experimental data across the physical sciences. Typically, this superposition is performed manually, or recently by one of a few automated algorithms. However, these methods are often heuristic in nature, are prone to user bias via manual data shifting or parameterization, and lack a native framework for handling uncertainty in both the data and the resulting model of the superposed data. In this work, we develop a data-driven, non-parametric method for superposing experimental data with arbitrary coordinate transformations, which employs Gaussian process regression to learn statistical models that describe the data, and then uses maximum a posteriori estimation to optimally superpose the data sets. This statistical framework is robust to experimental noise, and automatically produces uncertainty estimates for the learned coordinate transformations. Moreover, it is distinguished from black-box machine learning in its interpretability -- specifically, it produces a model that may itself be interrogated to gain insight into the system under study. We demonstrate these salient features of our method through its application to four representative data sets characterizing the mechanics of soft materials. In every case, our method replicates results obtained using other approaches, but with reduced bias and the addition of uncertainty estimates. This method enables a standardized, statistical treatment of self-similar data across many fields, producing interpretable data-driven models that may inform applications such as materials classification, design, and discovery.

Cited by

Related