2022/12/05 by Jan Weinreich, Guido Falk von Rudorff, Weinreich, Jan +3 · 1 citation
Chemistry · Computer Science · Materials Science · #Chemical Physics (physics.chem-ph) #Computational Drug Discovery Methods #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Machine Learning in Materials Science #Mass Spectrometry Techniques and Applications #Materials Science (cond-mat.mtrl-sci)
paper · pdf · doi:10.48550/arxiv.2212.04322
openalex publication_date 2022/12/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Large machine learning models with improved predictions have become widely available in the chemical sciences. Unfortunately, these models do not protect the privacy necessary within commercial settings, prohibiting the use of potentially extremely valuable data by others. Encrypting the prediction process can solve this problem by double-blind model evaluation and prohibits the extraction of training or query data. However, contemporary ML models based on fully homomorphic encryption or federated learning are either too expensive for practical use or have to trade higher speed for weaker security. We have implemented secure and computationally feasible encrypted machine learning models using oblivious transfer enabling and secure predictions of molecular quantum properties across chemical compound space. However, we find that encrypted predictions using kernel ridge regression models are a million times more expensive than without encryption. This demonstrates a dire need for a compact machine learning model architecture, including molecular representation and kernel matrix size, that minimizes model evaluation costs.