2024/09/03 by Erzhi Liu, Liu, Erzhi, Jerry Yao-Chieh Hu +7
Computer Science · #Artificial Intelligence (cs.AI) #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data
paper · pdf · doi:10.48550/arxiv.2409.01688
openalex publication_date 2024/09/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We introduce a refined differentially private (DP) data structure for kernel density estimation (KDE), offering not only improved privacy-utility tradeoff but also better efficiency over prior results. Specifically, we study the mathematical problem: given a similarity function f (or DP KDE) and a private dataset X ⊂ ℝd, our goal is to preprocess X so that for any query y∈ℝd, we approximate ∑x ∈ X f(x, y) in a differentially private fashion. The best previous algorithm for f(x,y) =‖ x - y ‖1 is the node-contaminated balanced binary tree by [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024]. Their algorithm requires O(nd) space and time for preprocessing with n=|X|. For any query point, the query time is d log n, with an error guarantee of (1+α)-approximation and ε-1 α-0.5 d1.5 R log1.5 n. In this paper, we improve the best previous result [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024] in three aspects: - We reduce query time by a factor of α-1 log n. - We improve the approximation ratio from α to 1. - We reduce the error dependence by a factor of α-0.5. From a technical perspective, our method of constructing the search tree differs from previous work [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024]. In prior work, for each query, the answer is split into α-1 log n numbers, each derived from the summation of log n values in interval tree countings. In contrast, we construct the tree differently, splitting the answer into log n numbers, where each is a smart combination of two distance values, two counting values, and y itself. We believe our tree structure may be of independent interest.