vix.ing · top · new · best · stats · spec

Offset-value coding in database query processing

2022/09/30 by Goetz Graefe, Graefe, Goetz, Do, Thanh
Computer Science · #Advanced Data Storage Technologies #Advanced Database Systems and Queries #Algorithms and Data Compression #Databases (cs.DB) #FOS: Computer and information sciences

paper · pdf · doi:10.48550/arxiv.2210.00034

openalex publication_date 2022/09/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Recent work shows how offset-value coding speeds up database query execution, not only sorting but also duplicate removal and grouping (aggregation) in sorted streams, order-preserving exchange (shuffle), merge join, and more. It already saves thousands of CPUs in Google's Napa and F1 Query systems, e.g., in grouping algorithms and in log-structured merge-forests. In order to realize the full benefit of interesting orderings, however, query execution algorithms must not only consume and exploit offset-value codes but also produce offset-value codes for the next operator in the pipeline. Our research has sought ways to produce offset-value codes without comparing successive output rows one-by-one, column-by-column. This short paper introduces a new theorem and, based on its proof and a simple corollary, describes in detail how order-preserving algorithms (from filter to merge join and even shuffle) can compute offset-value codes for their outputs. These computations are surprisingly simple and very efficient.

Related