vix.ing · top · new · best · stats

To pipeline or not to pipeline, that is the question

2020/02/03 by Harshad Deshmukh, Deshmukh, Harshad, Bruhathi Sundarmurthy +3 · 1 citation
Computer Science · #Advanced Data Storage Technologies #Advanced Database Systems and Queries #Blocking (statistics) #Computer science #Databases (cs.DB) #Distributed systems and fault tolerance #FOS: Computer and information sciences #Information retrieval #Key (lock) #Memory footprint #Operating system #Parallel computing #Pipeline (software) #Programming language #Query plan #Sargable #Search engine #Simple (philosophy) #Terminology #Theoretical computer science #Transfer (computing) #cs.DB

paper · pdf · doi:10.48550/arxiv.2002.00866

published in arXiv (Cornell University) (Cornell University)

arxiv created 2020/02/03 · openalex publication_date 2020/02/03 · arxiv updated 2020/02/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

In designing query processing primitives, a crucial design choice is the method for data transfer between two operators in a query plan. As we were considering this critical design mechanism for an in-memory database system that we are building, we quickly realized that (surprisingly) there isn't a clear definition of this concept. Papers are full or ad hoc use of terms like pipelining and blocking, but as these terms are not crisply defined, it is hard to fully understand the results attributed to these concepts. To address this limitation, we introduce a clear terminology for how to think about data transfer between operators in a query pipeline. We show that there isn't a clear definition of pipelining and blocking, and that there is a full spectrum of techniques based on a simple concept called unit-of-transfer. Next, we develop an analytical model for inter-operator communication, and highlight the key parameters that impact performance (for in-memory database settings). Armed with this model, we then apply it to the system we are designing and highlight the insights we gathered from this exercise. We find that the gap between pipelining and non-pipelining query execution, w.r.t. key factors such as performance and memory footprint is quite narrow, and thus system designers should likely rethink the notion of pipelining vs. blocking for in-memory database systems.

Citations

Related