2014/09/12 by Daniel Crankshaw, Crankshaw, Daniel, Peter Bailis +13 · 1 voice · 3 citations
Computer Science · Decision Sciences · #Advanced Data Storage Technologies #Advanced Database Systems and Queries #Data Quality and Management #Scientific Computing and Data Management #cs.DB
paper · pdf · doi:10.48550/arxiv.1409.3809
openalex publication_date 2014/09/12 · openalex created_date 2022/08/31 · openalex updated_date 2026/07/28
To support complex data-intensive applications such as personalized\nrecommendations, targeted advertising, and intelligent services, the data\nmanagement community has focused heavily on the design of systems to support\ntraining complex models on large datasets. Unfortunately, the design of these\nsystems largely ignores a critical component of the overall analytics process:\nthe deployment and serving of models at scale. In this work, we present Velox,\na new component of the Berkeley Data Analytics Stack. Velox is a data\nmanagement system for facilitating the next steps in real-world, large-scale\nanalytics pipelines: online model management, maintenance, and serving. Velox\nprovides end-user applications and services with a low-latency, intuitive\ninterface to models, transforming the raw statistical models currently trained\nusing existing offline large-scale compute frameworks into full-blown,\nend-to-end data products capable of recommending products, targeting\nadvertisements, and personalizing web content. To provide up-to-date results\nfor these complex models, Velox also facilitates lightweight online model\nmaintenance and selection (i.e., dynamic weighting). In this paper, we describe\nthe challenges and architectural considerations required to achieve this\nfunctionality, including the abilities to span online and offline systems, to\nadaptively adjust model materialization strategies, and to exploit inherent\nstatistical properties such as model error tolerance, all while operating at\n"Big Data" scale.\n