Predictable performance and high query concurrency for data analytics

Candea, George; Polyzotis, Neoklis; Vingralek, Radek

doi:10.1007/s00778-011-0221-2

Candea, George; Polyzotis, Neoklis; Vingralek, Radek

2011

Download

Formats

Format
BibTeX
MARC
MARCXML
DublinCore
EndNote
NLM
RefWorks
RIS

Files

Abstract

Conventional data warehouses employ the query- at-a-time model, which maps each query to a distinct physical plan. When several queries execute concurrently, this model introduces contention and thrashing, because the physical plans—unaware of each other—compete for access to the underlying I/O and computation resources. As a result, while modern systems can efficiently optimize and evaluate a single complex data analysis query, their performance suffers significantly and can be highly erratic when multiple complex queries run at the same time. We present in this paper Cjoin , a new design that substantially improves throughput in large-scale data analytics systems processing many concurrent join queries. In contrast to the conventional query-at-a-time model, our approach employs a single physical plan that shares I/O, computation, and tuple storage across all in-flight join queries. We use an “always on” pipeline of non-blocking operators, managed by a controller that continuously examines the current query mix and optimizes the pipeline on the fly. Our design enables data analytics engines to scale gracefully to large data sets, provide predictable execution times, and reduce contention. We implemented Cjoin as an extension to the PostgreSQL DBMS. This prototype outperforms conventional commercial systems by an order of magnitude for tens to hundreds of concurrent queries.

Details

Title Predictable performance and high query concurrency for data analytics

Author(s) Candea, George ; Polyzotis, Neoklis ; Vingralek, Radek

Published in The VLDB Journal

Volume 20

Issue 2

Pages 227-248

Date 2011

ISSN 0949-877X

Keywords

Join; Concurrency; Warehouse; Sharing; System

DOI https://doi.org/10.1007/s00778-011-0221-2

Other identifier(s) View record in Web of Science

Laboratories DSLAB

Record Appears in Scientific production and competences > I&C - School of Computer and Communication Sciences > IINFCOM > DSLAB - Dependable Systems Laboratory
Work produced at EPFL
Journal Articles
Published

Record creation date 2011-05-24

Actions

Preview

Select file: