Safe data parallelism for general streaming

Schneider S.; Hirzel M.; Gedik, B.; Wu, Kun-Lung

Safe data parallelism for general streaming

Files

Safe data parallelism for general streaming.pdf (1.48 MB)

Date

2015

Authors

Schneider S.

Hirzel M.

Gedik, B.

Wu, Kun-Lung

BUIR Usage Stats

1
views

25
downloads

Citation Stats

Abstract

Streaming applications process possibly infinite streams of data and often have both high throughput and low latency requirements. They are comprised of operator graphs that produce and consume data tuples. General streaming applications use stateful, selective, and user-defined operators. The stream programming model naturally exposes task and pipeline parallelism, enabling it to exploit parallel systems of all kinds, including large clusters. However, data parallelism must either be manually introduced by programmers, or extracted as an optimization by compilers. Previous data parallel optimizations did not apply to selective, stateful and user-defined operators. This article presents a compiler and runtime system that automatically extracts data parallelism for general stream processing. Data-parallelization is safe if the transformed program has the same semantics as the original sequential version. The compiler forms parallel regions while considering operator selectivity, state, partitioning, and graph dependencies. The distributed runtime system ensures that tuples always exit parallel regions in the same order they would without data parallelism, using the most efficient strategy as identified by the compiler. Our experiments using 100 cores across 14 machines show linear scalability for parallel regions that are computation-bound, and near linear scalability when tuples are shuffled across parallel regions.

Source Title

IEEE Transactions on Computers

Publisher

Institute of Electrical and Electronics Engineers

Keywords

Data processing, Distributed computing, Data handling, Data processing, Distributed computer systems, Parallel programming, Scalability, Semantics, Data parallelism, Data parallelization, Distributed runtime, Efficient strategy, Pipeline parallelisms, Stream processing, Stream programming, Streaming applications, Program compilers

Permalink

http://hdl.handle.net/11693/22636

Published Version (Please cite this version)

http://dx.doi.org/10.1109/TC.2013.221

Collections

Scholarly Publications - Computer Engineering

Language

English

Type

Article

Full item page

Safe data parallelism for general streaming

Files

Date

Authors

Editor(s)

Advisor

Supervisor

Co-Advisor

Co-Supervisor

Instructor

BUIR Usage Stats

Citation Stats

Series

Abstract

Source Title

Publisher

Course

Other identifiers

Book Title

Keywords

Degree Discipline

Degree Level

Degree Name

Citation

Permalink

Published Version (Please cite this version)

Collections

Language

Type

Safe data parallelism for general streaming

Files

Date

Authors

Editor(s)

Advisor

Supervisor

Co-Advisor

Co-Supervisor

Instructor

BUIR Usage Stats

Citation Stats

Share

Series

Abstract

Source Title

Publisher

Course

Other identifiers

Book Title

Keywords

Degree Discipline

Degree Level

Degree Name

Citation

Permalink

Published Version (Please cite this version)

Collections

Language

Type