Accelerating the HyperLogLog cardinality estimation algorithm

Bozkus, C.; Fraguela, B. B.

Accelerating the HyperLogLog cardinality estimation algorithm

Files

Accelerating the HyperLogLog Cardinality Estimation Algorithm.pdf (1.39 MB)

Date

2017

Authors

Bozkus, C.

Fraguela, B. B.

BUIR Usage Stats

2
views

13
downloads

Citation Stats

Abstract

In recent years, vast amounts of data of different kinds, from pictures and videos from our cameras to software logs from sensor networks and Internet routers operating day and night, are being generated. This has led to new big data problems, which require new algorithms to handle these large volumes of data and as a result are very computationally demanding because of the volumes to process. In this paper, we parallelize one of these new algorithms, namely, the HyperLogLog algorithm, which estimates the number of different items in a large data set with minimal memory usage, as it lowers the typical memory usage of this type of calculation from O(n) to O(1). We have implemented parallelizations based on OpenMP and OpenCL and evaluated them in a standard multicore system, an Intel Xeon Phi, and two GPUs from different vendors. The results obtained in our experiments, in which we reach a speedup of 88.6 with respect to an optimized sequential implementation, are very positive, particularly taking into account the need to run this kind of algorithm on large amounts of data. © 2017 Cem Bozkus and Basilio B. Fraguela.

Source Title

Scientific Programming

Publisher

Hindawi Limited

Keywords

Application programming interfaces (API), Program processors, Sensor networks, Cardinality estimations, Internet routers, Large amounts of data, Large datasets, Multi-core systems, Parallelizations, Sequential implementation, Software logs, Big data

Permalink

http://hdl.handle.net/11693/37065

Published Version (Please cite this version)

https://doi.org/10.1155/2017/2040865

Collections

Scholarly Publications - Computer Engineering

Language

English

Type

Article

Full item page

Accelerating the HyperLogLog cardinality estimation algorithm

Files

Date

Authors

Editor(s)

Advisor

Supervisor

Co-Advisor

Co-Supervisor

Instructor

BUIR Usage Stats

Citation Stats

Series

Abstract

Source Title

Publisher

Course

Other identifiers

Book Title

Keywords

Degree Discipline

Degree Level

Degree Name

Citation

Permalink

Published Version (Please cite this version)

Collections

Language

Type

Accelerating the HyperLogLog cardinality estimation algorithm

Files

Date

Authors

Editor(s)

Advisor

Supervisor

Co-Advisor

Co-Supervisor

Instructor

BUIR Usage Stats

Citation Stats

Share

Series

Abstract

Source Title

Publisher

Course

Other identifiers

Book Title

Keywords

Degree Discipline

Degree Level

Degree Name

Citation

Permalink

Published Version (Please cite this version)

Collections

Language

Type