Application of map/reduce paradigm in supercomputing systems

Demirci, Gündüz Vehbi

Application of map/reduce paradigm in supercomputing systems

buir.advisor	Aykanat, Cevdet
dc.contributor.author	Demirci, Gündüz Vehbi
dc.date.accessioned	2016-01-08T18:26:34Z
dc.date.available	2016-01-08T18:26:34Z
dc.date.issued	2013
dc.description	Cataloged from PDF version of article.	en_US
dc.description	Includes bibliographical references leaves 47-49.	en_US
dc.description.abstract	Map/Reduce is a framework first introduced by Google in order to rapidly develop big data analytic applications on distributed computing systems. Even though the Map/Reduce paradigm had a game changing impact on certain fields of computer science such as information retrieval and data mining, it did not have such an impact on the scientific computing domain yet. The current implementations of Map/Reduce are especially designed for commodity PC clusters, where failures of compute nodes are common and inter-processor communication is slow. However, scientific computing applications are usually executed on high performance computing (HPC) systems and such systems provide high communication bandwidth with low message latency where failures of processors are rare. Therefore, Map/Reduce framework causes performance degradation and becomes less preferable in scientific computing domain. Due to these reasons, specific implementations of Map/Reduce paradigm are needed for scientific computing domain. Among the existing implementations, we focus our attention on the MapReduce-MPI (MR-MPI) library developed at Sandia National Labs. In this thesis, we argue that by utilizing MR-MPI Library, the Map/Reduce programming paradigm can be successfully utilized for scientific computing applications that require scalability and performance. We tested MR-MPI Library in HPC systems with several fundamental algorithms that are frequently used in scientific computing and data mining domains. Implemented algorithms include all-pair-similarity-search (APSS), all-pair-shortest-path (APSP), and page-rank (PR). Tests were performed on well-known large-scale HPC systems IBM BlueGene/Q (Juqueen) and Cray XE6 (Hermit) to examine scalability and speedup of these algorithms.	en_US
dc.description.statementofresponsibility	Demirci, Gündüz Vehbi	en_US
dc.format.extent	x, 49 leaves, charts, tables	en_US
dc.identifier.itemid	B139371
dc.identifier.uri	http://hdl.handle.net/11693/15906
dc.language.iso	English	en_US
dc.rights	info:eu-repo/semantics/openAccess	en_US
dc.subject	Map/Reduce	en_US
dc.subject	Big Data	en_US
dc.subject	Data mining	en_US
dc.subject	Information Retrieval	en_US
dc.subject	Distributed Computing Systems	en_US
dc.subject.lcc	QA76.9.D343 D45 2013	en_US
dc.subject.lcsh	Big data.	en_US
dc.subject.lcsh	Data mining.	en_US
dc.subject.lcsh	Information retrieval.	en_US
dc.subject.lcsh	Supercomputers.	en_US
dc.subject.lcsh	Science--Data processing.	en_US
dc.title	Application of map/reduce paradigm in supercomputing systems	en_US
dc.type	Thesis	en_US
thesis.degree.discipline	Computer Engineering
thesis.degree.grantor	Bilkent University
thesis.degree.level	Master's
thesis.degree.name	MS (Master of Science)

Files

Original bundle

Now showing 1 - 1 of 1

Name:: 0006604.pdf
Size:: 528.81 KB
Format:: Adobe Portable Document Format

Download

Collections

Graduate School of Engineering and Science