Models and algorithms for parallel text retrieval

Cambazoğlu, Berkant Barla

Models and algorithms for parallel text retrieval

buir.advisor	Aykanat, Cevdet
dc.contributor.author	Cambazoğlu, Berkant Barla
dc.date.accessioned	2016-07-01T11:07:47Z
dc.date.available	2016-07-01T11:07:47Z
dc.date.issued	2006
dc.description	Cataloged from PDF version of article.	en_US
dc.description.abstract	In the last decade, search engines became an integral part of our lives. The current state-of-the-art in search engine technology relies on parallel text retrieval. Basically, a parallel text retrieval system is composed of three components: a crawler, an indexer, and a query processor. The crawler component aims to locate, fetch, and store the Web pages in a local document repository. The indexer component converts the stored, unstructured text into a queryable form, most often an inverted index. Finally, the query processing component performs the search over the indexed content. In this thesis, we present models and algorithms for efficient Web crawling and query processing. First, for parallel Web crawling, we propose a hybrid model that aims to minimize the communication overhead among the processors while balancing the number of page download requests and storage loads of processors. Second, we propose models for documentand term-based inverted index partitioning. In the document-based partitioning model, the number of disk accesses incurred during query processing is minimized while the posting storage is balanced. In the term-based partitioning model, the total amount of communication is minimized while, again, the posting storage is balanced. Finally, we develop and evaluate a large number of algorithms for query processing in ranking-based text retrieval systems. We test the proposed algorithms over our experimental parallel text retrieval system, Skynet, currently running on a 48-node PC cluster. In the thesis, we also discuss the design and implementation details of another, somewhat untraditional, grid-enabled search engine, SE4SEE. Among our practical work, we present the Harbinger text classification system, used in SE4SEE for Web page classification, and the K-PaToH hypergraph partitioning toolkit, to be used in the proposed models.	en_US
dc.description.statementofresponsibility	Cambazoğlu, Berkant Barla	en_US
dc.format.extent	xviii, 180 leaves	en_US
dc.identifier.itemid	BILKUTUPB100085
dc.identifier.uri	http://hdl.handle.net/11693/29882
dc.language.iso	English	en_US
dc.rights	info:eu-repo/semantics/openAccess	en_US
dc.subject	Search engine	en_US
dc.subject	Parallel text retrieval	en_US
dc.subject	Web crawling	en_US
dc.subject	Inverted index partitioning	en_US
dc.subject	Query processing	en_US
dc.subject	Text classification	en_US
dc.subject	Hypergraph partitioning	en_US
dc.subject.lcc	QA76.5 .C35 2006	en_US
dc.subject.lcsh	Parallel processing (Electronic computers).	en_US
dc.title	Models and algorithms for parallel text retrieval	en_US
dc.type	Thesis	en_US
thesis.degree.discipline	Computer Engineering
thesis.degree.grantor	Bilkent University
thesis.degree.level	Doctoral
thesis.degree.name	Ph.D. (Doctor of Philosophy)

Files

Original bundle

Now showing 1 - 1 of 1

Name:: 0003173.pdf
Size:: 2.19 MB
Format:: Adobe Portable Document Format
Description:: Full printable version

Download

Collections

Graduate School of Engineering and Science