Models and algorithms for parallel text retrieval
buir.advisor | Aykanat, Cevdet | |
dc.contributor.author | Cambazoğlu, Berkant Barla | |
dc.date.accessioned | 2016-07-01T11:07:47Z | |
dc.date.available | 2016-07-01T11:07:47Z | |
dc.date.issued | 2006 | |
dc.description | Cataloged from PDF version of article. | en_US |
dc.description.abstract | In the last decade, search engines became an integral part of our lives. The current state-of-the-art in search engine technology relies on parallel text retrieval. Basically, a parallel text retrieval system is composed of three components: a crawler, an indexer, and a query processor. The crawler component aims to locate, fetch, and store the Web pages in a local document repository. The indexer component converts the stored, unstructured text into a queryable form, most often an inverted index. Finally, the query processing component performs the search over the indexed content. In this thesis, we present models and algorithms for efficient Web crawling and query processing. First, for parallel Web crawling, we propose a hybrid model that aims to minimize the communication overhead among the processors while balancing the number of page download requests and storage loads of processors. Second, we propose models for documentand term-based inverted index partitioning. In the document-based partitioning model, the number of disk accesses incurred during query processing is minimized while the posting storage is balanced. In the term-based partitioning model, the total amount of communication is minimized while, again, the posting storage is balanced. Finally, we develop and evaluate a large number of algorithms for query processing in ranking-based text retrieval systems. We test the proposed algorithms over our experimental parallel text retrieval system, Skynet, currently running on a 48-node PC cluster. In the thesis, we also discuss the design and implementation details of another, somewhat untraditional, grid-enabled search engine, SE4SEE. Among our practical work, we present the Harbinger text classification system, used in SE4SEE for Web page classification, and the K-PaToH hypergraph partitioning toolkit, to be used in the proposed models. | en_US |
dc.description.statementofresponsibility | Cambazoğlu, Berkant Barla | en_US |
dc.format.extent | xviii, 180 leaves | en_US |
dc.identifier.itemid | BILKUTUPB100085 | |
dc.identifier.uri | http://hdl.handle.net/11693/29882 | |
dc.language.iso | English | en_US |
dc.rights | info:eu-repo/semantics/openAccess | en_US |
dc.subject | Search engine | en_US |
dc.subject | Parallel text retrieval | en_US |
dc.subject | Web crawling | en_US |
dc.subject | Inverted index partitioning | en_US |
dc.subject | Query processing | en_US |
dc.subject | Text classification | en_US |
dc.subject | Hypergraph partitioning | en_US |
dc.subject.lcc | QA76.5 .C35 2006 | en_US |
dc.subject.lcsh | Parallel processing (Electronic computers). | en_US |
dc.title | Models and algorithms for parallel text retrieval | en_US |
dc.type | Thesis | en_US |
thesis.degree.discipline | Computer Engineering | |
thesis.degree.grantor | Bilkent University | |
thesis.degree.level | Doctoral | |
thesis.degree.name | Ph.D. (Doctor of Philosophy) |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 0003173.pdf
- Size:
- 2.19 MB
- Format:
- Adobe Portable Document Format
- Description:
- Full printable version