Characteristics of Web-based textual communications
buir.advisor | Aykanat, Cevdet | |
dc.contributor.author | Küçükyılmaz, Tayfun | |
dc.date.accessioned | 2016-01-08T18:19:40Z | |
dc.date.available | 2016-01-08T18:19:40Z | |
dc.date.issued | 2012 | |
dc.description | Ankara : The Department of Computer Engineering and the Graduate School of Engineering and Science of Bilkent University 2012. | en_US |
dc.description | Thesis (Ph. D.) -- Bilkent University, 2012. | en_US |
dc.description | Includes bibliographical references. | en_US |
dc.description.abstract | In this thesis, we analyze different aspects of Web-based textual communications and argue that all such communications share some common properties. In order to provide practical evidence for the validity of this argument, we focus on two common properties by examining these properties on various types of Web-based textual communications data. These properties are: All Web-based communications contain features attributable to their author and reciever; and all Web-based communications exhibit similar heavy tailed distributional properties. In order to provide practical proof for the validity of our claims, we provide three practical, real life research problems and exploit the proposed common properties of Web-based textual communications to find practical solutions to these problems. In this work, we first provide a feature-based result caching framework for real life search engines. To this end, we mined attributes from user queries in order to classify queries and estimate a quality metric for giving admission and eviction decisions for the query result cache. Second, we analyzed messages of an online chat server in order to predict user and mesage attributes. Our results show that several user- and message-based attributes can be predicted with significant occuracy using both chat message- and writing-style based features of the chat users. Third, we provide a parallel framework for in-memory construction of term partitioned inverted indexes. In this work, in order to minimize the total communication time between processors, we provide a bucketing scheme that is based on term-based distributional properties of Web page contents. | en_US |
dc.description.provenance | Made available in DSpace on 2016-01-08T18:19:40Z (GMT). No. of bitstreams: 1 0006246.pdf: 1061769 bytes, checksum: 162c281b958bbddb4ba6f82b419c6237 (MD5) | en |
dc.description.statementofresponsibility | Küçükyılmaz, Tayfun | en_US |
dc.format.extent | xvii, 176 leaves | en_US |
dc.identifier.uri | http://hdl.handle.net/11693/15512 | |
dc.language.iso | English | en_US |
dc.rights | info:eu-repo/semantics/openAccess | en_US |
dc.subject | Web search engine | en_US |
dc.subject | result caching | en_US |
dc.subject | cache | en_US |
dc.subject | chat mining | en_US |
dc.subject | data mining | en_US |
dc.subject | index inversion | en_US |
dc.subject | inverted index | en_US |
dc.subject | posting list | en_US |
dc.subject.lcc | TK5105.888 .K83 2012 | en_US |
dc.subject.lcsh | World Wide Web. | en_US |
dc.subject.lcsh | Web search engines. | en_US |
dc.subject.lcsh | Data mining. | en_US |
dc.subject.lcsh | Indexing. | en_US |
dc.title | Characteristics of Web-based textual communications | en_US |
dc.type | Thesis | en_US |
thesis.degree.discipline | Computer Engineering | |
thesis.degree.grantor | Bilkent University | |
thesis.degree.level | Doctoral | |
thesis.degree.name | Ph.D. (Doctor of Philosophy) |
Files
Original bundle
1 - 1 of 1