Browsing by Subject "Markup languages"
Now showing 1 - 2 of 2
- Results Per Page
- Sort Options
Item Open Access Exploiting index pruning methods for clustering XML collections(Springer, Berlin, Heidelberg, 2010) Altıngövde, İsmail Şengör; Atılgan, Duygu; Ulusoy, ÖzgürIn this paper, we first employ the well known Cover-Coefficient Based Clustering Methodology (C3M) for clustering XML documents. Next, we apply index pruning techniques from the literature to reduce the size of the document vectors. Our experiments show that for certain cases, it is possible to prune up to 70% of the collection (or, more specifically, underlying document vectors) and still generate a clustering structure that yields the same quality with that of the original collection, in terms of a set of evaluation metrics. © 2010 Springer-Verlag Berlin Heidelberg.Item Open Access XML retrieval using pruned element-index files(Springer, Berlin, Heidelberg, 2010) Altıngövde, İsmail Şengör; Atılgan, Duygu; Ulusoy, ÖzgürAn element-index is a crucial mechanism for supporting content-only (CO) queries over XML collections. A full element-index that indexes each element along with the content of its descendants involves a high redundancy and reduces query processing efficiency. A direct index, on the other hand, only indexes the content that is directly under each element and disregards the descendants. This results in a smaller index, but possibly in return to some reduction in system effectiveness. In this paper, we propose using static index pruning techniques for obtaining more compact index files that can still result in comparable retrieval performance to that of a full index. We also compare the retrieval performance of these pruning based approaches to some other strategies that make use of a direct element-index. Our experiments conducted along with the lines of INEX evaluation framework reveal that pruned index files yield comparable to or even better retrieval performance than the full index and direct index, for several tasks in the ad hoc track. © 2010 Springer-Verlag Berlin Heidelberg.