Using query logs to establish vocabularies in distributed information retrieval

作者:

Highlights:

摘要

Users of search engines express their needs as queries, typically consisting of a small number of terms. The resulting search engine query logs are valuable resources that can be used to predict how people interact with the search system. In this paper, we introduce two novel applications of query logs, in the context of distributed information retrieval. First, we use query log terms to guide sampling from uncooperative distributed collections. We show that while our sampling strategy is at least as efficient as current methods, it consistently performs better. Second, we propose and evaluate a pruning strategy that uses query log information to eliminate terms. Our experiments show that our proposed pruning method maintains the accuracy achieved by complete indexes, while decreasing the index size by up to 60%. While such pruning may not always be desirable in practice, it provides a useful benchmark against which other pruning strategies can be measured.

论文关键词:Distributed information retrieval,Uncooperative environments,Indexing,Query logs

论文评审过程:Received 11 December 2005, Accepted 3 April 2006, Available online 7 July 2006.

论文官网地址:https://doi.org/10.1016/j.ipm.2006.04.003