Repository logo

Infoscience

  • English
  • French
Log In
Logo EPFL, École polytechnique fédérale de Lausanne

Infoscience

  • English
  • French
Log In
  1. Home
  2. Academic and Research Output
  3. Reports, Documentation, and Standards
  4. Query-Driven Indexing for Peer-to-Peer Text Retrieval
 
report

Query-Driven Indexing for Peer-to-Peer Text Retrieval

Skobeltsyn, Gleb  
•
Luu, Toan
•
Podnar, Ivana
Show more
2006

We present a query-driven algorithm for the distributed indexing of large document collections within structured P2P networks. To cope with the bandwidth consumption problem that has been identified as the major issue for the standard distributed approach using single term indexing, we leverage a distributed index that stores top-k document references only for carefully chosen indexing term combinations. In addition, since the number of possible term combinations extracted from a document collection can still be very large, we propose to use query statistics to index only such combination that are indeed frequently requested by the users. Thus, by avoiding the maintenance of superfluous indexing information, we achieve a substantial reduction in bandwidth and storage. A specific activation mechanism is applied to take into account changes in the query distribution, resulting in an efficient, constantly evolving query driven indexing structure. Moreover, our approach facilitates adjusting the indexing load according to the resources provided by the peers in the network. We claim that the size of the index and the generated indexing/retrieval traffic remains manageable even for web-size document collections at a price of a marginal loss in recall for rare queries. Our theoretical analysis and experimental results provide convincing evidence about the feasibility of the query-driven indexing strategy for large scale P2P text retrieval. Furthermore, our experiments confirm that the retrieval performance is only slightly lower than the one obtained with state-of-the-art, centralized query engines, such as Google or Yahoo.

  • Files
  • Details
  • Metrics
Loading...
Thumbnail Image
Name

Query-Driven Indexing for Peer-to-Peer Text Retrieval.pdf

Access type

restricted

Size

290.24 KB

Format

Adobe PDF

Checksum (MD5)

9b831d2b2da1b88c34ecf28a829466e8

Logo EPFL, École polytechnique fédérale de Lausanne
  • Contact
  • infoscience@epfl.ch

  • Follow us on Facebook
  • Follow us on Instagram
  • Follow us on LinkedIn
  • Follow us on X
  • Follow us on Youtube
AccessibilityLegal noticePrivacy policyCookie settingsEnd User AgreementGet helpFeedback

Infoscience is a service managed and provided by the Library and IT Services of EPFL. © EPFL, tous droits réservés