2009/07/14 by Martin Klein, Michael L. Nelson, Klein, Martin +2 · 1 voice · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Digital Libraries (cs.DL) #FOS: Computer and information sciences #H.3.3 #Information Retrieval (cs.IR) #Information Retrieval and Search Behavior #Machine Learning in Bioinformatics #Text and Document Classification Technologies #Web Data Mining and Analysis #cs.DL #cs.IR
paper · pdf · doi:10.48550/arxiv.0907.2268
10 pages, 11 figures, 5 tables, 40 references, accepted for publication at JCDL 2010 in Brisbane, Australia
openalex publication_date 2009/07/14 · arxiv published 2009/07/14 · arxiv created 2010/04/16 · arxiv updated 2010/04/16 · openalex created_date 2025/10/24 · openalex updated_date 2026/07/28
Missing web pages (pages that return the 404 "Page Not Found" error) are part of the browsing experience. The manual use of search engines to rediscover missing pages can be frustrating and unsuccessful. We compare four automated methods for rediscovering web pages. We extract the page's title, generate the page's lexical signature (LS), obtain the page's tags from the bookmarking website delicious.com and generate a LS from the page's link neighborhood. We use the output of all methods to query Internet search engines and analyze their retrieval performance. Our results show that both LSs and titles perform fairly well with over 60% URIs returned top ranked from Yahoo!. However, the combination of methods improves the retrieval performance. Considering the complexity of the LS generation, querying the title first and in case of insufficient results querying the LSs second is the preferable setup. This combination accounts for more than 75% top ranked URIs.