Publications > Search > Focused Crawling for Structured Data

Focused Crawling for Structured Data

Publication

Nov 4, 2014

Abstract

The Web is rapidly transforming from a pure document collection to the largest connected public data space. Semantic annotations of web pages make it notably easier to extract and reuse data and are increasingly used by both search engines and social media sites to provide better search experiences through rich snippets, faceted search, task completion, etc. In our work, we study the novel problem of crawling structured data embedded inside HTML pages. We describe Anthelion, the first focused crawler addressing this task. We propose new methods of focused crawling specifically designed for collecting data-rich pages with greater efficiency. In particular, we propose a novel combination of online learning and bandit-based explore/exploit approaches to predict data-rich web pages based on the context of the page as well as using feedback from the extraction of metadata from previously seen pages. We show that these techniques significantly outperform state-of-the-art approaches for focused crawling, measured as the ratio of relevant pages and non-relevant pages collected within a given budget.

Download

Venue:

ACM International Conference on Information and Knowledge Management (CIKM 2014)

Type:

Conference/Workshop Paper

Authors:

Robert Meusel
Roi Blanco

BibTeX

@inproceedings{ author = {Robert Meusel and Roi Blanco}, title = {Focused Crawling for Structured Data}, booktitle = {Proceedings of ACM International Conference on Information and Knowledge Management}, year = {2014} }

- Help
- About our ads

Focused Crawling for Structured Data

Publication

Abstract

ACM International Conference on Information and Knowledge Management (CIKM 2014)

Conference/Workshop Paper

Robert Meusel

Roi Blanco

BibTeX