Skip to main content

2004 | OriginalPaper | Buchkapitel

Mining Web Sites Using Wrapper Induction, Named Entities, and Post-processing

verfasst von : Georgios Sigletos, Georgios Paliouras, Constantine D. Spyropoulos, Michalis Hatzopoulos

Erschienen in: Web Mining: From Web to Semantic Web

Verlag: Springer Berlin Heidelberg

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config
loading …

This paper presents a new framework for extracting information from collections of Web pages across different sites. In the proposed framework, a standard wrapper induction algorithm is used that exploits named entity information that has been previously identified. The idea of post-processing the extraction results is introduced for resolving ambiguous fields and improving the overall extraction performance. Post-processing involves the exploitation of two additional sources of information: field transition probabilities, based on a trained bigram model, and confidence scores, estimated for each field by the wrapper induction system. A multiplicative model that is based on the product of those two probabilities is also considered for post-processing. Experiments were conducted on pages describing laptop products, collected from many different sites and in four different languages. The results highlight the effectiveness of the new framework.

Metadaten
Titel
Mining Web Sites Using Wrapper Induction, Named Entities, and Post-processing
verfasst von
Georgios Sigletos
Georgios Paliouras
Constantine D. Spyropoulos
Michalis Hatzopoulos
Copyright-Jahr
2004
Verlag
Springer Berlin Heidelberg
DOI
https://doi.org/10.1007/978-3-540-30123-3_6

Premium Partner