Bootstrapping Information Extraction from Semi-structured Web Pages
author:Charles Schafer, Google
published: Oct. 10, 2008, recorded: September 2008, views: 159
Slides
Related content
58:16
507 views - Tom Mitchell, 2007
01:01:41
2357 views - Kamal Nigam, 2006
24:28
147 views - Georgios Paliouras, 2006
01:19:12
337 views - Kalina Bontcheva, 2004
01:45:49
396 views - Ronen Feldman, 2005
07:44:46
654 views - Marie-Francine Moens, 2008
55:47
3587 views - William Cohen, 2006
01:47:58
121 views - Ronen Feldman, 1970
01:10:25
172 views - Roman Yangarber, 2007
02:28:57
1803 views - Bijan Parsia, 2006
Report a problem or upload files
If you have found a problem with this lecture or would like to send us extra material, articles, exercises, etc., please use our ticket system to describe your request and upload the data.Enter your e-mail into the 'Cc' field, and we will keep you updated with your request's status.
Description
We consider the problem of extracting structured records from semi-structured web pages with no human supervision required for each target web site. Previous work on this problem has either required significant human effort for each target site or used brittle heuristics to identify semantic data types. Our method only requires annotation for a few pages from a few sites in the target domain. Thus, after a tiny investment of human effort, our method allows automatic extraction from potentially thousands of other sites within the same domain. Our approach extends previous methods for detecting data fields in semi-structured web pages by matching those fields to domain schema columns using robust models of data values and contexts. Annotating 2-5 pages for 4-6 web sites yields an extraction accuracy of 83.8% on job offer sites and 91.1% on vacation rental sites. These results significantly outperform a baseline approach.
See Also:
Download slides:
ecmlpkdd08_carlson_bief_01.ppt (5.2 MB)
Launch in a standalone WM Player
Switch to Windows Media Player
Link this page
Would you like to put a link to this lecture on your homepage?Go ahead! Copy the HTML snippet !




Write your own review or comment: