Boilerplate Detection Using Shallow Text Features
published: Oct. 7, 2010, recorded: February 2010, views: 23538
Report a problem or upload filesIf you have found a problem with this lecture or would like to send us extra material, articles, exercises, etc., please use our ticket system to describe your request and upload the data.
Enter your e-mail into the 'Cc' field, and we will keep you updated with your request's status.
In addition to the actual content Web pages consist of navigational elements, templates, and advertisements. This boilerplate text typically is not related to the main content, may deteriorate search precision and thus needs to be detected properly. In this paper, we analyze a small set of shallow text features for classifying the individual text elements in a Web page. We compare the approach to complex, state- of-the-art techniques and show that competitive accuracy can be achieved, at almost no cost. Moreover, we derive a simple and plausible stochastic model for describing the boilerplate creation process. With the help of our model, we also quantify the impact of boilerplate removal to retrieval performance and show significant improvements over the baseline. Finally, we extend the principled approach by straight-forward heuristics, achieving a remarkable accuracy.
Download slides: wsdm2010_kohlschutter_bdu_01.pdf (2.8 MB)
Link this pageWould you like to put a link to this lecture on your homepage?
Go ahead! Copy the HTML snippet !
Reviews and comments:
To improve the audio quality of the video, please turn your speakers balance to left (= mono).
To test my algorithms, have a look at http://boilerpipe-web.appspot.com/ and http://code.google.com/p/boilerpipe/
More than one year has passed by now, and videolectures.net still has not managed to update the audio in the stream.
Until this is fixed, feel free to download/watch the video with corrected (mono) audio here:
How can I use this algorithm with jQuery ? Do you have any POC ? please let me know. thanks
how can i implement this algorithm in c language? Please Please let me know
thanks in advance
Write your own review or comment: