Boilerplate Detection Using Shallow Text Features
published: Oct. 7, 2010, recorded: February 2010, views: 23538
Slides
Related content
Report a problem or upload files
If you have found a problem with this lecture or would like to send us extra material, articles, exercises, etc., please use our ticket system to describe your request and upload the data.Enter your e-mail into the 'Cc' field, and we will keep you updated with your request's status.
Description
In addition to the actual content Web pages consist of navigational elements, templates, and advertisements. This boilerplate text typically is not related to the main content, may deteriorate search precision and thus needs to be detected properly. In this paper, we analyze a small set of shallow text features for classifying the individual text elements in a Web page. We compare the approach to complex, state- of-the-art techniques and show that competitive accuracy can be achieved, at almost no cost. Moreover, we derive a simple and plausible stochastic model for describing the boilerplate creation process. With the help of our model, we also quantify the impact of boilerplate removal to retrieval performance and show significant improvements over the baseline. Finally, we extend the principled approach by straight-forward heuristics, achieving a remarkable accuracy.
Link this page
Would you like to put a link to this lecture on your homepage?Go ahead! Copy the HTML snippet !
Reviews and comments:
To improve the audio quality of the video, please turn your speakers balance to left (= mono).
To test my algorithms, have a look at http://boilerpipe-web.appspot.com/ and http://code.google.com/p/boilerpipe/
Cheers,
Christian
More than one year has passed by now, and videolectures.net still has not managed to update the audio in the stream.
Until this is fixed, feel free to download/watch the video with corrected (mono) audio here:
http://www.l3s.de/~kohlschuetter/boil...
Best,
Christian
Hello Christian,
How can I use this algorithm with jQuery ? Do you have any POC ? please let me know. thanks
Hello Christian,
how can i implement this algorithm in c language? Please Please let me know
thanks in advance
kind regards,
Mohsin
Write your own review or comment: