SciELO - Scientific Electronic Library Online

 
vol.23 issue3Deep Semantic Role Labeling for Tweets using 5W1H: Who, What, When, Where, Why and HowEnriching Word Embeddings with Global Information and Testing on Highly Inflected Language author indexsubject indexsearch form
Home Pagealphabetic serial listing  

Services on Demand

Journal

Article

Indicators

Related links

  • Have no similar articlesSimilars in SciELO

Share


Computación y Sistemas

On-line version ISSN 2007-9737Print version ISSN 1405-5546

Abstract

SIDO, Jakub; KONOPIK, Miloslav  and  PRAžAK, Ondřej. English Dataset For Automatic Forum Extraction. Comp. y Sist. [online]. 2019, vol.23, n.3, pp.765-771.  Epub Aug 09, 2021. ISSN 2007-9737.  https://doi.org/10.13053/cys-23-3-3259.

This paper describes the process of collecting, maintaining and exploiting an English dataset of web discussions. The dataset consists of many web discussions with hand-annotated posts in the context of a tree structure of a web page. Each post consists of username, date, text, and citations used by its author. The dataset contains 79 different websites with at least 500 pages from each. Each web page consists of a tree structure of HTML tags with texts taken from selected web pages. In the paper, we also describe algorithms trained on the dataset. The algorithms employ basic architectures (such as a bag of words with an SVM classifier and an LSTM network) to set a baseline for the dataset.

Keywords : Information retrieval; web discussion.

        · text in English