Document and Corpus Level Inference For Unsupervised and Transductive Learning of Information Structure of Scientific Documents

Roi Reichart, Anna Korhonen · International Conference on Computational Linguistics · 2012

Inferring the information structure of scientific documents has proved useful for supporting information access across scientific disciplines. Current approaches are largely supervised and expensive to port to new disciplines. We investigate primarily unsupervised discovery of information structure. We introduce a novel graphical model that can consider different types of prior knowledge about the task: within-document discourse patterns, cross-document sentence similarity information based on linguistic features, and prior knowledge about the correct classification of some of the input sentences when this information is available. We apply the model to Argumentative Zoning (AZ) scheme and evaluate it on a fully unsupervised learning scenario and two transduction scenarios where the categories of some test sentences are known. The model substantially outperforms similarity and topic model based clustering approaches as well as traditional transduction algorithms. TITLE AND ABSTRACT IN FINNISH Dokumenttija korpustason inferenssiin perustuva ohjaamattomankoneoppimisen tekniikka tieteellisen julkaisujen rakenteen analyysissa Tieteellisten julkaisujen rakenteen analyysi voi tukea tietojen saatavuutta eri tieteenaloilta. Nykyiset koneoppimismetodit ovat pitkalti ohjattuja ja niiden soveltaminen uusille tieteenaloille on kallista. Tama artikkeli tutkii paaasiassa ohjaamatonta julkaisujen rakenteen analyysia. Lahtokohtana on uusi graafinen malli, joka pystyy integoimaan erilaista etukateistietoa tehtavasta: dokumenttien sisaisen diskurssin, dokumenttienvalisten samankaltaisuuden kielellisten ominaisuuksien suhteen, ja tietoa joidenkin lauseiden oikeasta luokittelusta, silloin kun tamankaltaista tietoa on saatavilla. Malli sovellettiin Argumentative Zoning (AZ) -analyysiin ja sen soveltuvuutta taysin ohjaamattomaan oppimiseen seka transduktio-oppimiseen, jossa joidenkin testilauseiden luokat on tiedossa, tutkittiin. Malli osoittautuu huomattavasti tarkemmaksi kuin samankaltaisuuteen ja klusterointiin perustuvat vertailumallit seka perinteiset transduktio-algoritmit.

Read the paper · More papers on PaperTik