Learning from data of varying quality for sentence role identification in MEDLINE abstracts

Natalia Aizenberg · Institutional Repositories DataBase (IRDB) · 2007

A most distinctive feature of scientific abstracts is their rhetorical structure, typically consisting of sections describing background information, objectives of the study, experimental methods, results and conclusions.There has been an increasing interest in recent years in identifying such structural roles, with particular motivations from the information retrieval point of view.In previous research done with respect to MEDLINE abstract, various sentence feature combinations were used in order to achieve successful performance, but one important issue has not yet been addressed: the unrepresentativeness of the major part of learning data; in this task, the learning set samples tend to originate from different sources baring many differences, while the application data source distribution does not necessarily obey that of the learning set.In this work we solve the issue mentioned previously by applying "example source" sensitive costs in the training process.

Read the paper · More papers on PaperTik