Text mining workflow for extraction of paragraphs from full articles describing drug-gene interactions to support Onco KEM software platform for personalized treatments

Fanny Perraudeau, David Morley, Mohammad Parsa Afshar, Mariana Guergova-Kuras · 2013

extracted the text from the PDF files of the full articles. The figures, tables and references of the articles were not extracted avoiding false negative paragraphs extraction. The keywords constituting the sentence were enriched with synonyms for both drugs and genes. The synonyms of the drugs were extracted from the CTD database. For the synonyms of the genes, the “gene_info” file downloaded from the FTP of the NCBI was used.

Read the paper · More papers on PaperTik