Text mining workflow for extraction of paragraphs from full articles describing drug-gene interactions to support Onco KEM software platform for personalized treatments
Fanny Perraudeau, David Morley, Mohammad Parsa Afshar, Mariana Guergova-Kuras · 2013
extracted the text from the PDF files of the full articles. The figures, tables and references of the articles were not extracted avoiding false negative paragraphs extraction. The keywords constituting the sentence were enriched with synonyms for both drugs and genes. The synonyms of the drugs were extracted from the CTD database. For the synonyms of the genes, the “gene_info” file downloaded from the FTP of the NCBI was used.