New Challenges For NLP Frameworks Programme
C. R. Ramakrishnan, William A. Baumgartner, Judith A. Blake, Gully Burns, Kevin Bretonnel Cohen, Harold Drabkin, Janan T. Eppig, Eduard H. Hovy, Chun‐Nan Hsu, Lawrence Hunter, Tommy Ingulfsen, Hiroaki Onda, Sandeep Pokkunuri, Ellen Riloff, Christophe Roeder, Karin M. Verspoor · 2010
We present a practical problem that involves the analysis of a large dataset of heterogeneous documents obtained by crawling the web for information related to web services. This analysis includes information extraction from natural-language (HTML and PDF) and machine-readable (WSDL) documents using NLP and other techniques, classifying documents as well as services (defined by sets of documents), and exporting the results as RDF for use in the back-end of a portal that uses Web 2.0 and Semantic Web technology. Triples representing manual annotations made on the portal are also exported back to our application to evaluate parts of our analysis and for use as training data for machine learning (ML). This application was implemented in the GATE framework and successfully incorporated into an integrated project, and included a number of components shared with our group’s other projects.