Parsing NCBI XML in Perl
2005
Parsing NCBI XML in Perl 5Perl remains the programming language of choice for many in bioinformatics.Perl has excellent support for processing and manipulating text, finding regular expression patterns, retrieving files via the Internet, and connecting to a wide variety of relational databases.This makes it an ideal language for parsing flat text files, such as GenBank Flat File records, integrating biological data from multiple sources, and performing sequence analysis.Building on these strengths, the bioinformatics community has developed BioPerl [72], a very successful open source module that includes numerous features, including the ability to retrieve biological data from remote data sources, run BLAST searches, and manipulate sequence data.Perl also has excellent support for XML, and is supported by a wide variety of third-party open source XML modules.This chapter provides an introduction to XML parsing in Perl, and introduces two standard interfaces: the Simple API for XML (SAX) and the Document Object Model (DOM).To explore SAX, we focus on the XML::SAX module, and to explore the DOM, we focus on the XML::LibXML module.To illustrate basic concepts, the chapter includes numerous examples for parsing XML documents from the National Center for Biotechnology Information (NCBI) at the U.S. National Institutes of Health.We also explore the NCBI EFetch service, and illustrate how to dynamically retrieve and parse sequence records from NCBI.This chapter assumes that you have a basic familiarity with Perl programming, and understand the fundamentals of object-oriented programming in Perl.If you do not have such background, you may want to check out one of the recommended Perl references [56; 66; 67; 73; 74]. How to Retrieve a Nucleotide Record: