E2M: Automatic Generation of MARC-Formatted Metadata by Crawling E-Publications

Siew-Phek T. Su, Yu Long, Daniel Cromwell · Information Technology and Libraries · 2002

This paper presents a system called E-pub to MARC (E2M), which automatically generates MARC-formatted metadata by crawling e-publications. The functions of its two key components, the Web Crawler and the MARC Converter, are introduced. The paper presents the methods and tools used for building the system. The process of crawling and gathering pertinent metadata stored in the e-publications and the transformation of the metadata into MARC-formatted records are described in detail. The complexity of the crawling and the record generation processes are also described. A comparison between the cataloging process of e-publication using the computer-aided E2M process and manual cataloging is presented to illustrate that the E2M process is a more cost effective and efficient method of organizing and proving access to e-publications. ********** The proliferation of scholarly electronic publications (e-publications) on the Internet has posed a challenging problem for catalog librarians. The process of manually cataloging information in this medium is not only time consuming but also human-resource intensive. In light of this challenge, the impetus of the project team was to develop a more efficient and effective method to catalog this type of material with the aid of a computer. Bibliographic control of Web resources and its related issues have been widely discussed and written about. (1) The issue of automating the e-publication cataloging process is an important one. However, little has been done in developing systems to automate the entire labor-intensive cataloging process using WebCrawler technology and techniques for automatic data conversion and loading. Although WebCrawlers have been used to extract information from Web pages, they are not programmed to extract the specific metadata needed for constructing catalog records and for loading them into bibliographic databases. For example, the two notable crawlers of the popular search engines, Google and Internet Archive, crawl the entire Web and extract keywords from Web pages to generate indexes for accessing relevant Web pages. (2) Meta-crawlers, such as MetaCrawler and Dogpile, integrate the search results obtained from different search engines. (3) Site-specific crawlers, such as WebSPHINX, allow users to specify site-specific crawling rules and perform so-called personal crawling. (4) The Hermes notification service system uses a component called wrapper to extract bibliographic data from HTML documents on publishers' Web sites and generate XML documents that contain bibliographic data. (5) The bibliographic data are typically the journal's table of contents (TOC). A commercially available tool for cataloging Web resources is the MARCit system. (6) The system provides a template for users to fill in such cataloging information as URL, author, title, and subject headings to convert the information to standard MARC-formatted records, which can be loaded into the local library management system. However, it does not have a WebCrawler component to automatically access and harvest the Web page metadata. In contrast to the above systems, the E2M system described in this paper deals with the entire e-publication cataloging process. It starts with the automatic extraction of metadata from Web pages and goes on to the conversion of metadata into MARC-formatted records. Next, these records are loaded into the local system for authority verification. The final stage of the process consists of exporting verified MARC records to the Online Computer Library Center (OCLC) catalog for sharing with the bibliographic information community. Project Domain The e-publications housed in the Extension Digital Information Source (EDIS) database of the Institute of Food and Agricultural Sciences (IFAS) at the University of Florida (UF) is used as the project domain. (7) The EDIS database is the official electronic database of IFAS' current extension service and research publications. …

Read the paper · More papers on PaperTik