Regular expression-based reference metadata extraction from the web

Xiaoyu Tang, Qingtian Zeng, Tingting Cui, Zeze Wu · 2010

Accurate reference metadata extraction becomes an intriguing task to researchers who want to collect data of scientific publications. In this paper, we introduce an approach to extracting the reference metadata based on regular expressions. A prototype system named “Goldrusher” is created which automatically extracts data from the website of Association for Computing Machinery (ACM). The experimental results show that, by using our regular expression-based method, we can effectively extract author names, article titles, journal titles, DIOs, etc.

Read the paper · More papers on PaperTik