Unlocking Scholarly Insights: Leveraging Machine Learning Approaches for Citation Analysis and Intent Classification.

Lyudmila Balakireva, Martin Klein · 2024

Publicly funded organizations, notably institutions like the Los Alamos National Laboratory (LANL), are deeply vested in acquiring robust productivity metrics to gauge the entirety of their research output. This study explores the use of Large Language Models (LLMs), including BERT-based models, LLaMa-30b-instruct, and Mixtral-8x7b-instruct, for classifying URL-referenced resources (e.g., software, datasets) and determining authorship intent. We highlight the challenges of identifying resource types and the potential of BERT and LLMs to overcome them. Our analysis shows a growing number of URL citations, reflecting the increasing importance of digital resources. Our findings also reveal that LANL authors contribute substantially to accessible science, comprising about 10% of dataset and software mentions in LANL publications.

Read the paper · More papers on PaperTik