Unsupervised Multitask Learning for Oil and Gas Language Models with Limited Resources

Maxime Marlot, Divya Nidhi Srivastava, Fui Kent Wong, Ming Xiang Lee · 2023

Abstract In this paper we explore the development of an oil and gas language model (LM) using an unsupervised multitask learning approach. A large language model (LLM)enables computers to understand and generate human language. The study addresses data scarcity and domain-specific language challenges, showcasing the model's performance on specific oil and gas tasks and qualitative testing. To do that, we collected a highly diversified dataset of 33,000 documents in energy and oil and gas domains to train and benchmark our model.Our findings demonstrate that even a small model, properly finetuned on domain-specific data, outperforms larger models trained on generic corpora, highlighting the benefits of finetuning LMs in technical domains. The paper contributes to advancing natural language processing (NLP) in the oil and gas industry, emphasizing the importance of addressing domain-specific nuances and limitations for improved NLP model performance and reliability.

Read the paper · More papers on PaperTik