Disentangling Disparate Communication Streams and Recovering Underlying Structure
Eisha Nathan · 2024
In this paper we present various approaches for the tasks of disentanglement (or automatic forum identification) and structure recovery (or prediction of message replies) in online discussion forums such as Reddit and Stack Exchange. For the disentanglement task we present a fine-tuned BERT model. By leveraging transfer learning from a pre-trained BERT (Bidi-rectional Encoder Representations from Transformers) model, given the message content, our approach achieves accuracies of 83–92% in predicting where a message originated. For the task of structure recovery, we propose a novel neural network architecture that leverages deep learning techniques to model the sequential relationship between messages to predict if one message follows another with accuracies over 80%. Additionally we present few-shot learning strategies for both tasks of disentanglement and structure recovery. Given the sparse availability of limited samples for training in real data, few-shot learning, or the ability to train without needed large amounts of labeled data, is necessary. Experimental results demonstrate the effectiveness of our proposed architecture and methods in capturing contextual dependencies and improving prediction accuracy in such diverse online forums for both tasks and our few-shot architectures achieve comparable accuracies to the case when large training data is available.