Analyzing the Linguistic Structure of Question Texts to Characterize Answerability in Quora
Suman Kalyan Maity, Aman Kharb, Animesh Mukherjee · IEEE Transactions on Computational Social Systems · 2018
Quora is one of the most popular community question & answer (Q&A) sites of recent times. However, with increasing question posts over time and the posts covering a wide range of topics (unlike focused Q&A sites like Stack Overflow), not all of them are getting answered. Measuringanswerability(i.e., whether a question shall get answered or not) involves collecting expensive human judgment data that can differentiate the characteristics of an answered question from an unanswered (akaopen) one. Factors to judge if a question would remain open include its subjectivity, openendedness, vagueness, ambiguity, and so on. It is difficult to collect such judgments for thousands of questions, requiring automatic framework to deal the issue of answerability of questions. In this paper, we quantify: 1)user-leveland 2)question-level linguistic activities—that cannicely correspond to many of the judgment factors noted earlier, can beeasily measured for each question postand thatappropriately discriminates an answered question from an unanswered one. Our central finding is thatthe way users use language while writing the question text can be a very effective means to characterize answerability.This characterization further helps us to predict early if a question remaining unanswered for a specific time period$t$will eventually be answered or not and achieve an accuracy of 76.26% ($t=1$month) and 68.33% ($t=3$months). Notably, features representing thelanguage use patternsof the users are most discriminative and alone account for an accuracy of 74.18%. We also compare our method with some of the similar works[1],[2]achieving a maximum improvement of ~39% in terms of accuracy.