Using Text Mining to Classify Lay Requests to a Medical Expert Forum and to Prepare Semiautomatic Answers
Wolfgang Himmel, Ulrich Reincke, Hans Wilhelm Michelmann · 2008
We developed a scoring procedure to automatically classify lay requests to an internet medical forum about involuntary childlessness. The requests should be classified according to their subject matter (32 categories) and the sender’s expectation (6 categories). Building upon this procedure, the experts’ answers to former requests could be the basis of an automatic answer for new incoming requests, finding their “nearest neighbors”. Our text mining approach comprised the following steps: a large start list of relevant words and the calculation of the Cramer’s V statistic for the association between relevant words and the 38 categories. We trained logistic regression models with high precision and recall. We then simulated a scenario in which a subset of requests (n=50) served as ‘new’ requests. To find the most nearest neighbors, we applied a formula, which gave high weight to singular value decompositions (SVDs) but also considered the automatically classified subject matter of this ‘new’ request and—to a lesser degree—the sender’s expectation. If we can implement this procedure into the real life of a health internet forum, the workload of medical experts could be lightened, visitors to a forum could receive a more timely answer to their question and they could become aware of former questions and answers that were rather similar to the one they had sent.