STATISTICAL AND MACHINE METHODS FOR AUTOMATICALLY EXTRACTING CAUSAL RELATIONSHIPS FROM TEXT (REVIEW)
Khairutin Shtanchaev · Известия Южного федерального университета. Технические науки · 2023
Until the 2000s, the concept of non-statistical methods was used to solve the problem ofautomatic extraction of causal relationships (CR). These methods used manually constructedlinguistic templates. Obviously, the CR that did not fit into the built templates could not bedefined. Non-statistical methods required constant manual control by experts, up to the evaluation.Almost all methods were aimed at extracting explicit CR. In some methods, attemptswere made to untie the extraction system from a specific subject area. To eliminate the abovedisadvantages, the methods developed in the future began to shift towards statistical dataprocessing and machine learning. In this article, statistical and machine methods of CR e xtractionare considered. A few valuable papers related to the new paradigm of CR extractionwere analyzed. The aim of the research was to evaluate new methods with the ability to identifytheir advantages and disadvantages. The great advantage of machine and statisticalmethods is independence from the subject area while maintaining the accuracy of extraction.Such methods are worse in accuracy, but they are not tied to a specific problem area. Themethods themselves, unlike non-statistical ones, which used linguistic and syntactic comparisonwith templates manually, are focused on finding these templates. Even though machineand statistical methods are mostly independent of the subject area and use large corpora oftext for teaching, they are intended mainly for the English language. There is also no standardizeddata set that would allow methods to be compared with each other. All works devotedto methods ignored the extraction of implicit CR.