Detecting Corporate Fraud: An Application of Machine Learning
Ophir Gottlieb, Curt Salisbury, Howard Howan Stephen Shek, Vishal Vaidyanathan · 2006
This paper explores the application of several machine learning algorithms to published corporate data in an effort to identify patterns indicative of securities fraud. Generally Accepted Accounting Principles (GAAP) represent a conglomerate of industry reporting standards which US public companies must abide by to aid in ensuring the integrity of these companies. Notwithstanding these principles, US public companies have legal flexibility to maneuver the way they disclose certain items in the financial statements making it extremely hard to detect fraud manually. Here we test several popular methods in machine learning (logistic regression, naive Bayes and support vector machines) on a large set of financial data and evaluate their accuracy in identifying fraud. We find that some variants of SVM and logistic regression are significantly better than the currently available methods in industry. Our results are encouraging and call for a more thorough investigation into the applicability of machine learning techniques in corporate fraud detection.