Can the Market See Through Fraud? Machine Learning-Based Fraud Prediction and Stock Mispricing
Ziming Wang, Yuanzhi Wen, Hongfeng Han · Emerging Markets Finance and Trade · 2026
Can the stock market see through fraud risk that is inferable from public text but not yet confirmed by regulators? Applying machine learning to the management discussion and analysis section of Chinese A-share annual reports, we train classifiers on linguistic features to estimate a firm-level fraud probability for each year. These predictions carry genuine informational content, as they are significantly associated with regulatory enforcement actions up to three years ahead. In asset pricing tests, firms with higher predicted fraud probabilities exhibit significantly larger absolute mispricing in the subsequent year, and the association is stable across alternative specifications and endogeneity controls. The effect is not uniform: when analyst coverage is sufficiently high, the pricing gap effectively vanishes, suggesting that professional intermediaries perform the cognitive work of translating textual fraud signals into prices, whereas for less-covered firms the same signals persist as undigested residuals. The Chinese market thus does not efficiently process publicly available textual information about fraud risk, a specific and quantifiable dimension of inefficiency that machine learning can identify.