Japanese-Mobile-Receipt-OCR-1.3K: A Comprehensive Dataset Analysis and Fine-tuned Vision-Language Model for Structured Receipt Data Extraction
Sabari Nathan · 2025
We introduce Japanese-Mobile-Receipt-OCR-1.3K, a curated dataset of 1,300 real-world Japanese receipt images captured via mobile phones and annotated with 34,727 text entries. We also present a fine-tuned vision-language model for end-to-end structured receipt extraction.Our dataset analysis quantifies linguistic and layout characteristics that challenge receipt understanding. These include a heavy-tailed token length distribution (mean ≈ 9.3 tokens, maximum 255), diverse text complexity across fields, and marked heterogeneity in character composition with substantial proportions of Kanji, Kana, and numerals. We further assess semantic coverage by quantifying numeric, monetary, and temporal expressions, and measuring named entity recognition coverage across common receipt fields. Leveraging these insights, we adapt a 3B-parameter vision-language backbone to produce structured JSON outputs capturing hierarchical field relationships and standardized numeric and currency formats.Extensive experiments show consistent improvements over strong baselines. Our approach achieves notable gains in field naming consistency, hierarchical structure accuracy, and numeric formatting, alongside reductions in word error rate and character error rate.This work delivers a complete pipeline-from dataset curation and statistical analysis to model adaptation and evaluation-establishing a robust benchmark and practical methodologies for Japanese receipt understanding.