Sifting US Census Records with Computer Vision and Machine Learning
Gregory Jansen · 2024
This paper shares the culmination of my work to computationally enhance researcher access to U. S. Census records, by targeting their personal transcription labor on those document pages that are most likely to contain relevant information. Much research on the United States population over time concerns demographic groups that may be identified, for example, through the race column on census population schedules, which are the handwritten forms on which census takers would record household information. This project was created to support the researcher efforts of Dr. Richard Marciano and the study of the community impact of the forced relocation of Japanese American households during the second world war. In particular through a detailed comparison between the Japanese American households and people recorded in 1940 and in 1950 Sacramento California. While the census forms have a different layout in each decade, the general design is tabular with rows and columns that may be used to visually segment the document. This paper, and the code notebooks that are published along with it, demonstrate a computer vision technique for segmenting population schedules to extract the individual cell images from their race column. Then the individual cell images are cleaned up and fed into two different neural network models, for identifying the handwritten race code within them. Finally, we created a user interface that allows a researcher to perform a visual review of uncertain results from the above process and thereby create a reliable dataset containing only those population schedule pages that pertain to their research. The Python code notebooks that were used to perform this analysis and the review process are linked within the paper and are freely available for reuse under a Creative Commons share-alike license.