Generating High-Quality Training Data for Automated Land-Cover Mapping
Umaa D. Rebbapragada, R. Lomasky, Carla E. Brodley, M. A. Friedl · 2008
This paper presents two machine learning techniques that greatly reduce the number of person-hours required to generate high-quality training data for land cover classification. The first technique uses active learning to guide the generation of training data by selecting only the most informative examples for labeling. The second technique identifies and mitigates the impact of mislabeled instances. Both techniques are tested on data from NASA's Moderate Resolution Imaging Spectroradiometer (MODIS), which has required thousands of person hours to label. Our results shows that the active learning method requires fewer labeled examples than random sampling to produce a high quality classifier. Our results on class noise mitigation show that if mislabelings occur, we can further improve classifier accuracy, and that weighting instances by their label confidence outperforms an analogous method that discards suspected mislabelings. If combined, these methods have the potential to make training data generation a more efficient and reliable process.