Domain adaptation for applications in computer vision with limited data
Mattias Billast · 2024
This dissertation explores the challenges and solutions of using domain adaptation in real-time computer vision applications with limited labeled data. Computer vision, initially based on traditional feature extraction methods, has progressed significantly with deep learning, achieving breakthroughs in areas like image classification and object detection. However, deep learning models often require large labeled datasets, which can be expensive and time-consuming to obtain, especially for custom applications. Domain adaptation offers a way to tackle this problem by using external data sources to improve model performance in a target domain. Its effectiveness depends on the domain gap—if the gap between source and target data is too large, adaptation becomes difficult. The dissertation focuses on two applications: maritime autonomous navigation and human motion prediction. In maritime navigation, the goal is to detect, track, and locate obstacles for autonomous vessels, but a lack of labeled data poses a challenge. By using domain adaptation techniques, data from external sources (such as public object detection datasets) is leveraged to improve model accuracy. The second application involves predicting the physical and cognitive ergonomics of operators performing repetitive tasks. This is done by analyzing human pose data and anticipating movements to prevent musculoskeletal issues. Data from a VR setup helps train the model, with domain adaptation used to improve its performance despite limited labeled data. Both applications require real-time performance with lightweight models. Domain adaptation techniques are used to enhance the models by incorporating external data, like maritime object detection datasets or VR controller data for human pose prediction. Overall, the thesis highlights the importance of domain adaptation in improving model accuracy with limited data, showing that external data sources can significantly enhance real-time computer vision applications, both in real-world and academic settings. The key contribution is that domain adaptation can utilize any useful external data to improve performance.