Learned Hybrid Video Coding for Human Perception and Multiple Machine Vision Tasks
Martin Benjak, Saifullah Khan, Yi‐Hsin Chen, Wen-Hsiao Peng, Jörn Östermann · 2025
In this work, we present a learned multi-task video codec that is optimized for human and machine vision. The codec consists of an encoder that maps images from the pixel domain to a latent representation and multiple decoders that map the latent to either an image for human consumption or multiple task-specific features for different machine vision tasks. This allows a single bitstream to be used for multiple tasks while also reducing the decoder complexity for machine vision tasks. Unlike most learned codecs, our method performs inter-coding at the latent level instead of the pixel domain. Experiments show that the proposed method achieves a compression performance for machine vision tasks comparable to other multi-task codecs designed for machine vision only, while also providing video reconstruction. The code is available at https://github.com/GreenAutoML4FAS/HybridMultiTaskCoding.