OSPC: OCR-Assisted VLM for Zero-Shot Harmful Meme Detection
Chenxi Zhu, Haotian Gao, Yuxiao Duan, Hao Guo, Minnan Luo, Xiang Hui Zhao · 2024
Harmful memes refer to the memes which contain social bias towards a certain group, such as gender, race, disabilities and so on. Detecting these harmful memes requires the model to have both visual and linguistic understanding of memes and some external knowledge about the culture and morality. This technical report summarises the fourth place solution of the Online Safety Prize Challenge, which enhances a visual language model (VLM) by integrating external Optical Character Recognition (OCR) tools, employs prompt engineering and manual resource management to optimize performance within limited resource constraints. This strategy achieved a score of 0.723/0.643(AUROC/ACC) in this challenge with one single submission