User guide for NIST Media Forensic Challenge (MFC) datasets

Haiying Guan, Andrew Delgado, Yooyoung Lee, Amy N. Yates, Daniel K. Zhou, Timothée Kheyrkhah, Jon Fiscus · 2021

More than 300 individuals from 150 organizations across 26 countries and regions use the NIST released Media Forensic Challenge (MFC) datasets for their research. The MFC datasets were created for use in the DARPA MediFor (Media Forensics) program. Since their release, multiple questions have been fielded regarding the dataset properties, including contents, metadata definitions, usage, data repurposing, etc. For example: what do the datasets contain? What are the definitions of the different kinds of metadata? How does one label the data with the reference information to build the training data for machine learning algorithms? How would one modify/extract the data for their own research purposes? This document serves as a user guide for the MFC datasets, including those used in the Nimble Challenge (NC). This guide includes: 1) a description about MFC datasets including background, evolution history, and the dataset summary by the evaluation tasks; 2) user access and permissions of MFC datasets; 3) an introduction to the MFC data by providing a simple example of a manipulation journal graph and its detailed corresponding MFC dataset reference files; 4) an introduction to a flexible subset selection approach, Selective Scoring, to sample the test probes from the entire test set for the particular task evaluation; 5) information to help users gain a deeper understanding of the metadata by presenting two commonly used approaches to illustrate the manipulation operation statistic histogram distributions, and 6) a general template of the NIST MFC evaluation dataset to facilitate the future dataset generation.

Read the paper · More papers on PaperTik