From Dataset Recycling to Multi-Property Extraction and Beyond
Tomasz Dwojak, Michał Pietruszka, Łukasz Borchmann, Jakub Chłędowski, Filip Graliński · 2020
This paper investigates various Transformer architectures on the WikiReading Information Extraction and Machine Reading Comprehension dataset. The proposed dual-source model outperforms the current state-of-theart by a large margin.Next, we introduce WikiReading Recycled-a newly developed public dataset, and the task of multipleproperty extraction.It uses the same data as WikiReading but does not inherit its predecessor's identified disadvantages.In addition, we provide a human-annotated test set with diagnostic subsets for a detailed analysis of model performance.