Long document similarity dataset, Wikipedia excerptions for movies collections

­ Anonymous · Zenodo (CERN European Organization for Nuclear Research) · 2022

Movies-related articles extracted from Wikipedia. For all articles, the figures and tables have been filtered out, as well as the categories and "see also" sections. The article structure, and particularly the sub-titles and paragraphs are kept in these datasets Movies The Wikipedia Movies dataset consists of 100,371 articles describing various movies. Each article may consist of text passages describing the plot, cast, production, reception, soundtrack, and more.

Read the paper · More papers on PaperTik