Zenodo Open Metadata snapshot - Training dataset for records and communities classifier building

Zenodo Team · Zenodo (CERN European Organization for Nuclear Research) · 2022

This dataset contains Zenodo's published open access records and communities metadata, including entries marked by the Zenodo staff as spam and deleted. The datasets are gzipped compressed JSON-lines files, where each line is a JSON object representation of a Zenodo record or community. Records dataset Filename: zenodo_open_metadata_{ date of export }.jsonl.gz Each object contains the terms: part_of, thesis, description, doi, meeting, imprint, references, recid, alternate_identifiers, resource_type, journal, related_identifiers, title, subjects, notes, creators, communities, access_right, keywords, contributors, publication_date which correspond to the fields with the same name available in Zenodo's record JSON Schema at https://zenodo.org/schemas/records/record-v1.0.0.json. In addition, some terms have been altered: The term files contains a list of dictionaries containing filetype, size, and filename only. The term license contains a short Zenodo ID of the license (e.g. "cc-by"). Communities dataset Filename: zenodo_community_metadata_{ date of export }.jsonl.gz Each object contains the terms: id, title, description, curation_policy, page which correspond to the fields with the same name available in Zenodo's community creation form. Notes for all datasets For each object the term spam contains a boolean value, determining whether a given record/community was marked as spam content by Zenodo staff. Some values for the top-level terms, which were missing in the metadata may contain a null value. A smaller uncompressed random sample of 200 JSON lines is also included for each dataset to test and get familiar with the format without having to download the entire dataset.

Read the paper · More papers on PaperTik