Danish Legislative Speech Corpus

Frederik Hjorth · Harvard Dataverse · 2021

This dataset contains text and speaker data for 1,886,747 snippets of legislative speech from Denmark's Parliament, Folketinget. The data set draws in part on the ParlSpeech V2 data set, in part on Folketinget's publicly available XML transcripts. In order to homogenize the lengths of text units, longer speeches are broken down into snippets of 3-5 sentences each. Version 3 extends the data coverage through 3 September 2026. Earlier versions are retained for archival purposes; most users will want the newest version, paradf-v3.rds. Recent XML transcripts may carry Folketinget's preliminary ("Foreløbig") status and may subsequently be revised.

Read the paper · More papers on PaperTik