Evaluating Diversity in Automatic Poetry Generation
Yanran Chen, Hannes Gröner, Sina Zarrieß, Steffen Eger · 2024
Natural Language Generation (NLG), and more generally generative AI, are among the currently most impactful research fields.Creative NLG, such as automatic poetry generation, is a fascinating niche in this area.While most previous research has focused on forms of the Turing test when evaluating automatic poetry generation -can humans distinguish between automatic and human generated poetry -we evaluate the diversity of automatically generated poetry (with a focus on quatrains), by comparing distributions of generated poetry to distributions of human poetry along structural, lexical, semantic and stylistic dimensions, assessing different model types (word vs. character-level, general purpose LLMs vs. poetry-specific models), including the very recent LLaMA3-8B, and types of fine-tuning (conditioned vs. unconditioned).We find that current automatic poetry systems are considerably underdiverse along multiple dimensions -they often do not rhyme sufficiently, are semantically too uniform and even do not match the length distribution of human poetry.Our experiments reveal, however, that style-conditioning and character-level modeling clearly increases diversity across virtually all dimensions we explore.Our identified limitations may serve as the basis for more genuinely diverse future poetry generation models. 1 L 0.62 12 34 20.18 20 2.84 de LLaMA3 con 0.76 10 47 21.69 21 4.14 en HUMAN 1.00 4 67 28.06 28 6.26 en DeepSpeare 0.57 15 33 23.85 24 2.85 en SA 0.92 12 52 27.36 27 5.38 en ByGPT5 S 0.80 12 44 25.30 25 5.09 en ByGPT5 L 0.77 11 47 24.97 25 4.87 en GPT2 S 0.69 13 55 24.11 24 4.48 en GPT2 L 0.72 13 56 24.74 24 4.94 en GPTNeo S 0.55 11 55 22.67 22 3.89 en GPTNeo L 0.48 13 34 21.93 22 3.16 en LLaMA2 S 0.87 15 75 28.60 27 7.52 en LLaMA2 L 0.67 12 54 23.95 24 4.50 en LLaMA3 0.59 14 60 23.20 23 4.23 en ByGPT5 con S 0.85 13 42 26.21 26 4.96 en ByGPT5 con L 0.84 14 42 25.85 25 4.84 en GPT2 con S 0.86 17 61 28.37 27 6.18 en GPT2 con L 0.83 16 70 27.8227 6.15 en GPTNeo con S 0.74 16 49 25.13 24 4.47 en GPTNeo con L 0.53 12 35 22.26 22 3.36 en LLaMA2 con S 0.70 17 74 33.55 32 7.83 en LLaMA2 con L 0.81 15 56 26.92 26 5.80 en LLaMA3 con 0.78 16 65 27.12 26 5.35