Query Generators for Datasets of Interchangeable Recipe Steps
Juni Kim · Zenodo (CERN European Organization for Nuclear Research) · 2023
Actions, like "cook chicken," "change tire," and "fold paper," have start and end states. Furthermore, a single action like "cook chicken" starting with a particular start state can have multiple possible end states like baked, fried, or boiled. Some of these possible end states may be unacceptable depending on the situation: boiling breaded chicken is less tasty than frying it. Training models to understand which actions are swappable with each other—which actions can be swapped with each other and still produce a possible and acceptable end state—would also allow them to reason about the relationship between the start state and the action an entity goes through. Models that understand how to evaluate the swappability of two steps can better understand, modify, and generate procedural text, like instructions. Currently existing models struggle to reason about these events because of the highly variable and context-sensitive nature of an action’s effect. To make progress toward understanding events, we focus on recipes as a subset of procedural text with an abundance of online resources. We present a seed query database containing possibly swappable steps for recipes and a generalizable to automatically generate new datasets from other instruction-based text. Our method clusters Sentence Transformers embeddings created from recipe titles so that recipes with similar end states are grouped together. Inside these clusters, we generate possible replacements from recipe steps that share a marginal amount of similarity. After clustering a set of 600 recipes by cosine similarity and analyzing differences between recipes in each individual cluster, we found splitting the set into 50 clusters to be most optimal. Applying our methodology onto these 50 clusters, we are able to generate approximately 80,000 queries. Our library is a stepping stone for the future development of models that are capable of modifying instructions without changing final results.