Self-Supervised Behavior Cloned Transformers are Path Crawlers for Text Games

Ruoyao Wang, Peter Jansen · 2023

In this work, we introduce a self-supervised behavior cloning transformer for text games, which are challenging benchmarks for multistep reasoning in virtual environments.Traditionally, Behavior Cloning Transformers excel in such tasks but rely on supervised training data.Our approach auto-generates training data by exploring trajectories (defined by common macro-action sequences) that lead to reward within the games, while determining the generality and utility of these trajectories by rapidly training small models then evaluating their performance on unseen development games.Through empirical analysis, we show our method consistently uncovers generalizable training data, achieving about 90% performance of supervised systems across three benchmark text games. 11 Released as open source: https://github.com/ cognitiveailab/pathfinding-rl PathCrawlingCrawl all paths that lead to reward. Evaluate GeneralityTrain agents on groups of paths.Evaluate agents on unseen text games to test path generality.Arithmetic Game: Your task is to solve the math problem.Then, pick up the item with the same quantity as the math problem answer, and place it in the box.90% on unseen games take 24 bananas put 24 bananas in answer box The math problem says: multiply 6 times 4 You take 24 bananas.You put 24 bananas in the answer box.Task-critical information 3 4 Train LLM with these paths Evaluate LLM on unseen games take math problem read math problem You are in the kitchen.You see a math problem, 3 pears, 24 bananas, ...You take the math problem. 2Game Variation 1: take math problem, read math problem, take 24 bananas, put 24 bananas in answer box Game Variation 2: take math problem, read math problem, take 12 apples, put 12 apples in answer box Game Variation 3: take math problem, read math problem, take 16 pears, put 16 pears in answer box ... Path Group: take(X), read(X), take(Y), put (Y, Z) T5Infer this micro-action sequence is a generalizable solution take(X), read(X), take(Y), put (Y, Z) Generalizable Micro-Action Sequence 15% on unseen gamesYou put 24 bananas in the answer box.Train LLM with these paths Evaluate LLM on unseen games take 24 bananas put 24 bananas in answer box You are in the kitchen.You see a math problem, 3 pears, 24 bananas, ...You take 24 bananas.

Read the paper · More papers on PaperTik