Extracting Common Inference Patterns from Semi-Structured Explanations

Sebastian Thiem, Peter A. Jansen · 2019

Complex questions often require combining multiple facts to correctly answer, particularly when generating detailed explanations for why those answers are correct.Combining multiple facts to answer questions is often modeled as a "multi-hop" graph traversal problem, where a given solver must find a series of interconnected facts in a knowledge graph that, taken together, answer the question and explain the reasoning behind that answer.Multihop inference currently suffers from semantic drift, or the tendency for chains of reasoning to "drift" to unrelated topics, and this semantic drift greatly limits the number of facts that can be combined in both free text or knowledge base inference.In this work we present our effort to mitigate semantic drift by extracting large high-confidence multi-hop inference patterns, generated by abstracting large-scale explanatory structure from a corpus of detailed explanations.We represent these inference patterns as sets of generalized constraints over sentences represented as rows in a knowledge base of semi-structured tables.We present a prototype tool for identifying common inference patterns from corpora of semi-structured explanations, and use it to successfully extract 67 inference patterns from a "matter" subset of standardized elementary science exam questions that span scientific and world knowledge. Error Class Sparsity in Explanation AnnotationFact 1 Friction occurs when two object's surfaces move against each other Fact 2As an object's smoothness increases, it's friction will decrease when it's surface moves against another surface.Issue These facts are not observed together in a single question's explanation, so they are not connected. Sparsity in Knowledge Base Fact 1If food is cooked then heat energy is added to that food.Fact 2A stove generates heat for cooking.Missing A campfire generates heat for cooking. IssueMissing facts in the knowledge base limit the generalization of patterns to new scenarios (e.g.campfire). Permissiveness in automatically populated edges Fact 1Melting means changing from a solid to a liquid by adding heat energy Fact 2Wax is an electrical energy insulator Issue Creating edges based on shared words (here, "energy") does not always generate meaningful connections.Permissiveness in automatically populated column links Fact 2 A tape measure * is used to measure distance.Fact 2 centimeters (cm) are a unit used for measuring is distance. IssueIdeally this edge should generalize to all kinds of measuring tools and units (e.g.X is used to measure Y, Z is a unit for measuring Y).The connection between tape measure * in Fact 1 and measure in Fact 2 makes generalization unlikely, and should be removed.Measure Count Graph Nodes: Nodes before merging 700 Nodes after merging 540 (77%) Graph Edges: Edges before curation 637 Edges after curation 771 (21%) Grid Row-to-Row Connections: Row-to-row connections before curation 1384 Row-to-row connections modified 631 (46%) Row-to-row connections removed 224 (16%) Grid Edge Constraints: Edge constraints before curation 2101 Edge constraints removed 133 (6%) Edge constraints marked optional 27 (1%)

Read the paper · More papers on PaperTik