A General Information Extraction Framework Based on Formal Languages
Markus L. Schmid · arXiv (Cornell University) · 2025
For a terminal alphabet $Σ$ and an attribute alphabet $Γ$, a $(Σ, Γ)$-extractor is a function that maps every string over $Σ$ to a table with a column per attribute and with sets of positions of $w$ as cell entries. This rather general information extraction framework extends the well-known document spanner framework, which has intensively been investigated in the database theory community over the last decade. Moreover, our framework is based on formal language theory in a particularly clean and simple way. In addition to this conceptual contribution, we investigate closure properties, different representation formalisms and the complexity of natural decision problems for extractors.