Speech-to-Tree: Cultivating Dependency Structures from Spoken English

Manish Prajapati, Hemant A. Patil · 2025

Dependency parsing has long been a cornerstone of natural language understanding, allowing machines to dissect and learn syntactic structures of sentences. However, traditional methods face certain limitations when applied to spoken language, mainly due to their reliance on a two-step process: first, they transcribe speech into text using Automatic Speech Recognition (ASR), then, they parse the resulting text. This pipeline introduces potential ASR errors, particularly in noisy or resource-scare environments, leading to degraded parsing accuracy and loss of prosodic information, such as intonation and stress. This information is critical to understanding the syntax of spoken language. Our work addresses these challenges by pioneering an end-to-end approach that directly generates dependency trees from the English language, eliminating the need for an intermediate transcription step. Our proposed method aligns speech with syntactic structures, preserving prosodic and acoustic cues, which improves parsing. This multi-modal strategy allows the model to capture the nuances of speech, such as silence and emphasis, which are not present in text-based parsing. Our proposed approaches ensure a seamless fusion of auditory and linguistic information, enabling a direct speech-to-tree (S2T) mapping, which eliminates the error-prone ASR bottleneck. This not only mitigates the impact of transcription inaccuracies, but also leverages the rich contextual information embedded in speech to produce more accurate syntactic representations. Our work paves the way for more natural and effective spoken language understanding, offering a transformative perspective on how machines can interpret embedded human speech structure in a single, unified step, with potential applications in real-time spoken dialogue systems, accessibility tools, and beyond.

Read the paper · More papers on PaperTik