Information Extraction from Hindi Texts

Kamlesh Dutta, Saroj Kaushik, Nupur Prakash · 2004

The paper presents an information extraction system that takes input from Hindi texts and improves the information content retrieved by using anaphor/pronoun resolution mechanism.The information extraction system developed consists of three major modules: The language Parser, Resolution System and Information Extractor.The language parser used is HPSG (Head-Driven Phrase Structure Grammar) based that provides both syntactic and semantic information to the anaphor resolution system.HPSG was chosen because it provides a set of constraint on the co-referential structures in the language, which bounds the search for an antecedent to a more precise location in the discourse.The semantic information included in its parsing may be helpful for removing ambiguity in anaphor/pronoun resolution.The anaphor resolution system uses few heuristic rules to resolve intrasentential references while centering theory is used for intersentential resolution EvaluationThe anaphor approach used is tested over 10 short stories and following accuracy was observed:Correct resolution: 63% Correct third person pronoun resolution: 69.2% Correct Definite pronoun resolution:

Read the paper · More papers on PaperTik