Information Extraction in Semantic, Highly-Structured, and Semi-Structured Web Sources

Víctor Manuel Alonso Rorís, Juan Manuel Santos-Gago, Roberto Pérez Rodríguez, Carlos Rivas Costa, Miguel Gómez Carballa, Luis Anido · Polibits · 2014

"The evolution of the Web from the original proposalmade in 1989 can be considered one of the most revolutionarytechnological changes in centuries. During the past 25 years theWeb has evolved from a static version to a fully dynamic andinteroperable intelligent ecosystem. The amount of data producedduring these few decades is enormous. New applications,developed by individual developers or small companies, can takeadvantage of both services and data already present on the Web.Data, produced by humans and machines, may be available indifferent formats and through different access interfaces. Thispaper analyses three different types of data available on theWeb and presents mechanisms for accessing and extracting thisinformation. The authors show several applications that leverageextracted information in two areas of research: recommendationsof educational resources beyond content and interactive digitalTV applications."

Read the paper · More papers on PaperTik