Parsed DMOZ data

Gaurav Sood · Harvard Dataverse · 2016

DMOZ is a large communally maintained open directory that categorizes web content. The data are posted in a complex XML format. The python scripts posted here were used to parse the data posted at: http://rdf.dmoz.org/ on June 12, 2016 to produce a csv file posted here. The structure of the file is "URL","Category 1","Category 2",.......... Given the categories are separated by commas, doing read_csv without the right options can be problematic Here's some code to read in the file: https://gist.github.com/soodoku/a97e6cf2800429d1c541ac2fb65e4c98

Read the paper · More papers on PaperTik