Data Encodings and File Formats
Field Cady · 2017
This chapter talks about most important file formats for a data scientist, such as CSV, JSON, XML, HTML, and Tar. This will include sample code for parsing them, discussions about when they are useful, and some thoughts about the future of data formats. The chapter discusses how data is laid out in the physical memory of a computer. This will involve peaking under the hood of the computer to look at performance considerations and give us a deeper understanding of the discussed file formats. Unicode is actually a family of encoding standards, all of them aiming to supplement basic ASCII with the massive range of other characters needed today and possibly in the future. The main version of Unicode available is UTF-8, and it is fast becoming the most popular encoding around. The chapter discusses this main version of Unicode.