Updating Compressed Column-Stores
Sándor Héman · Centrum Wiskunde & Informatica (CWI), the national research institute for mathematics and computer science in the Netherlands · 2015
During my computer science studies at university, my main interests were in the areas of hardware architecture, systems and algorithms/data structures.Inspired by the lectures of Martin Kersten (thank you for those!),I figured that the field of database systems would satisfy all my interests, so I applied at his research group within CWI.This is where I got introduced to Peter Boncz, initially as my copromotor.It quickly became clear to me that Peter's brain operates at stunning speeds, just like the database kernel (MonetDB/X100) he was researching.He thought me the ins and outs of hardware-conscious algorithm design through a torrent of "crazy" and unconventional, but at the same time clever and inspiring, ideas.The first one that caught my interest, was the use of data compression to help keep his extremely "data hungry" X100 database engine from starving.With a hint that I should forget about theoretically optimal constructs, like Huffman codes, and rather focus on raw CPU cycles and experimentation.From then on, I was part of the small, hard-working team that did research in the area of high-performance analytic database architecture.Together with Marcin Żukowski, Niels Nes and Peter himself, we spent many days (and nights!) brainstorming, hacking, benchmarking and writing papers.Through Arjen de Vries and Roberto Cornacchia, we also had several "flirts" with the field of information retrieval.I would like to thank the ones involved for these exciting and challenging moments.Eventually, however, after focusing on all kinds of data-and compute-intensive optimizations, I was starting to feel a bit of a tension.Yes, we had a fast engine.And, yes, we had a heavily scan optimized storage layout.However, we were silently ignoring "That-What-Must-Not-Be-Named": column-store updates.Still young and adventurous, I decided to take up the challenge, and attack the mysteries around transactional column-store update support.Soon after this "defining" moment in my PhD career, things took a bit of a turn, however.The year 2008 was the year of both an internship at HP Labs and the Vectorwise spin-off.First things first, however, as 2008 was also the year I got married to my beloved wife, Iris.Thank you Iris for all the love, support and care, during both the good and the darker times.Our wonderful honeymoon in California was followed directly by a research internship at HP Labs.I would like to thank Götz Graefe, Umesh Dayal and Meichun Hsu for providing me this opportunity in the first place, and Goetz in particular, for sharpening my knowledge in the areas of B + -trees and fault tolerance.Thanks also to Dimitris, Stavros and Mehul for showing me around "the Valley".After returning from the US, the Vectorwise spin-off company was already up and running.This meant that, besides getting the work on column-store updates published, it now had to be made production ready as well.Luckily, we were now in a position to hire more people, and I would like to thank Hui Li and Michał Szafranski for helping out with checkpointing, Gosia Wrzesinska for help around transaction management and Lefteris Sidirourgos for his help with publishing a paper during these busy times.An apology is in place towards Gosia, for the nightmares we both had around "join index updates" ;-) Thanks also to Giel de Nijs, not only as a colleague, but also as a "biking buddy".The CONTENTS vii early morning and evening rides provided a welcome break from office and urban life.For getting Vectorwise where it stands today, I would furthermore like to thank