Time-series Rule Discovery on Gene Expression Data
Isabell Schu, Gerhard Weikum, Hans‐Peter Lenhof · Max Planck Institute for Plasma Physics · 2006
Gene expression data capture which genes are activated or inhibited at a particular point in time. The mechanism of gene expression depends on transcription factors that work as promoters or enhancers of the gene expression. However, the control mechanism of gene expression is partly unknown. We aim at finding dependencies between genes of the form if gene A is active then gene B becomes active or inactive within a certain time with the goal to constitute new and important biological information about gene expression. To reach this goal we model the dependencies, described above, as association rules. We extend the most efficient algorithm for finding association rules, the A-priori algorithm [2], to discover rules with a certain time offset from the given gene expression data. We have to handle the problem of finding an immense number of rules and false positive rules, i.e. rules that are marked as ’good’ rules but are not relevant in biology. To keep track of all these rules we provide an interactive toolkit that guides the user in finding relevant rules with different types of visualizations that allows to control the quality of the rules. The user can improve the search results by changing the required parameters depending on the quality of the retrieved rules. Although we do not use any prior knowledge, we are able to extract known relations between genes from the public available gene expression data of the baker’s yeast (Saccharomyces cerevisiae) examined by Cho [12] and Spellman [42]. Applying our extended version of the A-priori algorithm we are able to find genes that are involved in the regulation of the same biological processes, for example the cell cycle and DNA replication. This indicates that many of the other identified rules may represent real regulatory interactions.