Scaling up Prioritized Grammar Enumeration for scientific discovery in the cloud

Tony Worm, Kenneth Chiu · 2014

Symbolic Regression (SR) is the data driven search for mathematical relations as performed by a computer. In essence, SR is a search over all possible equations to find those which best model the data on hand. Prioritized Grammar Enumeration (PGE) is a recently proposed algorithm which has been shown to have great efficacy and efficiency on the Symbolic Regression problem, using just a single compute core. PGE reformulates the SR problem as a search over a grammar, makes reductions in the magnitude of the search space, and introduces mechanisms for exploring that space efficiently. Notably, PGE provides reliability and reproducibility of results, a key aspect to any system used by scientists at large. In this paper, we enhance the PGE algorithm in several ways. First, we extend PGE to discover differential equations. Second, we incorporate multiple prioritization heaps into PGE, reducing point evaluations while maintaining efficacy. Finally, we decouple the PGE subroutines into a set of services, contain each with Docker, and deploy them onto the cloud. Our algorithm experiments cover a range of dynamical systems from a multitude of domains. and our cloud experiments explore a variety of architectural setups. Our results show PGE to have great promise and efficacy in automating the discovery of equations at the scales needed by tomorrow's scientific data problems.

Read the paper · More papers on PaperTik