An online learning approach for trend recognition
David Paulk · 2015
This work assesses the performance of the Gradient Descent and Exponentiated Gradient online learning algorithms on the Wikipedia Page Traffic Statistics Dataset. The two algorithms are trained to predict future Wikipedia page traffic during a 7-month period. Predictions are a weighted combination of feature attributes, which are various measurements of page traffic change. The algorithms improve their predictions as they learn patterns from the data sequence by updating a weight distribution over the features. Cumulative loss for predictions during the 7-month sequence is used as an evaluation metric. The cumulative loss measurements demonstrate that the Exponentiated Gradient algorithm predicts future Wikipedia page traffic more accurately Gradient Descent does with the computed features. The cumulative loss measurements also suggest that the Exponentiated Gradient algorithm becomes less confused than the Gradient Descent algorithm becomes when irrelevant features are present among relevant features.