Bayesian forward regularization replacing Ridge in online randomized neural network with multiple output layers
Junda Wang, Minghui Hu, Ning Li, Ponnuthurai Nagaratnam Suganthan · Pattern Recognition · 2025
Forward regularization (-F) with unsupervised knowledge was advocated to replace canonical Ridge regularization (-R) in online linear learners, as it achieved a lower relative regret boundary. However, we observe that -F cannot perform as expected in practice, even possibly losing to -R for online tasks. We identify two main causes for this: (1) inappropriate intervened regularization, and (2) non-i.i.d. nature and data distribution changes in online learning (OL), both of which result in unstable posterior distribution and optima offset of the learner. To improve these, we first introduce the adjustable forward regularization (- k F), a more general -F with controllable knowledge intervention. We also derive - k F’s incremental updates with variable learning rate, and study relative regret and boundary in OL. Inspired by the regret analysis, to curb unstable penalties, we further propose - k F-Bayes style with k synchronously self-adapted to revise the intractable tuning of - k F by considering parametric posterior distribution changes in non-i.i.d. online data streams. Additionally, we integrate the - k F and - k F-Bayes into a multi-layer ensemble deep random vector functional link (edRVFL) and present two practical algorithms for batch learning, avoiding past replay and catastrophic forgetting. In experiments, we conducted tests on numerical simulation, tabular, and image datasets, where - k F-Bayes surpassed traditional -R and -F, highlighting the efficacy of ready-to-work - k F-Bayes and the great potentials of edRVFL- k F-Bayes in OL and continual learning (CL) scenarios. • - k F provides a more flexible unsupervised knowledge intervention for online learners. • We derive - k F’s incremental updates and study relative regret in OL. • We propose - k F-Bayes to consider parametric posterior changes in non-i.i.d. streams. • We integrate the - k F and - k F-Bayes into edRVFL and present two algorithms for CL.