Approximate discriminative training of graphical models

Jeremy Jancsary · reposiTUm (TU Wien) · 2012

The grapheme-to-phoneme prediction task 107 37 Grapheme-to-phoneme: Results by competing training methods 107 38 The ConjugateGradient method for Gaussian inference 116 39 Relative speed-up of CG over several stages of refinement 116 40 A blocked Gibbs sampler suitable for our setting 117 41 Blocked Gaussian belief propagation 118 42 General form of the negative log-likelihood and its gradient 122 43 General form of the negative log-pseudolikelihood and its gradient 123 44 Pseudolikelihood: Gradient of the expected energy of a factor 124 45 Convex set of 2 × 2 matrices with bounded eigenvalues 124 46 Pseudolikelihood: Training efficiency 126 47 Encoding discrete labels via orthonormal bases 126 48 Quadratic fit of a repulsive pairwise discrete energy table 127 49 Chinese characters: Associativity of the learned pairwise potentials 128 50 The loss function matters: Bias of models trained for different losses 130 51 Derivative of the loss function w.r.t. a single model parameter 132 52 Increasing expressiveness of Gaussian models via conditioning 133 53 Illustration of how Regression Tree Fields (RTFs) work 133 54 Use of regression trees in standalone applications vs. RTF 134 55 Repetitive instantiation of factors in a regression tree field 135 56 Benefits of non-parametric pairwise factors: Increased PSNR 135 57 Benefits of splitting tree nodes for maximum increase in gradient norm 137 58 The OptimizeLossJointly for direct risk minimization 138 59 The OptimizeLikelihoodJointly algorithm for maximum PL 139 60 Benefits of joint versus separate maximization of pseudolikelihood 139 61 The Chinese Characters in-painting task 142 62 The Snakes discrete multi-label prediction task 143 63 The mixed discrete/continuous Joint Detection and Registration task 145 64 The continuous Face Colorization task 146 65 Complementarity of existing denoising methods 147 66 Visual improvement in denoising quality vs. previous state of the art 149 67 Illustration of what the model learned about images 150 68 Improvement in JPEG deb locking vs. the state of the art 151 69 Failure of common denoising methods to remove structured noise 152 List of Tables TightenBound: Impact of the set of spanning trees 66 TightenBound: Standard deviation of the approximation error 66 Accuracy of test set predictions on the Chinese Characters task 143 Accuracy of test set predictions on the Snakes task 144 Typical running time of competing denoising methods 149 Accuracy of test set predictions on the Denoising task 150 Accuracy of test set predictions on the JPEG Deblocking task 151Papers that directly contribute to this thesis J. Jancsary, S. Nowozin, and C.

Read the paper · More papers on PaperTik