A Margin-based Loss with Synthetic Negative Samples for Continuous-output Machine Translation
Gayatri Bhat, Sachin Kumar, Yulia Tsvetkov · 2019
Neural models that eliminate the softmax bottleneck by generating word embeddings (rather than multinomial distributions over a vocabulary) attain faster training with fewer learnable parameters.These models are currently trained by maximizing densities of pretrained target embeddings under von Mises-Fisher distributions parameterized by corresponding model-predicted embeddings.This work explores the utility of margin-based loss functions in optimizing such models.We present syn-margin loss, a novel marginbased loss that uses a synthetic negative sample constructed from only the predicted and target embeddings at every step.The loss is efficient to compute, and we use a geometric analysis to argue that it is more consistent and interpretable than other marginbased losses.Empirically, we find that synmargin provides small but significant improvements over both vMF and standard marginbased losses in continuous-output neural machine translation.