Rationalizing Text Matching: Learning Sparse Alignments via Optimal Transport
Kyle Swanson, Lili Yu, Tao Leí · 2020
Selecting input features of top relevance has become a popular method for building selfexplaining models.In this work, we extend this selective rationalization approach to text matching, where the goal is to jointly select and align text pieces, such as tokens or sentences, as a justification for the downstream prediction.Our approach employs optimal transport (OT) to find a minimal cost alignment between the inputs.However, directly applying OT often produces dense and therefore uninterpretable alignments.To overcome this limitation, we introduce novel constrained variants of the OT problem that result in highly sparse alignments with controllable sparsity.Our model is end-to-end differentiable using the Sinkhorn algorithm for OT and can be trained without any alignment annotations.We evaluate our model on the Stack-Exchange, MultiNews, e-SNLI, and MultiRC datasets.Our model achieves very sparse rationale selections with high fidelity while preserving prediction accuracy compared to strong attention baseline models.† * Denotes equal contribution.† Our code is publicly available at https://github. com/asappresearch/rationale-alignment.Can I find duplicate songs with different names?I have so many duplicate songs but they have different names.Is there an application I can use to find and delete the duplicates?How to find (and delete) duplicate files?I have a largish music collection and there are some duplicates in there.Is there any way to find duplicate files.At a minimum by doing a hash and seeing if two files have the same hash.… I'm happy using the command line if that is the easiest way.