Bilingual Word Embeddings with Bucketed CNN for Parallel Sentence Extraction

Jeenu Grover, Pabitra Mitra · 2017

We propose a novel model which can be used to align the sentences of two different languages using neural architectures.First, we train our model to get the bilingual word embeddings and then, we create a similarity matrix between the words of the two sentences.Because of different lengths of the sentences involved, we get a matrix of varying dimension.We dynamically pool the similarity matrix into a matrix of fixed dimension and use Convolutional Neural Network (CNN) to classify the sentences as aligned or not.To further improve upon this technique, we bucket the sentence pairs to be classified into different groups and train CNN's separately.Our approach not only solves sentence alignment problem but our model can be regarded as a generic bag-of-words similarity measure for monolingual or bilingual corpora.

Read the paper · More papers on PaperTik