Computing Movie Script Similarity with Neural Word Embeddings
Theodore Patsis · 2021
Typical applications of machine learning in the movie domain are box office revenue prediction and recommender systems using features derived from viewers' behavioral data, casting, budget, etc. None of these models have used a movie's script as their key feature, which is one of the only artifacts available early in the production process. We examine film similarity taking into account only the film's script, with a view to support decision making early in the movie production process. We describe two models for calculating movie script similarity, one based on a one-hot vector encoding, and the second using neural word embeddings (namely, word2vec). In a quantitative evaluation, we asked 3 human expert annotators to rank the output of each algorithm, and a random baseline, on a test set of 45 films stratified across many genres. The model using word2vec embeddings ranked first in 61 % of annotator judgments and ranked above the one-hot vector encoding in 70% of rankings.