Construction of document feature vectors using BERT
Hirotaka Tanaka, Rui Cao, Jing Bai, Wen Ma, Hiroyuki Shinnou · 2020
In recent years, the performance of various natural language processing systems has been greatly improved by pre-training models such as Bidirectional Encoder Representations from Transformers (BERT). The BERT model obtains word em-beddings in context by the multi-head attention of a Transformer. In document-classification tasks, document feature vector can be constructed from the BERT output. However, the length of the BERT input sequence is upper-limited, which reduces the amount of information obtainable from BERT by the standard method. This limitation is problematic when handling long documents. Against this background, we propose a method that constructs document feature vector in BERT by embedding all words in long documents.