Datetime Feature Recommendation by Word Embedding Methods Using Data Column Names
Satoshi Masuda, Tomohiro Takeda · International Journal of Software Engineering and Knowledge Engineering · 2025
The analysis of large volumes of data to derive new insights is commonly referred to as data science, and its widespread adoption is increasingly necessary. One critical step in data science workflows is feature extraction from data, known as feature engineering. This process heavily depends on expert experience, which has led to current research efforts aimed at its automation. In this paper, we introduce a novel approach to automate feature extraction from the textual information found in data column names. Specifically, we employ natural language processing and source code analysis techniques on existing source code and column names, with a particular focus on datetime features, to construct a knowledge database. Utilizing this knowledge database, we propose a system that recommends datetime features based on newly provided textual information. Departing from conventional methods such as one-hot encoding and Word2Vec word embeddings, our approach is informed by prior research and shifts toward Doc2Vec, which vectorizes at the document level. In our experiments, we validate the classification accuracy of the knowledge database and demonstrate its application in predictive tasks, such as housing price prediction, where it shows improved prediction accuracy. Our proposed approach, which utilizes a Doc2Vec model pre-trained on Wikipedia, results in enhanced prediction accuracy during vectorization.