Detecting Common Weakness Enumeration Through Training the Core Building Blocks of Similar Languages Based on the CodeBERT Model

Chansol Park, R. Young Chul Kim · 2023

The traditional code analysis approach to measure code quality, such as bad smells, coupling, cohesion, complexity, and common weakness enumeration, was a rule-based mechanism. In current, analyzing code through AI is being studied by many researchers. AI code analysis approach trains code patterns to AI model through labeled code datasets. The problem is that AI models require plenty of training datasets for better performance, but there are not enough datasets to train. To solve this problem, we suggest training an AI model with core building blocks of a target language and similar languages together. We can solve two problems the lack of datasets and the low performance of the AI model. Finally, we will compare the performance of the two models. One is a model trained with a dataset containing only one target language, and the other is a model trained with a dataset containing similar languages.

Read the paper · More papers on PaperTik