Identifying Terms in Open Source Software License Texts
Georgia M. Kapitsaki, Demetris Paschalides · 2017
Open source software is nowadays widely used and any open source software must carry a prominent license. However, the legal, natural language text of open source licenses is not always easy to interpret and an extensive manual analysis of the text may be required, in order to fully understand its content. Existing approaches present license content based on such manual interpretation. In this paper, we propose an automated license term extraction system (FOSS-LTE) for the identification of the license terms from a specific license text and the creation of a representation of these terms divided into rights, obligations and additional conditions. We present the process employed for the creation of the license term extraction system using NLP techniques and we evaluate its accuracy on a set of sentences from available licenses collected for this purpose.