Computational Intelligence in Software Cost Estimation: Evolving Conditional Sets of Effort Value Ranges
Efi Papatheocharous, Andreas S. · InTech eBooks · 2008
In the area of software engineering a critical task is to accurately estimate the overall project costs for the completion of a new software project and efficiently allocate the resources throughout the project schedule.The numerous software cost estimation approaches proposed are closely related to cost modeling and recognize the increasing need for successful project management, planning and accurate cost prediction.Cost estimators are continually faced with problems stemming from the dynamic nature of the project development process itself.Software development is considered an intractable procedure and inevitably depends highly on several complex factors (e.g., specification of the system, technology shifting, communication, etc.).Normally, software cost estimates increase proportionally to development complexity rising, whereas it is especially hard to predict and manage the actual related costs.Even for well-structured and planned approaches to software development, cost estimates are still difficult to make and will probably concern project managers long before the problem is adequately solved.During a system's life-cycle, one of the most important tasks is to effectively describe the necessary development activities and estimate the corresponding costs.This estimation, once successful, allows software engineers to optimize the development process, improve administration and control over the project resources, reduce the risks caused by contingencies and minimize project failures (Lederer & Prasad, 1992).Subsequently, a commonly investigated approach is to accurately estimate some of the fundamental characteristics related to cost, such as effort and schedule, and identify their interassociations.Software cost estimation is affected by multiple parameters related to technologies, scheduling, manager and team member skills and experiences, mentality and culture, team cohesion, productivity, project size, complexity, reliability, quality and many more.These parameters drive software development costs either positively or negatively and are considerably very hard to measure and manage, especially at an early project development phase.Hence, software cost estimation involves the overall assessment of these parameters, even though for the majority of the projects, the most dominant and popular metric is the effort cost, typically measured in person-months.Recent attempts have investigated the potential of employing Artificial Intelligence-oriented methods to forecast software development effort, usually utilising publicly available www.intechopen.comTools in Artificial Intelligence 2 datasets (e.g., Dolado, 2001;Idri et al., 2002;Jun & Lee, 2001;Khoshgoftaar et al., 1998;Xu & Khoshgoftaar, 2004) that contain a wide variety of cost drivers.However, these cost drivers are often ambiguous because they present high variations in both their measure and values.As a result, cost assessments based on these drivers are somewhat unreliable.Therefore, by detecting those project cost attributes that decisively influence the course of software costs and similarly define their possible values may constitute the basis for yielding better cost estimates.Specifically, the complicated problem of software cost estimation may be reduced or decomposed into devising and evolving bounds of value ranges for the attributes involved in cost estimation using the theory of conditional sets (Packard, 1990).These ranges may then be used to attain adequate predictions in relation to the effort located in the actual project data.The motivation behind this work is the utilization of rich empirical data series of software project cost attributes (despite suffering from limited quality and homogeneity) to produce robust effort estimations.Previous work on the topic has suggested high sensitivity to the type of attributes used as inputs in a certain Neural Network model (MacDonell & Shepperd, 2003).These inputs are usually discrete values from well-known and publicly available datasets.The data series indicate high variations in the attributes or factors considered when estimating effort (Dolado, 2001).The hypothesis is that if we manage to reduce the sensitivity of the technique by considering indistinct values in terms of ranges, instead of c r i s p d i s c r e t e v a l u e s , a n d i f w e e m p l o y a n e v o l u t i o n a r y technique, like Genetic Algorithms, we may be able to address the effect of attribute variations and thus provide a near-to-optimum solution to the problem.Consequently, the technique proposed in this chapter may provide some insight regarding which cost drivers are the most important.In addition, it may lead to identifying the most favorable attribute value ranges for a given dataset that can yield a 'secure' and more flexible effort estimate, again having the same reasoning in terms of ranges.Once satisfactory and robust value ranges are detected and some confidence regarding the most influential attributes is achieved, then cost estimation accuracy may be improved and more reliable estimations may be produced.The remainder of this work is structured as follows: Section 2 presents a brief overview of the related software cost estimation literature and mainly summarizes Artificial Intelligence techniques, such as Genetic Algorithms (GA) exploited in software cost estimation.Section 3 encompasses the description of the proposed methodology, along with the GA variance constituting the method suggested, a description of the data used and the detailed framework of our approach.Consequently, Section 4 describes the experimental procedure and the results obtained after training and validating the genetic evolution of value ranges for the problem of software cost estimation.Finally, Section 5 concludes the chapter with a discussion on the difficulties and trade-offs presented by the methodology in addition to suggestions for improvements in future research steps.