Cost-Benefit Analysis of Computer Resources for Machine Learning
Richard A. Champion · Antarctica A Keystone in a Changing World · 2007
Normalized road and dasymetric density across the San Francisco Bay Area.Road density is a measure of distance to the nearest road for each pixel.Normalization scales this to the range 0 to 1. Dasymetric density uses census-block and land-cover information to estimate population density per 30-m pixel.Population density values near the high end of the scale (near 1 person per 30-m pixel) suggest artifacts generated from inaccuracies in the land-cover information .................3 2. Calibration time as a function of the number of training points.The training time is approximately quadratic (indicated by the exponent 2.19) and explains more than 98 percent of the variance in the performance data ..............................................................3 3. Population density predicted from a probabilistic neural network using normalized road density and training-set sizes of 1,000 and 50,000 points.The average of the curves suggests how the goodness-of-fit varies with training sets of intermediate sizes.The blips in the curves for normalized road density greater than about 0.6 are likely due to noise in the data.For these data a large increase in the number of training points does not lead to a large improvement in goodness of fit ............................4 4. Scatterplot showing the geometrically uneven distribution of data points.The clustering of data near the origin indicates that the highest population densities are found in the areas containing the highest density of roads .............