Monitoring with Nagios and Trend Analysis with Cacti
Syed Ali · Apress eBooks · 2014
Monitoring is perhaps one of the most important pieces of infrastructure management. When systems go down, monitoring should alert the site reliability engineers (SREs) so they can investigate the service affected and try to bring the system back online. After that, a root cause analysis should be conducted and actions should be taken to prevent similar issues in the future. Ideally, monitoring will alert about issues before they cause a service outage. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.