Thermal Debugging Tool for Servers
Eduardo García-Espinosa, Enrique Gonzalez-Garcia, Adolfo Hernandez-Padilla, Raymundo Aguillon, Paulo Lopez‐Meyer · 2020
Among the different possible random causes of failure in server clusters, temperature variability inside the server's chassis is the most common/relevant variable. Random failures are difficult to isolate, which makes it difficult to find the root cause and debug it. At the silicon level, the experience shows that process, voltage, and temperature (PVT) variations are the most relevant variables that are used to replicate issues, since such variables increase the failure probability, which allows to understand and speed up the failure mechanism of a given problem. When random problems happen in server cluster environments at a customer data center, the debug in the validation laboratory becomes a challenging task since it is difficult to reproduce the conditions that triggers the problem. With this in mind, this paper shows a methodology that focus in one of these variables, temperature, to apply controlled and located variations at the system level to expedite the replication of thermal issues and to perform validation for commercial platforms.