Observing and Controlling Performance in Microservices
André Bento · Portuguese National Funding Agency for Science, Research and Technology (RCAAP Project by FCT) · 2019
Microservice based software architecture are growing in usage and one type of data generated to keep history of the work performed by this kind of systems is called tracing data.Tracing can be used to help Development and Operations (DevOps) perceive problems such as latency and request work-flow in their systems.Diving into this data is difficult due to its complexity, plethora of information and lack of tools.Hence, it gets hard for DevOps to analyse the system behaviour in order to find faulty services using tracing data.The most common and general tools existing nowadays for this kind of data, are aiming only for a more human-readable data visualisation to relieve the effort of the DevOps when searching for issues in their systems.However, these tools do not provide good ways to filter this kind of data neither perform any kind of tracing data analysis and therefore, they do not automate the task of searching for any issue presented in the system, which stands for a big problem because they rely in the system administrators to do it manually.In this thesis is present a possible solution for this problem, capable of use tracing data to extract metrics of the services dependency graph, namely the number of incoming and outgoing calls in each service and their corresponding average response time, with the purpose of detecting any faulty service presented in the system and identifying them in a specific time-frame.Also, a possible solution for quality tracing analysis is covered checking for quality of tracing structure against OpenTracing specification and checking time coverage of tracing for specific services.Regarding the approach to solve the presented problem, we have relied in the implementation of some prototype tools to process tracing data and performed experiments using the metrics extracted from tracing data provided by Huawei.With this proposed solution, we expect that solutions for tracing data analysis start to appear and be integrated in tools that exist nowadays for distributed tracing systems.