Observability for managers: how to know the business is running

For:
CTOs, CIOs, IT managers and SRE team leads assessing monitoring and the cost of an observability platform
Reading time:
10 min
Episode:
29 min

This episode is in Polish. Full Polish version with transcript

After you click, the video loads from YouTube (Google). Google may store data on your device and process it in the USA as well. More (PDF, in Polish) · Watch on YouTube

In brief

If your infrastructure is monitored and you still hear about an outage when the business calls you, CPU and memory monitoring is not enough. Observability gives the business, developers and operations the same data about what happens in the systems, so problems show up earlier and numbers settle disputes. You will learn when optimization and higher availability do not pay off, what to do with an expensive vendor platform and what a rollout on an open standard requires.

Key takeaways

  • Resource monitoring does not tell you whether the business is running, while observability lets you see a problem before it starts, so the team leaves firefighting mode and can warn the business itself.
  • Separate tools in each team have to be bought, learned and maintained, and during an outage one side sees it while the other may not.
  • Numbers show how many users the system is slow for and which machines can be downsized, so you know where to spend money and where to stop.
  • When the bill for a commercial platform is high, an alternative is OpenTelemetry with the Grafana stack, which your SRE team probably already knows, although the rollout requires changes on the application side.
  • Tools come first, then processes: alerts, anomaly detection and logging standards, which a platform team can roll out in the background.
  • Instead of full availability, you agree on a measurable commitment with the business and check which systems can occasionally be down.

Monitoring sees resources, observability sees system behavior

Monitoring is checking whether the CPU, the memory and maybe also the disk are full. A manager asks something else: whether their business is running. Monitoring reports an outage or, at best, warns before one.

The speakers have often heard companies say they have some monitoring or observability in place, and yet they almost always find the same situation. The business calls the manager, often quite angrily, to say the system has been down for half an hour. The operations team checks the CPU and RAM, sees that they are fine and replies that the system seems to be working.

Observability is not a trendy name for monitoring. Observability is a goal: a state in which you can infer from system data how the system behaves and what happens inside it. It covers systems and processes. With it, you see that something is happening, or even see it before it starts.

Logs, metrics and traces are successive levels of maturity

Everyone has logs, better or worse. Sometimes developers log in to machines and read the scrolling logs like in The Matrix. That is a skill, but at a larger scale not necessarily a useful one.

Metrics give real numbers, and you can collect a great many of them. A trace is a record of how a whole business process flows through the services and microservices in the organization. Traces show which calls between services work and which do not, their latency, where they fail and how response times are distributed. With them, you know where it pays to invest time in optimization.

A team in firefighting mode does not deliver its plan

A team that reacts to outages works in firefighting mode. It has to restore the system immediately, usually at the expense of quality. The problem comes back, because the team does what is needed right away and then usually has no time to fix it properly.

In this mode, nothing can be planned. The team announces something for the quarter, the quarter ends and the work has not even started. The team did a lot, but it was not valuable work.

With observability, the team invests in operational processes, has more planned time, and there are simply fewer outages. That leaves time to show developers that something runs slowly or is misused, or that the application performs poorly.

A comparison of two ways an operations team works. In firefighting mode, an angry call from the business leads to a quick fix right away, at the expense of quality. There is no time for a proper fix, so the problem comes back, the cycle repeats and the plan for the quarter stalls. With observability, the team sees the problem early, warns the business itself and has time for planned work, and there are fewer outages.
Diagram: firefighting mode and working with observability

When the team knows that an outage is coming, it can call the business itself: something is happening with the system, we are working on it. Even in a serious outage, that is a very different situation from an angry call from the business.

Separate tools in every team are an extra cost

Administrators, developers, ops and SRE teams work on their own pieces of the system. Everyone looks at their own part and everyone says that things are fine. During an outage, a DevOps engineer sits down with a developer, each has their own truth, and they cannot agree. One side sees the outage, the other may not see it.

Each team has its own tools. You have to pay for them, people have to learn them, and someone has to maintain them. A local debugger is no longer enough for developers. In recent years, they have come to expect numbers on how the system really behaves. If they maintain their tools themselves, they lose time on it, and software development and delivery slow down.

The direction is centralization. With unified tools, development, operations and monitoring work from the same numbers and speak one language. Few people will argue with numbers that the organization collects itself.

Numbers show what to optimize and where to downsize machines

A good level of observability lets you talk about what is happening rather than about impressions, because dashboards and numbers describe every situation. The manager gets an argument to use with every team.

In a large organization, someone will always complain that the system is slow. With numbers, you see how many people it is slow for, for example 2% of users. Then you decide whether it is worth spending money to improve that result. Szymon Warda: “Not every system is worth optimizing to the limit.” (translated)

Numbers also bring order to infrastructure decisions. For a new system, a team requests 20 CPUs and 100 GB of RAM, and then it turns out the system uses 50% of that. When a month of charts shows that CPU, memory and disk never go above a certain value, the conclusion is: let’s use smaller machines. That means very large savings, and the more infrastructure you have, the easier they are to get.

Commercial observability platforms are a significant cost

A vendor observability system costs 5–6 figures in dollars a year. For a mid-sized team, that is a fairly significant cost, and the amount depends on the size of the organization.

A large financial-sector client, unnamed here, already pays one of the big vendors a bill of 7 and 8 figures in dollars a year. The client saw what the platform Protopia recommends can do and is considering leaving that vendor. A few years ago, industry publications described organizations that paid 15 and 50 million dollars a year for observability systems.

The Grafana stack is the market standard

The platform Protopia recommends is built on the Grafana stack and integrates natively with Kubernetes. When you propose it to your developers, they will probably say they know it. Your SRE team knows which system you mean, because it uses it. If you look, you will find it in your organization, only probably less well maintained and not wired into processes.

These systems are the market standard in Kubernetes and at the cloud giants, the hyperscalers. The open-source tools in this stack scale very well, are cost-effective and are easy to maintain. You do not depend on a closed tool, so you draw on the knowledge of a huge user community, and you implement some good practices with ready-made components developed in the ecosystem. Open source has its own problems too.

A rollout requires work on the application side

Observability requires some work and changes on the application side, although these changes keep getting smaller. The more the team does, the more value it gets, and this is long-term work. It is worth doing once and doing it right, because OpenTelemetry has become the standard in observability. Big vendors also use it to integrate with applications.

Do not expect everything to work on its own. One client turned on automatic observability from a very well-known solution on its cluster, and the whole system went down. The mechanism hooked into a place it should not have, and the system stopped working immediately. Finding the cause took two days. “Magic” solutions often work this way: they promise too much and then fail to deliver, because they cannot.

The goal is a business that runs predictably. You do not want a situation where you upgrade an operations tool and your business is down for a whole day, and nobody knows why. Szymon Warda: “We know full well that rolling out any system has a cost.” (translated)

Alerts can warn of an outage, charts show it after the fact

A falling chart shows an outage that has already happened. Organizations sometimes have 4, 5, 6 or more monitors with nice charts. They look like a Christmas tree for show, and nobody uses them.

Alerts let you act earlier: they can notify you that something will happen in a day, in two days or at some later point. Observability also includes a change in the organization’s habits. First the technical part provides the capabilities, then the team introduces processes, for example alerting and anomaly detection, which catches signals that something is starting to break.

Systems sometimes produce tens of thousands of metrics, so you have to know how to filter them. Without that, it is drinking from a fire hose, and you can lose your teeth.

Developers see how their code runs in production

One unified tool makes shift left possible: developers see how the system behaves in production. The “it works on my machine” argument disappears, because there are numbers from production, including the SQL queries being executed. The data shows that something needs fixing, or that everything is fine and further performance work makes no sense, because the system runs very well there.

The feedback loop gets shorter. The team deploys a new version and a day later sees whether it works, does not work or runs slowly, instead of learning about it after a quarter. Measurement shows whether the new version is better or worse, so the team detects a regression much faster and knows exactly what happened.

The feedback loop after a deployment. The team deploys a new version to production and a day later has numbers from production, including the SQL queries being executed. The measurement shows whether the version is better or worse. On a regression, the team fixes the code and deploys another version. When the system works well, further performance work makes no sense.
Diagram: the feedback loop with numbers from production

In critical systems, where it makes sense, you can deploy a new version partially and watch how the traffic behaves. Such deployments do not pay off everywhere.

SLA, SLO and SLI define what the team commits to

Without numbers, you cannot measure whether an investment in improving a metric brings value. With numbers, you can set an SLA, SLOs and SLIs and state the service level the team commits to.

Some companies say they must have 100% uptime. That is unrealistic. To put it jokingly, every additional nine after the decimal point is another zero before the decimal point in cost, sometimes even more, sometimes not.

With such a commitment, you can measure whether the contract with the business is met. Potential outages and maintenance windows can be translated into the organization’s real finances. Sometimes it turns out that some systems can be down for a certain period without any problem, while others are more critical.

The same measurements remove friction between development teams that say their system fails because another one does. A team commits that its system runs within the agreed SLA, and development stops being about complaints that something does not work for someone.

A platform team rolls out standards in the background

The organization does not have to decide on its own, before it starts, what to monitor and how. There are universal standards and good practices to adopt instead of reinventing the wheel. Every organization is different to some degree, so the standards need adjusting.

If you have a platform team, it rolls out further standards in the background: how to log, which metrics to collect, which error levels to report and what goes into the logs. When GDPR came in, teams had a lot of work removing first names, last names and user-identifying data from logs. That was a one-time effort, and now these rules go into the standard.

Developers do not have to think about it: a new system gets the agreed components and libraries and reports in the agreed way. A costly piece of system development goes away. The standard is current and industry-wide.

Data from legacy, new, evolving and external systems can be collected into one pipeline and reported in a unified way. Instead of pieces in different tools, you see the whole organization and know how each part behaves and where.

Data from legacy, new, evolving and external systems flows through one pipeline to a platform on the Grafana stack, which reports it in a unified way. The business, developers and operations see the same numbers and the whole organization instead of pieces in different tools.
Diagram: data from every system in one pipeline

Speakers

  • Szymon Warda

    Szymon Warda

    Founder, Managing Partner, Technology Advisor.

    Szymon Warda co-founded Protopia and is a member of the Grafana Champions program. He co-hosts Patoarchitekci, where he has talked about distributed systems, observability and FinOps since 2019.

    All posts by this author
  • Mikołaj Szczerbicki

    Mikołaj Szczerbicki

    Head of Sales & Business Development.

    Mikołaj Szczerbicki is Head of Sales & Business Development at Protopia and co-hosts the Powered by Protopia podcast. He scopes and prices projects, so he asks about the cost, risk and timeline of AI, Azure and Kubernetes work.

    All posts by this author

FAQ

How do you measure systems of different teams that talk to each other?

Besides observability data, API Management can measure some aspects of the communication between systems, at a much higher level.

How does leaving firefighting mode affect the team?

It affects employee retention, because work is simply better. Working from outage to outage is not good management of the systems or of the team's time.

How long after the rollout does the organization work differently?

A quarter after the rollout, it is already a different organization that thinks ahead. The change comes from leaving firefighting mode and from standardizing the development processes.