Observability in practice: a Grafana stack rollout plan and demo

For:
CTOs, IT managers, architects and DevOps team leads who organize monitoring across several environments
Reading time:
12 min
Episode:
36 min

This episode is in Polish. Full Polish version with transcript

After you click, the video loads from YouTube (Google). Google may store data on your device and process it in the USA as well. More (PDF, in Polish) · Watch on YouTube

In brief

If you run several environments, your monitoring has probably grown over time without a plan, and during an outage admins and developers look at different data. A self-hosted Grafana stack collects metrics, logs and traces from the whole organization in one place and has no per-user licensing. From a small company upward, the stack rolls out in phases alongside the tools you already have, and your team mainly provides access.

Key takeaways

  • A self-hosted Grafana stack makes sense in a company bigger than a micro business, while a small team with one monolith is better served by an off-the-shelf tool.
  • The rollout starts with an inventory and a single entry point in Grafana, because a big-bang change in which people do not move properly to the new tools ends in chaos.
  • Your team gives access to CI/CD and to the Kubernetes cluster (Protopia can also set up the cluster), then points to the systems, and every system handed over ends with a training session.
  • Modern applications are instrumented with a few libraries, and purchased systems run through an agent with no code changes.
  • eBPF profiling shows individual method calls without integrating with the application and with lower overhead than other APM systems, but it works only on Linux.
  • Zabbix or ELK can stay because Grafana works with them, and the rollout can go in parts, although the recommended option is the complete one, with training.

A self-hosted Grafana stack makes sense beyond a micro business

A self-hosted Grafana stack starts to make a lot of sense in a small or medium-sized organization, because at that size the overhead of data collection and the volume of data are already significant. The stack has no limits tied to the scale of use. In the smallest companies, the math is different:

Szymon Warda: “If you have a team of 10–20 developers and one monolithic system, get something off the shelf. It will simply be cheaper to maintain.” (translated)

An off-the-shelf tool then gives a good start, a fast rollout and a fast return on investment.

An example: an organization with Zabbix, with ELK for security, with a hybrid environment (some on-premises, some cloud), with several clouds and different programming languages. For such a company a self-hosted Grafana stack makes sense, and some observability components already run there.

The choice depends on the size of the organization. A team of 10–20 developers with one monolithic system is better served by an off-the-shelf tool, because it is cheaper to maintain. For a small or medium-sized organization with a hybrid environment, several clouds and different programming languages, a self-hosted Grafana stack without per-user licenses makes sense.
Diagram: a self-hosted Grafana stack or an off-the-shelf tool

All monitoring goes to one place

When a company has several environments, its monitoring has most likely grown over time with nobody organizing it. Each system then has its own island: a developer looks at system A, an admin at system B, sometimes even for the same application. The worst case is when admins report an outage and developers answer that everything works on their side, because they do not see the error.

The first goal is to revisit monitoring and put it in order: collect information from all systems in one place, so that everyone looks at the same data and draws the same conclusions. Clients often say they have legacy systems and not everything runs on Kubernetes. Kubernetes is the easiest platform to monitor, but it is not the only option. The stack also covers virtual machines and applications on physical servers, and it brings monitoring of the cloud, on-premises and hybrid setups with more than one cloud into one view.

Monitoring also takes in client applications, databases, queues, cache and even DNS servers, because teams and systems affect each other. The cause of an outage can be an overloaded DNS server, a disk array or an identity server, so data has to come from all these places.

Prometheus probably runs at your company too, but you likely do not use it to the full. The number of systems it can collect metrics from runs into the thousands, and its community has prepared connectors for almost everything. When a team quickly sets up its own system, the result is generally an empty shell: it collects data from 10–20% of what could be collected and gives no picture of the whole. During an outage, nobody then knows what the root cause was.

Per-user licensing discourages broad access

Off-the-shelf systems such as Datadog or Dynatrace are usually licensed per user, which discourages opening them to the whole organization. They are relatively easy to deploy, but cost is where the problems start, and they are severe. When 100 or 200 people should use the tool, the seat count starts to hurt.

At that point, the speakers fairly often see organizations limit who has access. During an outage, support teams are then less sure what they can do, and someone who cannot see something sets up yet another monitoring system. The second common licensing model counts processors or monitored systems. The Grafana stack has no per-user costs, so it can reach much further across the organization.

The rollout starts with an inventory

The rollout runs in several phases and does not need much involvement from your team. The first phase is an inventory: who uses which tool and what actually runs in the organization. A big-bang change looks good on paper. But if you suddenly take their tools away from admins and from first- and second-line support and do not move them over properly, the organization ends up in chaos.

In the second step, Grafana becomes the single entry point to all metrics and monitoring data in the organization. It replaces the wiki page with a list of URLs that only some people know about. At this stage, forgotten systems come to light, including those the company pays for and does not use. For each one, a decision follows: remove it, replace it or do something else with it.

Most guides on the internet say to build the platform at this point and send data to it. First, however, comes Grafana Alloy: collectors that gather data from virtual machines, Kubernetes, the cloud and other sources. Alloy improves the quality of this data, checks what it looks like and buffers it. A very common situation in companies: developers set up Kafka to ship logs, although Kafka is not quite suited to that.

Once the data from system A looks similar to the data from system B, the monitoring platform goes up. Metrics in Prometheus come first: they show the state of the whole environment, its trends and its daily cycles. On these numbers you can then build alerts and forecasts of when something will stop working.

The base stack has four components: Alloy collects the data, Prometheus handles metrics and is the market standard, Loki collects logs and Tempo collects traces. Loki takes in terabytes of data, is time- and cost-efficient and needs practically no maintenance.

Kubernetes, virtual machines, physical servers and the cloud send data to Grafana Alloy collectors. Alloy improves the quality of the data, watches over it and buffers it. It then passes metrics to Prometheus, logs to Loki and traces to Tempo. These three systems form the base stack, and Grafana is the single entry point to all the data.
Diagram: the base stack from collectors to Grafana

Your team gives access to the cluster and CI/CD

Protopia sets up the stack. Your company provides access to a Kubernetes cluster (or Protopia sets up the cluster itself) and access to CI/CD, for example GitHub or Bitbucket. The whole stack is built from code to make it resilient. That is where your team’s involvement in this part ends.

From there, the work goes in phases. You point to the systems, PoCs follow, and Protopia sets up collectors, prepares load tests to check that the system will work, and connects everything to the main Grafana instance.

Every system handed over for use comes with a training session. People with a full backlog have no time to learn on their own, so it is easier to run a one-day, 4-hour training session and show what they can get out of the tool. This encourages employees to use the new stack instead of the old tools.

Your team gives access to the Kubernetes cluster, or Protopia sets up the cluster, and access to CI/CD. Protopia builds the stack from code. Then your team points to the systems. For each system, Protopia runs a PoC, sets up collectors, prepares load tests and connects the system to Grafana. The cycle ends with a training session and returns to the next system.
Diagram: what your team provides and what Protopia does

The stack is known and used by the CNCF and by organizations around the world. A new hire is expected to know these tools already, so you do not have to teach them the tools themselves, and the company builds on a standard it can extend.

Developers get instrumentation standards

Next come the standards the organization requires from application developers. Protopia interviews several groups of developers, learns how they work and where they start from, prepares recommended standards and, if needed, splits the path to them into phases.

Monitoring works without integration with the applications and without developer support, but more value comes only when the applications take part. For modern technologies such as Java, .NET, Python or JavaScript, the effort is half an hour to 3 hours per system: you add a few libraries and it works. Developers have probably tried already, and the libraries may be under-configured, so small configuration changes make a big difference.

For systems that you bought, no longer maintain or no longer develop, there is auto-instrumentation. Auto-instrumentation means starting an application through an agent, with no changes to its code. It does not give everything and it is not perfect, but it gives a lot: telemetry, tracing and metrics. It is very safe, and sometimes the only change is the target image the application runs in.

Team knowledge goes into alerts

Standards are also written for admins and DevOps teams. When developers, admins and DevOps look at the same deep data, you can ask when a system is about to stop working and what the main reason for its problems is.

Tribal knowledge shared in a meeting evaporates quickly, so the knowledge of these people goes into alerts and predictive metrics, and over time the system knows more. Much of this knowledge does not have to be gathered from people. Most popular systems, such as PostgreSQL or Redis, have ready-made or official alert sets, frequently promoted by the vendors themselves. Such sets often have dozens, if not hundreds, of rules. Some Kubernetes deployments load alert rules by default, and you only configure who gets notified and when. This knowledge is available because a great many organizations use the stack.

Drilldown in Grafana shows errors, traffic and response time

The most important screen in the demo is Drilldown on tracing data. It shows the number of errors, the number of requests and the distribution of their duration. RED stands for Request, Error, Duration. It is one of the three basic views of any system. The operator does not need to know where the data comes from and how it flows: they get a ready way to look at the system.

The errors tab separates errors your customers see from errors the system swallows. The latter do not surface, but they are worth dealing with. Next, you see how many errors each system produces, and root cause errors: the statistical source of errors for each service over the last half hour. The cause usually sits lower in the call chain than the system that failed.

The comparison sets the current behavior of the system against a previous period. Chasing every error probably does not pay off, because something always happens, so what counts is what has changed. Traces show the path of a request through the whole call chain: where something failed, which services it affected and whether the user saw it. The durations tab points to the main source of latency and automatically picks out the requests that really were slow.

eBPF profiling extends the base stack

Use of eBPF-based profiling is growing fast in very large organizations. Monitoring shows days and hours, logs and traces show minutes. Profiling goes below 20 milliseconds, for example when developers analyze performance or when something has failed. It works in real time, only on Linux, and best on Kubernetes. It is possible on virtual machines too, but it is not recommended, because it does not go as smoothly as it should.

eBPF profiling vs. other APM systems
Feature eBPF profiling Other APM systems
Performance overhead 1–3% usually 5–10%
Enabling nothing to enable often must be enabled
Sampling depth down to individual method calls often much worse
Application awareness the application does not know about it the application must know about it

A developer understands a detailed call stack, an operator does not. The Explain flame graph button sends the flame graph data to a language model, in the demo to OpenAI. The model explains what happens in a given piece of code, for example that it fails or uses a lot of memory, and suggests changes based on best practices. The answer is not perfect, but it is very useful: the operator does not need to know the code, and a developer who maintains a system they do not know well gets a concrete answer. The stack does not tie you to one model and works with Anthropic and with any system compatible with the OpenAI API.

Grafana connects to databases

With controlled access, databases can be connected to Grafana. A familiar outage scenario: the database throws an exception and someone has to find a DB admin to check what is going on. This feature has to be built and secured. When a trace shows an exception because a command did not execute, one button runs a read-only query on that database. The path from problem to answer then shrinks to seconds instead of hours spent looking for a person with permissions.

The rollout can go in parts

Monitoring changes can roll out in phases and in parts: start with cleanup or with better visibility, then add further elements as needed. The recommended option is the complete one: the rollout, real onboarding and a training package, so that people know how to get value out of what has been built.

Grafana also works with many existing systems, including Datadog, so the move to one central place can be smooth. Tools deeply rooted in the company, such as Zabbix or ELK, do not have to go, because Grafana and Prometheus work with them. The change is evolutionary: the goal is to win people over with value they can see, without forcing anything.

Speakers

  • Szymon Warda

    Szymon Warda

    Founder, Managing Partner, Technology Advisor.

    Szymon Warda co-founded Protopia and is a member of the Grafana Champions program. He co-hosts Patoarchitekci, where he has talked about distributed systems, observability and FinOps since 2019.

    All posts by this author
  • Mikołaj Szczerbicki

    Mikołaj Szczerbicki

    Head of Sales & Business Development.

    Mikołaj Szczerbicki is Head of Sales & Business Development at Protopia and co-hosts the Powered by Protopia podcast. He scopes and prices projects, so he asks about the cost, risk and timeline of AI, Azure and Kubernetes work.

    All posts by this author

FAQ

Does the Grafana stack lock the company into one vendor?

The stack is open source and is not a small product of a single company.

Why not send all the data straight to a central system?

If you send data of mixed quality straight to a central system, you get a low-quality result: garbage in, garbage out.

Is installing a tool enough for observability?

Setting up Prometheus and handing it over to people is not enough. They also need to know how to use it properly.