Monitoring & CI/CD

What is Prometheus

What it is

A metrics collection system โ€” it gathers numbers about how servers and services are doing.

How we use it

It feeds Grafana with data and can send an alert itself when something goes wrong โ€” before the client notices.

Where it helps a business

  • You want to see in advance that disk space or memory is running out, not learn it from a crash.
  • There are many services, and a system rather than a person should watch each one.
  • You need history: how many errors and how much load there was over the week.

How we use it

Prometheus paired with Grafana is part of our infrastructure stack for projects with several servers. In our single-server cases simpler tools are enough: container health checks, a watchdog with automatic restarts and Telegram messages. For example, the Tech Poly VPN monitoring checks the node's port every 5 minutes and alerts the administrator only after several failures in a row.

Common problems

  • An alarm for every trifle. Without thresholds and repeated checks, people stop reading the alerts.
  • Metrics without reaction. Collecting numbers is pointless if nobody gets a message about a problem.

When you do not need it

On one server with a couple of services, a metrics system is needless complexity. Start with health checks and alerts.

โ† All terms

Need a website, a bot or automation?

Terms explained โ€” now let's get to work: tell us about the task and we'll turn it into a clear work plan.