User guideWayseer 0.28.3Contents
Recipes
Many sources already export to Prometheus, and the Prometheus module can show them without a module of their own. Each recipe here is a config for one exporter: its targets become entities of the right kind, and the queries under series and flows become their metrics and traffic. Paste the parts you need into your Prometheus instance.
Every recipe assumes each exporter is a scrape target with a job label, as below, and names the job in kinds. If your job is named differently, change it in kinds and nowhere else: the queries keep job and instance as they are. A target's entity is named job/instance, such as mysql/db-1:9104, and runs on the host its instance names; a job whose kind is host, such as Netdata's, is that host itself.
The metric names are prefixed by the exporter, so two recipes can share one instance. The queries are yours: the language model never writes or sees them.
MySQL and MariaDB
For mysqld_exporter, scraped as the job mysql:
modules:
- kind: prometheus
name: prom
options:
url: http://localhost:9090
kinds: {mysql: database}
series:
- query: sum by (job, instance) (rate(mysql_global_status_queries[5m]))
metric: mysql.queries
unit: per_second
description: statements the server ran
entity: {labels: [job, instance], kind: database}
- query: max by (job, instance) (mysql_global_status_threads_connected)
metric: mysql.connections
unit: count
description: clients connected
entity: {labels: [job, instance], kind: database}
- query: max by (job, instance) (mysql_global_status_innodb_buffer_pool_bytes_data)
metric: mysql.buffer_pool
unit: bytes
description: data held in the InnoDB buffer pool
entity: {labels: [job, instance], kind: database}
- query: max by (job, instance) (mysql_slave_status_seconds_behind_master or mysql_slave_status_seconds_behind_source)
metric: mysql.replication_lag
unit: seconds
description: how far the replica is behind its source
entity: {labels: [job, instance], kind: database}
Each server is a database, with its statements, connections and buffer pool in Grid, Detail and Timeline. Replicas also have their lag; a source has none. MySQL 8.0.22 and later call the lag seconds_behind_source, and the query reads either name.
It can't show the schemas and tables inside each server, or replication as traffic between them: the exporter gives a replica's source by name but no rate of what it reads. To see a server's tables, use the SQL module beside this one.
Redis and Dragonfly
For redis_exporter, scraped as the job redis. Dragonfly answers the same INFO fields, so the recipe works for it through the same exporter; Dragonfly's own /metrics uses other names.
modules:
- kind: prometheus
name: prom
options:
url: http://localhost:9090
kinds: {redis: database}
flows:
- query: sum by (master_host, instance) (rate(redis_slave_repl_offset[1m]))
from: {label: master_host, kind: host}
to: {label: instance, kind: host}
unit: bytes
series:
- query: max by (job, instance) (redis_connected_clients)
metric: redis.clients
unit: count
description: clients connected
entity: {labels: [job, instance], kind: database}
- query: sum by (job, instance) (rate(redis_commands_processed_total[5m]))
metric: redis.commands
unit: per_second
description: commands the server ran
entity: {labels: [job, instance], kind: database}
- query: max by (job, instance) (redis_memory_used_bytes)
metric: redis.memory
unit: bytes
description: memory the server holds
entity: {labels: [job, instance], kind: database}
Each server is a database with its clients, commands and memory. A replica reports its primary in master_host, so the flow is replication between their hosts, in bytes per second, in Flow. It joins hosts the module found from targets, so it needs the replica's master_host to be a name the primary's host goes by: a replica that names its primary by IP adds no flow, and the module's health note counts it.
It can't show keys or keyspaces as entities, or a cluster's slots. Lag shows only as flow that stops: a replica that falls behind reads less.
MongoDB
For Percona's mongodb_exporter, scraped as the job mongodb. Replication lag needs the exporter's --compatible-mode; without it, that query is empty and the rest still show.
modules:
- kind: prometheus
name: prom
options:
url: http://localhost:9090
kinds: {mongodb: database}
series:
- query: max by (job, instance) (mongodb_ss_connections{conn_type="current"})
metric: mongodb.connections
unit: count
description: clients connected
entity: {labels: [job, instance], kind: database}
- query: sum by (job, instance) (rate(mongodb_ss_opcounters[5m]))
metric: mongodb.operations
unit: per_second
description: inserts, queries, updates, deletes and commands
entity: {labels: [job, instance], kind: database}
- query: max by (job, instance) (mongodb_ss_mem_resident) * 1048576
metric: mongodb.memory
unit: bytes
description: resident memory of the server
entity: {labels: [job, instance], kind: database}
- query: max by (name) (mongodb_mongod_replset_member_replication_lag)
metric: mongodb.replication_lag
unit: seconds
description: how far the member is behind the primary
entity: {labels: [name], kind: host}
Each server is a database with its connections, operations and memory. The exporter reports lag by member, named as the replica set knows it, such as mongo-2:27017, so lag is a metric of each secondary's host, matched by hostname. The primary has none.
It can't show databases and collections as entities, or replication as traffic: the exporter gives no rate between members.
Kafka
For kafka_exporter, scraped as the job kafka. One exporter speaks for the whole cluster, so its target is a cluster.
modules:
- kind: prometheus
name: prom
options:
url: http://localhost:9090
kinds: {kafka: cluster}
flows:
- query: sum by (topic, consumergroup) (rate(kafka_consumergroup_current_offset[5m]))
from: {label: topic, kind: queue, make: true}
to: {label: consumergroup, kind: service, make: true}
unit: messages
series:
- query: sum by (topic) (rate(kafka_topic_partition_current_offset[5m]))
metric: kafka.messages
unit: per_second
description: messages written to the topic
entity: {labels: [topic], kind: queue}
- query: max by (topic) (kafka_topic_partitions)
metric: kafka.partitions
unit: count
description: partitions of the topic
entity: {labels: [topic], kind: queue}
- query: sum by (consumergroup) (kafka_consumergroup_lag_sum)
metric: kafka.lag
unit: count
description: messages the group has still to read
entity: {labels: [consumergroup], kind: service}
- query: max by (consumergroup) (kafka_consumergroup_members)
metric: kafka.members
unit: count
description: consumers in the group
entity: {labels: [consumergroup], kind: service}
The exporter names a topic and a consumer group in one series, so the flow makes both: each topic is a queue, each consumer group a service, and the messages a group reads from a topic per second are a flow in Flow and Topology. A group named like a service another module found, such as a Kubernetes deployment, joins that service instead of making one. Topics carry their write rate and partitions, and groups their lag and members.
Only topics some group reads are made, so a topic nobody reads has no entity and its series show nowhere. Brokers, partitions and which broker leads each partition aren't shown: the exporter gives partitions as numbers on a topic, not as things. A Kafka module would be needed for them, and for topics no group reads.
Netdata
For Netdata's Prometheus export, with each agent scraped as the job netdata at /api/v1/allmetrics?format=prometheus. Each agent is a host.
modules:
- kind: prometheus
name: prom
options:
url: http://localhost:9090
kinds: {netdata: host}
series:
- query: sum by (instance) (netdata_system_cpu_percentage_average{dimension!="idle"})
metric: netdata.cpu
unit: percent
description: CPU in use
entity: {labels: [instance], kind: host}
- query: max by (instance) (netdata_system_ram_MiB_average{dimension="used"}) * 1048576
metric: netdata.memory
unit: bytes
description: memory in use
entity: {labels: [instance], kind: host}
- query: sum by (instance) (netdata_system_net_kilobits_persec_average{dimension="received"}) * 1000
metric: netdata.received
unit: bits_per_second
description: network traffic in
entity: {labels: [instance], kind: host}
- query: -sum by (instance) (netdata_system_net_kilobits_persec_average{dimension="sent"}) * 1000
metric: netdata.sent
unit: bits_per_second
description: network traffic out
entity: {labels: [instance], kind: host}
Each host has its CPU, memory and network traffic. Netdata keeps traffic sent as a negative number, so the last query turns it round. The names assume Netdata's defaults: averages, and dimension names rather than IDs.
If one Netdata parent streams many children and Prometheus scrapes only the parent with format=prometheus_all_hosts, every child's series carries its hostname in instance. They show on hosts other modules found by that name, but the parent's one target makes no host for each child.
It can't show what Netdata knows below the host, such as disks, network interfaces, containers and applications, as entities: they are only labels on its series, and a series here is a metric of one entity.
When a source earns its own module
A recipe adds nothing to the app's size and works with any Prometheus. A source earns a module of its own only when it has structure Prometheus can't carry, such as Kafka's brokers and partitions, a MongoDB replica set or a database's tables, and its client is light enough to build into the app. Until then, each recipe above says what it leaves out.