Skip to content

Commit a7658dd

Browse files
committed
Document the client local metrics endpoint
1 parent f7433ce commit a7658dd

5 files changed

Lines changed: 144 additions & 3 deletions

File tree

‎src/components/NavigationDocs.jsx‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -867,6 +867,7 @@ export const docsNavigation = [
867867
{ title: 'gRPC Daemon Socket', href: '/client/grpc-socket' },
868868
{ title: 'HTTP/JSON Daemon Socket', href: '/client/json-socket' },
869869
{ title: 'Environment Variables', href: '/client/environment-variables' },
870+
{ title: 'Local Metrics Endpoint', href: '/client/local-metrics' },
870871
{ title: 'MDM Integration', href: '/client/mdm-integration' },
871872
{
872873
title: 'Settings',

‎src/pages/client/local-metrics.mdx‎

Lines changed: 117 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,117 @@
1+
import {Note} from "@/components/mdx";
2+
3+
export const description = 'Scrape NetBird client connection health with Prometheus through the local /metrics endpoint on the client daemon.'
4+
5+
# Local Metrics Endpoint
6+
7+
The NetBird client daemon can expose a Prometheus `/metrics` endpoint on the local machine. It reports whether the client is connected to Management and Signal, how many peers it knows and how many of those are connected, the latency to each directly connected peer, and how long connection establishment, sync processing, and logins take.
8+
9+
The endpoint is opt-in and off by default. When enabled, it binds to `127.0.0.1:9191` unless you set another address. It is independent of [client metrics push](/manage/client-metrics): nothing here is sent anywhere, the data is only served to whatever scrapes the endpoint.
10+
11+
<Note>
12+
Available since NetBird <strong>v0.78.0</strong>.
13+
</Note>
14+
15+
## Enabling the endpoint
16+
17+
```bash
18+
netbird up --enable-local-metrics
19+
```
20+
21+
To listen on a different port:
22+
23+
```bash
24+
netbird up --enable-local-metrics --local-metrics-address 127.0.0.1:9300
25+
```
26+
27+
To disable it again:
28+
29+
```bash
30+
netbird up --enable-local-metrics=false
31+
```
32+
33+
The setting is stored in the [profile](/client/profiles) config, so it survives restarts and applies per profile. Switching profiles starts, stops, or rebinds the endpoint to match the profile you switch to.
34+
35+
<Note>
36+
`netbird up` ignores configuration flags when the client is already connected. Run `netbird down` first, then `netbird up` with the flags.
37+
</Note>
38+
39+
Once the client is up, check the endpoint:
40+
41+
```bash
42+
curl http://127.0.0.1:9191/metrics
43+
```
44+
45+
## Exposing it beyond localhost
46+
47+
The endpoint has no authentication. Anything that can reach it can read your peer names, the latency to each of them, and your connectivity state. Keep it on a loopback address and let the scraper run on the same host.
48+
49+
If a scraper genuinely has to reach it from elsewhere, binding to a non-loopback address requires root on Linux and macOS, and administrator privileges on Windows. The daemon runs with those privileges itself, so an unprivileged local user must not be able to publish this to the network:
50+
51+
```bash
52+
sudo netbird down
53+
sudo netbird up --enable-local-metrics --local-metrics-address 0.0.0.0:9191
54+
```
55+
56+
The daemon logs a warning for every non-loopback bind. Put the endpoint behind a firewall rule, or a reverse proxy that adds authentication, if you do this.
57+
58+
## Scraping with Prometheus
59+
60+
```yaml
61+
scrape_configs:
62+
- job_name: netbird-client
63+
static_configs:
64+
- targets: ['127.0.0.1:9191']
65+
```
66+
67+
## Metrics reference
68+
69+
### Connection state
70+
71+
These are read from the daemon at scrape time and are always present while the daemon runs.
72+
73+
| Metric | Type | Labels | Description |
74+
| --- | --- | --- | --- |
75+
| `netbird_management_connected` | gauge | — | `1` when connected to the management service, `0` otherwise. |
76+
| `netbird_signal_connected` | gauge | — | `1` when connected to the signal service, `0` otherwise. |
77+
| `netbird_peers` | gauge | — | Number of peers known to this client, matching the total in `netbird status`. Includes peers that are currently offline. |
78+
| `netbird_peers_connected` | gauge | `connection_type` | Number of connected peers, split into `p2p` and `relay`. |
79+
| `netbird_peer_latency_seconds` | gauge | `peer` | Round-trip latency to a directly connected peer, labeled with its FQDN. Relayed connections carry no latency measurement and produce no series. |
80+
81+
### Connection establishment and management interactions
82+
83+
These are recorded as the client runs, so they appear only once the engine is up and the corresponding event has happened at least once. They reset when the daemon restarts.
84+
85+
| Metric | Type | Labels | Description |
86+
| --- | --- | --- | --- |
87+
| `netbird_peer_connection_stage_duration_seconds` | histogram | `stage`, `connection_type`, `attempt_type` | Duration of peer connection establishment stages. |
88+
| `netbird_sync_duration_seconds` | histogram | — | Duration of processing a sync message from the management service. |
89+
| `netbird_sync_phase_duration_seconds` | histogram | `phase` | Duration of an individual sync processing phase, for example `routes_apply`, `filtering`, or `added_peers`. |
90+
| `netbird_login_duration_seconds` | histogram | `success` | Duration of logins to the management service, split by whether the login succeeded. |
91+
92+
Label values for `netbird_peer_connection_stage_duration_seconds`:
93+
94+
| Label | Values |
95+
| --- | --- |
96+
| `stage` | `signaling_to_connection`, `connection_to_wg_handshake`, `total` |
97+
| `connection_type` | `ice_p2p`, `ice_turn`, `relay` |
98+
| `attempt_type` | `initial` for the first connection to a peer, `reconnection` for a later one |
99+
100+
<Note>
101+
Every label is a bounded enum except `peer` on `netbird_peer_latency_seconds`, which carries one series per directly connected peer. On a client with many direct connections this is the one metric whose cardinality grows with your network.
102+
</Note>
103+
104+
## Grafana dashboard
105+
106+
NetBird ships a [client dashboard](/selfhosted/observability/dashboards#client) that graphs everything above. Import [`client.json`](https://github.com/netbirdio/netbird/blob/main/infrastructure_files/observability/grafana/dashboards/client.json) into Grafana and point it at the Prometheus datasource that scrapes your clients.
107+
108+
## MDM
109+
110+
Both settings can be enforced through [MDM](/client/mdm-integration):
111+
112+
| Key | Type | Description |
113+
| --- | --- | --- |
114+
| `enableLocalMetrics` | boolean | Turn the local `/metrics` endpoint on or off. |
115+
| `localMetricsAddress` | string | Listen address for the endpoint, for example `127.0.0.1:9191`. |
116+
117+
An MDM-supplied address is applied as-is and is not subject to the privilege check, since the policy already comes from an administrator.

‎src/pages/client/mdm-integration.mdx‎

Lines changed: 8 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -60,7 +60,7 @@ stale residue on the device.
6060

6161
## Policy keys reference
6262

63-
The same 16 keys apply on every platform. Names are camelCase in the
63+
The same 20 keys apply on every platform. Names are camelCase in the
6464
managed-configuration payload; the Windows ADMX template renders the
6565
PascalCase variant in the Group Policy Editor — both are recognized.
6666

@@ -82,6 +82,8 @@ PascalCase variant in the Group Policy Editor — both are recognized.
8282
| `disableUpdateSettings` | boolean | Block every configuration change from UI or CLI on this device (read-only mode). |
8383
| `disableProfiles` | boolean | Hide the profile menu in the GUI and reject profile CRUD via CLI. |
8484
| `disableNetworks` | boolean | Hide the Networks / Exit Node menus in the GUI and reject the related RPCs. |
85+
| `enableLocalMetrics` | boolean | Expose the client's [local Prometheus `/metrics` endpoint](/client/local-metrics). |
86+
| `localMetricsAddress` | string | Listen address of the local `/metrics` endpoint (default `127.0.0.1:9191`). |
8587
| `splitTunnelMode` | string | `allow` or `disallow` — split-tunnel policy mode (Android only at the client level; harmless on desktop). |
8688
| `splitTunnelApps` | string | Comma-separated list of package names that the split-tunnel mode applies to (Android only). |
8789

@@ -100,6 +102,11 @@ PascalCase variant in the Group Policy Editor — both are recognized.
100102
See [Enforcing the Exit Node on Managed Devices](/use-cases/remote-access/exit-nodes#enforcing-the-exit-node-on-managed-devices)
101103
for the full recipe, including why the policy must be deployed
102104
before users touch the exit node switch.
105+
- `localMetricsAddress` is applied as pushed. Users are held to a
106+
loopback address unless they are root, because the endpoint is
107+
unauthenticated, but a policy already comes from an administrator and
108+
is not subject to that check. A non-loopback address publishes peer
109+
names and connectivity state to anything that can reach the port.
103110
- `splitTunnelMode` and `splitTunnelApps` are wired into Android's
104111
`VpnService.Builder.addAllowedApplication()` flow; on Windows and
105112
macOS the daemon parses the keys but ignores them. They are safe to

‎src/pages/manage/client-metrics.mdx‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -52,3 +52,10 @@ The `NB_METRICS_PUSH_ENABLED` environment variable on the client takes precedenc
5252

5353
You can additionally set `NB_METRICS_INTERVAL` to a duration value (e.g., `30m`, `1h`) to override how often metrics
5454
are pushed.
55+
56+
## Scraping metrics locally
57+
58+
Independently of metrics push, a client can expose the same connection stage, sync, and login metrics, plus live
59+
connectivity and per-peer latency gauges, on a local Prometheus endpoint that you scrape yourself. The data never
60+
leaves the machine unless your own scraper collects it. See
61+
[Local Metrics Endpoint](/client/local-metrics).

‎src/pages/selfhosted/observability/dashboards.mdx‎

Lines changed: 11 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,8 @@
1-
export const description = 'Ready-made Grafana dashboards for NetBird Management, Signal, Relay, and the Enterprise Commercial License management stack.'
1+
export const description = 'Ready-made Grafana dashboards for NetBird Management, Signal, Relay, the client, and the Enterprise Commercial License management stack.'
22

33
# Grafana dashboards
44

5-
NetBird ships ready-to-use Grafana dashboards for the Management, Signal, and Relay services, plus an extended Management dashboard for [Enterprise Commercial License](/selfhosted/enterprise/getting-started) deployments. They are maintained in the `netbirdio/netbird` repository under [`infrastructure_files/observability/grafana/dashboards`](https://github.com/netbirdio/netbird/tree/main/infrastructure_files/observability/grafana/dashboards) and import directly into Grafana.
5+
NetBird ships ready-to-use Grafana dashboards for the Management, Signal, and Relay services and for the client, plus an extended Management dashboard for [Enterprise Commercial License](/selfhosted/enterprise/getting-started) deployments. They are maintained in the `netbirdio/netbird` repository under [`infrastructure_files/observability/grafana/dashboards`](https://github.com/netbirdio/netbird/tree/main/infrastructure_files/observability/grafana/dashboards) and import directly into Grafana.
66

77
## Available dashboards
88

@@ -12,6 +12,7 @@ NetBird ships ready-to-use Grafana dashboards for the Management, Signal, and Re
1212
| Management (Enterprise) | [`management-enterprise.json`](https://github.com/netbirdio/netbird/blob/main/infrastructure_files/observability/grafana/dashboards/management-enterprise.json) |
1313
| Signal | [`signal.json`](https://github.com/netbirdio/netbird/blob/main/infrastructure_files/observability/grafana/dashboards/signal.json) |
1414
| Relay | [`relay.json`](https://github.com/netbirdio/netbird/blob/main/infrastructure_files/observability/grafana/dashboards/relay.json) |
15+
| Client | [`client.json`](https://github.com/netbirdio/netbird/blob/main/infrastructure_files/observability/grafana/dashboards/client.json) |
1516

1617
### Management
1718

@@ -29,6 +30,10 @@ Covers active peers, peer connection durations, message forwarding throughput an
2930

3031
Covers connected peers (total / active / idle), peer authentication latency, peer store latency, and inbound/outbound relay traffic bandwidth.
3132

33+
### Client
34+
35+
Covers management and signal connectivity, known and connected peers split by connection type, per-peer latency, peer connection establishment stages, and sync and login durations. Unlike the service dashboards it reads from the NetBird client itself, which exposes these metrics through an opt-in local endpoint — see [Local Metrics Endpoint](/client/local-metrics) for enabling it and for the metric reference.
36+
3237
## Importing a dashboard
3338

3439
1. In Grafana, go to **Dashboards → New → Import**.
@@ -49,10 +54,14 @@ The Management, Signal, and Relay dashboards expose these template variables:
4954

5055
Your deployment may use only a subset of these variables; unused ones can be left at the default `All`. The Enterprise dashboard uses a different variable set — see [its variables table](/selfhosted/enterprise/grafana-dashboard#dashboard-variables).
5156

57+
The Client dashboard uses `datasource`, `job`, and `instance`, where `instance` selects the client to look at.
58+
5259
<Note>
5360
The Management dashboard expects HTTP request metrics to carry an `exported_endpoint` label rather than `endpoint`. If your Prometheus relabeling drops or renames this label, edit the dashboard panel queries accordingly.
5461
</Note>
5562

5663
## Scraping for the dashboards
5764

5865
The dashboards assume a Prometheus scrape configuration that reaches the `/metrics` endpoint of each NetBird service. See [Service endpoints](/selfhosted/observability#service-endpoints) for defaults and per-service overrides.
66+
67+
The Client dashboard scrapes the clients themselves rather than a server component. Its endpoint is off by default and binds to localhost when enabled, so it needs a scraper on each client host — see [Local Metrics Endpoint](/client/local-metrics).

0 commit comments

Comments
 (0)