Metrics and alerts
Scrape strazad's metrics with a token, read what each one counts, and alert on the counters that mean a person must act.
- Who
- You, as the operator who runs strazad
- Where
- The strazad config file and your Prometheus server
- Profile
- Standalone and enterprise
On this page
strazad counts what it decides, what it records and what it delivers, and serves those counts at /metrics on its main listener in the Prometheus text format. This page sets a token on that endpoint, points Prometheus at it, says what each metric counts, and names the conditions worth an alert. /metrics labels its counts by MCP server, sink, route, status, effect and lane. No metric names a session or a user. The endpoint carries only Straza’s own series, with no Go runtime or process metrics beside them, and the approver listener on port 8443 does not serve it.
Set a metrics token
/metricsanswers anyone who reaches the main listener until you setserver.metricsToken. With a token set, every scrape must send it as a bearer token. Generate a long random value and put it in the config file on every strazad host:server: metricsToken: REPLACE_WITH_THE_GENERATED_VALUEopenssl rand -hex 32prints a value that fits. The environment variableSTRAZA_METRICS_TOKENsets the same thing, and on the Helm chart the key goes inconfigYaml. Restart strazad after the change, because it reads the token at boot.curl -s http://127.0.0.1:8420/metrics curl -s -H "Authorization: Bearer $STRAZA_METRICS_TOKEN" http://127.0.0.1:8420/metrics | head -3You should see
The first request answers 401 with
{"error":"metrics token required (server.metricsToken)"}. The second prints the start of the metrics, beginning with# HELP straza_approvals_unroutable_total.A token does not make the endpoint fit for a public ingress. Scrape it over the private network, as TLS and exposure explains, and read Hardening for the rest of the server’s exposure settings.
Point Prometheus at every replica
Each strazad process keeps its own counts, and they start at zero when it starts. Scrape every replica at its own address, so a counter never jumps between two processes behind a load balancer. This scrape job reads the token from a file on the Prometheus host:
scrape_configs: - job_name: straza scheme: https authorization: type: Bearer credentials_file: /etc/prometheus/straza-metrics-token static_configs: - targets: ["straza-1.internal:8420", "straza-2.internal:8420"]Drop
scheme: httpsfor a main listener without TLS on a trusted network. On Kubernetes, discover the pods.A fresh server shows its counters without labels at zero and its two storage gauges. A labelled series appears with the first event it counts, so
straza_pdp_decisions_totalis absent until strazad has decided something. After one request to/version, the request counter reads:straza_http_requests_total{route="GET /version",status="200"} 1Know what each metric counts
The route label is the route pattern strazad matched, such as
GET /version,POST /v1/decideor/mcp, and a dash when no route matched. Two gauges are measured when strazad starts and then once an hour.Metric What it counts straza_pdp_decisions_totalBy effect. Decisions strazad makes itself,allowordeny. They come fromPOST /v1/decide, which a client calls for a rule that needs a person’s approval or a server check, and from the MCP gateway. A hook that decides from its local snapshot does not count here.straza_pdp_decision_secondsBy effect. How long those decisions took, as a histogram from 20 microseconds to 100 milliseconds.straza_failclosed_totalBy lane. Internal failures answered as a deny, withhookfor/v1/decideandgatewayfor/mcp. A failed approval service, fingerprint or classifier counts, and so does a record the full audit queue refused underblock. Policy denies and approval outcomes never count.straza_gateway_throttled_totalBy app. Tool calls the gateway refused because one session called one MCP server faster than the server’sstraza.limits.rpsallows. A server whose manifest sets no rate has no limit. The drafting tools of the built-instrazaserver have their own limits: two calls a second for one session, and one submitted draft every ten seconds for one person.straza_gateway_catalog_oversize_totalCatalog builds with more tools than apps.catalog.warnSize, 100 by default. The warning in the log names the role.straza_approvals_unroutable_totalApproval requests denied at once because nobody could decide them: no approver role resolved to a person and no sponsor could answer. straza_audit_lost_totalAudit records strazad took in and could not confirm as written to the database. It starts at zero in each process. straza_audit_dropped_totalAudit records dropped because the in-memory audit queue was full, under governance.auditBackpressure: drop-with-counter.straza_audit_refused_totalRecords a client uploaded that POST /v1/audit/batchrefused, because the session they name is unknown, belongs to another user or device, or is past the cap of 1000 sessions in one upload. The client has deleted them.straza_sink_deadletter_totalBy sink. Deliveries a sink parked, because the receiver refused the event or a failure outlived its retries. Parked events are kept for a replay.straza_sink_duplicates_totalBy sink. Deliveries the receiver answered with 409, which means it already held the event.straza_sink_replayed_totalBy sink. Parked events that an operator’s replay delivered and took off the parked list.straza_http_requests_totalBy routeandstatus. Every answered request on both listeners./healthz,/readyzand/metricsare left out.straza_http_request_secondsBy route. How long requests took, with the same three routes left out. A streamed answer counts when the stream ends.straza_http_errors_totalBy routeandstatus. Server-side failures: answers of 500 and above, recovered panics, and internal JSON-RPC errors on/mcp, whose status readsrpcand the code. Answers in the 400 range are the caller’s mistake and do not count.straza_transcript_store_bytesBytes the recorded transcripts take. On Postgres it is the size of the transcript table, and on SQLite the size of the whole database file. straza_data_disk_free_bytesFree bytes on the filesystem that holds the data directory, or -1 when strazad cannot tell. Alert on what needs a person
Each row below is a condition that the metric’s own help text calls an alert, or a counter with one error log line per increment. The log line carries the cause, so an alert sends you to the log first.
Alert when What it means, and what you do straza_audit_lost_totalorstraza_audit_dropped_totalrisesA decision ran and its audit record never reached the chain. Read the strazad log for audit record lost, as Doctor and logs explains, and look at the database.straza_audit_refused_totalrisesA client’s uploaded records were refused, and the log line is their only trace. Read the warn line that starts audit batch:, which names the session and the reason.straza_failclosed_totalrisesAn internal failure turned a call into a deny. Find the error line fail-closed: internal failure answered as a denyand itscorrelation_id.straza_http_errors_totalrisesstrazad answered a request with a server-side failure. Each count has one error line in the log, with its correlation_idand route.straza_approvals_unroutable_totalrisesAn approval rule names approvers who resolve to nobody. Read the deny reason or the warn line, then give the approver role a holder or fix the sponsor. straza_sink_deadletter_totalrisesA sink receiver refuses events or has been down past its retries. Run strazactl sinks list, fix the receiver, thenstrazactl sinks replay <sink>, as Sinks and SIEM shows.straza_sink_duplicates_totalgrows steadilyThe receiver is not removing duplicates by event id. Check the receiver’s pipeline. One 409 after a redelivery is normal. straza_gateway_throttled_totalkeeps climbingAn agent calls one MCP server in a tight loop. Find the session in the gateway’s deny records, which name the rate. straza_transcript_store_bytespassesgovernance.transcriptBytesWatermarkRecording is filling the disk. strazad also logs transcript storage above watermarkevery hour. Narrow the recording rule, shortengovernance.captureRetention, grow the volume, or raise the watermark.straza_data_disk_free_bytesruns lowThe data directory’s disk is filling. Free or grow the disk. In the enterprise profile the database disk belongs to Postgres, so watch it there too. These rules turn four of the rows into Prometheus alerts. The windows and the disk threshold are examples, so set them to your own on-call habits.
groups: - name: straza rules: - alert: StrazaAuditRecordsLost expr: increase(straza_audit_lost_total[10m]) > 0 or increase(straza_audit_dropped_total[10m]) > 0 annotations: summary: strazad lost or dropped audit records. Read its log for "audit record lost". - alert: StrazaFailClosed expr: increase(straza_failclosed_total[10m]) > 0 annotations: summary: An internal failure was answered as a deny. Find the error line by its correlation_id. - alert: StrazaSinkParked expr: increase(straza_sink_deadletter_total[15m]) > 0 annotations: summary: The sink {{ $labels.sink }} parked events. Fix the receiver, then replay them. - alert: StrazaTranscriptsAboveWatermark expr: straza_transcript_store_bytes > 10737418240 annotations: summary: Recorded transcripts are above 10 GiB.
Settings that move these numbers
A few settings decide when a counter moves, so read them together with your alerts. The Configuration page lists every key.
| Setting | Default, and what it changes |
|---|---|
governance.auditBackpressure |
Default drop-with-counter on standalone, block on enterprise. Under drop-with-counter a database outage shows in the lost and dropped counters. Under block, decisions wait and then fail closed, which shows in straza_failclosed_total. |
governance.transcriptBytesWatermark |
Default 10 GiB. The size at which strazad starts logging the watermark warning. |
apps.catalog.warnSize |
Default 100. The catalog size that counts as oversize, and -1 turns the check off. |
governance.auditIngestBacklogLimit |
Default 50000. While this many records wait to be published, client uploads get 429 and retry. They show in straza_http_requests_total with the route POST /v1/audit/batch and the status 429. |
governance.auditIngestPerSessionRPS |
Default 5. Uploads per second one session may make at one strazad, refused with 429 above it. |
events.auditStreamMaxAge and events.auditStreamMaxBytes |
Default twice governance.captureRetention, so 1440h, and 2 GiB. How long and how much the audit stream keeps. A sink that stays down longer than that misses the records that aged out, so act on a parked sink before then. |
straza.limits.rps in a server’s manifest |
No default. The rate per session and server above which the gateway throttles. |
Patterns across many records, such as a burst of denies from one session, are the audit sentinel’s job. It records its verdicts on the audit chain rather than as metrics, and it is off until you set governance.sentinel.enabled, as The sentinel shows.
Walked on v1.1.0-117-g106081a8 on 2026-10-06. An Alpine 3.20 container with a standalone strazad, where the /metrics answer of a fresh boot, the answer after one request to /version, and both requests of the token step against a server started with STRAZA_METRICS_TOKEN were run with curl on the default address. The Prometheus scrape job and the alert rules were written from the metric names and the Prometheus configuration page and were not loaded into a Prometheus server.