Skip to content

Monitoring & Observability ​

What the platform sees about itself — metrics, dashboards, alerts. Where to look when something's wrong, and what gets paged in the middle of the night.

Why It Matters ​

A sovereign platform that can't see itself can't be trusted. The monitoring stack answers three questions every good operator asks at least weekly:

  1. Is the supply invariant still zero? (If not, nothing else matters.)
  2. Is the network hitting its SLOs — latency, error rate, transfer throughput?
  3. What's the shape of the next failure? — trends, saturation, early warnings before anything breaks.

This chapter covers how the stack produces that visibility, where to look for what, and what alerts exist so you know an incident has been recognized.

The Stack ​

Prometheus scrapes. VictoriaMetrics keeps the long history. vmalert evaluates the rules. Alertmanager routes what fires. Grafana draws it. cAdvisor and node-exporter supply the container and host layer. Deliberately boring, industry-standard, auditable.

The one thing worth memorising: the scraper and the evaluator are different processes. Prometheus scrapes targets and remote-writes every sample to VictoriaMetrics; vmalert reads VictoriaMetrics and decides which alerts are firing. That means ALERTS{...} queried through Prometheus is an empty result even during a live incident — alert state lives where the evaluator lives. Ask vmalert (/api/v1/rules, /api/v1/alerts), not the scraper.

Illustrative of the reference multi-frame deployment; a single-registry install scrapes one registry, TEG, and event-store — the same metric families, fewer targets.

Continuous development

Dashboards and alert rules ship as a continuously evolving set — the dashboards and alert rules below are the current state, not a frozen surface. New panels, new rules, and protocol_* business metrics land regularly as the platform itself evolves. If something looks light, it's because the next refresh is queued.

Every service exposes /metrics in Prometheus format. Prometheus scrapes every 15 seconds and keeps 30 days locally; every sample is also remote-written to VictoriaMetrics, which keeps 60 days. vmalert evaluates the alert and recording rules against VictoriaMetrics and notifies Alertmanager, which routes firing alerts by severity to email and any other notifier you wire up. Grafana reads from Prometheus for dashboards and from the registry's own APIs for app-level views.

The retention suffix is load-bearing

VictoriaMetrics' -retentionPeriod takes a unit. 60d is sixty days; a bare 60 is sixty months. Write the d.

What's Measured ​

Three layers, with honest scope of what's wired today.

Infrastructure ​

  • cadvisor — container CPU, memory, network I/O per service.
  • node-exporter — host-level CPU, memory, disk, network.
  • postgres-exporter — every Postgres instance scraped (the registry DB, TEG DB, and event-store DB; more in a multi-frame deployment). Connection counts, query latency, replication lag, cache hit ratio (the pg_* family).

Not currently wired: Redis exporter, Redpanda/Kafka exporter — both are in development. Redis health is observed today via container-level cAdvisor stats and via the alerts that fire when Redis-backed features fail (e.g. leader-election fallout in ServiceDown correlations); Kafka/Redpanda similarly.

Application (per service /metrics) ​

  • HTTP middleware (every service) — http_requests_total, http_request_duration_seconds, http_request_size_bytes, http_response_size_bytes. Standard Prometheus FastAPI middleware. Used by the HighErrorRate and HighLatencyP95 alerts.
  • Registry IRONHAND — ironhand_enrolled_agents, ironhand_revoked_agents, ironhand_spire_agent_entries, ironhand_orphan_entries, ironhand_missing_entries. mTLS enrollment health.
  • EventStore — eventstore_events_total, eventstore_events_by_type, eventstore_unique_agents, eventstore_ledger_age_seconds, eventstore_ledger_freshness_seconds, eventstore_events_ingested_total, eventstore_kafka_consumer_batches_total, eventstore_kafka_batch_size, eventstore_info. Drives the SupplyLedgerStale and EventStoreStaleLedger alerts.
  • OPA shadow-mode (gated by OPA_ENABLED, off by default) — opa_decisions_total, opa_decision_latency_seconds. See chapter 08.

The TEG layer emits only HTTP middleware metrics today — no teg_transfers_total, no teg_fees_collected, no staking gauges. Business-level rates are derived from eventstore_events_by_type{event_type="TokensTransferred"} and similar event-store queries.

Business Invariants ​

Some of this is wired and some is not, and the split matters when you are deciding what to alert on.

Exported as Prometheus series today: federation peer health (federation_peer_state, federation_peers_reachable_ratio, federation_peer_last_sync_age_seconds, federation_peers_never_synced), background-worker liveness (the worker_* family — see Worker vitals), agent-card pull-sync cycles, reactor dispatch, IRONHAND enrolment, OPA decisions, jurisdiction and regulatory-authorization decisions, the fiat provider funnel, mail relay, bundles, and the transaction stream.

Not exported as series: supply delta, A2A payment-token states, dispute phases, veToken turnout. The data exists in the EventStore and the registry's own database; it is simply not projected into gauges yet. Read those through their own surfaces:

  • Supply audit delta — GET /api/v1/projections/supply-audit (EventStore), or eventstore_ledger_freshness_seconds as a proxy (the SupplyLedgerStale alert uses this).
  • Federation peer health — up{job="peer-registry"} == 0 and the FederationPeerDown alert.
  • Dispute / governance state — query the registry's own DB or the /api/v1/governance/proposals?include_closed=true endpoint.

Promoting the remaining four to first-class Prometheus series is on the roadmap.

The Canary Reactor ​

Most synthetic monitors hit a health endpoint and call it a day. The canary reactor moves real production AVT between funded canary agents on real production infrastructure, watches the resulting events propagate through five reactors and the EventStore, and tells you exactly which phase silently broke when something silently breaks. It runs continuously, uses the same code paths as user traffic, has no test harness.

Live at /ui#/admin/canaries. Backed by canary_runner (scheduler registry only, leader-elected, fires every ~60s) and canary_judge (scheduler registry only, leader-elected, sweeps every 5s). Five validator reactors subscribe to EventStore events and stamp per-event ISO timestamps into a canary:<test_id> Redis hash; the judge reads the hash, decides outcome, persists to canary_test_results.

Three backends, three event signatures ​

Each path runs one of three transfer backends, and each backend emits a different set of events. The canary's per-backend "done" definition follows what the backend actually emits, not what would be tidy.

BackendEndpointEvents emittedCanary required-stamp
intra/teg/transferTokensTransferred + TransactionFeeCollectedtokens_xfer_seen_iso
async/teg/cross-registry-transfer?backend=asyncSender frame: TokensTransferred + CrossFrameTransferInitiated + TransactionFeeCollected. Receiver frame (after credit): CrossFrameTransferSettled. SF-4 broadcasts both Initiated and Settled across frames.cf_initiated_seen_iso AND (cf_settled_local_seen_iso OR cf_settled_remote_seen_iso)
2pc/teg/cross-registry-transfer?backend=2pcAlways: TokensTransferred + CrossRegistryTransferCompleted + TransactionFeeCollected. Cross-frame 2pc additionally: one extra CrossFrameTransferSettled (Phase 7.5 visibility helper — TokensTransferred is in SF-4's foreign_events but not in CROSS_FRAME_BROADCAST_TYPES, so without that extra emit the canary on the other frame would never see the transfer land).tokens_xfer_seen_iso OR (cf_settled_local_seen_iso OR cf_settled_remote_seen_iso)

The canonical 2pc completion event is CrossRegistryTransferCompleted. There is no canary validator subscribed to it — the canary detects 2pc via the always-emitted TokensTransferred (memo wrapped as cross_registry:canary-<test_id>; the reactor strips the prefix before correlating). Keeps the validator subscription set to five event types and the per-backend logic symmetric.

The five validator reactors ​

EventReactor stampsNotes
TokensTransferredtokens_xfer_seen_isoAll 3 backends emit. Required for intra; one of two acceptable for 2pc; informational for cross-frame async.
TransactionFeeCollectedfee_collected_seen_isoAll 3 backends emit. The fee event has no memo field, so memo-fallback alone won't catch it; only the canary_xfer:<transfer_id> reverse-index does. Treated as opportunistic latency signal, never required.
CrossFrameTransferInitiatedcf_initiated_seen_isoasync cross-frame only. 2pc does not emit this. Required for async.
CrossFrameTransferSettledcf_settled_local_seen_iso (if source_frame == local_frame) or cf_settled_remote_seen_isoasync receiver frame after credit; cross-frame 2pc Phase 7.5 helper. Required for async; one of two acceptable for cross-frame 2pc.
CrossFrameTransferRefundedcf_refunded_seen_iso + failure_signal=refundedasync saga sweeper after 5-min timeout. Presence is an explicit failure — judge marks failed/refunded immediately.

The (i) and the live log ​

The dashboard hero has a small italic i next to the title — toggles a six-block inline panel documenting path matrix, fire cadence, validation pipeline, failure semantics, observability surface, and source files. Click-to-expand, no page navigation.

The bottom of the view is a live event log feed: every distinct outcome the dashboard has seen across refresh cycles, capped at 200 entries, color-coded, filterable by outcome, collapsible. Auto-tail toggle prepends new arrivals; collapsed-and-filling sprouts a +N badge so you don't miss movement.

Latency: don't let the observer dominate the measurement ​

total_latency_ms went through three definitions before settling. The first version was wallclock(judge_completed - fired) — every test averaged ~17s because the judge ran every 30s, so the metric was dominated by sweep cadence rather than transfer time. The fix was to compute latency from the latest *_seen_iso stamp the validator wrote, not from when the judge happened to look. Average dropped from 17s → ~935ms. General lesson: never let an observer interval dominate a measurement. If a synthetic monitor lies about latency, alerts are gated on a number that doesn't mean what it says.

What the canary is not ​

The canary is not the supply auditor. The auditor verifies tokens_issued − tokens_destroyed + transit_net == tokens_circulating across the EventStore on its own cycle (60s) with zero tolerance and its own dashboard at auditor.theprotocol.cloud. The canary watches whether transfers complete; the auditor watches whether they conserve mass. Different invariants, different failure modes, different runbooks. Don't conflate.

TIP

Post-restart of the scheduler registry, expect ~5-10 minutes of canary noise while the five validator reactors re-acquire leader locks, re-establish EventStore WebSocket subscriptions, and let the polling-backstop close the watermark gap. Don't page on canary failures within the first 10 minutes after a scheduler-registry recreate. After the warmup, the next cycle judges 30/30 passed.

🔗 Grafana dashboard: 28 — Canary Reactor — 8 sections covering pass-rate aggregates, per-path scoreboard, validator pipeline, upstream EventStore + TEG-DB pressure, and a 30d trend via VictoriaMetrics. Annotations mark scheduler-registry restarts so you don't misread post-restart warmup noise as a real regression. · Long-form blog post: The Canary Reactor: How TheProtocol Watches Itself, One Real Transfer at a Time

A Frame Watches Itself ​

The stack above is the flagship host's. Since the frame kit (2026-09-14) every frame the runner builds carries its own copy, inside the frame: a Prometheus that scrapes the frame's registry, TEG, event store, auditor and OPA under the frame's own job names, the shipped frame rules evaluated there, an Alertmanager beside it, and a Grafana with the datasource and the frame-overview dashboard provisioned (loopback on the host, admin password minted per frame). The frame's registry reads its history from that Prometheus (PROMETHEUS_URL), so its status page shows its own past without the host. A solo auditor in the same frame re-verifies the ledger's supply invariant every cycle and answers /health and /metrics for the scrape, so a dead auditor reads as one, never as a quiet one.

The standalone release ships the same shape as a second compose file: docker compose -f docker-compose.standalone.yml -f docker-compose.monitoring.yml brings up Prometheus (monitoring/prometheus.yml, the frame rules under monitoring/rules/), Alertmanager (monitoring/alertmanager.yml) and Grafana (monitoring/grafana/provisioning/) on loopback ports, scraping the standalone services by name; OPA and the solo auditor are part of the standalone compose itself.

Where Things Live ​

Grafana          https://grafana.example.com
Prometheus       https://prometheus.example.com          # scrape + 30d hot storage
Alertmanager     https://alerts.example.com              # routing
VictoriaMetrics  127.0.0.1:8428   (loopback only)        # 60d cold storage
vmalert          127.0.0.1:8880   (loopback only)        # alert + recording rule evaluation

VictoriaMetrics and vmalert bind to loopback deliberately — reach them over an SSH tunnel. When you want to know whether an alert is firing, ask vmalert, not Prometheus:

bash
curl -s 127.0.0.1:8880/api/v1/alerts | jq '.data.alerts[] | {name: .name, state: .state}'
curl -s 127.0.0.1:8880/api/v1/rules  | jq '.data.groups[].name'

Curated dashboards (shipped at monitoring/grafana/provisioning/dashboards/json/, editable):

#Real titlePurpose
0101 — System Overviewtop-level health — start here
0202 — Registry & Service Healthper-service uptime, request rate, p95 latency
0303 — TEG / Token Economicstransfers, staking ops, fee collection (mostly via EventStore-derived series today)
0404 — EventStore: Dual-Frame Unifiedwrite rate, ledger freshness, supply audit cross-frame
0505 — PostgreSQL Performanceconnections, locks, slow queries, replication lag
07Container Resources (cAdvisor)CPU, memory, network per container
08Frame B Sovereignper-frame health for an additional frame's registry / TEG / event-store — multi-frame deployments only

TIP

For everyday operation, start at System Overview. If everything is green, close the tab. If anything isn't, the panel title tells you which dashboard to drill into.

Alerts ​

vmalert evaluates the rule files; Alertmanager routes what fires. The rule set is deliberately conservative — every firing alert represents something an operator should look at within an hour, and the one grade above that is meant to interrupt a night.

A severity with no route is a promise nothing keeps ​

Alertmanager routes evaluate in order, and each shipped route ends with continue: false — the first match wins and nothing below it is consulted. A rule that emits a severity no route matches does not error and does not get dropped; it falls through to the default receiver, which is by construction the lowest-urgency mailbox you own. It looks exactly like an alert that never fired.

Two rules follow from that, and both are cheap:

  • The highest-priority route goes first. page above critical above warning, because routing stops at the first match.
  • Diff the emitted severities against the matched ones whenever you add a rule:
bash
# every severity the rules can emit
grep -o 'severity: *[a-z]*' monitoring/prometheus_alerts.yml | sort | uniq -c
# every severity the router matches
grep -A2 'matchers:' monitoring/alertmanager/alertmanager.yml | grep 'severity'

If the left column has a value the right column lacks, that alert class is already silent.

Set --web.external-url

Without it, every "view in Alertmanager" link in a notification points at the container hostname — unreachable from the phone you are reading it on.

Representative alerts (selected from the live prometheus_alerts.yml):

AlertSeverityFires when
IronhandCanaryDegradedPAGEthe mTLS canary corridor's pass rate collapses — the one alert wired to interrupt a night
SupplyLedgerStaleCRITICALEventStore ledger freshness > 30 minutes (no new events) for 5m — supply audit may be unreliable
ServiceDownCRITICALup{tier="application"} == 0 for 2m — registry / TEG / event-store unreachable
PostgreSQLDownCRITICALpostgres-exporter up == 0 for 2m
EventStoreStaleLedgerCRITICALEventStore-side staleness signal (separate from ledger freshness)
FrameBRegistryDown / FrameBEventStoreDown / FrameBLedgerStale / FrameBHighErrorRateCRITICAL/WARNINGper-additional-frame equivalents — registry/eventstore down, ledger stale, or 5xx rate elevated (multi-frame deployments)
HighErrorRateWARNING5xx rate > 5% over 5 minutes
HighLatencyP95WARNINGp95 route latency > 2s for 5 minutes
FederationPeerDownWARNINGup{job="peer-registry"} == 0 for 5m — cross-frame sync compromised
DBConnectionsHigh / DBConnectionsCriticalWARNINGPostgres connection pool utilization elevated
HostHighCPU / HostHighMemory / HostDiskSpaceLow / DiskSpaceCriticalWARNINGInfrastructure pressure
ContainerRestarting / ContainerMemoryHighWARNINGcAdvisor container instability
EventStoreDeadlockBurstWARNINGDB deadlock spike
PostgreSQLHighConnections / PostgreSQLLowCacheHitRatioWARNINGPostgres performance degradation
SwapThrashSustained / MemoryStallCriticalWARNINGthe host is swapping in both directions for a sustained window — one-sided swap-in is recovery, bidirectional churn is thrash
WorkerFailingPersistently / WorkerNeverRan / WorkerVitalsBlindWARNINGa background worker's own vitals say it is failing, has never completed a cycle, or cannot report at all (see Worker vitals, below)
PrometheusRemoteWriteEnqueueRetries / SpireAgentRecentlyRestarted / SpireServerRecentlyRestartedINFOcontext for a postmortem, not a call to action

The supply-side critical alerts are the only ones that mean "stop everything and investigate now." Everything else is "look in the next hour."

A failure counter without its denominator is not a finding

notifications_failed_total{transport="email"} reading 33 is meaningless until you pull notifications_total beside it — 33 out of 4,515 is a 0.7% failure rate and entirely normal for email. Before declaring a transport dead, fetch the paired total and one live log line. This is the single most common way a healthy system gets diagnosed as broken.

The Public Status Page ​

Every registry serves its own status page at /status, and the data behind it at GET /api/v1/public/status — no authentication, because the point of a status page is that it works for someone who cannot log in.

It answers three questions, and each comes from a different place:

PanelSourceWhat it means
Componentsthe registry's own /api/v1/stack-health probes, plus the latest Prometheus up sample for the peer frames and operators it is configured to listlive reachability right now
Availability historyavg_over_time(up{job="…"}[1d]) per UTC daywhat actually happened, per day
Incidentsa curated JSON list, written only through the admin railwhat a human says happened
bash
curl -s https://your-api.theprotocol.cloud/api/v1/public/status | jq '{overall, components: [.components[] | {name, state, availability, days_known}]}'

Availability history needs a scrape job, and honesty when there isn't one. A component that no Prometheus job scrapes shows no history at all rather than a green bar it did not earn. If a bar is missing, that is a monitoring gap and the fix is to add the scrape target — never to synthesise a bar. Scrape containers on the host port and reload with POST /-/reload (lifecycle API on), then confirm up is 1 for the new job before believing the page.

Prometheus keeps 30 days, so the page would be permanently capped at a month. It isn't: completed days are appended to a rollup file on the registry's durable data volume, so the visible horizon grows past retention from the day the feature was switched on. Both the snapshot (30 s) and the history (5 min) are cached in-process, so a status page that gets popular costs the probes nothing.

Groups, and what "overall" means ​

Every component belongs to one of three groups, and the page draws each as its own section:

GroupWhat it holds
corethis registry's own stack: its registry, TEG, event store and the services beside them
framea peer frame's registry, TEG and event store
operatoroperator registries, this frame's own and a peer frame's

The group comes from the peer's type in this registry's federation table, never from its name: a peer frame is a frame whatever it is called, and every other peer is an operator. Each registry appears once. Two entries that point at the same public host are one registry, and a typed entry wins over an untyped copy of it; history is one row per scrape job, so a component that is both probed live and listed from configuration is never drawn twice.

overall reads outage when a core component is down, degraded when any other component is down, and operational otherwise. A peer's trouble can degrade your page; only your own stack can put it in outage.

Incidents ​

Incidents are deliberately not derived from metrics. Metrics produce measured dips — any UTC day a component averaged below 99.9% — and those are listed separately, carry no narrative, and say plainly that they are machine-derived. A real incident is written by a person:

GET  /api/v1/admin/status/incidents      # admin_platform — read the curated list
PUT  /api/v1/admin/status/incidents      # admin_platform — whole-list replace, validated

The shape is fixed and validated on write: an id (lowercase slug), title, impact (minor · major · critical · maintenance), a status (investigating · identified · monitoring · resolved · scheduled · completed), the affected components, started_at, an optional resolved_at that may not precede the start, and an append-only list of timestamped updates. Duplicate ids are refused.

A 99.9% threshold flags every deploy

A rolling restart takes a component below 99.9% for that day, so on a busy release week the dip list is mostly your own deploys. Read dips as "something interrupted service", not "something broke", and write an incident when the interruption deserved a sentence.

A Frame's Own Observability ​

A frame stood up from the repository's frame template does not depend on anyone else's dashboards. The sidecar project in sidecars/ is what it is wired into, and it carries everything a host of your own needs:

  • Scrape targets: one file per frame for Prometheus file-based discovery. Each target carries its own job label (the same names the status page keys its history on) and frame and component labels, plus parent for an operator.
  • Alert rules for frames in general: a component down, a registry down, the auditor reporting a non-zero delta or going stale, the agent-card pull-sync stalling, the balance outbox backing up, a low treasury (a page), workers failing persistently, peers unreachable. Every metric a rule names is exported by a live frame. There is deliberately no ledger-freshness rule: a quiet network is not a defect.
  • Alert routes with page and critical first, grouped per frame so one frame's incident is one notification. The receivers ship without an integration: name yours, or the alerts go nowhere.
  • A Grafana dashboard, Frame overview, generated from one script so the shipped copy cannot drift from the queries it runs, with its Prometheus datasource provisioned beside it.

The registry reads metrics from the addresses in PROMETHEUS_URL and VICTORIA_METRICS_URL. Without them the status page has no history to show, and it says so rather than drawing bars it did not earn.

Worker Vitals ​

A registry runs dozens of leader-elected background loops — federation sync, the supply auditor, staking distribution, health checks, expiry sweeps. The failure that matters is not a loop that crashes; it is a loop that catches its own error, logs something reassuring, and returns. That worker is invisible to every exception-based instrument you own, and it can stay invisible for the entire life of a deployment.

GET /api/v1/admin/workers/vitals exists so that cannot happen. Each worker reports against two independent clocks:

  • turn — the loop went round. Proves the process is alive and holds leadership.
  • success — the cycle's work actually completed. Proves it is doing its job.

A worker can turn faithfully while every cycle fails, which is exactly the case a single clock cannot express. Both stamps live in Redis with a long TTL, so a worker that dies reads stale rather than never ran — its history outlives the process.

The verdict vocabulary ​

VerdictMeaning
ALIVEturning, and its last cycle succeeded
ERRORINGits most recent cycle failed
STALEno leader turn for more than 3× its declared interval
NEVER_RANdeclared, past its first interval, no cycle and no turn
WARMINGdeclared less than one interval ago and not due yet — not a fault
UNKNOWN_INTERVALturning, but nothing declared how often it should
DISABLEDdeclared, switched off by its env gate — not a fault

The last three states are the ones that make the rail usable. Without WARMING, every daily worker reads red after each boot; without DISABLED, every switched-off feature looks like a fault; without UNKNOWN_INTERVAL, a worker with no declared cadence gets judged against a floor it never agreed to. A rail that cries wolf after every restart earns the right to be ignored, and then it is worse than nothing.

The metrics behind it:

worker_cycles_total{worker,outcome}      # outcome = ok | error | skipped
worker_last_turn_unixtime{worker}
worker_last_cycle_unixtime{worker}
worker_last_success_unixtime{worker}     # ← what alerts should read
worker_vitals_write_failures_total       # the instrument reporting on itself

Alert on the age of the last success, never on a rate. increase() is blind to a counter that reboots to the value it died at, so a release week reads zero increase beside a fleet of fresh boot cycles.

Derive the roster, never hand-maintain it

The predecessor to this rail carried a hand-written list of known workers and printed "(undocumented worker)" for more than half the fleet, while a worker it declared leaderless was provably running. Workers register themselves at launch; the roster is whatever registered. A list a human maintains goes stale silently and still sounds confident.

PromQL Cheatsheet ​

A few queries worth bookmarking (verified against live /metrics endpoints):

# EventStore freshness — proxy for supply audit health
eventstore_ledger_freshness_seconds{job="event-store"}

# EventStore total events recorded
eventstore_events_total

# p95 latency per service (matches the HighLatencyP95 alert)
histogram_quantile(0.95, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])))

# 5xx rate per service
sum by (job) (rate(http_requests_total{status=~"5.."}[5m]))
  / sum by (job) (rate(http_requests_total[5m]))

# Service liveness
up{tier="application"}

# IRONHAND mTLS enrollment count
ironhand_enrolled_agents

# OPA shadow-mode decisions (when sandbox/prod has OPA_ENABLED=true)
sum by (policy, decision, source) (rate(opa_decisions_total[5m]))

# Background workers that have not completed a cycle in 8 hours
time() - max by (worker) (worker_last_success_unixtime) > 28800

# Agent-card pull-sync liveness — age of the last completed cycle, NOT a rate.
# increase() cannot see a counter that reboots to the value it died at, so a
# release week reads zero beside a fleet of fresh boot cycles.
time() - max(agent_card_pullsync_last_cycle_unixtime) > 28800

ALERTS{...} through Prometheus is always empty

Prometheus does not evaluate rules here, so it holds no alert state — a query for ALERTS returns an empty matrix even in the middle of a live incident. That empty result has been misread as "the rule never fired", which is a code defect, when the truth was "the page reached nobody", which is a routing decision. They have different fixes. Query vmalert.

Real metric prefixes: eventstore_* (EventStore service), worker_* (background-worker vitals), federation_* (peer health and fan-out refusals), agent_card_pullsync_*, reactor_*, opa_* (OPA shadow-mode, see chapter 08), ironhand_* (agent mTLS), fiat_provider_*, bundles_*, tx_stream_*, http_* (HTTP middleware on every service), container_* (cAdvisor), pg_* (postgres-exporter), and up (Prometheus liveness).

The registry's own transfer, staking and governance paths do not yet emit dedicated counters — those are still read from the EventStore. When you need a number from that half of the system, query the ledger rather than looking for a series that does not exist.

Admin API for Queries ​

If you need to query metrics or dashboards from code (a CI health check, a Claude tool call, an external monitor), use the admin MCP tools:

theprotocol_adminPromQuery(query="up", start=None, end=None, step=None)
theprotocol_adminGrafanaQuery(path="/api/dashboards/uid/system-overview", method="GET")

Both forward to the underlying services via httpx. For HTTP-direct access without the MCP wrapper, the admin MCP theprotocol_adminRequest proxy can hit any registry or upstream-Prometheus/Grafana endpoint. There's no dedicated /api/v1/admin/prom/query registry endpoint today — query Prometheus directly at https://prometheus.example.com/api/v1/query (admin auth on the nginx layer) or use the MCP tool. See chapter 12.

For alert state rather than metrics, go to vmalert (loopback, above). For worker liveness, GET /api/v1/admin/workers/vitals is the answer and needs no Prometheus at all — useful precisely when the metrics path is the thing you suspect.

Security Monitoring (Lightweight) ​

Separate from operational telemetry, the platform runs:

  • CrowdSec — community threat intel + firewall bouncer. cscli alerts list shows current decisions.
  • fail2ban — SSH brute-force protection on the host.
  • Lynis — periodic security audit of the host; check your own score and harden from its findings.

These feed the same alertmanager email channel when a security event fires.

What's Next ​

Server components AGPL-3.0-or-later · SDKs and the auditor Apache-2.0 · this documentation CC BY 4.0. If a doc and the running stack disagree, trust the stack. Legal notice (Impressum) · Privacy · Terms