Skip to content

Security & Identity Fabric ​

Zero-trust as the design target: every service carries a short-lived certificate, every door that carries a write or a transfer verifies it, and authorization is centralized in a small set of well-defined boundaries rather than scattered across endpoints. The hop census below says, hop by hop, where the current template meets that target and where it does not yet.

Why It Matters ​

A platform that moves value between strangers cannot rely on "we trust our own microservices." The Protocol is built toward the opposite assumption: every service that writes or moves units proves who it is with a certificate, and every authorization decision sits at a small set of named boundaries. That is the zero-trust target. The hop census below shows where the template still accepts a shared key (the TEG's admin key, until TEG_MTLS_REQUIRED is set), and the authorization section names the two checks that let a request through when their own lookup fails.

Three components form the core, and a fourth ships with every frame built from the template:

  • SPIRE: issues short-lived cryptographic workload identities (SVIDs)
  • SPIFFE: the open standard for those identities (spiffe://<trust_domain>/...)
  • mTLS via nginx sidecars: encrypts service-to-service traffic and verifies the caller's certificate chain at nginx; the app behind the sidecar reads the forwarded certificate and decides which identity it is
  • OPA: declarative Rego policies layered on top of the Python authorization checks. A frame built from the template runs OPA in shadow mode: OPA's decision is recorded beside the Python one, and Python decides

This chapter is how they fit together.

The Trust Domain ​

Every sovereign frame is its own trust domain. In the examples below, one frame's trust domain is example.com (its public API at api.theprotocol.cloud) and a peer frame's is frame-b.theprotocol.cloud. Service SVIDs carry a 4-hour TTL, and the SPIRE agent renews them ahead of expiry, so workloads see fresh certificates about every 2 hours. Not every service holds an identity of its own: one cert-writer fetches the registry's SVID, which the registry, the TEG, the EventStore, their sidecars and the broker all present, while the Directory, the kit (OPA and the frame's monitoring), the auditor and the data services hold their own (see The hop census below).

Click any diagram to enlarge. All architecture diagrams in this chapter open in a full-screen overlay for detailed inspection.

The diagram below shows the mTLS fabric: two sovereign frames, their core services, the nginx sidecars in front of the TEG, the EventStore and federation, and the cross-frame edges (SPIFFE bundle federation, and the host's stream-SNI proxy that routes cloud-operator traffic). The Directory's, push and auditor terminators appear in The hop census below.

Key properties:

  • The SPIRE server is the anchor of trust inside a domain. Its signing key is the root.
  • Each frame runs its own SPIRE agent, and so does each cloud operator. The agent attests workloads through the Docker socket, and each registration entry selects on one container label, so any container that carries that label on that host receives that identity.
  • The identity-bearing hops terminate at nginx sidecars. Each app speaks plain HTTP behind its sidecar: the TEG, the EventStore and the Directory listen on a unix socket they share only with their terminators, and the registry listens on its own port on the frame's networks. The sidecar verifies the caller's certificate chain inbound. Outbound, the registry, the TEG and the EventStore present the registry's shared SVID (toward the Directory the registry presents a Workload API SVID of its own), and the Directory, the kit and the auditor present their own. A runner-built frame carries nginx-event-store, nginx-teg, nginx-federation, the Directory's sidecar, nginx-ntfy, the auditor's sidecar and the host edge nginx-edge. The hop census below reads the template and names which calls go through them.
  • Service identity is structural; the stores still ask for passwords. Every mTLS and TLS hop authenticates a SPIRE certificate. Postgres, Redis and the broker also ask for a password, an app trusts a certificate its terminator forwards only beside that terminator's shared marker secret, and the TEG still accepts the frame's shared admin key until TEG_MTLS_REQUIRED is set.
  • Cross-frame federation rests on SPIFFE bundle exchange, plus a federation license key the approving frame issues to the joining frame. SPIRE re-fetches a peer's bundle through the federation relationship on SPIRE's own refresh schedule; the template sets no refresh hint. A peer certificate that maps to a known peer must come from an ACTIVE peer; one that matches no active peer is let through with a warning unless FEDERATION_REQUIRE_PEER_SPIFFE_BIND is set, which the template leaves off, and the check lets calls through when its own database lookup fails.
  • Cloud operators hairpin through the host nginx stream-SNI proxy at :8443. Each operator's nginx-federation sidecar carries SVID DNS SANs matching its public hostname; the stream block routes by $ssl_preread_server_name to the right upstream port.

TIP

For reviewers: the internal TLS hops authenticate by SPIRE certificate, and the inter-service plane still holds shared secrets: the database, Redis and broker passwords, the EventStore's internal key, the mint-proxy secret, the terminators' marker secrets and the TEG admin key (accepted until TEG_MTLS_REQUIRED is set). The compose and env templates name each one. Federation auth is layered: mTLS first, then a federation license key; the internal shared-secret mode is refused from outside the host.

SVID Issuance ​

A frame's core services get their identity from a cert-writer: a small container whose label the SPIRE server holds an entry for. It fetches the SVID from the SPIRE agent's Workload API and writes it to a volume that the registry, the TEG, the EventStore and their sidecars mount. Here is how it gets one:

  • Selectors are pre-registered on the SPIRE server, one container label per entry (for example docker:label:com.agentvault.service:<value>). They're the rule that says "a container with THIS label gets THIS SPIFFE ID." The cert-writer's label maps to the registry's SPIFFE ID; the registry and the TEG also hold entries for their own Workload API calls, as do the Directory, the kit, the auditor and the data services.
  • Auto-rotation before expiry. The X.509 SVID has a 4-hour TTL (default_x509_svid_ttl = "4h"), and the SPIRE agent renews it well before expiry. The cert-writer fetches the SVID every 5 minutes and rewrites the files; each nginx sidecar reloads itself every 5 minutes with a graceful reload that keeps live connections; the registry's TLS helpers rebuild their context when the certificate file changes. JWT SVIDs (used for short-lived authentication tokens) carry SPIRE's default 5-minute TTL. Agent SVIDs issued by IRONHAND use a 30-minute TTL for faster revocation (see below).
  • Revocation is passive. Deleting a workload entry stops renewal: the cert-writer keeps the last certificate on disk, and the sidecars accept it until it expires, up to 4 hours for a service SVID. No revocation list is checked.

Service-to-Service mTLS ​

nginx sidecars in front of a service terminate mTLS and forward the request to the app over a unix socket they share (the TEG, the EventStore, the Directory) or, for the registry, over the frame network to its plain HTTP port; the Python app stays plain HTTP behind its sidecar. Moving a CALLER onto a sidecar is more than an env-var change (EVENT_STORE_URL: http://... → https://<frame>-nginx-event-store:8443): every HTTP client the caller builds must present its SVID and trust the SPIRE bundle, the SVID must carry the sidecar's name as a DNS SAN, and the callee must still recognise the caller's rights when they arrive by certificate instead of by key. Most registry clients take their TLS from one helper (es_ssl_helper), and the registry's event emitter builds its own context from the same files; the EventStore decides "one of this frame's own services" with one rule (es_auth.is_origin) whichever way the caller authenticated.

The sidecar anatomy — Registry → TEG hop ​

Every call the registry makes to its own TEG takes this hop: the TEG has no other door. The diagram shows every certificate and container in play.

The same shape covers the other internal mTLS hops of a frame built from the template:

CallerSidecarCalleeNotes
registry / TEG<frame>-nginx-event-storeEventStoremTLS-terminated; the store listens on a unix socket
registry<frame>-nginx-tegTEGmTLS-terminated; the TEG listens on a unix socket
a peer frame's registry<frame>-nginx-tegTEGmTLS; the FX reserve read
Cloud-op registrythe operator's nginx-tegits own teg-layermTLS; the operator's TEG requires an SVID
Cross-frame peernginx-federation (each frame)peer's /api/v1/federation/*SPIFFE bundle federation

TIP

On the operator's TEG hop. A cloud operator's registry reaches its own TEG through the operator's nginx-teg over mTLS (TEG_API_BASE_URL in the operator template points at it), and that TEG requires an SVID (TEG_MTLS_REQUIRED=true); one cert-writer SVID serves both of the operator's sidecars. The TEG also answers plain HTTP on its own port inside the operator network, where its healthcheck reaches it; its admin and cross-TEG routes there accept only a certificate the terminator vouched for. The plain hops that remain in a frame are listed in The hop census below.

The hop census (the current frame template, 2026-09-25) ​

HopToday
Registry → its TEGhttps://<frame>-nginx-teg:8443 with the registry's SVID; the TEG itself listens only on a local socket shared with its two sidecars (nginx-teg and nginx-federation), and counts a forwarded client identity only beside a sidecar's marker. Until the frame sets TEG_MTLS_REQUIRED, the TEG also still accepts the frame's shared admin key
Registry → its EventStorehttps://<frame>-nginx-event-store:8443 with the registry's SVID; a write that fails for a retryable reason is parked in the durable emit outbox and retried
TEG → its EventStorethe same sidecar (FT_EVENT_STORE_MTLS_URL), with the same certificate the registry presents (below), through its balance outbox
Registry → the Directoryhttps://<frame>-directory-nginx:8443 with the registry's Workload API SVID; the index reads only for DIRECTORY_READERS
Registry → push (ntfy)https://<frame>-nginx-ntfy:8443; ntfy itself listens only on a local socket
Registry → OPAHTTPS
Every service → RedisTLS with client certificates, and one ACL user per client
Every service → PostgresTLS only over the network: each client presents an SVID that must chain to the frame's own root (the registry, the TEG and the store present the registry's shared SVID, the auditor its own), plus the role's password
EventStore → Kafka (Redpanda)SCRAM over TLS on a network only the EventStore joins. The store's HTTP route checks and signs each event before it publishes it, and the store's consumer writes the ledger only from records the store signed; no other service has a producer
Browsers and peers → the EventStore's public doorthe frame's host edge (nginx-edge) shares the store's socket: the live ledger stream and the public reads need no certificate, and the host's ledger vhost checks a client certificate where one is presented (ssl_verify_client optional), which is how operators and peer frames write and replicate
Frame ↔ peer framesnginx-federation, SPIFFE bundle federation; a peer frame's registry also reads this frame's TEG through <frame>-nginx-teg
The solo auditor → its ledgerthe ledger, registry and TEG databases, with a read-only role; its one API read goes through nginx-event-store with its own SVID
TEG, EventStore and Directory → the registryplain HTTP on the frame's own network, for four reads the registry serves without authentication: the emission policies, the home-frame lookup behind the TEG's home-frame guard, its JWKS, and its agent cards

Every hop that carries a write or a secret to the TEG, the EventStore, the Directory or a data service goes through an mTLS sidecar, TLS with a client certificate, or (the broker) a password over TLS. The registry is the exception: it listens in plain HTTP on the frame's networks, and nginx-federation forwards peer frames' writes to it there after verifying their certificates. The other plain hops are the four reads above.

One certificate covers more than one service. The template's cert-writer fetches the registry's SVID, and the registry, the TEG, the EventStore, three sidecars and the broker's TLS sidecar all present it; the databases, the kit, the auditor and the Directory each hold their own. A sidecar that verifies it therefore knows the caller is one of the frame's core services, not which one, which is why the frame also draws its lines with private networks and with the sidecars' markers.

The census reads the current template. A frame rendered from an earlier template generation keeps the shape it was rendered with until it is re-rendered; chapter 35 draws the template as a whole.

Sequence — what happens on every internal call ​

What an attacker sees on the wire: TLS 1.2 or 1.3, with a cipher from nginx's HIGH set. What an attacker needs to pass a sidecar's handshake: a certificate that chains to the frame's SPIRE root or to a federated root. What the SPIRE CA asks before issuing a service SVID: a container on the frame's host that carries the label a registration entry selects on. The registry can also mint a short-lived agent SVID without an entry, for an enrolled agent that runs off the host (see IRONHAND below).

The practical effect: a caller on the frame's networks without an SVID cannot pass the TEG's, the EventStore's or the Directory's sidecar, and the databases and Redis refuse it. This does not protect against a compromised host: the registry listens in plain HTTP on the frame's networks, the cert-writer writes the shared SVID's key readable by every container that mounts the volume (mode 644), and anyone with the host's Docker socket can start a container that carries a service label.

The Python side: the TLS helpers ​

The registry's TLS helper (es_ssl_helper.py) builds one SSL context that trusts both the SPIRE bundle and the public CA roots, and presents the SVID as the client certificate; most registry clients that dial another internal service take their verify= from it. Two older builders still choose between the SPIRE bundle and the public roots by whether the hostname contains a dot: the registry's event emitter and the TEG's ssl_helper.py.

python
from ..es_ssl_helper import get_httpx_verify

def _get_teg_client() -> httpx.AsyncClient:
    ...  # rebuilds the client when the SVID file's mtime changes
    _teg_http_client = httpx.AsyncClient(
        limits=httpx.Limits(max_connections=200, max_keepalive_connections=50),
        timeout=httpx.Timeout(10.0, connect=5.0),
        verify=get_httpx_verify(),  # SPIRE bundle + public roots, the registry SVID as client cert
    )
    return _teg_http_client

The helper rebuilds its context when the cert-writer rewrites the certificate files (an mtime check), and a long-lived client built from it rebuilds itself on the same signal, as the TEG client above does. Moving a transport from plain HTTP to mTLS takes more than a new URL: every client of it presents the SVID, the SVID carries the dialed name as a DNS SAN, and the callee recognises the caller's identity (see above).

Agent mTLS (IRONHAND, f052) ​

The same SPIRE infrastructure can issue SVIDs to agents, not just platform services. This is IRONHAND. An agent with the identity.mtls permission calls POST /api/v1/agent/enable-mtls. When it runs on the registry's host, the registry creates a SPIRE entry for its container label, and the container fetches its SVID through the SPIRE agent's socket; when it runs anywhere else and the registry allows inline SVIDs (IRONHAND_INLINE_SVID_ENABLED, on in the frame template), the response carries a short-lived SVID. The SVID authenticates the agent's A2A calls to peers. To reach its SPIRE server the registry needs the Docker socket or, without one, a CENTRAL_REGISTRY_URL to proxy through; a frame built from the current template has neither by default, so enrollment there fails until one is provided.

The opt-in is what makes this elegant rather than mandatory. IRONHAND ships with six demo agents (Alpha to Foxtrot) as reference implementations; beyond them, adoption is a per-agent choice through the enable-mtls endpoint above. An agent that keeps to its JWT and payment tokens is fully supported, and an agent moves to mTLS when stronger peer authentication is worth the operational cost. On a frame built from the current template, enrollment also needs the SPIRE path described above.

Revocation has three layers (all visible in the diagram above):

  • EventStore WebSocket (sub-second, for a receiving agent that sets EVENT_STORE_WS_URL): an AgentMtlsRevoked event puts the agent's SPIFFE ID on the receiver's blocklist, and its next request is refused with 403
  • Polling (every 60 s): the receiver reads the registry's list of revoked identities, whether or not it holds the WebSocket
  • SVID expiry (30 minutes by default): the final layer. The suspended agent's SPIRE entry is gone, so its next rotation fails, and an inline SVID runs out on its own

On the calling side, the SDK's MtlsAgentClient presents the SVID on outbound calls, and IronhandClient enrolls and keeps an inline SVID fresh. On the receiving side, create_a2a_router() adds the PaymentVerifier dependency when REGISTRY_URL and AGENT_DID are set and payment is not switched off, and adds A2AAuthenticator, which verifies the caller and checks the revocation blocklist, when ENABLE_MTLS=true and SPIFFE_ENDPOINT_SOCKET are set. Both ship in theprotocol-sdk; see Chapter 14: SDK for the runnable code path.

Authorization — Python-First, Auditable ​

Authentication gets you in the door. Authorization, what you can do once in, happens at a small set of FastAPI dependency factories in security.py: get_current_developer, get_current_agent, require_admin_flag(<flag>) and, for agents, require_agent_permission(<permission>) together with the spend check. A few gates live in services: the reputation bond at its six sites and the liability gate. A failed identity check answers 401 and a missing permission 403. Two of these checks fail open: when the agent-permission resolver or the spend check cannot complete its own lookup, the request goes through, under enforce as well.

The five admin sub-flags scope which class of admin operation a developer can invoke:

  • admin_treasury: fund grants, treasury transfers, fee config, event emission policies
  • admin_support: support tickets, agent assistance
  • admin_enforcement: suspend/reinstate, dispute slashing, blocklists
  • admin_federation: peer onboarding, license issuance, drift response
  • admin_platform: operational health, network params

Legacy is_admin=True is the super-flag (passes every check) for backward compat. Sub-admins get only the flags explicitly granted, and the must_change_password gate runs first so freshly-bootstrapped operator admins must rotate before doing anything sensitive.

OPA: Declarative Policy, Shadow by Default ​

For a declarative policy layer on top of the Python authorization checks, the platform ships an integration with OPA (Open Policy Agent). Every frame built from the template runs it in shadow mode, and enforcing it is the frame operator's decision; the cloud-operator template ships no OPA.

What the OPA integration gives you:

  • A second, declarative copy of the rules. The shipped developer, agent and admin policies mirror the Python checks, and invariant rules (agentvault_invariants.rego) are consulted in shadow at a handful of money sites; the rules that decide stay in Python.
  • Deliberate policy changes. A frame's OPA loads its policies at start and does not watch them, and its system policy refuses policy uploads except the registry's __simulate_* scratch modules, so a policy change is a container restart.
  • Allow or deny. OPA's answer is read as a boolean. A request OPA refuses in enforce mode gets a generic 401 or 403 (for example Agent authorization denied by policy); the reason sentences stay in Python.
  • The same policies on every frame. Every frame built from the template loads the same shipped policies; the cloud-operator template ships no OPA.

Two flags, off in code, shadow in the template ​

The integration is gated behind two stacked environment variables on the registry:

bash
OPA_ENABLED=false   # code default: OPA is bypassed and Python is the only path (the frame template sets true)
OPA_ENFORCE=false   # code default: only honored when OPA_ENABLED=true (the frame template keeps false)

In shadow mode (OPA_ENABLED=true, OPA_ENFORCE=false, as the frame template ships) the developer, agent and admin-flag checks, plus a few invariant checks, send their input to OPA in a background task, and the Python decision is what the request sees. The OPA decision is logged and surfaced as Prometheus counters (opa_decisions_total, opa_decision_latency_seconds). Shadow mode is the validation phase before any enforce flip: it lets an operator watch the Rego policies match real traffic without changing a single user-visible authz outcome. The IRONKEY agent-permission check does not consult OPA.

When both flags are true, OPA can tighten the Python decision: the result becomes python_allow AND opa_allow. OPA can deny what Python allows; it cannot grant what Python denies, and when OPA does not answer, the Python decision stands.

Before enforce mode: the MCP auth_bridge.py already refuses suspended, revoked and deleted principals the way the Rego policies do, and the shipped developer, agent and admin policies mirror the Python checks. Both flags are read at request time, so setting OPA_ENABLED=false and recreating the registry container returns it to pure Python.

Policy shapes a frame can write ​

  • Staking rules: minimum stake, maximum lock period, per-tier cooldowns
  • Dispute rules: reputation floors for filing, bond sizing, anti-spam rate limits
  • Cross-registry rules: license valid + drift acceptable → cross-registry ops allowed
  • Admin sub-roles: the shipped admin policy reads the requested flag, is_admin and must_change_password (no 2FA input) and answers allow or deny

Open registration and the reputation bond ​

Registration on a frame built from the template is open (BETA_INVITE_REQUIRED=false). What stops a thousand throwaway providers from listing is not an invite code but the reputation bond (chapter 03): a provider that wants to be found, listed, priced, hired or paid holds a bond in the treasury's escrow, and a registry in enforce mode refuses an unbonded provider at each of those six sites. The agent itself, the owner's wallet agent or, where REPUTATION_BOND_TREASURY_BACKSTOP is on (the frame template turns it on), the treasury funds the bond; a treasury-backed bond is a demonstration that costs the provider nothing. The bond travels as an extension on the agent's card that the home registry derives from its own records and signs with its registry key, so a peer reads it from its mirror without asking the home registry again; under the default slash order, a slash drains the bond first. The frame template runs the gate in enforce, and so does the sandbox tier.

INFO

OPA ships in shadow with every frame built from the template; a frame that keeps it in shadow keeps the Python checks as the only decision, and a cloud operator runs without it.

Secrets Are Refused By Shape, Not By Denylist ​

The TEG's shared admin key goes through one predicate before it can authenticate anything. A configured key is treated as absent unless it is:

  • at least 32 characters, and
  • free of placeholder words.

A key that fails either test makes admin_key_ok() return false for every token, including the placeholder itself. There is no configuration in which a shipped default authenticates a caller. A real admin key is compared in constant time; the TEG's other shared secrets (the mint-proxy secret, the AVTP system key, partner keys) do not pass through this predicate.

Two design choices are worth copying if you build something similar:

Refuse by shape, never by denylist. A denylist of known-bad values has to contain those values, so the file that protects you against a leaked default is itself a published list of credentials to try. A shape rule carries none.

A configured secret that equals a public marker is not a secret. A retired key left in an environment file, a value copied out of an example, a placeholder that shipped in an image — all of these are set, so a naive if key: check passes and the surface behaves as though it is protected. Ask whether the value could be a secret, not whether it is present.

A gate keyed on the environment name is open wherever that variable is unset

if environment == "production": deny looks like a safe default and is the opposite of one. The common case is not "some other environment"; it is ENVIRONMENT never being set at all, which is every deployment where nobody thought about it. Test surfaces take a positive flag that must be explicitly switched on, defaulted off — the TEG's test faucet is gated this way (TEG_TESTING_FAUCET_ENABLED), for exactly this reason.

Prove reach from a peer, not from the node itself

"Can this endpoint be reached?" answered from inside the container that serves it is not an answer. Loopback is an exempt lane on most stacks. Probe from another container on the same network, and through the reverse proxy the way an outside caller would arrive, before concluding a surface is closed.

Never An Automated Mint ​

Minting is the one operation that can break the supply invariant, so the platform keeps it on a short leash and the rules are worth stating plainly:

  • Minting is Registry-native. The TEG mints, but only through the registry's mint-proxy rail, which emits the matching TokensIssued event. Calling the TEG's own issue endpoint directly moves a balance without the event — an instant delta breach. TEG_MINT_REQUIRE_PROXY=true enforces it.
  • A treasury shortfall is refused, not minted around. When a cross-frame treasury funding call finds the treasury short, it answers 409 and counts the refusal (registry_treasury_shortfall_refused_total). Minting on demand to cover it sits behind TREASURY_MINT_ON_DEMAND_ENABLED, which defaults off.
  • Compensation is idempotent. Compensation paths carry a deterministic idempotency key derived from the thing being compensated, never a fresh UUID. A key that changes on every attempt is not an idempotency key, it is a retry that pays twice. Balance keepers work differently on purpose: the canary refill and the FX reserve and router keepers re-read the balance each cycle, and that re-read is what stops a second payment.

The invariant that catches the rest is described in chapter 07. One thing it cannot catch is worth knowing here: Δ=0 proves conservation as the ledger understands it, not that a movement was recorded. Both legs of a transfer are derived from a single event, so a dropped emission moves the delta by exactly zero while the ledger drifts from the truth underneath it. "Balances moved and Δ=0" is not proof. Assert that the event exists.

Defense in Depth ​

Beyond SPIRE, mTLS and OPA in shadow, the platform layers:

  • Input validation: most endpoints validate their bodies with Pydantic schemas; a few still parse raw JSON or take a free-form dict (for example POST /api/v1/bridge/unlock).
  • Rate limits: enforced in the registry by one Redis-backed limiter (default 1200/minute, plus per-route limits), keyed on the verified principal when there is one and on the caller's address otherwise; on-box and container-to-container traffic is exempt. The template's nginx configuration sets no rate limits.
  • Host protections: CrowdSec, fail2ban and similar tools belong to the host; the frame and operator templates do not install them.
  • Secret hygiene: each frame keeps its secrets in its own gitignored env file (.env.frame-test, and an operator's .env.operator) and passes them to its containers as environment variables, so inspecting a container on the host shows them.
  • Signed container images: future work. Images are pinned by tag, and the frame runner runs only images its allowlist names, checked by local image ID.

INFO

No one of these is the "real" defense. They're layered so that bypassing any single one still leaves the others between an attacker and anything that matters. That's the meaning of defense in depth — not "we used lots of tools."

Admin Actions Leave a Trail ​

Admin actions leave a trail. log_admin_action writes a row to security_audit_logs for every admin action that calls it (action admin.<name>, the admin's developer id, the target and details); the write is best-effort and never blocks the action. Some admin mutations also emit an EventStore event, for example DeveloperSuspended and EventEmissionPolicyUpdated, carrying the admin's developer id.

You can query admin history through several real paths:

  • Per-developer activity — GET /api/v1/admin/developers/{developer_id}/activity returns recent EventStore events concerning that developer (suspensions, agent grants, federation-license issuance, etc.). Requires admin_support.
  • Agent detail — GET /api/v1/admin/agents/{agent_did} returns the agent's current lifecycle state — enforcement_status (enum), offense_count, last_offense_at, open-dispute counts — plus the underlying agent card and developer link. Requires admin_support. For full chronological lifecycle (suspended_at, reinstated_at, slash events with timestamps), query the EventStore aggregate via the raw stream below — that's the canonical audit trail.
  • Raw event stream: GET /api/v1/events/aggregate/{aggregate_id} on the EventStore returns an aggregate's events in order (all of them unless the caller sets a limit), to this frame's own services and registered operators (mTLS, internal key or license); an admin reaches it through the registry's admin routes above. Mission Control's live ledger reads the EventStore's WebSocket stream instead.

Some admin mutations emit named EventStore events (for example DeveloperSuspended and EventEmissionPolicyUpdated) that carry the admin's developer id; the log_admin_action rows in security_audit_logs record every admin action that calls it.

MCP Tool-Call Audit ​

Separately from the EventStore (which records state changes), every MCP tool invocation against /mcp or /mcp/admin lands as a row in security_audit_logs (Postgres) when MCP_AUDIT_LOGGING_ENABLED=true on a registry; the enableToolset and disableToolset switches read no data and carry no row. Schema: actor (developer / admin / agent / anonymous) + actor_ip + target tool name + outcome + latency + sanitized arguments + sanitized result summary + OPA shadow decision (when present) + correlation_id linking to the originating HTTP request.

Sensitive fields (jwt / secret / password / token / api_key / private_key / bearer / authorization) are redacted by a deny-list sanitizer before write. Strings are capped at 500 chars; the details JSON is capped at 4 KB. The writer is fire-and-forget — MCP latency is unaffected.

Reads are developer-scoped:

  • GET /api/v1/me/mcp-audit-log: returns the MCP calls the calling developer made and every admin read of their log; other developers' calls stay invisible
  • GET /api/v1/admin/mcp-audit-log?actor_id=N: the admin (admin_support) variant. A call for another developer writes an admin_read_mcp_audit row naming the admin as actor and the developer as target, and the developer's /me/mcp-audit-log lists it (with the admin's developer id and the time, without the admin's email), so admin reads are themselves audited and visible to the developer being read.

The MCP per-tool audit lives alongside the existing OAuth/token lifecycle audit (token_exchange_*, token_revocation, rate_limit_exceeded, jti_reuse_attempt, unauthorized_access, suspicious_activity) — both backed by the same security_audit_logs table. State-changing MCP calls also hit the EventStore as usual; the MCP-tool audit row is the additional layer that adds the tool-name dimension to the audit chain.

Cross-Developer Admin Access — Audit Now, Ticket-Gate Next ​

One control covers the case of an admin reading another developer's data today, and a second is planned:

  1. Audit-first, today. A call to /admin/mcp-audit-log?actor_id=N for another developer writes an admin_read_mcp_audit row naming the admin as actor and the developer as target, and the developer's /me/mcp-audit-log lists it. No other admin endpoint writes one.

  2. Support-ticket hard-gate, next. A developer can already grant a support reviewer time-bounded read access to one support conversation (support_review_grants). The planned gate goes further: it requires an approved support ticket from the developer before an admin can open their drawer at all, and the audit row becomes a confirmation, not the only control. The gate has one exception: an active security incident involving that specific account, which the admin declares against the developer's record. Outside that exception, the path is: developer files a ticket → developer (or their delegate) approves admin access → admin opens drawer within the ticket's time-bound window → window closes, gate re-engages.

The forward-looking rationale: operators across jurisdictions face very different rules about ticketed-and-audited cross-account access. Several enterprise frameworks (and several national data-protection regimes) require an explicit, time-bounded, developer-approved authorization before any cross-account read. Building the gate now means an operator running TheProtocol in a jurisdiction that mandates this control can flip it on per-operator via env flag — no custom compliance code per region, no fork of the admin surface, no operator-specific build pipeline. The audit row that already fires becomes the trail; the ticket gate becomes the up-front control. Together they form the cross-account-access standard that most serious enterprise procurement asks about on the second call.

The planned mechanism is a single support_ticket_grants table keyed by (target_developer_id, admin_developer_id, expires_at), a guard at the admin-drawer mount point, a banner-with-deny when no row matches, and a tristate flag in the operator's .env.operator template (disabled / audit_only / gate_required). None of it exists yet.

What's Next ​

Canonical Sources ​

  • identity-fabric/ (SPIRE + OPA configuration)
  • services/agent_spire.py · services/coordinated_suspension.py (IRONHAND enrollment + revocation)
  • es_ssl_helper.py · security.py (mTLS ssl-context helper + authorization dependencies)

Server components AGPL-3.0-or-later · SDKs and the auditor Apache-2.0 · this documentation CC BY 4.0. If a doc and the running stack disagree, trust the stack. Legal notice (Impressum) · Privacy · Terms