Identity Enforcement
Whether a target actually enforces the identity scheme it declares — measured from behaviour, never assumed from the declaration.
Declaring an identity scheme in a manifest is one line of JSON. Enforcing it is a running verifier, key resolution against a directory, and a decision on every request. The gap between the two is invisible to the merchant — an integration "works" identically whether signatures are checked or ignored — which is exactly why it's the most valuable thing to measure.
How it's measured
The harness sends one valid, correctly signed request as a baseline, then a series of probes — single-variable mutations of that request. Enforcement is read as divergence: does the target treat the mutated request differently from the valid one? Divergence, not rejection — a scheme that grants verified traffic a trust upgrade rather than rejecting invalid traffic still has observable consequences, and counts.
| Probe dimension | The question it asks |
|---|---|
| Presence | Is an unsigned request treated differently? |
| Integrity | Is a tampered signature treated differently? |
| Binding | Is a signature over the wrong context treated differently? |
| Freshness | Is an expired signature treated differently? |
| Replay | Is a replayed nonce treated differently? |
| Key resolution | Is a signature from an unknown key treated differently? |
Active probes run only against reference stores, stores you own, or explicit opt-ins — never a cold third-party production endpoint.
The five verdicts
Deliberately not pass/fail: a store that accepts everything is not failing a scheme, it is ignoring it.
| Verdict | Means |
|---|---|
| Enforced | Valid requests accepted; every invalid variant treated differently. |
| Partial | Some invalid variants treated differently; others accepted unchanged. |
| Not enforced | No gating observed — signed and unsigned requests treated identically. |
| Inconclusive | The valid signed request did not succeed, so enforcement could not be assessed. Never rendered as failure. |
| Not run | Handshake only — the negative probes have not been run. |
Two rules keep the readout fair. "Not enforced" is only ever reported for a target that declares the scheme — a store that never claimed it has no identity verdict at all. And an unsigned-allowed store is not a downgrade: allowing unsigned traffic is the Permissive posture, the recommended rollout mode, and is reported as a characteristic alongside the verdict, never folded into it.
Postures vs verdicts
A posture is the stance a store is configured to take. A verdict is what the harness measures the behaviour to be. The readouts worth acting on are the disagreements — a store that believes it is Strict but measures Partial, or declares a scheme and measures Not enforced.
| Posture | Valid | Missing | Invalid |
|---|---|---|---|
| Strict | allow | reject | reject |
| Permissive | allow | allow | reject |
| None | allow | allow | allow |
Strict is the production stance; Permissive is the rollout stance; None exists on reference stores so you can see what non-enforcement looks like from the agent side. Invalid signatures are rejected under every published profile — that column is the load-bearing one.
Where you see it
- Conformance page — the verdict with the full probe-by-probe readout.
- Agent identity rail — when a session runs signed, the handshake and per-call signing surface live, and you can run the probe set from there.
- Signature Verification — the same checks from the other vantage: your agent's signatures verified by our store.