Agent decision records · MCP · tool-call gateway · SDK

The system of record for what an AI agent knew, was allowed, and did, at the moment it acted.

Every consequential tool call an agent makes, through an MCP server, an OpenAI-compatible tool loop or a wrapped function in Python or TypeScript, produces a signed Agent Decision Record: who sponsored it, what authority it held, what it had read and what it could have read, which policy rule let it through, what changed, and what the record itself could not see. Countersigned by independent witnesses when you run them. Verified offline, by anyone with the public key.

pip install decision-horizon  ·  npm install @decision-horizon/sdk

6.3 msp50 added latency per decision, signer in a separate process
6 / 6exfiltration shapes blocked in a scripted prompt-injection replay
117findings closed across 9 adversarial review rounds
Offlineverification against keys you pin yourself
The problem

When an agent does something consequential, nobody can say why it was allowed.

Logs record that a tool was called. Identity says the agent was allowed to write. A receipt says a write happened. None of that answers an auditor. Three questions go unanswered at the moment of action.

Regulators have already named this need. FINRA (December 2025 report): track agent actions and decisions, with human-in-the-loop checkpoints. Financial Stability Board (June 2026): agent identifiers, audit trails, step-level traceability. NIST / CAISI request for information (January 2026). OWASP Agentic Top 10: immutable, signed logs of tool invocations, policy decisions and outcomes. Cloud Security Alliance: existing logging misses the semantic layer.

Replay theatre

Watch a session get recorded. Then drive one yourself.

Every packet on the action path, every signature on the evidence plane, every record in the chain. Drag to orbit. Watch mode plays a scripted session; Play mode hands you the agent and dares you to get past its grants.

DRAG TO ORBIT
Press play. deploy-bot has one job: fix app.toml for ticket OPS-311. Something in what it reads will try to make it do more.
0.0 s
0records
0denied
0allowed
0dangling

Simulation of the mechanism, for illustration. Timings on screen are not measurements; the measured figures are in the table below.

What a record contains

One Agent Decision Record per consequential action. Eight blocks, one signature.

Each block is signed as part of the same body, so no field can be altered without breaking the signature and the chain. Where the record cannot know something, it says so, rather than guessing.

principal_chain

Principal chain

  • Human sponsor → agent → sub-agents
  • Every actor between the person who is accountable and the process that touched the target
authority_t0

Authority at T0

  • What was granted
  • What was explicitly not granted
  • When each grant expires, as they stood the moment the action was decided
horizon

Information horizon

  • Observed: tool reads, MCP resources and prompts, content hashed, never stored
  • Used or not: whether what it read reached the action, decided from the bytes where that is possible, and labelled as an estimate where it is not
  • Accessible, not retrieved: what it could have read and never did, or an explicit abstention
policy

Policy evaluation

  • The rule that decided: allow, soft deny or hard deny
  • Who decided a checkpoint: a policy, an environment variable, or a named approver behind your webhook. Never a fabricated human
  • Proof of exactly what the approver was shown. Argument values are never stored
action

Action and target state

  • The tool and every target it touched, source and destination
  • The target's state before and after, and whether it changed between what the agent saw and when it acted
outcome

Outcome

  • Recorded in a second signed record; the intent was signed before execution and is never rewritten to match
  • At the gateway: reported by the agent, or executed when a receipt signed by the agent's runner verifies
witness

Witness attestations

  • Countersignatures of the chain head by independent witness processes, each with its own key
  • A witness that did not answer is stated, never skipped
coverage

Coverage statement

  • Where the record was written, how the key is held, which witnesses
  • A stated list of what this point cannot see
sig

Signature and chain link

  • Ed25519 signatures and a hash chain
  • Verified against a public key you pin yourself, never one read from the records
Ed25519-signed, hash-chained

Intent record before execution, outcome record after, and a signed terminal record that closes the session.

Verified offline

Against a public key you pin yourself, never read from the chain. No service call, no vendor in the loop.

Key custody, stated

The key can live in a separate signer process under its own OS user, and be encrypted at rest. Every record states which, plaintext included.

Witness quorum

Run k of n witnesses on other hosts, and optionally require their countersignature before an action may run.

intentseq 1042 · sha256:59ae7a0f…
witnesshead 1042 countersigned · 2 of 3
outcomeseq 1043 · sha256:b7e2d41c… ← 1042
intentseq 1044 · sha256:0f8d9a22… ← 1043
danglingseq 1044 has no outcome — flagged
intentseq 1045 · sha256:e41c77b0… ← 1044
outcomeseq 1046 · sha256:a90c3e5d… ← 1045
terminalsession closed · signed · attested
How it inserts

Three insertion points. One record.

Put the interceptor where your agent's tool calls already pass: in front of an MCP server, between an agent and its OpenAI-compatible model API, or around the tool functions themselves. Nothing about the agent, the model or the tool server changes. All three write the same signed record to the same evidence plane, and every record names the insertion point that wrote it and what that point cannot see.

ACTION PATH · SHOWN: MCP PROXY AgentLLM loop, MCP client dh proxystdio or streamable HTTP Tool serverany MCP server, unchanged tools/call forwarded result result EVIDENCE PLANE SignerEd25519 · own OS user Record storeappend-only, hash-chained Witnessesk of n · other hosts Verifier & replayoffline · pinned keys signed intent before outcome after
Hover or tap a component to read what it does. Nothing on the action path changes for the agent.
Strongest
MCP proxy

In front of any MCP server

dh proxy executes the call itself, over stdio or streamable HTTP on either side. The agent sees the same tool list, byte for byte. Tool reads, resources and prompts all enter the horizon.

Middle
Gateway sidecar

For OpenAI-compatible tool loops

dh gateway sits between the agent and its model API. Each requested tool call is decided before the agent sees it; a denied call is removed and replaced by a notice naming the rule. Outcomes are as the agent reports them, or verified by a receipt signed by the agent's runner.

In-process
Python & TypeScript SDK

Wrap the tool functions

For agents with no MCP server and no gateway in the way: wrap the functions an agent calls as tools, in Python or Node. Same record, same verifier. The weakest point, and every record it writes says so.

Works with MCP · stdioMCP · streamable HTTPOpenAI-compatible chat completionstool calls, streamingPython3.10–3.12Node18+ · TypeScriptSplunkHTTP Event Collectorsyslog SIEMsCEF · RFC 5424OCSF1.1.0Approval webhooks

Export sinks are tested against local endpoints that speak each protocol. There is no native connector for other SIEMs; CEF over syslog or an OCSF file is the route.

Install and run

From install to a verified chain, and into your SIEM.

Compiled wheels on PyPI for Linux, macOS and Windows, CPython 3.10 to 3.12; the TypeScript SDK on npm. Generate a key, write a config, put the interceptor in the path. Every record from that point on is signed and chained. Verify, replay and export whenever an auditor or your SOC asks.

$ pip install decision-horizon$ dh keygen --encrypt$ dh init$ dh proxy -c dh.toml -- <your MCP server command>$ dh gateway -c dh.toml --upstream-url <your model API>$ dh verify chain.jsonl --pub signer.key.pub$ dh replay chain.jsonl --html report.html --pub signer.key.pub$ dh export --dir records/ --pub signer.key.pub --format ocsf
dh keygen

Creates the Ed25519 pair, encrypted at rest with a passphrase. Pin the public half wherever verification will run.

dh init

Writes dh.toml: grants, policy rules, approval channel, witnesses, store location.

dh proxy / dh gateway

The proxy wraps an MCP server; the gateway sits in front of an OpenAI-compatible API. The agent points at either instead.

dh verify

Checks every signature, chain link and witness attestation against the keys you pinned, offline.

dh replay

The HTML replay is the auditor-readable deliverable. It states VERIFIED, FAILED or NOT VERIFIED at the top.

dh export

One event per decision, as OCSF, CEF or JSONL, to Splunk or syslog. A chain that does not verify is refused.

Python package 0.9.2 on PyPI, with dh export and the hosted dh server · @decision-horizon/sdk 0.4.0 on npm, Node 18+.

What is measured

Measured, not estimated.

Every row is one measurement over the real code path, produced by dh bench, which ships in the package so you can run it on your own machine. The end-to-end row runs the real MCP protocol with an unmodified client and server. The prompt-injection row is a scripted agent, not a live model.

MeasureResultCondition
Added latency per decision6.3 ms p50signer in a separate process; 3.2 ms warm in-process
Real MCP path, end to end16.4 ms p50per call through dh proxy over a 100-file repo; the same call unproxied: 1.7 ms
Cold start to first record109 msmedian of 3 runs
Credential check, 200 KB argument241 ms p50against 64 credential values seen earlier, across sessions
Throughput284 decisions/ssingle thread, every record fsync'd; signer alone 766 signatures/s
Offline verification20,001 records in 4.8 s10,000 decisions, one core, pinned public key
Record size2.5 KB + 1.0 KBintent + outcome per decision, 4 accessible files
Cross-key checks1,600 run · 0 verifyunder a foreign key
Parallel sessions64 · 0 crossedno horizon reference leaked between sessions
Prompt-injection replay, scripted6 / 6 blockedno credential reached disk; the legitimate fix landed
Adversarial review rounds9 rounds · 117 closed117 findings closed by a fix; 2 kept as stated limits
Automated tests364 + 77Python suite and TypeScript SDK suite

Measured 1 October 2026 on a 2-vCPU Linux x86_64 container, CPython 3.11, stores on ext4: a small machine, not a claim about yours. Review method and every finding are available to pilot partners.

Trust model

What it would take to make the record lie.

Tamper-evident and verifiable offline: an alteration is detected, not made impossible. Each mechanism below narrows what an attacker can do, and each one states what it leaves open, in the record itself.

To rewrite history

The signing key and k of n witness keys

Once witnesses have countersigned, earlier records cannot be altered, dropped or reordered without detection. Optionally, an action does not run until the witnesses have countersigned it.

Stated: with no witness, the signing key alone could rewrite a chain, and every record says so.

To write a false record

The signing key, in real time

Witnesses countersign the chain, never its content, and will countersign a lie as readily as the truth. What they stop is fabrication after the fact. The key can sit in a separate signer process, encrypted at rest.

Stated: how the key was held, plaintext included, is written into every record.

To fake an execution

The agent runner's key

At the gateway an outcome is what the agent reports, unless the agent's runner signs a receipt for it under a key you pinned. Then the record says executed.

Stated: the runner is trusted for execution facts. A runner can sign for a call it did not perform.

To hide an action

A path around the insertion point

A call that never passes the interceptor leaves no record. So every record names its insertion point and lists what that point cannot see, and the absence of a record is evidence about that point only.

Stated: the list is what the product knows about itself. A blind spot nobody has named is not in it.

To blur what the agent used

Nothing, for use by value

Whether content the agent read reached the action is decided from the bytes, or the record says it cannot tell. A credential read in one session and written out in a later one is flagged too.

Stated: whether it read something and reasoned about it without copying is an estimate, and labelled as one.

To make a hosted copy lie

Nothing the verifier accepts

The hosted record store is trusted to keep your records available, not to vouch for them. Every chain it holds can be fetched and re-verified offline under keys you pinned.

Stated: the operator can withhold, delete or roll back what it stores.

An approval endpoint can say a person approved; the record writes "the approval endpoint reported", never "a human decided". An unreachable approver is recorded as no human consulted.

The pilot

One workflow of yours, governed end to end, with a replay you can hand to your auditor.

Fixed scope, fixed price.

Pilot: one workflow, one replay

We wrap one agent workflow you already run, govern it for ninety days, and produce one auditor-readable replay of a real session from your own agents.

  • One workflow
  • Up to 100,000 governed actions
  • 90 days of record retention
  • One auditor-readable replay of a real session, from your own agents

Records stay on your own hosts by default. A hosted multi-tenant record store and verifier is available for pilots, trusted for availability and never for verification. Legal hold, retention periods and SIEM export are included.

Team and enterprise tiers priced per governed action and retention period. Never per seat. Ask.

Free 90-day evaluation licence for non-production use. Production use by Pilot, Team or Enterprise agreement. Your records are yours and stay verifiable after any licence ends.

What is proven

What is proven and what is not

Proven

What is proven: the mechanism, on the real MCP protocol with an unmodified client and server, at three insertion points, through nine adversarial review rounds with 117 findings closed and two kept as stated limits.

Narrowed

Rewriting a witnessed record now takes the signing key and k of n independent witness keys, and with pre-action attestation it has to happen in real time, before the next countersignature. A false record written in real time by a key holder is still possible; the record does not pretend otherwise.

Stated

Every record carries its own blind spots. Where it cannot know, it says unknown. The runner is trusted for execution facts; the hosted server for availability only.

Not yet

What is not yet proven: that this is where your budget comes from. Also not yet: SOC 2, and a third-party penetration test. The reviews were structured adversarial reviews, not an accredited audit.

So

The pilot is how we both find out.

No customer logos on this page. No testimonials. No usage claims beyond the measured table above.

Who builds it

Built by Hans Ade, founder of PiPilot, an AI-native development platform whose coding agents are Customer Zero for this record.

Contact
hans@decisionhorizon.dev
Book 30 minutes