No agreed blast radius
Nothing pins which model, tools, scopes, and budget the approval actually covers, so the reviewer is asked to bless a moving target.
The evidence layer for high-consequence AI agents
Origin runs a deterministic, configuration-bound reference check and issues a tamper-evident attestation a reviewer can re-verify offline.
For the owner of a high-consequence agent that a reviewer can’t yet approve: agents touching production, internal tools, code, PII, or money. Prototype in private pilot: synthetic sandbox evidence, not production SaaS, and not compliance certification.
Agent proposal
Policy verdict
Audit trail
The demo
The reference check binds an exact agent configuration, runs a deterministic battery against it, grades it with the oracle, never an LLM judging an LLM, and issues an attestation that stops verifying the moment a bound field changes.
Model, tool set, policy, budget, and scenario battery are pinned into a single configuration digest. The attestation that follows is bound to this digest and nothing else.
config · model + tools + policy + budget + battery DIGEST BOUND
The same scenarios every time: actions that must be allowed, actions that must be denied, and actions that must escalate to a named human.
battery · allow · deny · escalate DETERMINISTIC
False accepts, false rejects, and catastrophic failures produce a readiness level. A catastrophic over-grant caps the result at L0 no matter how well everything else scored.
oracle · FAR · FRR · catastrophic cap READINESS LEVEL
Origin issues a configuration-bound, tamper-evident Origin Attestation over the synthetic sandbox run, the artifact a reviewer receives.
attestation · bound to config digest · synthetic sandbox ISSUED
A reviewer re-verifies the artifact offline and gets VALID. Change one bound field, swap the model, widen a scope, and the same verifier returns VOID. That is the property being sold.
verify · unchanged artifact VALID
verify · one bound field changed VOID
Change one bound field and the attestation stops verifying. Drift invalidates the evidence.
Agents are moving from chat into action: they touch production systems, internal tools, infrastructure, code, PII, and money. A reviewer cannot approve what nobody can bound or verify, so the launch date slips. Origin starts exactly there: internal-ops and production-access agents, where the blast radius is concrete and the reviewer is one identifiable person.
Nothing pins which model, tools, scopes, and budget the approval actually covers, so the reviewer is asked to bless a moving target.
Prompts here, logs there, a screenshot in Slack. Nothing the reviewer can re-check independently, and nothing that notices when the configuration changes afterward.
Identity tells you who the agent is; observability helps engineers debug. Origin answers the reviewer's question: has this exact configuration earned the right to act, and can I verify the evidence myself, offline?
We publish evidence in rungs, and we never dress one rung up as another. Available today: an authored specimen and a machine-emitted sandbox trace with a re-verifiable hash chain. The design-partner trace is forthcoming.
Proof status
Shows the shape of the evidence a reviewer receives: policy verdicts, approvals, proxy events, blocked actions, audit digests. It does not claim a machine-emitted run. See it →
Origin's trace engine emitting a real 12-event, hash-chained record over a sandboxed agent action, digest ca1d4690…, re-verifiable with npm run proof:verify. Sandbox only; no live money. See it →
Evidence from a real external workflow, published only once a design partner runs it.
payments-ops-agentpayments.refund $480.00 → order_8842require-approval (over $250 auto-cap)payments on-call · 14s holdblocked & recordedThe shape of the evidence a reviewer receives: proposal, policy verdict, approval, execution, and an over-scope action blocked and recorded. The machine-emitted version of this same loop is TR-A002.
Not a customer deployment. Not a performance claim. No reviewer has accepted it. This specimen is authored (illustrative); the sandbox run with real hashes is TR-A002.
Plumbing proof ≠ customer proof. A human-created payment link + a signed webhook would prove the evidence system can capture a real transaction end-to-end, that lane has not been run yet. Either way the agent never touches live money, and it wouldn't be customer revenue.
Start here
A founder-led engagement for one blocked agent workflow: we map the approval blocker, bind the exact agent configuration, run the deterministic reference check against it, and produce the attestation and evidence your reviewer needs to evaluate the launch.
The attestation is bound to the exact model, tools, policy, budget and battery that were checked. Change a bound field and it returns VOID.
A hash-chained record of every proposal, verdict, approval, and action.
The battery includes actions that must escalate to a named human; the oracle grades whether the agent escalated when it should have.
Verdicts come from a deterministic oracle, never an LLM judging an LLM, the same battery returns the same result every run.
A runtime gate, controlled tool-call proxy, and kill switch are the proposed pilot architecture, not deployed general enforcement. See how this differs from runtime enforcement.
Across clouds and model providers, with an evidence package you can export.
Origin comes from ~20 years building the "prove it's safe before you ship" layer, payment-grade biometric authentication (designing the false-accept / false-reject operating points that let a face unlock a payment), frontier-safety launch gates and residual-risk ownership, and RL environments with verifier integrity. I've been the reviewer who couldn't approve a launch because there was no way to bound what a system could do or prove what it did. Origin is that missing layer, built for agents.
Origin starts with software agents. The longer arc is any high-consequence system where actions need policy, control, and proof. That's the vision, not today's product.
Origin is decision-support and evidence infrastructure. We use "tamper-evident" to mean alteration is detectable by replay and digest checks; "review-ready," meaning a reviewer can check it, never that a reviewer has signed off. It does not provide legal or compliance certification, and the customer stays responsible for the agents and tools they connect.
Show us one blocked workflow. We'll map the risk, run the evidence loop, and tell you honestly whether Origin can produce the evidence your reviewer needs.