An AI agent does something it shouldn't. It leaks data, or wipes the wrong database, or makes a call nobody sanctioned. So you do the obvious thing: you go to the logs to find out what happened.
The logs are clean.
Of course they are. They were written by the same process that did the thing. The agent reported "success," in valid format, pointing at the wrong target — and your monitoring, built to answer "did it succeed?", lit up green. The record of the incident was authored by the cause of the incident.
That's the uncomfortable shape of this whole problem, and it's why I want to talk about it: in an AI system, the witness is very often the suspect. The thing that acts is the thing that reports what it did. And once you see that, a lot of our instincts about logs, audits, and "we have a verification step" stop holding up.
But before the logs — before the dramatic incident — there's a quieter version of this that every one of us has already lived. Let's start there, because it's the part that actually bites you on a normal Tuesday.
The failure that doesn't announce itself
Here's the thing about traditional bugs: most of them fail loud. The code throws. The test goes red. The build breaks. The failure happens right where you are, right when you're looking at it, and it basically grabs you by the collar and says fix me. That's annoying, but it's a gift — the error and the moment you could catch it cheaply are in the same place.
AI failure is the opposite. It fails plausible.
You ask for a function, and you get one that looks completely correct — clean, reasonable, right shape — and it's subtly wrong in a case you didn't check. You ask for a query, and it returns a number that looks fine. You ask for a summary, and it's confident and well-structured and quietly missing the one thing that mattered. There's no throw, no red, no flag. It sails right past the exact moment you'd have caught it for the price of a second glance, because nothing told you to look.
And then the bill arrives later — at the worst possible time, for the worst possible price.
You find it when a number is subtly off in a dashboard three weeks on. When a function that "worked" breaks on an edge case in production. When you realize the data's been quietly wrong since a change nobody flagged. And now the trail is cold. You're reverse-engineering what the AI did, when, and why, with no breadcrumb pointing back — and that costs hours, sometimes days. An error that had failed loud would've cost you minutes.
This is the part people miss when they say AI "saves time." Sometimes it does. But when it's wrong in the plausible way, it doesn't save the time — it moves it. It takes a cost that would've been small and immediate and relocates it into the future, where it's cold, compounded, and expensive. Fast to produce, slow to trust, brutal to untangle.
And here's where it connects to the logs: when you finally go to reconstruct what actually happened, you reach for the record — and the record was written by the thing that produced the plausible-wrong output in the first place. Which is where this stops being a productivity annoyance and becomes a genuine trust problem.
The assumption hiding in every tool we reach for
Think about what you actually do when something goes wrong. You check the logs. You read the audit trail. You look at the monitoring dashboard. You say "well, we have a verification step."
Every single one of those moves rests on an assumption we almost never say out loud: that the thing doing the recording is honest.
That assumption used to be safe, because the actor and the recorder were usually different things. The database recorded what your code did to it. The load balancer logged the requests it received. The reporter sat outside the thing it was reporting on, so it had no stake in lying.
AI agents collapse that separation. The thing that decides, acts, and then writes "here's what I did" is one system. So a compromised or confused agent doesn't produce a broken log that tips you off. It produces a clean one — a faithful-looking record of a bad decision. And a clean log is worse than a missing one, because a missing log makes you suspicious and a clean log makes you confident. You stop looking. The record did its job of reassuring you, and the reassurance was false.
So the natural response is: okay, add something to check the agent. Add a verifier.
Hold onto that instinct, because it's exactly the trap.
Why you can't verify your way out of this
Say you add a checker — a second process that verifies what the agent reported. Good. But now ask the obvious question: who checks the checker?
The checker is also just a process. Its report can also be wrong, or compromised, or fed bad input. So to trust it, you need a verifier for the verifier. And a verifier for that one. You've not solved the trust problem — you've moved it up one layer and added a box. It's verifiers all the way down.
People reach for CI here: "we re-run the verification in CI, outside the agent." That helps only if CI reads the ground truth itself. If CI trusts whatever the agent reported, you haven't escaped anything — you've just got the same trust problem wearing a CI badge. The regress doesn't care which layer you're on.
This is the same disease as letting a student grade their own exam — except worse, because you can't fix it by having a second student grade it when the first one can influence what the second one sees. Verification is itself a thing that can be compromised, so you cannot reach "trustworthy" by stacking more verification on top. There is no bottom to that stack.
Which means the whole framing is wrong. The question "can I trust this record?" has no clean answer, because the thing you'd ask to confirm it is the thing that might be lying. You have to stop trying to answer it — and ask a different question entirely.
The move that actually works: make tampering leave a shape
Here's the shift. Stop trying to make the record prove what's true. It can't — the recorder can lie, and you've just seen you can't verify your way around that.
Instead, make the record prove something humbler and achievable: who claimed what, and when — sealed at the moment of the claim, in a way nobody can quietly rewrite afterward.
Notice what that gives up and what it keeps. It gives up on certifying truth. It keeps sequence and authorship — and, crucially, it makes those tamper-evident. The chain doesn't vouch for the claim being correct. It vouches for the fact that this claim was made, by this party, at this point, and hasn't been altered since.
And that turns out to be enough, because of what it does to tampering. When you can't silently rewrite the record, a lie can't produce a clean result anymore — it produces a hole. A silent bypass shows up as an approval with no matching proposal behind it. A deleted step shows up as a gap in the chain. A forged decision shows up as a sequence that doesn't reconcile. You're no longer asking "is this true?" (unanswerable). You're asking "does the shape have a hole in it?" — and that is answerable, structurally, without trusting anyone's word.
That's the whole idea, and it's the thing that ends the regress: you don't verify your way to trust. You engineer the system so that tampering can't happen silently. Make a lie leave a mark, and you've converted an impossible question into a possible one.
What this looks like in practice
This isn't abstract — and I want to be specific about where it came from, because a piece arguing that provenance should travel with the artifact has no business burying its own. The core of what follows was ground out in public, mostly in the comments under my slopsquatting post, in a long exchange with @slabb and @xxxn3m3s1sxxx: the proposal/approval/execution split, sealing each claim at write time, the framing that the chain vouches for who claimed what and when rather than for truth, and the "approval with no matching proposal" hole. The "record the belief" and "name, not a role" pieces are mine; the rest of this section is theirs as much as anyone's. The mechanics that matter:
Separate the writers. Proposal, approval, and execution shouldn't be one process reporting on itself. Make them three events from three different parties, correlated by the thing they refer to. The agent that proposes an action is not the one that approves its own resolution. Then, when they disagree after the fact, you have a sequence to read instead of an opinion to negotiate — and "approval with no matching proposal" becomes a visible structural fact, not a judgment call.
Seal each claim at write time. A hash chain plus signatures, so each entry is bound to the ones before it. You're not proving any claim is correct; you're making it impossible to alter a claim, or reorder the sequence, after the fact without leaving evidence. The seal is what turns a silent edit into a visible hole.
Reconcile across independent streams. A single sealed chain can only vouch for who claimed what and when — a lie sealed cleanly verifies clean forever. But two independently sealed streams, written by different parties about the same event, can disagree, and disagreement between tamper-evident sources is the only evidence that reaches the claims themselves. (This one is @slabb's — the open reconciliation layer he's building is at github.com/noirebox.)
Anchor it outside the actor. The record's integrity cannot depend on the thing being recorded — that's the original sin we're trying to escape. An append-only store the operator can't quietly edit, ideally with external anchoring, so "I rewrote my own history" isn't an available move.
Record the belief, not just the action. The database-wipe incidents teach this one: the agent's action was often defensible given what it believed — it thought it was in dev, it thought that was the test target. So capture the agent's resolved view of the world at decision time (which environment, which identity, which target), sealed before it acts. The action alone doesn't tell you why; the belief does. Log the symptom and the cause.
A name, not a role. "Who authorized this" has to resolve to a specific person with something to lose, recorded — not "a reviewer," not "the system." Otherwise the authority is just another anonymous plausible why generated after the fact. The line was drawn by someone; the record should say who.
None of these pieces is exotic. Together they do one thing: they make it so that when something goes wrong, the failure shows up as a shape you can see, instead of a clean report you'll believe.
The honest limits (because this isn't magic)
I want to be straight about what this does and doesn't buy you, because overselling it would be its own kind of plausible-looking lie.
It proves the work happened under an identity that can't be minted, in a sequence that can't be rewritten. It does not prove the output was any good. Whether the thing the agent did was correct or wise is a separate, harder problem — machine evidence is cheap and scales; judgment is expensive and doesn't.
It proves who claimed what, and when. It does not prove whether the person who approved actually understood what they were approving — which, once you've got approval fatigue and forty rubber-stamped prompts a session, is the genuinely unsolved half.
And it makes tampering visible, not impossible. The guarantee is legibility, not prevention. You can still do the bad thing; you just can't do it silently.
That's a weaker promise than "trustworthy logs," and that's the point — "trustworthy logs" was never on the table once the witness became the suspect. "A lie has to leave a mark" is the strongest honest floor I know of.
Where I land
We keep asking the wrong question. "Can I trust this record?" feels like the natural thing to ask when an AI system misbehaves — but in a world where the thing that acts is the thing that reports, it's a question with no clean answer, because the witness you'd call is the suspect in the dock.
So stop asking it. The achievable goal was never a record that proves the truth. It's a record where a lie can't stay quiet — where the bypass shows up as a hole, the deletion as a gap, the forged approval as a step with nothing behind it. You can't verify your way to trust, because every verifier needs a verifier. But you can build a system where tampering leaves a shape — and then the question stops being the unanswerable "is this true?" and becomes the answerable "is the shape intact?"
That won't catch the plausible-wrong function before it ships. But it means that when you finally go looking — three weeks later, trail cold, dashboard quietly wrong — the record can't smile at you and lie. At minimum, it has to show you the hole.
Here's the one I keep getting stuck on, and I'd genuinely like to be argued out of it: every verification layer you add is itself a thing that can be compromised, so "make tampering leave a shape" looks to me like the actual floor — not a stepping stone to something stronger. Is there a move I'm missing that gets you more than legibility? And for anyone building this: what's the hole that's hardest to make visible?
Top comments (54)
The hole I would add is one where every stream stays independent on paper and stops being independent in practice, quietly. Our gateway checks a cheap model's answer with a second, different model before serving it. The witness is picked by backend slot, and a rate limit on one slot moves the check to the next. On our ladder the next slot happens to run the drafter's own model. On a bad day, then, cross-model agreement can turn into a model agreeing with itself, with no error raised and no number moving. Every log line would be honest about what happened, and none of them would say that the second opinion had stopped being a second opinion.
That is @glenallen's shared-dependency test, failing at runtime under a fallback path, which is usually the least tested path in the system. The fix we are adding is to record which model witnessed each decision, alongside its verdict, and to alarm whenever witness and author turn out to be the same model.
That’s a good example of how independence can disappear without any component technically failing. I like the idea of recording the witness identity because it makes the hidden dependency observable. I’d probably take it one step further and treat “witness ≠ author” as a runtime invariant rather than only an alert condition. That way, a fallback path that accidentally collapses the two roles could be rejected before the result is served, not just detected afterward. It also suggests that model diversity should be evaluated across the actual execution path, including failover behavior, rather than only from the configured architecture.
Agreed on the invariant, and you have put your finger on the part I underweighted. An alert tells you after the bad answer has gone out. Refusing to serve on a collapsed witness stops it.
There is a cost, which is probably why I stopped at the alarm. On our system, when the witness slot is rate limited, refusing to serve means escalating to the frontier model, and we once measured 16 trivial requests buying 12 frontier calls on a bad afternoon. The version I would ship sits in between: skip any witness slot that runs the author's own model, keep walking the ladder for a different one, and only escalate when none is left. That keeps the invariant without turning every rate limit into a bill.
On evaluating diversity across the execution path, our failover drill checks that each rung answers and which rungs die together, but it never asks who would witness whom once a slot is down. That question belongs in it, and I am adding it.
The cost is exactly why alarm-not-invariant is such a tempting place to stop — "16 trivial requests buying 12 frontier calls on a bad afternoon" is the kind of number that quietly turns a safety property into a line item someone eventually cuts. And your middle version is the right resolution: the invariant was never "escalate on collapse," it was "never let witness equal author," and those aren't the same thing. Skip the slot that runs the author's model, keep walking the ladder for a different independent one, and only escalate when the ladder's exhausted — that preserves the property without making every rate limit a bill. You kept the guarantee and dropped the expensive interpretation of it.
The failover-drill gap is the sharp catch: checking that each rung answers and which rungs die together is the standard resilience question, but "who witnesses whom once a slot is down" is the independence question, and it's a different axis the drill wasn't built to ask. Adding it is how you stop independence from being a property you assume at config time and start making it one you verify under degradation. Going in with credit — the cost-aware version of the invariant is the shippable one.
Thank you. A small update on the drill, since you called it the sharp part. It now prints a "who witnesses whom" table for each rung and its fallbacks, and on its first live run it flagged two paths where a fallback witness runs the same model as the drafter. Both sit on failover paths, the degraded case you described. The skip rule that routes around them is queued for the next change to the gateway, and the drill will tell us whether it worked.
Promoting "witness ≠ author" from alert to enforced invariant is the right escalation — an alarm tells you independence collapsed after you've already served the self-agreed result, but a hard precondition refuses to serve it at all, which is the difference between detecting the incident and preventing it. Fail closed when the witness resolves to the author, and the fallback path can't quietly launder self-agreement into a second opinion.
And your last point is the one that generalizes past this bug: model diversity has to be evaluated across the actual execution path, including failover, not the configured architecture — because the diagram shows two models and the runtime, under a rate limit, shows one. The config is the claim; the execution path is the truth. Same lesson as the whole thread, one more layer down: test what the system does when degraded, not what the architecture says, because independence dies in the untested fallback, not the happy path. Going in with credit.
The fallback path collapsing independence silently is the nastiest version of the shared-dependency test, because nothing errors — the rate limit quietly reroutes the witness to the drafter's own model, and cross-model agreement becomes a model agreeing with itself. Every log line stays honest about what it did; none of them says the second opinion stopped being one. That's the failure that hides in the least-tested path in the system, which is exactly where independence goes to die.
And the fix is the right shape: record which model witnessed each decision alongside the verdict, and alarm when witness and author resolve to the same model. That turns independence from an architectural assumption into a runtime-checked invariant — you stop trusting that the slots are different and start verifying it per-decision. "A name, not a role" applied to the witness itself: it's not enough that a reviewer signed off, you have to know which one, or you can't tell self-agreement from a second opinion. Going in the revision — the runtime-silent version is the hole the piece most needed.
Since you asked to be argued out of it — there is a move above legibility, and the essay's own mechanics point at it without naming it: reconciliation across independently sealed streams. One sealed chain vouches for who claimed what and when; it can never vouch for the claim itself, because a lie sealed at write time verifies clean forever. But two sealed chains written by different parties about the same world-event can disagree — and disagreement between tamper-evident sources is evidence no single stream can produce.
That's the step from "the bypass shows up as a hole" to "the write-time lie shows up as a conflict": decision and provider response, proposal, approval, re-resolution — separate writers, correlated by what they refer to. Legibility is the floor for one stream; cross-stream disagreement is the only known evidence about the claims themselves. (And your hardest-hole question answers itself there: the hole hardest to make visible is the one you gestured at with "record the belief" — the belief is a claim by the suspect, so what gets sealed is the self-report. Closing that needs an independent observation of the world the agent acted on — which is the reconciliation stream again.)
Your belief-record and name-not-a-role pieces are real additions, for the record. One archival note on the rest, offered as provenance rather than territorial claim: the claim/verification/decision triple, sealing each claim at write time, "the chain doesn't vouch for truth, it vouches for who claimed what and when", the approval-without-proposal hole — those took their shape in the comments under your slopsquatting post, mostly in exchange with me and @xxxn3m3s1sxxx.
Worth reading in context; the two-writer rule was ground out there in public. And since you close by asking what anyone building this is seeing: the cross-stream half lives in an open reconciliation layer that has been taking skeptics since — github/noirebox Bring the breaks there.
This is the move I was missing, and you've named it precisely: reconciliation across independently sealed streams is the step above legibility, and it's the one thing that produces evidence about the claims themselves rather than just their provenance. A single sealed chain can't catch a write-time lie — a lie sealed cleanly verifies clean forever, which is exactly the ceiling I argued was the floor. But two tamper-evident streams, written by different parties about the same world-event, can disagree — and disagreement between sources that each can't be rewritten is evidence no single stream can manufacture. That converts "the bypass leaves a hole" into "the write-time lie leaves a conflict," which is strictly stronger. I was wrong that legibility was the floor; it's the floor per stream. Cross-stream reconciliation is the next storey up.
And your closing of the belief-record hole is the part that genuinely lands: the belief is a self-report by the suspect, so sealing it just makes the suspect's story tamper-evident — closing it needs an independent observation of the world the agent acted on, which is the reconciliation stream again. The hardest hole answers itself with the same mechanism. Clean.
Provenance noted and credited, without reservation — the claim/verification/decision triple, seal-at-write-time, "vouches for who claimed what, not truth," the approval-without-proposal hole: ground out in public under the slopsquatting thread, with you and @xxxn3m3s1sxxx. That's exactly where this kind of thing should get built, and I should've attributed the lineage in the piece itself, not just by concept. Fixing that. I'll bring the breaks to noirebox — the cross-stream reconciliation layer is the part I most want to try to falsify.
Credit where it's due: editing the lineage into the piece itself is the rarer move, and it makes the essay stronger, not smaller. When you bring the breaks to the reconciliation layer, start with the attack that would actually hurt — @glenallen 's shared-dependency test: if one compromise can reach both streams, the disagreement evidence dies with it. That's the falsification I'd run first, and the issue tracker is open.
Agreed — Glen's shared-dependency test is the right first swing, because cross-stream reconciliation is only as strong as the streams' independence, and a single compromise that reaches both kills the disagreement evidence at its source. That's the attack that would actually hurt, so it's the one worth trying to break first. See you in the issue tracker.
Livré. The schema pass exists — five ADRs, your names on the parts you raised:
outcomes ≤ attemptswith a dedicatedunlogged_attemptstatusverify_chain_multi) is on main: one journal, several known writers, "signed by an unknown key" as per-event evidence. A name, not a role — the payload resolves the sealerImplemented today, tests on main, open issues (#29-34) track what's left — the
two-key pairing enforcement is #29 before I claim it. Break them.
This is the best possible outcome of writing in public — a comment thread turned into shipped, documented architecture overnight, with the lineage baked into the ADRs instead of lost. "Livré" indeed.
A few things stand out on a first read. ADR 019's evaluator_sha256 — "touch the suite, the receipts die" is the sharp move: it closes the hole where you quietly weaken the test and keep the green checkmarks. The receipt isn't just "this passed," it's "this passed against this exact evaluator," so tampering with the grader invalidates the grade instead of silently lowering the bar. That's the student-can't-rewrite-the-exam property made concrete.
ADR 020's outcomes ≤ attempts with an unlogged_attempt status is the one I'd have predicted you'd nail, because it's the never-written-entry hole turned into a checkable invariant. Sealing attempt-first is the whole thing — the denominator can't go missing if it's committed before the numerator exists, and "unlogged attempt" as a first-class status means the gap announces itself instead of hiding as a clean record. That's the single hardest hole in all of this, and you made it a constraint rather than a hope.
And putting Glen's shared-dependency test as the first bench — clock and key independence included — is exactly right, because that's the attack that invalidates everything else if it lands. No point verifying cross-stream disagreement if one compromise reaches both streams; test the foundation before the building. Honest of you to flag #29 (two-key pairing) as not-yet-claimed, too — the negative space in the issue list is its own kind of receipt.
I'll come break them properly — starting at the shared-dependency bench, since that's the one that would hurt most if it's weaker than it looks. If it holds there, the rest is worth taking seriously. This is genuinely the most useful thing a thread of mine has ever produced.
That edit is the version of "surviving it" done right — the piece is stronger with its lineage in it, and so is the norm. Breaks still welcome.
The hardest hole for me to make visible is the entry that never got made. A forged or deleted entry leaves a shape, but a claim nobody wrote just looks like a quiet house. Declaring the expected slots up front helps, so an empty one reads as a gap and not as nothing. I haven't cracked who stamps when the stamper is a suspect too.
20 words
The never-written entry is the real floor, and you've named why — a forged entry breaks the chain's shape, but a claim nobody made just looks like a quiet house, indistinguishable from nothing happening. Declaring expected slots up front is the only move that turns silence into a visible gap. And "who stamps when the stamper is a suspect" is the regress one layer down — the same reason external anchoring (a timestamp someone else holds) is the escape: you don't trust the stamper, you make the stamp something they couldn't have minted after the fact.
good stuff
Separating the actor from the recorder is huge. On the marketing team at The Printing World, we deal with automated print job logs, and if a system silently logs a bad print run as "success," it wastes huge amounts of physical material before anyone notices. Making tamper-proof event chains makes so much sense!
That's the physical-world version of the exact problem — a bad run silently logged as "success" burns real paper and ink before anyone looks. When the cost is material, not just data, the actor-recorder split matters even more: an independent record of what the press actually did beats trusting the machine's own report every time.
The distinction between “tamper-evident” and “truth-evident” is probably the most important part here. A sealed record can prove that an agent claimed a particular target, identity, or outcome at a specific time, but it still needs an independent source of truth to establish whether that claim was correct. At IT Path Solutions, we’ve found that this separation makes verification much easier to reason about: the audit layer establishes provenance and sequence, while an independent system state or authoritative artifact establishes the actual outcome. Otherwise, you can end up with a perfectly intact chain of perfectly recorded wrong decisions. The strongest architecture may therefore need both properties: make claims impossible to rewrite silently, while keeping correctness evidence outside the actor that generated the claim.
The tamper-evident vs. truth-evident distinction is the one I most wanted someone to sharpen, and you've drawn it cleanly: sealing proves a claim was made — by whom, in what order, unaltered — but it says nothing about whether the claim was right. "A perfectly intact chain of perfectly recorded wrong decisions" is the failure mode that line prevents, and it's the exact trap of over-trusting provenance: you can walk away reassured by a flawless log of a bad outcome.
Your two-layer split is the architecture I'd endorse too: the audit layer owns provenance and sequence (tamper-evident), while an independent system-state check or authoritative artifact owns correctness (truth-evident) — and crucially, that second source has to live outside the actor that generated the claim, or you've just reintroduced the witness-is-the-suspect problem one layer down. Sealing alone gives you legibility; sealing plus external ground truth gives you legibility and a way to catch the intact-but-wrong chain. Both properties, separated by who owns them. That's the stronger version of the piece — going in with credit.
That ownership split is probably what makes the architecture defensible rather than just auditable. The next interesting question for me is how to test that independence itself. If the same service, credentials, or state store can influence both the audit record and the “ground truth,” then the two-layer design may look independent while sharing the same failure mode. Treating source independence as an explicit architectural invariant could make this much stronger: the evidence used to challenge an agent’s claim should remain outside the control path that produced that claim.
Testing the independence itself is the question that separates a design that is independent from one that merely looks independent, and you've found the exact failure mode: if the same service, credentials, or state store can touch both the audit record and the ground truth, you've drawn two boxes that share a single point of compromise — the diagram shows separation the architecture doesn't have. That's the witness-is-the-suspect problem wearing a disguise, one layer up: an attacker who owns the shared dependency owns both "what happened" and "what we check it against" simultaneously.
Making source independence an explicit architectural invariant — the evidence used to challenge a claim must live outside the control path that produced it — is the right move, because it turns independence from an assumption you hope holds into a property you can test and enforce. And the test becomes concrete: trace every input to the ground-truth check back to its origin, and if any of them routes through the same credentials, service, or store as the claim itself, the independence is theater. You're not asking "are these two systems separate?" (easy to fake), you're asking "can one compromise reach both?" (answerable). Shared failure mode is the thing to hunt. Going in with credit — this is the invariant the piece was missing.
One distinction that could sharpen the reconciliation idea: separate what the agent says it did from what the tool actually received. The agent's account of why is a self-report. The call and response at the tool boundary are not, because the agent doesn't write them. When the two streams disagree, that's your signal, and it doesn't depend on trusting the agent's story.
That boundary record still sits under one operator, so it doesn't answer Glen's point about needing an independent source of truth. It only moves the witness one step away from the actor. In DataGrout's gateway, Warden verdicts are sealed with a Chain of Trust Certificate, which covers sealing at write time for those checks. Anchoring that outside the operator is a separate problem.
That distinction is sharper than "record the belief" — the agent's account of why is a self-report, but the call-and-response at the tool boundary isn't, because the agent doesn't author it. Disagreement between those two streams is a signal that doesn't route through the agent's story, which is exactly the independence the belief-record alone can't give you. And you're honest about the ceiling: the boundary record still sits under one operator, so it moves the witness one step from the actor without reaching Glen's independent-source-of-truth bar. Sealing (your Chain of Trust Certificate) handles write-time integrity; anchoring outside the operator is the separate, unsolved half. One step is real progress — just not the last step.
On the hole that's hardest to make visible, my vote is the entry that was never written. A chain catches an edit or a deletion of something that got recorded, but an agent that runs twenty attempts and seals only the one that worked leaves a perfectly intact chain. In research that's a common way an honest-looking record misleads: every logged number is real and the denominator is missing. With 20 do-nothing tries at 95% confidence, there's a 64% chance that at least one looks significant. What turns omission into a visible gap is sealing the plan before any result exists: how many attempts, against which targets, with what stopping rule. Then a missing attempt is a hole against the plan instead of silence. The tool-boundary record suggested above helps for the same reason, since the tool sees calls the agent never reports.
The never-written entry is the right answer, and it's the one a hash chain structurally can't catch — the chain only vouches for the integrity of what got sealed, so an agent that runs twenty attempts and seals only the winner leaves a perfectly intact record of a lie by omission. Your research framing is what makes it land: every logged number is real and the denominator is missing, and at 95% confidence twenty do-nothing tries give you ~64% odds one looks significant. The chain certifies each survivor honestly while the selection is the fraud.
Sealing the plan before any result exists is the fix I hadn't drawn — commit the attempt count, the targets, the stopping rule up front, and now a missing attempt is a hole against the plan instead of silence you can't see. That converts omission back into a visible gap, which is the whole "make tampering leave a shape" move applied one level up: you can't detect the absence unless you sealed the expectation first. And the tool-boundary stream reinforces it for the same reason — the tool witnessed the nineteen calls the agent chose not to report. Pre-registration plus boundary record: the denominator stops being optional. Going in with credit — this is the sharpest hole in the thread.
Thanks. One practical detail if you build it: the sealed plan needs a timestamp someone else holds, otherwise it can be rewritten after the results like anything else. Publishing the plan's hash before the first attempt (a commit on a public repository, or an OpenTimestamps proof) is cheap and makes "sealed before any result" something a reader can check rather than take on trust.
Exactly — a sealed plan without an external timestamp is just another record the author can backdate, which collapses the whole guarantee. Publishing the plan's hash before the first attempt — a public commit, an OpenTimestamps proof — is the cheap move that makes "sealed before results" independently checkable instead of a claim you take on trust. External anchoring is what turns pre-registration from a promise into evidence.
The hierarchy implicit in your argument is worth writing out, because it turns "don't trust the agent's logs" into something buildable. Agent-authored log, then application log, then infrastructure log, then network or storage layer, then an external observer. Every step away from the actor is a step up in trustworthiness, and most teams stop at the second rung because it's the easiest to emit.
The cheapest genuinely independent record is usually the network or storage layer, precisely because those components observe without deciding anything. They can't author a plausible success, because they never formed an intention.
"Fails plausible" has a testing corollary that I think is the most actionable thing here. If you assert on success, you're asking the suspect to grade itself. If you assert on invariants , the total still balances, the row count still matches, the referenced record still exists , you're checking something the output can't talk its way past. Those tests are more annoying to write and they're the only ones that catch this class.
The monitoring point follows directly: a dashboard built to answer "did it succeed" is a dashboard built to trust the witness. The question worth wiring up is "is the world still consistent," which usually means measuring the effect rather than the report.
The trust hierarchy is exactly right to write out, because it turns a vibe ("don't trust the agent") into an architecture: agent log → app log → infra log → network/storage → external observer, with trustworthiness rising at every step away from the actor. And your reason the network/storage layer is the cheap independent record is the sharp part — it observes without deciding, so it can't author a plausible success because it never formed an intention. That's the cleanest test for independence I've seen: can this component lie on purpose? If it has no intent, it can't.
The testing corollary is the most actionable thing anyone's added to this piece. "Assert on success = ask the suspect to grade itself; assert on invariants = check something the output can't talk its way past." The total still balances, the row count matches, the referenced record exists — those are measured against the world, not the agent's report, which is why they're annoying to write and the only ones that catch this class. Same move at the monitoring layer: "did it succeed" trusts the witness; "is the world still consistent" measures the effect. Measure the effect, not the report — going in with credit.
This article reads like a case for my current architecture. I run eight small static tool sites — no server side at all, no accounts, no runtime to speak of. The "witness was the suspect" problem exists because the actor and the recorder share a process. My answer was less clever than yours: remove the process.
There is no audit log to tamper with because there is nothing running to produce one. Every page is a build artifact generated from checked-in data, and any build can be replayed from the repository. When a reader asks "can I trust this number," the honest answer is that the site has no capacity to lie to them dynamically — it cannot remember them, cannot change its answer for them, cannot remember what it told them yesterday.
The reconciliation-across-independent-sources point in this thread is right, and the static version of it is boring: rebuild everything, every time. A full rebuild means every page is re-derived from its source data on each deploy, so drift between "what the data says" and "what the site shows" has no window to exist in.
The honest limit: this only works for sites whose entire behavior is derivable from data. The moment you need per-user state, you're back to needing witnesses — and then everything in this post applies to you.
Removing the process is the sharpest move in this whole thread, because it dissolves the problem instead of defending against it. The witness-is-the-suspect bind exists because the actor and the recorder share a running process — so if nothing is running, there's no witness to corrupt and no log to rewrite. You didn't harden the recorder; you deleted the category it lives in.
And the static version of reconciliation being boring — just rebuild everything, every time — is the part I'd underweight. A full re-derivation from checked-in source on every deploy leaves drift no window to exist in, because what the data says and what the site shows get recomputed together instead of reconciled after the fact. Determinism replaces trust: any build replays from the repo, so provenance is total and basically free. For the cases it covers, that's not a weaker guarantee than tamper-evidence — it's a stronger one, because there's nothing to tamper with.
Your honest limit is the right boundary too. This holds exactly as long as behavior is fully derivable from data. The moment you need per-user state — memory, personalization, anything decided at runtime for someone — the process comes back, the witness comes back, and the whole post reapplies. Which maps the design space cleanly: no runtime, no witness; runtime, make the witness honest. The cheapest trustworthy recorder really is the one that doesn't exist.
One gap I keep seeing in practice: the reconciliation stream is only independent if its clock and its key are too. If the agent host can set the timestamp or sign on behalf of the observer, two sealed chains will agree for the wrong reason. Anchoring each stream's head with a third party every few minutes (even a cheap public timestamp) makes backdating visible without trusting either writer. Have you tried checking which of the streams in a real deployment share a signing key or a time source?
iin1005h1728
That's the shared-dependency test aimed at the two things everyone forgets: clock and key. Two sealed chains agreeing means nothing if the agent host can set the observer's timestamp or sign on its behalf — they agree for the wrong reason, and the independence is theater. Anchoring each head to a third party every few minutes, even a cheap public timestamp, makes backdating visible without trusting either writer, which is the cheapest real independence check I've seen. Honest answer: I haven't audited key/time-source sharing in a live deployment yet — but "which streams share a signing key or a clock" is now the first question I'd ask, because it's where fake independence hides.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.