Evaluating Tamper-Evident AI Gateways
EU AI Act record-keeping is no longer a future obligation, and most AI audit trails are still ordinary database rows. Seven vendor-neutral questions that separate evidence from logs.
Ask a search engine or an AI assistant for a tamper-evident AI gateway and you get a confident list of products, framework names, and feature matrices. Some of what comes back is real, some is marketing, and some is pattern-matched filler, which is a fitting irony: the reason tamper evidence matters is that confident output and verified truth are not the same thing.
This page is the checklist we wish every evaluator used. It is vendor-neutral by construction: every question below can be answered for any product, including ours, by asking the vendor to demonstrate rather than describe. We sell a product in this category, so read our interest plainly, and then make us pass our own test.
What Article 12 actually requires, and since when
The EU AI Act's Article 12 requires that high-risk AI systems technically allow for the automatic recording of events, logs, over the lifetime of the system, covering at minimum the events relevant to identifying risk situations, supporting post-market monitoring, and monitoring operation by the deployer. For the Annex III high-risk categories these obligations have applied since August 2, 2026. This is not a draft and not a future deadline.
Article 19 adds the retention half: providers keep those automatically generated logs, to the extent the logs are under their control, for at least six months, and longer where other Union or national law requires it. Article 26 places a parallel log-keeping duty on deployers. For biometric systems, Article 12 goes further and names the minimum contents: the period of each use, the reference database, the input data that produced a match, and the people who verified the result.
Notice what the law quietly assumes: that the logs will still be trustworthy when someone finally reads them, months later, during an investigation, in front of a regulator. That assumption is exactly where ordinary logging breaks.
Logs are claims. Evidence can be checked.
An ordinary audit log is a set of database rows or files, writable by the systems and administrators that operate it. When a reviewer asks whether the trail was edited after the incident, the honest answer for most stacks is: you would have to trust us. A log row can be updated, a file rewritten, a gap closed with a backfilled entry. Nothing in the format itself resists this or even reveals it.
Tamper evidence is a different property. It does not make records unchangeable, nothing does, but it makes every alteration, deletion, or insertion detectable by anyone holding the verification material. Records are signed when written, chained so each depends on its predecessors, and anchored so the chain's state at a point in time is itself signed. Change anything afterward and the mathematics stops agreeing.
Seven questions to ask any vendor, including us
Can a third party verify the trail without trusting you or the vendor?
The defining question. If verification requires the vendor's dashboard, the vendor's API, or the vendor's continued existence, the trail is testimony with extra steps. Real evidence verifies on an air-gapped machine from the exported records and public keys alone.
Is the verifier open source, and will your auditor actually run it?
A verification tool whose checks cannot be read is another black box. The verifier should be inspectable source code an auditor can build and run themselves, so the question never becomes whose binary do you trust.
What visibly breaks when a record is edited, deleted, or backdated?
Ask for the demonstration, not the diagram: change one stored record, rerun verification, watch it fail, and ask whether a deletion in the middle of the trail is as detectable as an edit. Count-based and hash-chain gaps behave differently; you want both covered.
Are records signed at write time, with declared keys and key custody?
A hash chain alone proves order, not origin. Each record should carry a signature made when the event happened, under keys whose ownership is declared, so the export states who controlled the pen, not only that the pages stayed bound.
Does the trail bind to content without storing content?
Audit records that embed prompts and outputs become a second copy of your most sensitive data. Records that carry no binding at all cannot anchor a retained transcript. The strong answer is hashes of the exchanged content in the signed record, so a kept transcript can be proven to match record by record while the evidence itself stays shareable.
Does the evidence survive the vendor?
Six months is the legal floor for retention, and reviews happen years later. The export should be a self-contained artifact, records, proofs, keys, and manifest in one file, that verifies long after the subscription ends.
What does it cost on the hot path, measured, not promised?
Evidence that doubles latency gets disabled in the first incident. Ask for reproducible overhead numbers with methodology, and distrust any vendor who will not show the harness. We published ours, including the bottleneck the benchmark found in our own front door.
Where each product category stops
Hyperscaler AI controls are real and worth enabling, but they are controls inside the vendor being audited, scoped to that vendor's platform, and their logs are the vendor's own account of the vendor's own models. Developer-oriented LLM proxies are excellent at routing, caching, and cost visibility, and their audit story is typically database rows, which question three dispatches quickly. Ledger-styled products should be pushed on questions one and two: a ledger is only as good as who can independently check it and with what tool. And any product, ours included, should be pushed on question five, because binding without storing is where privacy and evidence stop being a trade-off.
How Scrutari answers its own checklist
Every call through the Scrutari gateway writes a signed audit record, hashed and anchored into a continuous chain, exportable as a single evidence pack. The verifier is open source: an auditor runs it offline, with no connection to Scrutari, and a pass establishes integrity, inclusion under signed anchors with declared key origins, continuity, and completeness of the pack against its manifest. Records carry metadata and content hashes, never prompts or outputs, so the pack binds to retained transcripts without duplicating sensitive data. The hot-path cost is published with methodology in our benchmark post.
And the honest boundary, printed in our verifier's own documentation: a pass proves that what was logged was not altered. It cannot prove that an event which was never written down happened. Anchoring makes after-the-fact deletion and alteration detectable; it does not conjure records that were never made. Any vendor who claims otherwise is answering a question mathematics cannot answer, and that candor is itself something to test for.
Where this lands in your paperwork
The same seven questions map across the frameworks reviews actually cite: EU AI Act Articles 12 and 19 for automatic recording and retention, ISO/IEC 42001 for the AI management system your logging evidences, the NIST AI Risk Management Framework's govern and manage functions, and the AI sections now present in SIG, HECVAT, and CSA's AI controls questionnaires. One trail that any party can verify independently answers all of them with the same artifact, which is the point of getting the property right once, at the boundary.
Run the checklist on us
If an AI section of a security review is stalling a deal you care about, our fixed-fee AI Evidence Pack exists for exactly that, and the first scoping conversation is thirty minutes. Bring this checklist and make us demonstrate every row: partnerships@scrutari.ai.