Blog

Baking in Evidence

Every control you operate should produce the same evidence

Evidence tends to be in whatever shape the tool that produced it happened to emit; a JSON blob here, a PDF there, a spreadsheet far more often than it should. So how do we ensure it is more than just output, that it consistently proves what happened? Read on...

Continuing from the previous post on attestations and evidence, this post looks deeper into defining policy as code, evaluating and producing a consistent evidence format that shows why a control passed or failed.

Problems with Traditional Evidence

If you were to pick a release at random in a large regulated organisation and look at how controls were closed, you might find something like this.

ControlFINOS MitigationFormatClosureDescription
Never alone appliedCode ReviewJSONautomatedEvidence of a PR being merged and branch protection in place
Sufficient testing was undertakenTest Execution and Sign-OffPDFautomatedJUnit reports demonstrating some tests were executed
SBOM availableComponent Inventory.xlsxmanualProvenance unknown — we don’t know who generated this or how

None of these are obviously pinned to the same artifact. They could relate to anything. How do I know, for example, the tests that were run were against the artifact we’re deploying?

Each one has a different evidence format which might be highly technical in nature. A typical PR log will have Git SHAs and branch names, something a non-technical auditor might not recognise. It’s probably not obvious that the list of SHAs relates to the artifact being deployed. Perhaps worse, none of these describe the rules by which compliance can be determined. The list of PRs doesn’t include a description that “no author can be the reviewer”. That part is hidden in the runtime as branch protection rules. The point is, no one can really tell by just looking at the evidence. We need more.

None of them describe what was checked, against which version of which rules, using what inputs. Every control invented its own shape of evidence. An auditor has to learn all of them and still ends up taking the result on trust.

Modern Evidence

When evidence just shows the output of a process, an auditor has no choice but to look into that process. When it’s tooling, they can inspect the configuration (for example, look at your branch protection rules) but at scale that can become time consuming and unreliable, especially when you need to demonstrate it from six months ago. For manual processes, it quickly becomes impractical.

To be able to independently verify and re-check a decision and to be able to do that three years in the future, your evidence will need:

  • an identifiable artifact (a fingerprint of the subject of the claim, or similar)
  • the facts that were used and where each one came from
  • the rules that were applied, verbatim and versioned
  • how the facts met the rules (or otherwise)
  • the result and when it was reached

Read that list again and notice what isn’t in it. Nothing specific to a control. Nothing about code review, test coverage or CVEs.

It’s because there are no control-specific items in the list that we can create a single evidence format that will satisfy an auditor for any control.

The Bakery Example

The bakery example is useful as the production line is very similar to the software supply chain. Imagine the chef who is baking a batch of cakes. She wants to demonstrate the batch is allergen free so she can sell to the local millennials.

She’s using the approach from the previous post and:

  • recording the list of ingredients used, oven temperature and cooking times
  • defining a policy to ensure no nuts are in the ingredients list
  • expanding that policy to ensure baking temperatures and times are sufficient to cook the ingredients properly

She is capturing the policy as Rego and the ingredients and cooking details as attestations. The evaluation of the policy against the attestations produces a pass/fail decision in the format below.

A unified evidence document headed Allergen Free Cake Batch. A green shield and COMPLIANT badge at the top, a strip showing evaluated at, artifact, attestation set hash and policy hash, then three sections: attestations listed in a table with their source and timestamp, the policy as human-readable clauses plus its provenance and the verbatim Rego source and the evaluation showing each clause, the attestation it used, the expression evaluated and its result.

Attestations: What Was Observed

// attestations.json
{
  "artifact": { "id": "cake-batch-2026-05-07-001", "type": "batch-id" },
  "attestations": [
    { "id": "att-002", "type": "allergens", "field": "allergens",
      "value": ["milk", "eggs"], "source": "ingredient classifier",
      "observed_at": "2026-05-07T09:12:05Z" },
    { "id": "att-003", "type": "bake", "field": "temp_c",
      "value": 180, "source": "oven telemetry",
      "observed_at": "2026-05-07T09:44:10Z" },
    { "id": "att-004", "type": "bake", "field": "minutes",
      "value": 32, "source": "oven telemetry",
      "observed_at": "2026-05-07T10:16:10Z" }
  ]
}

Policy: What Should Have Been True

Within the evidence, the full policy is captured. Both as Rego (technically, the policy-as-code part) but also as a human-readable table of clauses. Rego can get pretty gnarly and having a human readable summary is powerful for the non-technical audience. The fact that this is code and not hearsay or departmental memory means trust levels go up between Audit and the teams operating the control.

In plain language, an auditor can see the version of the policy that was actually used, eliminating the frantic scramble to demonstrate how a system would have behaved six months ago.

# bakery.rego
package bakery

# Facts, projected from the attestation set. Each one remembers which attestation it came from.
fact[att.field] := {"value": att.value, "from": att.id} if {
	some att in input.attestations
}

default nut_free := false

nut_free if not "nuts" in fact.allergens.value

default temp_ok := false

temp_ok if {
	fact.temp_c.value >= 175
	fact.temp_c.value <= 200
}

default time_ok := false

time_ok if {
	fact.minutes.value >= 25
	fact.minutes.value <= 40
}

# The clause table: what each rule means, what it drew on and what was actually evaluated.
clauses := [
	{"clause": "Nut allergens absent", "predicate": "nut_free", "holds": nut_free,
		"inputs": [fact.allergens.from],
		"expression": sprintf(`not "nuts" in %v`, [fact.allergens.value])},
	{"clause": "Bake temperature in range", "predicate": "temp_ok", "holds": temp_ok,
		"inputs": [fact.temp_c.from],
		"expression": sprintf("%v >= 175 and %v <= 200", [fact.temp_c.value, fact.temp_c.value])},
	{"clause": "Bake time in range", "predicate": "time_ok", "holds": time_ok,
		"inputs": [fact.minutes.from],
		"expression": sprintf("%v >= 25 and %v <= 40", [fact.minutes.value, fact.minutes.value])},
]

The fact rule projects the set of attestations into something the clauses can read and helps each fact remember the attestation it came from. That from field is what later lets the evidence cite a specific input value (att-003 above) rather than just asserting that the temperature was 180. It will produce the following.

{
  "ingredients": { "from": "att-001", "value": ["flour","eggs","milk","sugar","butter"] },
  "allergens":   { "from": "att-002", "value": ["milk","eggs"] },
  "temp_c":      { "from": "att-003", "value": 180 },
  "minutes":     { "from": "att-004", "value": 32 },
  "decorations": { "from": "att-005", "value": ["sprinkles"] },
  "signed_by":   { "from": "att-006", "value": "Baker A (emp-12345)" }
}

The clauses table is the policy describing itself. Each entry carries the human-readable statement, the attestations it drew on and a rendered expression with the real values substituted in. In a few lines, it allows the document to “show the workings”.

Evaluation: Show Your Workings

The third section does the heavy lifting of showing how attestations were evaluated against the policy. I like to think of Rego as a series of expressions that use attestations as inputs to evaluate.

There’s one hidden piece of Rego that shows how to do this. The evidence.rego below is infrastructure and only needs writing once. The “policy” itself is in bakery.rego above and would be specific to each control, but the evidence generation is generic and reusable.

# evidence.rego
package evidence

import data.bakery

verdict(true) := "Pass"

verdict(false) := "Fail"

default compliant := false

compliant if every c in bakery.clauses {
	c.holds
}

status := "Compliant" if compliant

status := "Non-compliant" if not compliant

steps := [object.union(object.remove(c, ["holds"]), {
	"step": i + 1,
	"result": verdict(c.holds),
}) |
	some i, c in bakery.clauses
]

decision := {
	"status": status,
	"steps": steps,
}

Run it via OPA:

opa eval --data bakery.rego --data evidence.rego --input attestations.json 'data.evidence.decision'

…to produce a decision:

// data.evidence.decision
{
  "status": "Compliant",
  "steps": [
    { "step": 1, "clause": "Nut allergens absent", "predicate": "nut_free",
      "inputs": ["att-002"],
      "expression": "not \"nuts\" in [\"milk\", \"eggs\"]", "result": "Pass" },
    { "step": 2, "clause": "Bake temperature in range", "predicate": "temp_ok",
      "inputs": ["att-003"],
      "expression": "180 >= 175 and 180 <= 200", "result": "Pass" },
    { "step": 3, "clause": "Bake time in range", "predicate": "time_ok",
      "inputs": ["att-004"],
      "expression": "32 >= 25 and 32 <= 40", "result": "Pass" }
  ]
}

Given the JSON, we can render it however we like.

When It Fails

If we run the same control over a batch where the oven drifted to 210°C, we’d fail the control.

An evidence document headed Allergen Free Cake Batch with a red shield and a NON-COMPLIANT badge, showing its evaluation section. Step two reads 210 >= 175 and 210 <= 200 with a Fail result; steps one and three pass and the overall row reads Non-compliant.

The expression at step 2 was evaluated as 210 >= 175 and 210 <= 200 which fails with the att-003 input of 210. It fails the temp_ok rule and the overall status is non-compliant.

Notice that the evidence itself is versioned and independently verifiable. Everything is hashed, so the hash of the attestation set above is different from the happy path above because the facts are different. The document is bound to the inputs that produced it. This will make the most paranoid auditor happy.

Advantages

One format, n controls. Auditors learn the evidence format once. This dramatically changes the tone of conversation with audit.

Anyone can check it. The inputs and policy are both hashed, so a third party can take the same facts, run the same rule and see whether they get the same answer. That’s a much stronger position than “trust the tool that closed the control”.

The rule is on the page. When the policy is printed in the evidence, an auditor can challenge the rule rather than the outcome. Is 175–200°C actually the right range? That’s a far more useful conversation than relitigating whether the control was closed. Today that rarely happens because the rules are buried in a tool or in someone’s memory.

It works as a fitness function. You cannot produce this document from a screenshot. You cannot produce it if the thing collecting the evidence also decides the outcome. Adopting the format forces controls towards declarative, opinionless attestations and external evaluation. Mandating the output shape turns out to be a more effective lever than mandating the architecture.

It’s an interface. The evidence schema becomes the contract between the teams implementing controls and the system of record that stores them. Contracts can be versioned and tested.

You Still Need Trust

It comes with caveats though. It doesn’t automatically make an attestation trustworthy.

Guaranteeing that attestations are made from trusted authorities is still a responsibility of the system in which the controls operate. For example, allowing development teams to make their own attestations invites the question of “can we trust them?”, whereas running from a centralised secure pipeline may give more confidence. Digital signatures, minimising trusted authorities, applying zero-trust principles and fingerprinting are all techniques we can use to improve the trustworthiness of attestations.

Decisions and enforcement are still separate jobs. OPA can say whether the artifact satisfies the policy. Something else has to enforce a consequence. This is the difference between a Policy Decision Point (PDP) and a Policy Enforcement Point (PEP).

Where to Start

Pick one control. Write down the facts it consumes and where each one comes from. Write the rule in Rego — badly, if necessary, the first version only has to be honest, not complete. Make the evaluation emit its own workings rather than a verdict, render the document and put it in front of someone in audit.

Then refuse to accept a different shape for the next control.

The format is the easy part. The work is that every control needs an attestation schema and a policy as code before it can produce any of this. But you design the document once and every control you move onto it makes the next conversation with audit easier.

Get that right and there’s less debate, so you can all focus on the things that matter — actually reducing risk. Audit the system, not the output.

Discussion