Skip to content
Evidence and reports

Files you can hand to an auditor.

One folder per attestation, self-contained, verifiable offline, and honest about its gaps. This page describes the format the workspace reads.

The report folder

suncly-reports/<attestation-id>/ holds four kinds of file:

  • result.json: the evidence bundle, format suncly-result/1.
  • transcripts/<run-id>.json: one evidence document per recorded run, byte for byte as stored.
  • report.md: the human-readable report as Markdown.
  • report.html: the same report, self-contained, no network needed.

result.json

The bundle without the inline transcripts (they are the files next to it). Top-level fields, with the data model's names:

FieldContents
agentid (derived from the card URL), name, owner, risk_level.
card_versioncard_hash, raw_json exactly as fetched, fetched_at.
parsed_cardThe card as JSON and as the parsed model: skills, interfaces, capabilities, modes.
card_urlThe URL that was fetched.
contract, test_casesThe approved contract (version, status, approved_by, approved_at) and its test cases with input and criteria.
attestationStatus, trigger, timestamps, budget_limit, cost_total, signature, signing_key_id.
runsEvery recorded run (verdict, judge_layer, latency_ms, attempt, transcript_ref) with the document_hash of its transcript file.
decisionsEvery decision record, oldest first. The first is the Policy engine's.
resultsPer test case: pass_count, fail_count, inconclusive_count. Never combined.
planned_runs, not_executedHow many runs were planned and which never ran, with the reason: budget, runner_crashed or withheld.
card_recheckWhat the final re-fetch found: unchanged, changed, unavailable or not performed.
not_testedThe “What was NOT tested” list, category and detail.
sandbox_declared, drafter_nameThe sandbox declaration and the source of the test plan.
signer_public_key, signature_payloadThe public key (base64url) and the exact payload that was signed.
proposalsThe open-question proposals in effect for this run.

Transcript files

Each file is one evidence document with two parts:

  • transcript: the target URL and binding, timestamps, latency_ms, the outcome (responded_task, responded_message, unreachable, timeout, protocol_error), the final task state and final response, every request and response exchanged with headers and bodies after redaction, and a redaction summary naming the rules that fired.
  • judgement: the verdict, the judge layer, a summary, and every check with its name, its passed value (true, false, or null when Layer 1 cannot decide) and a detail such as “expected TASK_STATE_COMPLETED, observed TASK_STATE_INPUT_REQUIRED”.

The file is written once in canonical JSON (RFC 8785). Its SHA-256 is what the signature covers.

What the signature covers

{
  "payload_version": 1,
  "attestation_id": "…",
  "card_hash": "sha256:…",
  "contract": {"id": "…", "version": 1},
  "results": [
    {"test_case_id": "…", "skill_id": "order-status", "kind": "skill",
     "pass": 5, "fail": 0, "inconclusive": 0}
  ],
  "transcript_hashes": {"<run-id>": "sha256:…"},
  "decision": {"outcome": "flag", "policy_version": "unconfigured"}
}

The payload is canonicalized with RFC 8785 and signed with the deployment's Ed25519 key; the signature is stored as ed25519:<base64url> and signing_key_id is a SHA-256 fingerprint of the public key. Failed and invalidated attestations are signed with "decision": null. The signature does not cover a later human decision, the report prose, or anything about the agent's future behaviour.

Verifying

suncly verify suncly-reports/<attestation-id> [--public-key <base64url>]

Checks, in order: signature present; public key matches signing_key_id; signature valid over the payload; card_hash recomputed from raw_json equals the payload and the record; attestation and contract ids match; every transcript file hashes to its signed value; the recorded decision equals the signed one; the signed per-test-case counts match the recorded runs. The workspace runs the same checks in the browser, with the CLI as the reference.

What was NOT tested

Every report carries this section. The categories the current version can emit:

  • skill without test case: a declared skill has no test case (no examples, or left out by the contract file).
  • runs never executed: how many planned runs did not run, and why.
  • inconclusive runs: how many runs ended inconclusive; they count neither as pass nor as fail.
  • declared capability not exercised: streaming, push notifications, extended card, extensions.
  • interface not used: other bindings or URLs the card lists.
  • probes: no probe test cases exist yet (stage 4).
  • semantic correctness: Layer 1 checks structure, state, output modes and latency; meaning needs Layer 2 (stage 4).
  • production endpoint: tests ran against the declared sandbox only.
  • card observation: fields the A2A specification marks REQUIRED that the card omits.

Loading into the workspace

Open Import, then drop result.json together with the files from transcripts/ (or select the whole report folder). The workspace validates the format, keeps the transcript text byte for byte so hashes can be re-checked, and stores the bundle in this browser only. Nothing is uploaded anywhere. Remove a bundle from the workspace settings at any time.

Pilot

Run the first evaluation with us.

Suncly is in pilot. If your team approves A2A agents by hand today, tell us about one agent and one sandbox, and we will run the first evaluation together.

Or write to team@suncly.com.