The report folder
suncly-reports/<attestation-id>/ holds four kinds of file:
result.json: the evidence bundle, formatsuncly-result/1.transcripts/<run-id>.json: one evidence document per recorded run, byte for byte as stored.report.md: the human-readable report as Markdown.report.html: the same report, self-contained, no network needed.
result.json
The bundle without the inline transcripts (they are the files next to it). Top-level fields, with the data model's names:
| Field | Contents |
|---|---|
agent | id (derived from the card URL), name, owner, risk_level. |
card_version | card_hash, raw_json exactly as fetched, fetched_at. |
parsed_card | The card as JSON and as the parsed model: skills, interfaces, capabilities, modes. |
card_url | The URL that was fetched. |
contract, test_cases | The approved contract (version, status, approved_by, approved_at) and its test cases with input and criteria. |
attestation | Status, trigger, timestamps, budget_limit, cost_total, signature, signing_key_id. |
runs | Every recorded run (verdict, judge_layer, latency_ms, attempt, transcript_ref) with the document_hash of its transcript file. |
decisions | Every decision record, oldest first. The first is the Policy engine's. |
results | Per test case: pass_count, fail_count, inconclusive_count. Never combined. |
planned_runs, not_executed | How many runs were planned and which never ran, with the reason: budget, runner_crashed or withheld. |
card_recheck | What the final re-fetch found: unchanged, changed, unavailable or not performed. |
not_tested | The “What was NOT tested” list, category and detail. |
sandbox_declared, drafter_name | The sandbox declaration and the source of the test plan. |
signer_public_key, signature_payload | The public key (base64url) and the exact payload that was signed. |
proposals | The open-question proposals in effect for this run. |
Transcript files
Each file is one evidence document with two parts:
transcript: the target URL and binding, timestamps,latency_ms, theoutcome(responded_task, responded_message, unreachable, timeout, protocol_error), the final task state and final response, every request and response exchanged with headers and bodies after redaction, and a redaction summary naming the rules that fired.judgement: the verdict, the judge layer, a summary, and every check with its name, itspassedvalue (true, false, or null when Layer 1 cannot decide) and a detail such as “expected TASK_STATE_COMPLETED, observed TASK_STATE_INPUT_REQUIRED”.
The file is written once in canonical JSON (RFC 8785). Its SHA-256 is what the signature covers.
What the signature covers
{
"payload_version": 1,
"attestation_id": "…",
"card_hash": "sha256:…",
"contract": {"id": "…", "version": 1},
"results": [
{"test_case_id": "…", "skill_id": "order-status", "kind": "skill",
"pass": 5, "fail": 0, "inconclusive": 0}
],
"transcript_hashes": {"<run-id>": "sha256:…"},
"decision": {"outcome": "flag", "policy_version": "unconfigured"}
}The payload is canonicalized with RFC 8785 and signed with the deployment's Ed25519 key; the signature is stored as ed25519:<base64url> and signing_key_id is a SHA-256 fingerprint of the public key. Failed and invalidated attestations are signed with "decision": null. The signature does not cover a later human decision, the report prose, or anything about the agent's future behaviour.
Verifying
suncly verify suncly-reports/<attestation-id> [--public-key <base64url>]Checks, in order: signature present; public key matches signing_key_id; signature valid over the payload; card_hash recomputed from raw_json equals the payload and the record; attestation and contract ids match; every transcript file hashes to its signed value; the recorded decision equals the signed one; the signed per-test-case counts match the recorded runs. The workspace runs the same checks in the browser, with the CLI as the reference.
What was NOT tested
Every report carries this section. The categories the current version can emit:
- skill without test case: a declared skill has no test case (no examples, or left out by the contract file).
- runs never executed: how many planned runs did not run, and why.
- inconclusive runs: how many runs ended inconclusive; they count neither as pass nor as fail.
- declared capability not exercised: streaming, push notifications, extended card, extensions.
- interface not used: other bindings or URLs the card lists.
- probes: no probe test cases exist yet (stage 4).
- semantic correctness: Layer 1 checks structure, state, output modes and latency; meaning needs Layer 2 (stage 4).
- production endpoint: tests ran against the declared sandbox only.
- card observation: fields the A2A specification marks REQUIRED that the card omits.
Loading into the workspace
Open Import, then drop result.json together with the files from transcripts/ (or select the whole report folder). The workspace validates the format, keeps the transcript text byte for byte so hashes can be re-checked, and stores the bundle in this browser only. Nothing is uploaded anywhere. Remove a bundle from the workspace settings at any time.