Glossary
The words behind an evaluation, defined once.
Each term starts with a one-sentence definition you can quote, then what it means inside Suncly. Where a term names something planned rather than built, it says so.
- A2A protocol (Agent2Agent)
- The A2A protocol is an open protocol in which an AI agent publishes an Agent Card describing itself and its skills, and answers requests as Tasks or Messages over bindings such as JSON-RPC.
- Suncly evaluates agents through A2A version 1.0 over JSON-RPC: it sends each test input as a Message and follows the Task to a terminal or interrupted state. The protocol defines how skills are described; it has no mechanism for checking that the agent performs them, which is the gap Suncly fills.
- See also:Agent CardSkill
- Agent Card
- An Agent Card is the JSON document an A2A agent publishes, usually at /.well-known/agent-card.json, naming the agent, its version, its interfaces, its capabilities and the skills it offers.
- Suncly keeps the card byte for byte, hashes it, and treats the skills list as the claims under test. A signed card proves who published it and that it was not altered; it says nothing about behaviour.
- See also:card_hashSkill
- Skill
- A skill is one capability an Agent Card declares, with an id, a name, a description, tags and optionally example inputs.
- Suncly drafts one test case per declared example of each skill. A skill with no examples gets no test case and is reported under what was not tested, because Suncly never invents input.
- See also:Test caseWhat was NOT tested
- AI agent evaluation
- AI agent evaluation is the practice of testing what an AI agent actually does against what it claims, before the agent is approved for use.
- In Suncly an evaluation is an attestation: a contract of test cases, approved by a person, run repeatedly against a sandbox endpoint, judged deterministically, signed, and reported with its coverage gaps. It is behavioural evaluation, not protocol conformance testing and not runtime monitoring.
- See also:AttestationBehavioural evaluation
- Behavioural evaluation
- Behavioural evaluation sends real inputs to an agent and judges the responses, as opposed to reading its documentation or checking its protocol conformance.
- Suncly's judge checks that each response is a well-formed A2A Task or Message, reached the expected final state, answered within the latency limit, carries content, uses the declared output modes, and satisfies required fields and schema. Whether an answer is correct in meaning needs the model-based judge of a later stage.
- Attestation
- An attestation is one execution of an approved contract against an agent: every test case run a configured number of times, each run judged, the results signed.
- Its status is queued, running, completed, failed, cancelled or invalidated. Only a completed attestation carries a decision. Failed (budget or card re-fetch) and invalidated (card changed) attestations are signed with no decision, and no decision is never an approval.
- See also:ContractRunDecision: approve, flag, block
- Contract
- A contract is a versioned, human-approved set of test cases for one version of an Agent Card.
- Nothing runs until a person approves the contract and their identifier is recorded as approved_by. Approved contracts are immutable; an edit creates a new version, and a changed card needs a new contract.
- See also:Test casecard_hash
- Test case
- A test case is one input to send to the agent and the criteria its response must satisfy.
- Criteria in the current version: expected final task state, latency limit, response present, declared output modes, required fields as JSON pointers, an optional JSON Schema, and model checks that stay inconclusive until the model judge exists.
- See also:Run
- Run
- A run is one execution of one test case, identified by a deterministic key of attestation, test case and attempt number.
- Retries reuse the key, so a crashed run never counts twice. Each run keeps its redacted transcript, verdict, latency, cost and timestamps.
- See also:Verdict: pass, fail, inconclusiveTranscript
- Verdict: pass, fail, inconclusive
- A verdict is the judge's outcome for one run: pass when every check passed, fail when any check failed, inconclusive when a check could not be decided.
- Inconclusive is never counted as a pass. Counts are reported per test case and never combined into a score.
- See also:InconclusiveJudge (Layer 1 and Layer 2)
- Inconclusive
- Inconclusive is the verdict for a run whose outcome could not be decided: the agent was unreachable, the transcript could not be read, or a criterion needs a judge that does not exist yet.
- Suncly reports inconclusive runs separately and lists them under what was not tested. They never count as passes.
- Judge (Layer 1 and Layer 2)
- The judge assigns a verdict to each run. Layer 1 is deterministic; Layer 2 is a model pinned by version with a fixed rubric, used only for criteria Layer 1 cannot decide.
- Only Layer 1 exists in the current version. Every report says that semantic correctness was not tested.
- Decision: approve, flag, block
- A decision is the Policy engine's recorded outcome for a completed attestation: approve, flag for human review, or block.
- Without a configured policy the only outcome is flag, so a human reviews every result. Automatic decisions are signed; a later human decision is a second record and is never written over the first.
- See also:Policy engineRisk level
- Policy engine
- The Policy engine aggregates results per test case, applies the customer's policy for the agent's risk level, records the decision and signs the attestation.
- Thresholds are a customer configuration that does not exist yet; Suncly ships no numbers of its own.
- Risk level
- The risk level (low, medium, high) records how much a human must be involved in approving an agent.
- By design, low-risk agents can be approved automatically on pass, medium-risk agents need a human on any drop, and high-risk agents need human sign-off every time. In the current version the level is recorded but every decision is flag.
- Probe
- A probe is a test case that exercises behaviour the card does not declare: undeclared capabilities, injected instructions, or failure handling.
- Probes are a planned stage. Every report lists them under what was not tested until they exist.
- Sandbox declaration
- The sandbox declaration (--sandbox) is the caller's statement that the endpoint under test is a sandbox or dry-run endpoint.
- Nothing runs without it, so tests cannot book, pay or delete anything real. Suncly cannot verify a sandbox; the report records the declaration as a declaration, and the production endpoint is listed as not tested.
- card_hash
- card_hash is the SHA-256 of the Agent Card's RFC 8785 canonical form without the signatures field.
- It pins an evaluation to an exact version of the claims. An unchanged hash reuses the approved contract, which makes two evaluations comparable. An unchanged card is not proof of an unchanged agent.
- Signed evidence
- Signed evidence is the attestation's Ed25519 signature over the card hash, the contract version, the per-test-case counts, the hash of every transcript and the decision.
- suncly verify re-checks a report folder offline with the public key. A valid signature proves the evidence was not altered after signing and which deployment produced it; it does not prove the agent is correct or will behave the same tomorrow.
- See also:TranscriptAttestation
- Transcript
- A transcript is the complete record of one run: every request and response, the final task state, latency, outcome and the checks applied.
- Transcripts are redacted inside the Runner before they are stored: the credential, sensitive headers and known token patterns are removed. Each transcript's hash is in the signed payload.
- What was NOT tested
- What was NOT tested is the mandatory report section listing every gap in the evidence.
- Categories: skills without a test case, runs never executed, inconclusive runs, declared capabilities not exercised, interfaces not used, probes and semantic checks that do not exist yet, the production endpoint, and card fields the specification requires but the card omits.
- Budget
- The budget is the maximum number of attempts an attestation may make against the agent, retries included.
- When it is reached no new run starts, the attestation ends failed, the runs never executed are listed, and no decision is made.
- Runner
- The Runner is the isolated process that calls the agent. It is the only component that holds the agent credential.
- It refuses any host other than the target, reads the credential from one environment variable, redacts transcripts before returning them, and runs only against a declared sandbox.
Pilot
Run the first evaluation with us.
Suncly is in pilot. If your team approves A2A agents by hand today, tell us about one agent and one sandbox, and we will run the first evaluation together.
Or write to team@suncly.com.