Skip to content
Sample data · fictional agent · produced by the real code path

One agent, three evaluations, nothing hidden.

Harbor Returns Agent is fictional. The evaluations below were produced by running Suncly against a local mock of it, with the contract approved as a fictional reviewer. The hashes, verdicts, transcripts and signatures are real output of the software; the agent, the orders and the people are not real. Nothing here is a customer evaluation.

Sample data, kept separate from real evaluations

Sample data. A fictional agent evaluated by the real Suncly code path on a developer machine. Not a customer evaluation. Loading it into the workspace marks every record as a sample.

Three evaluations of one agent

Same card hash, so the same approved contract runs again. start-return now stops at TASK_STATE_INPUT_REQUIRED on every run. The second order-status example fails two runs in five: the agent is inconsistent, which five repetitions make visible and one manual prompt would not.

Sample dataattestation f888d7d2 · trigger manual

Harbor Returns Agent

Card version 2.3.1 · owner Harbor Commerce platform team (fictional) · started 04 Oct 2026, 21:04 UTC

Evaluation status

Completed

Every planned run was judged, the card was unchanged at the end, a decision was recorded and the attestation was signed.

Approval status

Flag: human review · policyRisk medium

Automatic decision by the Policy engine. No policy is configured, so a human must review this result. Approve and block are never produced automatically in this version.

Needs a human decision

The Policy engine recorded flag. A reviewer has to decide on this evidence.
Pass8runs with every check passed
Fail7runs with a failed check
Inconclusive5never counted as a pass
Runs recorded20 / 20every planned run recorded

Inconsistent behaviour on 1 test case

The same input passed on some runs and failed on others: “Has order 48213 shipped yet?”. One manual prompt would have shown only one of those outcomes.

1 test case undecided

Every run was inconclusive: “How much will I get back if I return order 48213?”. Open the run evidence to see which check Layer 1 could not decide.

Results per test case

Pass, fail and inconclusive are shown side by side. Suncly never combines them into a score.

Results per test case: pass, fail and inconclusive counts
SkillTest inputSignalPassFailInconclusive
Order statusorder-status“Where is order 48213?”Passing500
Order statusorder-status“Has order 48213 shipped yet?”Inconsistent320
Start a returnstart-return“Start a return for order 48213, item 2”Failing050
Refund estimaterefund-estimate“How much will I get back if I return order 48213?”Undecided005
Totals across 4 test cases. Counts, not a score.875

What was not tested

Straight from the report. An approval on this evidence is an approval with these gaps.

  • skill without test caseskill 'cancel-order' has no test case; invariant 1 (one test case per declared skill) is not satisfied for it
  • inconclusive runs5 run(s) ended inconclusive; they count neither as pass nor as fail
  • declared capability not exercisedthe card declares streaming; no test case exercises it
  • probesno probe_undeclared, probe_injection or probe_failure test cases exist yet (stage 4)
  • semantic correctnessLayer 1 checks structure, state, output modes and latency; whether the content of each answer is correct needs Layer 2 (stage 4)
  • production endpointtests ran against the declared sandbox or dry-run endpoint http://127.0.0.1:53448/rpc; the production endpoint itself was not tested (DR-006)

What changed between evaluation 1 and 2

The card hash is identical, so the same approved contract ran both times. The counts are comparable; the agent is not the same.

Earlier

attestation d9b37b41

04 Oct 2026, 21:04 UTC

CompletedFlag: human review · policy

Later

attestation f888d7d2

04 Oct 2026, 21:04 UTC

CompletedFlag: human review · policy
  • Agent CardUnchangedSame card_hash. An unchanged card is not proof of an unchanged agent; the results below are.
  • Test conditionsSame contractContract version 1, approved by m.lind@harbor.example.
  • Evaluation statuscompleted
  • Test cases2 regressed · 0 improved2 unchanged, 0 mixed, 0 without runs on one side, 0 added, 0 removed
Per test case comparison of pass, fail and inconclusive counts
SkillTest inputEarlier (pass / fail / inconclusive)Later (pass / fail / inconclusive)Change
Order statusorder-status“Where is order 48213?”5 / 0 / 05 / 0 / 0
Unchanged
Order statusorder-status“Has order 48213 shipped yet?”5 / 0 / 03 / 2 / 0
Regressed
Start a returnstart-return“Start a return for order 48213, item 2”5 / 0 / 00 / 5 / 0
Regressed
Refund estimaterefund-estimate“How much will I get back if I return order 48213?”0 / 0 / 50 / 0 / 5
Unchanged

Computed in this workspace from the two signed bundles. Suncly's evidence store has no comparison operation; what counts as a “drop” for the Policy engine is an open question (OQ-PO2).

The reviewer's decision (sample)

Human decisions are not yet recorded inside Suncly (planned, stage 5). This is what a reviewer records in the workspace and exports with the signed report.

Block · reviewerHuman decision, outside the signed evidenceSample

start-return no longer completes from a single message: all five runs stop at input-required, while evaluation 1 completed five of five against the same card. order-status is inconsistent on the shipping question (3 pass, 2 fail with upstream timeouts). refund-estimate remains inconclusive pending a semantic check. Block until the returns flow completes without extra input; re-evaluate with the same contract.

m.lind@harbor.example · attestation f888d7d2-c378-4dd3-9b2f-71d6f266b92b

What remains untested, in the reviewer's words

  • cancel-order: the card declares no examples, so there is no test case. Ask the team for examples or write them into the contract file.
  • Whether refund amounts are correct: needs the model-based judge (stage 4). Until then the check is inconclusive by design.
  • Streaming: declared by the card, never exercised.
  • Prompt injection and behaviour outside the declared skills: probes are not implemented.
  • The production endpoint: tests ran against the declared sandbox only.

Work with these in the review workspace

The workspace is where your own report folders go. Load the sample to try the overview, the agent history, the comparison and the review note. Everything loaded from here carries a sample badge.

Pilot

Run the first evaluation with us.

Suncly is in pilot. If your team approves A2A agents by hand today, tell us about one agent and one sandbox, and we will run the first evaluation together.

Or write to team@suncly.com.