Agents that misbehave on purpose.
The Suncly Lab is where tests get built. Eleven mock A2A agents ship with the package; each models a behaviour the evaluation has to handle. They run on your machine and are sandboxes by construction.

The eleven
| Name | What it does | What Suncly reports |
|---|---|---|
honest | Does what its card says. | every run passes |
honest-async | Completes tasks asynchronously; the client has to poll GetTask. | every run passes after polling GetTask |
unreachable | Serves a card whose endpoint refuses connections. | no run passes; runs are inconclusive |
lying | Declares skills and output modes it does not honour. | every run fails the output_modes check |
flaky | Succeeds on odd calls and fails on even calls. | odd calls pass, even calls fail |
slow | Answers correctly, but slowly. | every run fails the latency limit |
direct-message | Answers with a direct Message instead of a Task. | handled without error; not a pass |
interrupted | Always asks for more input. | recorded as fail; never a pass |
leaky | Echoes the Authorization header back in its answer. | the credential appears in no transcript, report or log |
card-changer | Changes its card while an attestation runs. | the attestation ends invalidated |
no-examples | One of its skills declares no examples. | the skill without examples is listed as not tested |
Run a mock agent
Start one in a second terminal, then evaluate it from its card URL. Any of the eleven names works in place of honest.
python -m suncly.mock_agents honest --port 8701suncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox
The sample generator
frontend/scripts/make-sample.py runs the real attestation code path against a fictional local agent and writes the three sample bundles the site shows. It fails if the fake credential appears in any transcript.
.venv/bin/python frontend/scripts/make-sample.pyPlanned
- Prompt-injection, undeclared-behaviour and failure-handling probes (stage 4)
- Model-based judging for criteria Layer 1 cannot decide (stage 4)
Example contract file
Export the draft, edit it, run with the file. The example below is the sample agent's contract: four test cases, one skill without a test case, structural criteria only. The file format is documented in the command reference.
{
"suncly_contract_file": 1,
"card_hash": "sha256:c25be2adb553bc2efce2a1506c0b19194c8299c030e1f0b265890fbbeb96850b",
"agent_name": "Harbor Returns Agent",
"skills_without_test_case": [
"cancel-order"
],
"test_cases": [
{
"skill_id": "order-status",
"kind": "skill",
"input": {
"text": "Where is order 48213?"
},
"criteria": {
"final_state": "TASK_STATE_COMPLETED",
"latency_limit_ms": 2000,
"response_present": true,
"output_modes": [
"text/plain"
],
"required_fields": [
"/artifacts/0/parts/0/text"
],
"model_checks": [],
"accept_direct_message": false
}
},
{
"skill_id": "order-status",
"kind": "skill",
"input": {
"text": "Has order 48213 shipped yet?"
},
"criteria": {
"final_state": "TASK_STATE_COMPLETED",
"latency_limit_ms": 2000,
"response_present": true,
"output_modes": [
"text/plain"
],
"required_fields": [
"/artifacts/0/parts/0/text"
],
"model_checks": [],
"accept_direct_message": false
}
},
{
"skill_id": "start-return",
"kind": "skill",
"input": {
"text": "Start a return for order 48213, item 2"
},
"criteria": {
"final_state": "TASK_STATE_COMPLETED",
"latency_limit_ms": 2000,
"response_present": true,
"output_modes": [
"text/plain"
],
"required_fields": [
"/artifacts/0/parts/0/text"
],
"model_checks": [],
"accept_direct_message": false
}
},
{
"skill_id": "refund-estimate",
"kind": "skill",
"input": {
"text": "How much will I get back if I return order 48213?"
},
"criteria": {
"final_state": "TASK_STATE_COMPLETED",
"latency_limit_ms": 2000,
"response_present": true,
"output_modes": [
"text/plain"
],
"required_fields": [
"/artifacts/0/parts/0/text"
],
"model_checks": [
"the amount equals the order total minus the restocking fee"
],
"accept_direct_message": false
}
}
]
}