Skip to content
Suncly Lab

Agents that misbehave on purpose.

The Suncly Lab is where tests get built. Eleven mock A2A agents ship with the package; each models a behaviour the evaluation has to handle. They run on your machine and are sandboxes by construction.

Eleven small forms in a row at golden hour. Most cast true shadows; one casts a wrong shadow, one casts two, one casts none.

The eleven

The eleven mock agents, from BEHAVIOURS in src/suncly/mock_agents/behaviours.py
NameWhat it doesWhat Suncly reports
honestDoes what its card says.every run passes
honest-asyncCompletes tasks asynchronously; the client has to poll GetTask.every run passes after polling GetTask
unreachableServes a card whose endpoint refuses connections.no run passes; runs are inconclusive
lyingDeclares skills and output modes it does not honour.every run fails the output_modes check
flakySucceeds on odd calls and fails on even calls.odd calls pass, even calls fail
slowAnswers correctly, but slowly.every run fails the latency limit
direct-messageAnswers with a direct Message instead of a Task.handled without error; not a pass
interruptedAlways asks for more input.recorded as fail; never a pass
leakyEchoes the Authorization header back in its answer.the credential appears in no transcript, report or log
card-changerChanges its card while an attestation runs.the attestation ends invalidated
no-examplesOne of its skills declares no examples.the skill without examples is listed as not tested

Run a mock agent

Start one in a second terminal, then evaluate it from its card URL. Any of the eleven names works in place of honest.

python -m suncly.mock_agents honest --port 8701suncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox

The sample generator

frontend/scripts/make-sample.py runs the real attestation code path against a fictional local agent and writes the three sample bundles the site shows. It fails if the fake credential appears in any transcript.

.venv/bin/python frontend/scripts/make-sample.py

Planned

  • Prompt-injection, undeclared-behaviour and failure-handling probes (stage 4)
  • Model-based judging for criteria Layer 1 cannot decide (stage 4)

Example contract file

Export the draft, edit it, run with the file. The example below is the sample agent's contract: four test cases, one skill without a test case, structural criteria only. The file format is documented in the command reference.

{
  "suncly_contract_file": 1,
  "card_hash": "sha256:c25be2adb553bc2efce2a1506c0b19194c8299c030e1f0b265890fbbeb96850b",
  "agent_name": "Harbor Returns Agent",
  "skills_without_test_case": [
    "cancel-order"
  ],
  "test_cases": [
    {
      "skill_id": "order-status",
      "kind": "skill",
      "input": {
        "text": "Where is order 48213?"
      },
      "criteria": {
        "final_state": "TASK_STATE_COMPLETED",
        "latency_limit_ms": 2000,
        "response_present": true,
        "output_modes": [
          "text/plain"
        ],
        "required_fields": [
          "/artifacts/0/parts/0/text"
        ],
        "model_checks": [],
        "accept_direct_message": false
      }
    },
    {
      "skill_id": "order-status",
      "kind": "skill",
      "input": {
        "text": "Has order 48213 shipped yet?"
      },
      "criteria": {
        "final_state": "TASK_STATE_COMPLETED",
        "latency_limit_ms": 2000,
        "response_present": true,
        "output_modes": [
          "text/plain"
        ],
        "required_fields": [
          "/artifacts/0/parts/0/text"
        ],
        "model_checks": [],
        "accept_direct_message": false
      }
    },
    {
      "skill_id": "start-return",
      "kind": "skill",
      "input": {
        "text": "Start a return for order 48213, item 2"
      },
      "criteria": {
        "final_state": "TASK_STATE_COMPLETED",
        "latency_limit_ms": 2000,
        "response_present": true,
        "output_modes": [
          "text/plain"
        ],
        "required_fields": [
          "/artifacts/0/parts/0/text"
        ],
        "model_checks": [],
        "accept_direct_message": false
      }
    },
    {
      "skill_id": "refund-estimate",
      "kind": "skill",
      "input": {
        "text": "How much will I get back if I return order 48213?"
      },
      "criteria": {
        "final_state": "TASK_STATE_COMPLETED",
        "latency_limit_ms": 2000,
        "response_present": true,
        "output_modes": [
          "text/plain"
        ],
        "required_fields": [
          "/artifacts/0/parts/0/text"
        ],
        "model_checks": [
          "the amount equals the order total minus the restocking fee"
        ],
        "accept_direct_message": false
      }
    }
  ]
}