One protocol in. Any model behind it.
Suncly evaluates agents through the A2A protocol. It does not integrate with model providers or agent frameworks directly: if an agent publishes an A2A 1.0 Agent Card and answers over JSON-RPC, Suncly can test it, whatever built it. This page states exactly what is inspected and tested, and where coverage differs.
Protocol coverage
What the Runner speaks and what the Judge checks. Verified against src/suncly/runner and src/suncly/core.
| Item | Status | What Suncly does with it |
|---|---|---|
| A2A 1.0 over JSON-RPC (HTTP) | Available | SendMessage, then GetTask polling until a terminal or interrupted state. The first supportedInterfaces entry with JSONRPC and 1.0 is used. |
| Agent Card at /.well-known/agent-card.json | Available | Fetched over https (plain http for loopback sandboxes), size-limited, hashed with RFC 8785. Fields the spec requires but the card omits are reported. |
| Direct Message replies | Available with limits | Handled without error. A direct Message counts as completing a task only if the criteria say accept_direct_message. |
| Interrupted states (input or auth required) | Available with limits | Recorded as the final observed state and judged against the expected state. The Runner never invents input to continue. |
| gRPC and HTTP+JSON bindings | Planned | Not spoken. If the card lists them, the report lists them under interfaces not used. |
| Streaming, push notifications, extended cards, extensions | Planned | Not exercised. A card that declares them gets a 'declared capability not exercised' line in the report. |
| Authentication to the agent | Available with limits | One Authorization header value from an environment variable read only by the Runner. Other security schemes the card declares are kept as opaque fields and not negotiated. |
Model providers
Suncly is model-agnostic. It never calls a model provider in this version, and it does not need to know which model an agent uses. An agent built on any of these, or on none of them, is tested the same way through its A2A interface.
- Claude
- OpenAI GPT and Codex
- Grok
- Gemini
- Llama
- Mistral
- DeepSeek
- Qwen
- Cohere
- Open-source and fine-tuned models
Later stages add a model-based drafter and a model-based judge. Those will use your own model keys and a pinned judge model; the choice of provider is yours.
Agent frameworks and coding tools
What matters is whether the thing you want to evaluate exposes an A2A endpoint.
| Item | Status | What Suncly does with it |
|---|---|---|
| Any framework that serves an A2A 1.0 Agent Card and JSON-RPC endpoint | Available | Point suncly attest at the card URL of its sandbox deployment. |
| Agents that only expose a chat or vendor-specific API | Planned | Not reachable until they are wrapped in an A2A server. Suncly adds no other transports in this version. |
| Coding assistants and IDE agents | Planned | They are not A2A agents by themselves. To evaluate one, run it behind an A2A server in a sandbox and declare that sandbox. |
Integrations and interfaces
Where results go, and how Suncly is driven.
Command-line interface
Availablesuncly attest, demo, verify, keys init, db migrate, db check, doctor. Documented exit codes; --json for scripts.
HTTP API
PlannedPOST /attestations, GET /attestations/{id}, POST /contracts/{id}/approve, GET /agents/{id}/evidence, calling the same core library.
Stage 5. api.py is a placeholder; payloads and authentication are open questions (OQ-P1, OQ-P3).
CI gate
Available with limitsThe CLI runs in a pipeline today with --approve-as, --json and distinct exit codes. Exit code 0 means completed and signed, never approved.
A CI adapter that turns a decision into a pass or fail gate is stage 5. Because decisions are flag-only, no pipeline should gate on them yet.
Registry integration
PlannedAdapters that write approval status into your agent registry.
Stage 6. Target registries are not chosen (OQ-R5).
Bundled sandbox agents
Eleven mock A2A agents ship with the package for the demo and the end-to-end tests. Each one models a behaviour the evaluation has to handle correctly. They are sandboxes by construction and run on your machine.
| Mock agent | Behaviour | Expected result |
|---|---|---|
| honest | Does what its card says. | every run passes |
| honest-async | Completes asynchronously; the client has to poll GetTask. | every run passes after polling |
| lying | Declares text/plain output but answers with JSON. | every run fails output_modes |
| flaky | Succeeds on odd calls, fails on even calls. | an exact mix of pass and fail |
| slow | Answers correctly after the latency limit. | every run fails the latency limit |
| unreachable | Its endpoint refuses connections. | no run passes; runs are inconclusive |
| direct-message | Replies with a Message instead of a Task. | handled; not a pass |
| interrupted | Always asks for more input. | fail on final state; never a pass |
| leaky | Echoes the Authorization header back. | the credential appears nowhere |
| card-changer | Changes its card during the run. | the attestation ends invalidated |
| no-examples | One skill declares no examples. | that skill is listed as not tested |
Start one yourself: python -m suncly.mock_agents honest --port 8701, then evaluate it with suncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox.
Run the first evaluation with us.
Suncly is in pilot. If your team approves A2A agents by hand today, tell us about one agent and one sandbox, and we will run the first evaluation together.
Or write to team@suncly.com.