What Suncly accesses
Suncly touches three things, and nothing else on your network:
- The Agent Card URL you give it. Fetched over https (plain http is accepted only for loopback addresses, where local sandboxes run), with a timeout and a size limit (1 MB by default). The body is kept byte for byte as
card_version.raw_json. - The agent endpoint named in the card. The Runner refuses any request whose host is not the target agent's host; one test enforces this in the one place that makes HTTP calls. It speaks A2A 1.0 over JSON-RPC:
SendMessage, thenGetTaskuntil a terminal or interrupted state. - One credential, if the sandbox needs one. The full value of the Authorization header, read from the environment variable
SUNCLY_AGENT_AUTHORIZATION.
Suncly never calls a registry, a gateway, an identity provider or Suncly's own servers: there are none in this version.
Where credentials are handled
The Runner is the only component that holds the agent credential. It runs as a separate process started for each run; the job it receives on stdin carries no credential (a test checks the job schema has no such field), and the process reads the variable itself. A test proves that exactly one source module names the variable.
Before a transcript leaves the Runner it is redacted as data, over every string in every request and response: the credential value and its token part, the values of sensitive header keys (authorization, cookie, set-cookie, x-api-key, api-key, x-auth-token, proxy-authorization), documented token patterns (bearer tokens, JWTs, common provider key prefixes) and key=value secrets. Each rule that fired is named in the transcript's redaction summary. If redaction itself fails, the transcript is withheld and the run is not recorded. Nothing from the Runner's stderr is surfaced.
Credentials never appear in command lines, request bodies, report files or logs. The bundled leaky mock agent, which echoes the header back, is part of the test suite to prove it.
What goes to external model providers
Nothing, in this version. The test plan is drafted deterministically from the card's declared examples, and every verdict is deterministic. Suncly makes no call to any model provider and needs no model key.
Later stages add a model-based drafter and a model-based judge (Layer 2). They are designed to use your own model keys, with the judge model pinned by version so that a verdict measures the agent, not drift in the judge. Which provider is yours to choose. Until those stages exist, every report lists semantic correctness under what was not tested.
What evidence is stored, and where
Everything is written on the machine that runs Suncly, under the Suncly home folder (~/.suncly by default, or SUNCLY_HOME) and the reports folder (./suncly-reports by default):
store/: the seven entity records (agent, card version, contract, test cases, attestation, runs, decisions) as files, or in Postgres whenDATABASE_URLis set.transcripts/: one evidence document per run: the redacted transcript and its judgement. Write-once; its SHA-256 is in the signed payload.keys/: the deployment's Ed25519 signing key. The private key is never printed; the public key travels in every result.json.suncly-reports/<attestation-id>/:result.json, the transcript files,report.mdand a self-containedreport.html.
The store is append-only: run and decision records cannot be updated or deleted, and the Postgres migration enforces this with triggers. Corrections are new records.
Retention and deletion
Suncly defines no retention period and has no delete operation, because the evidence store is append-only by design. Evidence stays where Suncly wrote it until you remove the files or drop the database. Because the software runs on your machine or in your network, you decide where the folders live, who can read them, and when they go.
Owner input pending
Deployment options
- A developer machine. Python 3.12 or newer;
pip install -e .; the file store. - A CI runner. The CLI with
--approve-as,--jsonand documented exit codes; the report folder archived as a build artefact. - A server inside your network. The same CLI with
DATABASE_URLpointing at your Postgres (any PostgreSQL 13 or newer; the schema is applied withsuncly db migrate). Transcripts stay on that machine's disk until an object-storage adapter exists.
There is no hosted Suncly service, no account system and no data sent to Suncly in this version.
Sandbox requirements
Nothing runs unless the caller declares the endpoint a sandbox or dry-run endpoint (--sandbox). The Runner refuses undeclared targets a second time. Tests send the card's own example inputs, and later probes will deliberately send injected instructions and failure conditions; against production that could book, pay or delete something real. Suncly cannot verify that an endpoint is a sandbox; the declaration is yours and the report records it as a declaration.
The non-negotiable rules
Seven decision records the system is built around. Each one has a test named after it.
- Idempotent runs. Every run has a deterministic key; retries never double count.
- Evidence is immutable. Records are never edited; corrections are new records.
- Secrets never leave the Runner. Transcripts are redacted before storage.
- The judge model is pinned. A new judge model is a configuration change, not drift (applies once Layer 2 exists).
- Budget caps live in the Orchestrator. The component that starts runs is the one that stops them.
- Tests hit a sandbox or dry-run endpoint. Never production.
- Reports state what was NOT tested. No score hides the gaps.
Evaluation limitations
- A pass means every deterministic check passed on the sandbox endpoint during this evaluation. It is not a guarantee of future behaviour, and it says nothing about the production endpoint, which is never called.
- The signature proves the evidence was not altered after signing and which deployment produced it. It does not prove the agent is correct.
- Decisions are flag-only in this version. Suncly never approves or blocks automatically; a human records the decision outside Suncly, and this workspace helps route it.
- Suncly cannot verify that an endpoint is a sandbox. Your --sandbox declaration is recorded as a declaration.
- Tests come from the examples the card declares. A skill without examples is reported as not tested, because Suncly never invents input.
- Prompt-injection, undeclared-behaviour and failure-handling probes are planned, not implemented. Every report lists them under what was not tested.
The full architecture, with each component's “must never” list and failure behaviour, is in the repository documents listed on the documentation page.