1. Install
You need Python 3.12 or newer and git.
git clone https://github.com/Kristjanh2/Suncly.gitcd Sunclypython -m venv .venv.\.venv\Scripts\Activate.ps1 # macOS/Linux: source .venv/bin/activatepip install -e .suncly --version
2. See it work
suncly demoThe demo starts two bundled mock A2A agents on your machine, an honest one and a lying one, and evaluates both. Mock agents are sandboxes by construction, so the demo declares them as such and approves the drafted contract on your behalf as suncly-demo; suncly attest never does either. For each agent you see the card fetched and hashed, the contract approved, live progress per run with its verdict, the per-test-case counts, the decision line, the signature, the “What was NOT tested” list and the report folder.
Decision: flag. No policy is configured, so a human must review this result.
The honest agent passes every run. The lying agent's card declares text/plain output but it answers with JSON, so every run fails the output_modes check. Both end with the decision flag: without a configured policy, Suncly never approves or blocks anything. Reports land in ./suncly-reports/<attestation-id>/.
3. Evaluate your sandbox agent
Any A2A 1.0 agent over JSON-RPC works. To try the flow with a bundled agent first, start one in a second terminal:
python -m suncly.mock_agents honest --port 8701Then, in the first terminal:
suncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox--sandboxis required. Without it Suncly refuses to run anything. Suncly cannot verify that an endpoint is a sandbox; the flag is your declaration.- Suncly fetches the card, drafts one test case per declared example of each skill, and shows you the draft. Nothing runs until you approve it and enter your identifier, which is recorded as
approved_by. In a script, pass--approve-as <identifier>instead. - A skill that declares no examples gets no test case: Suncly never invents input. The draft says so, and the report lists the skill under “What was NOT tested”.
- Each test case runs
--runstimes (default 5), each from a separate Runner process, within a budget of attempts (default twice the planned runs). Both numbers are shown before anything runs. - Every run gets a deterministic verdict; the Policy engine records
flag; the attestation is signed with your deployment key; the report folder is written. - The same card gets the same approved contract next time. A changed card gets a new draft that needs a new approval.
Exit code 0 is not an approval
4. Use a credential
If the sandbox needs an Authorization header, put its full value in the environment before running. Only the Runner process reads it, and it is redacted from every transcript before anything leaves the Runner.
$env:SUNCLY_AGENT_AUTHORIZATION = "Bearer <token>"# macOS/Linux: export SUNCLY_AGENT_AUTHORIZATION="Bearer <token>"
5. Read the report
| File | Contents |
|---|---|
report.html | The report for a reviewer. Self-contained; opens offline. |
report.md | The same content as Markdown. |
result.json | The evidence bundle: attestation, runs, decisions, card version, contract, results, the signed payload and the public key. |
transcripts/<run-id>.json | One redacted transcript per run, with its Layer 1 checks. |
The report always states what was NOT tested: skills without a test case, runs never executed, inconclusive runs, declared capabilities no test exercised, interfaces not used, probes and semantic checks that need later stages, and the production endpoint itself.
6. Verify the attestation
suncly verify suncly-reports/<attestation-id>The verifier checks the signature over the signed payload, that the card hash matches the stored card, that every transcript file matches its signed hash, that the recorded decision is the signed one, and that the per-test-case counts match the recorded runs. Change one byte of a transcript or of result.json and it tells you which check failed and why. The public key travels in result.json; pass --public-key to verify against a key you obtained out of band.
7. Edit the test cases
Export the draft, edit it, and run with the file:
suncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox --export-draft contract.json# edit contract.json: add required_fields, a response_schema, a tighter latency_limit_mssuncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox --contract contract.json --approve-as <you>
The file format is in the command reference. The imported contract becomes a new version; its approval is recorded like any other.
8. Load it into the workspace
The review workspace reads report folders in this browser: drop result.json and the transcripts folder, and you get the overview, the agent's history, the evidence per run, in-browser verification and a comparison with the previous evaluation. Nothing is uploaded; the workspace stores bundles in this browser only.
9. Use Postgres instead of the file store
By default evidence lives in files under ~/.suncly/store. To use Postgres, set DATABASE_URL and apply the migration:
$env:DATABASE_URL = "postgresql://user:password@host:5432/suncly" # never commit this valuesuncly db migratesuncly db check
That is the only change. Transcripts stay on local disk under ~/.suncly/transcripts until an object storage adapter exists.
10. When something goes wrong
suncly doctor http://127.0.0.1:8701/.well-known/agent-card.jsonsuncly doctor checks the Python version, the deployment key, the store configuration and whether the card is reachable. Every error Suncly prints says what happened, why, and what to do next; add --debug for a traceback.
Every setting has a default, can be set in ~/.suncly/config.toml, and can be overridden by an environment variable. --home moves the whole state folder, which is useful for isolated runs:
suncly --home ./tmp-home demo --reports-dir ./tmp-reports