Skip to content
Getting started

From nothing to a signed evaluation in five minutes.

The same commands work on Windows PowerShell, macOS and Linux; where a command differs, both forms are given. Repository access comes with the pilot.

1. Install

You need Python 3.12 or newer and git.

git clone https://github.com/Kristjanh2/Suncly.gitcd Sunclypython -m venv .venv.\.venv\Scripts\Activate.ps1   # macOS/Linux: source .venv/bin/activatepip install -e .suncly --version

2. See it work

suncly demo

The demo starts two bundled mock A2A agents on your machine, an honest one and a lying one, and evaluates both. Mock agents are sandboxes by construction, so the demo declares them as such and approves the drafted contract on your behalf as suncly-demo; suncly attest never does either. For each agent you see the card fetched and hashed, the contract approved, live progress per run with its verdict, the per-test-case counts, the decision line, the signature, the “What was NOT tested” list and the report folder.

Decision: flag. No policy is configured, so a human must review this result.

The honest agent passes every run. The lying agent's card declares text/plain output but it answers with JSON, so every run fails the output_modes check. Both end with the decision flag: without a configured policy, Suncly never approves or blocks anything. Reports land in ./suncly-reports/<attestation-id>/.

3. Evaluate your sandbox agent

Any A2A 1.0 agent over JSON-RPC works. To try the flow with a bundled agent first, start one in a second terminal:

python -m suncly.mock_agents honest --port 8701

Then, in the first terminal:

suncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox
  • --sandbox is required. Without it Suncly refuses to run anything. Suncly cannot verify that an endpoint is a sandbox; the flag is your declaration.
  • Suncly fetches the card, drafts one test case per declared example of each skill, and shows you the draft. Nothing runs until you approve it and enter your identifier, which is recorded as approved_by. In a script, pass --approve-as <identifier> instead.
  • A skill that declares no examples gets no test case: Suncly never invents input. The draft says so, and the report lists the skill under “What was NOT tested”.
  • Each test case runs --runs times (default 5), each from a separate Runner process, within a budget of attempts (default twice the planned runs). Both numbers are shown before anything runs.
  • Every run gets a deterministic verdict; the Policy engine records flag; the attestation is signed with your deployment key; the report folder is written.
  • The same card gets the same approved contract next time. A changed card gets a new draft that needs a new approval.

Exit code 0 is not an approval

It means the attestation completed and was signed. The output says so. All codes are in the command reference.

4. Use a credential

If the sandbox needs an Authorization header, put its full value in the environment before running. Only the Runner process reads it, and it is redacted from every transcript before anything leaves the Runner.

$env:SUNCLY_AGENT_AUTHORIZATION = "Bearer <token>"# macOS/Linux: export SUNCLY_AGENT_AUTHORIZATION="Bearer <token>"

5. Read the report

FileContents
report.htmlThe report for a reviewer. Self-contained; opens offline.
report.mdThe same content as Markdown.
result.jsonThe evidence bundle: attestation, runs, decisions, card version, contract, results, the signed payload and the public key.
transcripts/<run-id>.jsonOne redacted transcript per run, with its Layer 1 checks.

The report always states what was NOT tested: skills without a test case, runs never executed, inconclusive runs, declared capabilities no test exercised, interfaces not used, probes and semantic checks that need later stages, and the production endpoint itself.

6. Verify the attestation

suncly verify suncly-reports/<attestation-id>

The verifier checks the signature over the signed payload, that the card hash matches the stored card, that every transcript file matches its signed hash, that the recorded decision is the signed one, and that the per-test-case counts match the recorded runs. Change one byte of a transcript or of result.json and it tells you which check failed and why. The public key travels in result.json; pass --public-key to verify against a key you obtained out of band.

7. Edit the test cases

Export the draft, edit it, and run with the file:

suncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox --export-draft contract.json# edit contract.json: add required_fields, a response_schema, a tighter latency_limit_mssuncly attest http://127.0.0.1:8701/.well-known/agent-card.json --sandbox --contract contract.json --approve-as <you>

The file format is in the command reference. The imported contract becomes a new version; its approval is recorded like any other.

8. Load it into the workspace

The review workspace reads report folders in this browser: drop result.json and the transcripts folder, and you get the overview, the agent's history, the evidence per run, in-browser verification and a comparison with the previous evaluation. Nothing is uploaded; the workspace stores bundles in this browser only.

9. Use Postgres instead of the file store

By default evidence lives in files under ~/.suncly/store. To use Postgres, set DATABASE_URL and apply the migration:

$env:DATABASE_URL = "postgresql://user:password@host:5432/suncly"   # never commit this valuesuncly db migratesuncly db check

That is the only change. Transcripts stay on local disk under ~/.suncly/transcripts until an object storage adapter exists.

10. When something goes wrong

suncly doctor http://127.0.0.1:8701/.well-known/agent-card.json

suncly doctor checks the Python version, the deployment key, the store configuration and whether the card is reachable. Every error Suncly prints says what happened, why, and what to do next; add --debug for a traceback.

Every setting has a default, can be set in ~/.suncly/config.toml, and can be overridden by an environment variable. --home moves the whole state folder, which is useful for isolated runs:

suncly --home ./tmp-home demo --reports-dir ./tmp-reports
Pilot

Run the first evaluation with us.

Suncly is in pilot. If your team approves A2A agents by hand today, tell us about one agent and one sandbox, and we will run the first evaluation together.

Or write to team@suncly.com.