Repeatable evaluation. Accountable approval.
Suncly turns the approval of an AI agent into a recorded procedure: a reviewed test plan, repeated runs in a sandbox, deterministic verdicts, signed evidence, and an explicit statement of what was not tested.
- Available
- Implemented in the current version and covered by tests.
- Available with limits
- Implemented, with a limitation a buyer must know about.
- Planned
- On the roadmap. No date is promised, and nothing on this site treats it as current.
What is tested, how often, and how it is judged.
Reads the A2A Agent Card and hashes it
AvailableFetches the card over https, keeps it byte for byte, and computes card_hash as SHA-256 of its RFC 8785 form without the signatures field.
Why it matters: Every evaluation is pinned to an exact version of what the agent claimed.
Where this lives in the code
src/suncly/domain/card.py, core/cards.py; tests/unit/test_card.py
Tests derived from declared skills
Available with limitsOne test case per declared example of each skill, up to a cap. Criteria: completes, answers, uses the declared output modes, within a latency limit.
Why it matters: Nobody writes the first test plan by hand, and every declared skill with an example is covered.
Limit: Drafting is deterministic and uses no model. A skill without examples gets no test case and is listed as not tested. Model-drafted tests and probes are stage 2 and 4.
Where this lives in the code
src/suncly/core/contract_builder.py (DeterministicDrafter); tests/unit/test_contract_builder.py
Customer-defined criteria and known-answer checks
Available with limitsExport the draft, edit the contract file, run with it: expected final state, required fields (JSON pointers), a JSON Schema for the response, latency limit, declared output modes.
Why it matters: Your acceptance criteria become repeatable checks instead of a reviewer's memory.
Limit: Checks are structural. A criterion about the meaning of an answer (model_checks) makes the run inconclusive until Judge Layer 2 exists (stage 4).
Where this lives in the code
src/suncly/domain/criteria.py, domain/contract_file.py; docs/API.md 'Contract file'
Repeated runs that expose inconsistent behaviour
AvailableEach test case runs N times (default 5) from a separate Runner process, with a deterministic run key per attempt so retries never double count.
Why it matters: An agent that works four times out of five shows up as four passes and one fail, not as a pass.
Where this lives in the code
src/suncly/core/orchestrator.py; tests/e2e/test_mock_agents.py (flaky agent)
Deterministic verdicts: pass, fail, inconclusive
Available with limitsLayer 1 checks the response is a well-formed A2A Task or Message, reached the expected final state, answered within the latency limit, carries content, uses the declared output modes, and satisfies required fields and schema.
Why it matters: Verdicts are reproducible and explained check by check. Inconclusive is a third outcome, never counted as a pass.
Limit: No model-based judgement yet. Whether an answer is correct in meaning is listed under what was not tested.
Where this lives in the code
src/suncly/core/judge.py; tests/unit/test_judge.py
Prompt-injection, undeclared-behaviour and failure-handling probes
PlannedTest cases of kind probe_injection, probe_undeclared and probe_failure, drafted alongside skill tests and approved the same way.
Why it matters: Would show how the agent behaves outside what its card declares.
Not yet: Stage 4. The data model and the contract file format already carry the kinds; no probe is drafted or run today, and every report says so.
Where this lives in the code
docs/ROADMAP.md stage 4; src/suncly/core/coverage.py lists the gap
What is kept, how it is signed, and how it is read later.
Signed, verifiable evidence
AvailableEvery attestation is signed with the deployment's Ed25519 key over the card hash, contract version, per-test-case counts, the hash of every transcript and the decision. suncly verify re-checks a report folder offline.
Why it matters: Evidence can be handed to an auditor and checked without trusting the person who handed it over.
Where this lives in the code
src/suncly/core/signing.py, core/verify.py; tests/unit/test_signing_policy_verify.py
Explicit inconclusive outcomes and coverage gaps
AvailableEvery report has a 'What was NOT tested' section: skills without a test case, runs never executed, inconclusive runs, declared capabilities not exercised, interfaces not used, probes and semantic checks that do not exist yet, and the production endpoint.
Why it matters: An approval on partial evidence is visible as such.
Where this lives in the code
src/suncly/core/coverage.py; tests/unit/test_non_negotiable_rules.py::test_dr_007_reports_state_what_was_not_tested
Exportable evaluation reports
AvailableOne folder per attestation: result.json (the evidence bundle and signed payload), the redacted transcripts, report.md and a self-contained report.html that opens offline.
Why it matters: Results travel as files you can attach to a ticket, archive, or load into this workspace.
Where this lives in the code
src/suncly/adapters/report/; tests/unit/test_adapters.py::test_report_folder_is_self_contained_and_offline
Append-only evidence store, file or Postgres
Available with limitsRuns and decisions are append-only; the database rejects updates and deletes. The same store interface runs on local files or Postgres (set DATABASE_URL and run suncly db migrate).
Why it matters: Evidence cannot be quietly edited after the fact.
Limit: Transcripts stay on local disk until an object-storage adapter exists.
Where this lives in the code
src/suncly/adapters/file_store.py, adapters/postgres/; tests/stores/test_store_contract.py, tests/db/
Regression comparisons between evaluations
Available with limitsThe same card hash reuses its approved contract, so two attestations of the same agent run the same tests and can be compared test case by test case.
Why it matters: You see what changed since the last evaluation, with the test conditions held constant.
Limit: The comparison is computed in this workspace from the two signed bundles. Suncly's store has no comparison or baseline operation; what counts as a drop is an open policy question (OQ-PO2).
Where this lives in the code
frontend/lib/evidence/compare.ts over result.json; core/cards.py (card hash reuse)
Who decides, on what, and how that is recorded.
Human review of the test plan
AvailableNothing runs until a person approves the drafted contract and their identifier is recorded as approved_by. Approved contracts are immutable; an edit is a new version.
Why it matters: The test plan is a reviewed artefact with a name on it, not a side effect of a run.
Where this lives in the code
src/suncly/core/contract_builder.py (ContractService.approve); tests/unit/test_non_negotiable_rules.py
Automatic approval decisions
Available with limitsThe Policy engine records one decision per completed attestation and signs it. Today the only outcome it can produce is flag, because no policy configuration exists.
Why it matters: Every completed evaluation ends with an explicit, signed decision record and a human in the loop.
Limit: approve and block are never produced automatically; thresholds per risk level are stage 5 and a customer configuration. Risk levels are recorded but do not yet change the outcome.
Where this lives in the code
src/suncly/core/policy_engine.py; tests/unit/test_architecture.py::test_no_code_path_constructs_an_approve_or_block_decision
Recording the reviewer's decision in the evidence store
PlannedA second decision record with the reviewer's identity, resolving the flag.
Why it matters: Would close the loop inside Suncly.
Not yet: No CLI command or endpoint records it yet (OQ-P2, stage 5). This workspace lets a reviewer record and export a decision note, kept outside the signed evidence and clearly marked as such.
Where this lives in the code
docs/API.md 'Operations without an interface'
Sandboxes, credentials, budgets and limits.
Sandbox or dry-run endpoints only
Available with limitsNothing runs unless the caller declares the endpoint a sandbox with --sandbox. The Runner refuses undeclared targets a second time.
Why it matters: Tests cannot book, pay or delete anything real.
Limit: Suncly cannot verify that an endpoint is a sandbox. The flag is your declaration, and the report records it as such.
Where this lives in the code
src/suncly/core/attestation.py, runner/process.py; tests/e2e/test_cli.py
Credentials held only by the Runner, redacted from evidence
AvailableThe agent's Authorization value comes from one environment variable, read only inside the Runner process. Transcripts are redacted before they leave it: the credential, sensitive headers, known token patterns.
Why it matters: A reviewer can read every transcript without seeing a secret, and the component that holds secrets is small enough to audit.
Where this lives in the code
src/suncly/runner/credentials.py, runner/redaction.py; tests/unit/test_architecture.py::test_only_the_runner_reads_the_credential
Budget and execution controls
AvailableA budget in attempts per attestation (default twice the planned runs), a per-run timeout, retries under the same run key, and a concurrency limit. When the budget is reached the attestation ends failed and the report lists every run never executed.
Why it matters: No runaway cost, and no silent partial result dressed up as a complete one.
Where this lives in the code
src/suncly/core/orchestrator.py (_Budget); tests/e2e/test_mock_agents.py::test_budget_stop_ends_failed_and_reports_runs_never_executed
How Suncly fits into pipelines and registries.
Command-line interface
Availablesuncly attest, demo, verify, keys init, db migrate, db check, doctor. Documented exit codes; --json for scripts.
Why it matters: Runs anywhere Python 3.12 runs, including inside your network, with no data leaving it.
Where this lives in the code
src/suncly/cli/; docs/API.md
HTTP API
PlannedPOST /attestations, GET /attestations/{id}, POST /contracts/{id}/approve, GET /agents/{id}/evidence, calling the same core library.
Why it matters: Would let this workspace start runs and read evidence directly.
Not yet: Stage 5. api.py is a placeholder; payloads and authentication are open questions (OQ-P1, OQ-P3).
Where this lives in the code
src/suncly/api.py; docs/API.md
CI gate
Available with limitsThe CLI runs in a pipeline today with --approve-as, --json and distinct exit codes. Exit code 0 means completed and signed, never approved.
Why it matters: A pipeline can run the evaluation and archive the report on every release.
Limit: A CI adapter that turns a decision into a pass or fail gate is stage 5. Because decisions are flag-only, no pipeline should gate on them yet.
Where this lives in the code
src/suncly/cli/exit_codes.py; src/suncly/adapters/ci.py (placeholder)
Registry integration
PlannedAdapters that write approval status into your agent registry.
Why it matters: Would make approval visible where agents are discovered.
Not yet: Stage 6. Target registries are not chosen (OQ-R5).
Where this lives in the code
src/suncly/adapters/registry.py (placeholder)
Risk is recorded today. It decides nothing yet.
Every agent carries a risk level, recorded on first sight (default high). The default approval policy below is the design; its thresholds are a customer configuration that does not exist yet, so every decision is flag.
| Risk level | Example | Approval, as designed | Today |
|---|---|---|---|
| low | read-only lookup | automatic on pass | flag; a human decides |
| medium | writes to internal systems | automatic on pass, human on any drop | flag; a human decides |
| high | payments, personal data | human sign-off every time | flag; a human decides |
A human is always required for the first contract approval, for new or changed skills, for borderline or dropping results, and for every evaluation of a high-risk agent. Suncly ships no threshold numbers of its own.
Questions buyers ask.
Is Suncly available today?
Yes, as a command-line tool in pilot. From a card URL it drafts a test plan, records your approval, runs the tests repeatedly against your sandbox, judges every run, records a flag decision, signs the attestation and writes a report. The HTTP API, model-based judging, probes and automatic approve or block decisions are planned stages, not current features.
Which agents can it evaluate?
Any agent that publishes an A2A 1.0 Agent Card and answers over JSON-RPC, whatever model or framework is behind it. Other A2A bindings, streaming and push notifications are not exercised yet and are listed in the report as not tested.
Does Suncly call my production agent?
No. Nothing runs unless you declare the endpoint a sandbox or dry-run endpoint, and the report records that declaration. Suncly cannot verify a sandbox; the declaration is yours.
Where do my credentials go?
Into one environment variable that only the Runner process reads. The Runner redacts the credential, sensitive headers and known token patterns from every transcript before it leaves the process. A test proves no other module reads the variable.
Does Suncly send anything to a model provider?
Not in this version. Drafting is deterministic and judging is deterministic. Later stages will use your own model keys for a model-based drafter and judge.
What does a decision look like?
Every completed evaluation gets one signed decision record from the Policy engine. Without a configured policy the outcome is always flag, so a human reviews every result. Approve and block are never produced automatically yet.
Can I compare two evaluations?
An unchanged card reuses its approved contract, so two evaluations run the same tests and their counts are comparable. The workspace shows the comparison test case by test case. It is computed from the two signed bundles; Suncly's store has no comparison operation of its own yet.
Does it produce a score?
No. Counts per test case, never combined, and inconclusive runs are counted separately. A score would hide what was not tested.
What does it cost?
No price is published. Suncly is in pilot and a payment layer is planned; access is arranged with the team.
Run the first evaluation with us.
Suncly is in pilot. If your team approves A2A agents by hand today, tell us about one agent and one sandbox, and we will run the first evaluation together.
Or write to team@suncly.com.