Skip to content
Product

Repeatable evaluation. Accountable approval.

Suncly turns the approval of an AI agent into a recorded procedure: a reviewed test plan, repeated runs in a sandbox, deterministic verdicts, signed evidence, and an explicit statement of what was not tested.

Available
Implemented in the current version and covered by tests.
Available with limits
Implemented, with a limitation a buyer must know about.
Planned
On the roadmap. No date is promised, and nothing on this site treats it as current.
Evaluation

What is tested, how often, and how it is judged.

  • Reads the A2A Agent Card and hashes it

    Available

    Fetches the card over https, keeps it byte for byte, and computes card_hash as SHA-256 of its RFC 8785 form without the signatures field.

    Why it matters: Every evaluation is pinned to an exact version of what the agent claimed.

    Where this lives in the code

    src/suncly/domain/card.py, core/cards.py; tests/unit/test_card.py

  • Tests derived from declared skills

    Available with limits

    One test case per declared example of each skill, up to a cap. Criteria: completes, answers, uses the declared output modes, within a latency limit.

    Why it matters: Nobody writes the first test plan by hand, and every declared skill with an example is covered.

    Limit: Drafting is deterministic and uses no model. A skill without examples gets no test case and is listed as not tested. Model-drafted tests and probes are stage 2 and 4.

    Where this lives in the code

    src/suncly/core/contract_builder.py (DeterministicDrafter); tests/unit/test_contract_builder.py

  • Customer-defined criteria and known-answer checks

    Available with limits

    Export the draft, edit the contract file, run with it: expected final state, required fields (JSON pointers), a JSON Schema for the response, latency limit, declared output modes.

    Why it matters: Your acceptance criteria become repeatable checks instead of a reviewer's memory.

    Limit: Checks are structural. A criterion about the meaning of an answer (model_checks) makes the run inconclusive until Judge Layer 2 exists (stage 4).

    Where this lives in the code

    src/suncly/domain/criteria.py, domain/contract_file.py; docs/API.md 'Contract file'

  • Repeated runs that expose inconsistent behaviour

    Available

    Each test case runs N times (default 5) from a separate Runner process, with a deterministic run key per attempt so retries never double count.

    Why it matters: An agent that works four times out of five shows up as four passes and one fail, not as a pass.

    Where this lives in the code

    src/suncly/core/orchestrator.py; tests/e2e/test_mock_agents.py (flaky agent)

  • Deterministic verdicts: pass, fail, inconclusive

    Available with limits

    Layer 1 checks the response is a well-formed A2A Task or Message, reached the expected final state, answered within the latency limit, carries content, uses the declared output modes, and satisfies required fields and schema.

    Why it matters: Verdicts are reproducible and explained check by check. Inconclusive is a third outcome, never counted as a pass.

    Limit: No model-based judgement yet. Whether an answer is correct in meaning is listed under what was not tested.

    Where this lives in the code

    src/suncly/core/judge.py; tests/unit/test_judge.py

  • Prompt-injection, undeclared-behaviour and failure-handling probes

    Planned

    Test cases of kind probe_injection, probe_undeclared and probe_failure, drafted alongside skill tests and approved the same way.

    Why it matters: Would show how the agent behaves outside what its card declares.

    Not yet: Stage 4. The data model and the contract file format already carry the kinds; no probe is drafted or run today, and every report says so.

    Where this lives in the code

    docs/ROADMAP.md stage 4; src/suncly/core/coverage.py lists the gap

Evidence

What is kept, how it is signed, and how it is read later.

  • Signed, verifiable evidence

    Available

    Every attestation is signed with the deployment's Ed25519 key over the card hash, contract version, per-test-case counts, the hash of every transcript and the decision. suncly verify re-checks a report folder offline.

    Why it matters: Evidence can be handed to an auditor and checked without trusting the person who handed it over.

    Where this lives in the code

    src/suncly/core/signing.py, core/verify.py; tests/unit/test_signing_policy_verify.py

  • Explicit inconclusive outcomes and coverage gaps

    Available

    Every report has a 'What was NOT tested' section: skills without a test case, runs never executed, inconclusive runs, declared capabilities not exercised, interfaces not used, probes and semantic checks that do not exist yet, and the production endpoint.

    Why it matters: An approval on partial evidence is visible as such.

    Where this lives in the code

    src/suncly/core/coverage.py; tests/unit/test_non_negotiable_rules.py::test_dr_007_reports_state_what_was_not_tested

  • Exportable evaluation reports

    Available

    One folder per attestation: result.json (the evidence bundle and signed payload), the redacted transcripts, report.md and a self-contained report.html that opens offline.

    Why it matters: Results travel as files you can attach to a ticket, archive, or load into this workspace.

    Where this lives in the code

    src/suncly/adapters/report/; tests/unit/test_adapters.py::test_report_folder_is_self_contained_and_offline

  • Append-only evidence store, file or Postgres

    Available with limits

    Runs and decisions are append-only; the database rejects updates and deletes. The same store interface runs on local files or Postgres (set DATABASE_URL and run suncly db migrate).

    Why it matters: Evidence cannot be quietly edited after the fact.

    Limit: Transcripts stay on local disk until an object-storage adapter exists.

    Where this lives in the code

    src/suncly/adapters/file_store.py, adapters/postgres/; tests/stores/test_store_contract.py, tests/db/

  • Regression comparisons between evaluations

    Available with limits

    The same card hash reuses its approved contract, so two attestations of the same agent run the same tests and can be compared test case by test case.

    Why it matters: You see what changed since the last evaluation, with the test conditions held constant.

    Limit: The comparison is computed in this workspace from the two signed bundles. Suncly's store has no comparison or baseline operation; what counts as a drop is an open policy question (OQ-PO2).

    Where this lives in the code

    frontend/lib/evidence/compare.ts over result.json; core/cards.py (card hash reuse)

Approval

Who decides, on what, and how that is recorded.

  • Human review of the test plan

    Available

    Nothing runs until a person approves the drafted contract and their identifier is recorded as approved_by. Approved contracts are immutable; an edit is a new version.

    Why it matters: The test plan is a reviewed artefact with a name on it, not a side effect of a run.

    Where this lives in the code

    src/suncly/core/contract_builder.py (ContractService.approve); tests/unit/test_non_negotiable_rules.py

  • Automatic approval decisions

    Available with limits

    The Policy engine records one decision per completed attestation and signs it. Today the only outcome it can produce is flag, because no policy configuration exists.

    Why it matters: Every completed evaluation ends with an explicit, signed decision record and a human in the loop.

    Limit: approve and block are never produced automatically; thresholds per risk level are stage 5 and a customer configuration. Risk levels are recorded but do not yet change the outcome.

    Where this lives in the code

    src/suncly/core/policy_engine.py; tests/unit/test_architecture.py::test_no_code_path_constructs_an_approve_or_block_decision

  • Recording the reviewer's decision in the evidence store

    Planned

    A second decision record with the reviewer's identity, resolving the flag.

    Why it matters: Would close the loop inside Suncly.

    Not yet: No CLI command or endpoint records it yet (OQ-P2, stage 5). This workspace lets a reviewer record and export a decision note, kept outside the signed evidence and clearly marked as such.

    Where this lives in the code

    docs/API.md 'Operations without an interface'

Operations

Sandboxes, credentials, budgets and limits.

  • Sandbox or dry-run endpoints only

    Available with limits

    Nothing runs unless the caller declares the endpoint a sandbox with --sandbox. The Runner refuses undeclared targets a second time.

    Why it matters: Tests cannot book, pay or delete anything real.

    Limit: Suncly cannot verify that an endpoint is a sandbox. The flag is your declaration, and the report records it as such.

    Where this lives in the code

    src/suncly/core/attestation.py, runner/process.py; tests/e2e/test_cli.py

  • Credentials held only by the Runner, redacted from evidence

    Available

    The agent's Authorization value comes from one environment variable, read only inside the Runner process. Transcripts are redacted before they leave it: the credential, sensitive headers, known token patterns.

    Why it matters: A reviewer can read every transcript without seeing a secret, and the component that holds secrets is small enough to audit.

    Where this lives in the code

    src/suncly/runner/credentials.py, runner/redaction.py; tests/unit/test_architecture.py::test_only_the_runner_reads_the_credential

  • Budget and execution controls

    Available

    A budget in attempts per attestation (default twice the planned runs), a per-run timeout, retries under the same run key, and a concurrency limit. When the budget is reached the attestation ends failed and the report lists every run never executed.

    Why it matters: No runaway cost, and no silent partial result dressed up as a complete one.

    Where this lives in the code

    src/suncly/core/orchestrator.py (_Budget); tests/e2e/test_mock_agents.py::test_budget_stop_ends_failed_and_reports_runs_never_executed

Interfaces and integrations

How Suncly fits into pipelines and registries.

  • Command-line interface

    Available

    suncly attest, demo, verify, keys init, db migrate, db check, doctor. Documented exit codes; --json for scripts.

    Why it matters: Runs anywhere Python 3.12 runs, including inside your network, with no data leaving it.

    Where this lives in the code

    src/suncly/cli/; docs/API.md

  • HTTP API

    Planned

    POST /attestations, GET /attestations/{id}, POST /contracts/{id}/approve, GET /agents/{id}/evidence, calling the same core library.

    Why it matters: Would let this workspace start runs and read evidence directly.

    Not yet: Stage 5. api.py is a placeholder; payloads and authentication are open questions (OQ-P1, OQ-P3).

    Where this lives in the code

    src/suncly/api.py; docs/API.md

  • CI gate

    Available with limits

    The CLI runs in a pipeline today with --approve-as, --json and distinct exit codes. Exit code 0 means completed and signed, never approved.

    Why it matters: A pipeline can run the evaluation and archive the report on every release.

    Limit: A CI adapter that turns a decision into a pass or fail gate is stage 5. Because decisions are flag-only, no pipeline should gate on them yet.

    Where this lives in the code

    src/suncly/cli/exit_codes.py; src/suncly/adapters/ci.py (placeholder)

  • Registry integration

    Planned

    Adapters that write approval status into your agent registry.

    Why it matters: Would make approval visible where agents are discovered.

    Not yet: Stage 6. Target registries are not chosen (OQ-R5).

    Where this lives in the code

    src/suncly/adapters/registry.py (placeholder)

Risk levels and policy

Risk is recorded today. It decides nothing yet.

Every agent carries a risk level, recorded on first sight (default high). The default approval policy below is the design; its thresholds are a customer configuration that does not exist yet, so every decision is flag.

Default approval policy by risk level, as designed
Risk levelExampleApproval, as designedToday
lowread-only lookupautomatic on passflag; a human decides
mediumwrites to internal systemsautomatic on pass, human on any dropflag; a human decides
highpayments, personal datahuman sign-off every timeflag; a human decides

A human is always required for the first contract approval, for new or changed skills, for borderline or dropping results, and for every evaluation of a high-risk agent. Suncly ships no threshold numbers of its own.

FAQ

Questions buyers ask.

Is Suncly available today?

Yes, as a command-line tool in pilot. From a card URL it drafts a test plan, records your approval, runs the tests repeatedly against your sandbox, judges every run, records a flag decision, signs the attestation and writes a report. The HTTP API, model-based judging, probes and automatic approve or block decisions are planned stages, not current features.

Which agents can it evaluate?

Any agent that publishes an A2A 1.0 Agent Card and answers over JSON-RPC, whatever model or framework is behind it. Other A2A bindings, streaming and push notifications are not exercised yet and are listed in the report as not tested.

Does Suncly call my production agent?

No. Nothing runs unless you declare the endpoint a sandbox or dry-run endpoint, and the report records that declaration. Suncly cannot verify a sandbox; the declaration is yours.

Where do my credentials go?

Into one environment variable that only the Runner process reads. The Runner redacts the credential, sensitive headers and known token patterns from every transcript before it leaves the process. A test proves no other module reads the variable.

Does Suncly send anything to a model provider?

Not in this version. Drafting is deterministic and judging is deterministic. Later stages will use your own model keys for a model-based drafter and judge.

What does a decision look like?

Every completed evaluation gets one signed decision record from the Policy engine. Without a configured policy the outcome is always flag, so a human reviews every result. Approve and block are never produced automatically yet.

Can I compare two evaluations?

An unchanged card reuses its approved contract, so two evaluations run the same tests and their counts are comparable. The workspace shows the comparison test case by test case. It is computed from the two signed bundles; Suncly's store has no comparison operation of its own yet.

Does it produce a score?

No. Counts per test case, never combined, and inconclusive runs are counted separately. A score would hide what was not tested.

What does it cost?

No price is published. Suncly is in pilot and a payment layer is planned; access is arranged with the team.

Pilot

Run the first evaluation with us.

Suncly is in pilot. If your team approves A2A agents by hand today, tell us about one agent and one sandbox, and we will run the first evaluation together.

Or write to team@suncly.com.