Test what your AI agents actually do.
Zeian executes realistic tasks against your agent and its MCP tools, scores the trajectory and the outcome, and fails the build when behavior regresses — before users find out.
suites/booking-flow.yaml
suite: booking-flow
agent: ./agents/travel
tools: [flights, bookings, refunds]
tasks:
- id: cancel-with-credit
input: cancel flight BK-2291
expect:
outcome: booking_cancelled
trajectory:
calls: [bookings.get, refunds.void]
max_calls: 5
forbid: [refunds.void ×2]
run
$ zeian run suites/booking-flow
✓ search-flights pass · 4 calls · $0.021
✓ book-cheapest pass · 7 calls · $0.033
✗ cancel-with-credit fail · trajectory
✓ rebook-same-day pass · 9 calls · $0.041
4/5 pass · 1 regression vs baseline
trace · cancel-with-credit
01 flights.search → ok 812ms
02 bookings.get → ok 204ms
03 refunds.void → ok 388ms
04 refunds.void → duplicate call
outcome correct. trajectory unsafe.
score 0.41 — retry loop on 4xx
verdict
task cancel-with-credit
status failed
reason tool called twice
baseline 4 calls → now 5
cost $0.018 → $0.031 (+72%)
→ blocks merge in CI
Example output. Illustrative only.
Agents pass demos and fail in production.
- Unit tests cover functions, not behavior
- An agent can pass every test and still call the wrong tool on a real task.
- Tool contracts drift
- MCP servers ship new schemas, prompts change, models update. Nothing fails loudly until a task breaks.
- Failures are silent
- Wrong tool calls, dropped arguments, retry loops. Most teams find out from users.
- Eval scores miss the path
- A right answer through a broken trajectory still costs money and trust.
Four steps from task suite to verdict.
01
Define
Write task suites that mirror the real work your agent handles. Each task declares inputs, allowed tools, and what done means.
02
Execute
Zeian runs the agent against each task with its real MCP tools attached, not stubs.
03
Evaluate
Every run is scored on outcome, tool-call sequence, arguments, error handling, cost, and latency.
04
Compare
Results are diffed against a golden baseline so drift shows up as a failing check, not a surprise.
Your tools are part of the product. Test them like it.
Zeian is designed to test the whole loop: the model's decisions, the MCP tools it calls, and the outcomes those calls produce.
01
Connect
Point Zeian at the same MCP servers your agent uses in production. No mocks, no re-implementation.
02
Exercise
Run every tool against realistic inputs, including the awkward ones your demos skip.
03
Assert
Check response schemas, argument handling, timeouts, and error paths against expectations.
A broken tool looks like a dumb agent. Zeian separates the two so you fix the layer that actually failed.
Score the path, not just the answer.
Zeian collects the full trace of each run: every tool call, every argument, every retry. Evaluation turns that trace into a verdict you can act on.
- Did it call the right tools, in the right order?
- Were the arguments correct and complete?
- Did it handle errors and retries sanely?
- Did it stay on task, or wander and loop?
Trace excerpt
# task: cancel-with-credit
01 flights.search → ok · 812ms
02 bookings.get → ok · 204ms
03 refunds.void → ok · 388ms
04 refunds.void → duplicate call, flagged
verdict: task outcome correct, trajectory unsafe
Baselines turn vibes into gates.
Promote a good run to a golden baseline. Every later run diffs against it, so a model update, prompt tweak, or MCP change that breaks behavior shows up immediately.
Baseline
Capture a known-good trajectory per task: calls, arguments, outcome, cost, latency.
Diff
Compare each new run against the baseline. Added calls, dropped calls, changed arguments, worse outcomes.
Gate
Fail the run when the diff matters. Ship the change only when the agent still does the job.
Run your agent suite on every pull request.
Zeian is built to run inside GitHub Actions. Suites execute on each PR, results post back as a check, and a regression blocks merge just like a failing unit test.
name: agent-evals
on: [pull_request]
jobs:
zeian:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: zeian-labs/run@v1
with:
suite: suites/production.yaml
baseline: main- Results land as a PR check with a per-task breakdown.
- Trajectory diffs link straight to the failing step.
- The check fails on regression, so broken changes can't merge quietly.
One pipeline from task definition to merge gate.
The shape Zeian is being built around. Each stage is explicit so a failure has an address.
Target architecture. See docs/architecture.md for what exists today.
What a suite looks like when it runs.
Each task reports its outcome and its trajectory. The regression line is the part your unit tests never told you.
$ zeian run suites/booking-flow.yaml --agent mcp://localhost:8123suite booking-flow · 6 tasks✓ search-flights pass 4 tool calls · 8.2s✓ book-cheapest-option pass 7 tool calls · 21.4s✗ cancel-with-credit fail refunds.void called twice✓ rebook-same-day pass 9 tool calls · 18.9s✓ price-check-parity pass 5 tool calls · 6.1s~ multi-passenger skip missing fixtureresult 5/6 passed · 1 regression vs baseline
Example output. Illustrative only.
Coming soon.
Zeian is in early development. Pricing will be published here before any paid tier exists. The repository and this site stay the source of truth.
Start here
Build agents you can actually verify.
Zeian is in early development. The fastest way to follow the work is the repository and the docs.
