Zeian Labs · Early development

Test what your AI agents actually do.

Zeian executes realistic tasks against your agent and its MCP tools, scores the trajectory and the outcome, and fails the build when behavior regresses — before users find out.

suites/booking-flow.yaml

suite: booking-flow

agent: ./agents/travel

tools: [flights, bookings, refunds]

tasks:

- id: cancel-with-credit

input: cancel flight BK-2291

expect:

outcome: booking_cancelled

trajectory:

calls: [bookings.get, refunds.void]

max_calls: 5

forbid: [refunds.void ×2]

run

$ zeian run suites/booking-flow

✓ search-flights pass · 4 calls · $0.021

✓ book-cheapest pass · 7 calls · $0.033

✗ cancel-with-credit fail · trajectory

✓ rebook-same-day pass · 9 calls · $0.041

4/5 pass · 1 regression vs baseline

trace · cancel-with-credit

01 flights.search → ok 812ms

02 bookings.get → ok 204ms

03 refunds.void → ok 388ms

04 refunds.void → duplicate call

outcome correct. trajectory unsafe.

score 0.41 — retry loop on 4xx

verdict

task cancel-with-credit

status failed

reason tool called twice

baseline 4 calls → now 5

cost $0.018 → $0.031 (+72%)

→ blocks merge in CI

Example output. Illustrative only.

01 The problem

Agents pass demos and fail in production.

Unit tests cover functions, not behavior
An agent can pass every test and still call the wrong tool on a real task.
Tool contracts drift
MCP servers ship new schemas, prompts change, models update. Nothing fails loudly until a task breaks.
Failures are silent
Wrong tool calls, dropped arguments, retry loops. Most teams find out from users.
Eval scores miss the path
A right answer through a broken trajectory still costs money and trust.
02 How Zeian works

Four steps from task suite to verdict.

  1. 01

    Define

    Write task suites that mirror the real work your agent handles. Each task declares inputs, allowed tools, and what done means.

  2. 02

    Execute

    Zeian runs the agent against each task with its real MCP tools attached, not stubs.

  3. 03

    Evaluate

    Every run is scored on outcome, tool-call sequence, arguments, error handling, cost, and latency.

  4. 04

    Compare

    Results are diffed against a golden baseline so drift shows up as a failing check, not a surprise.

03 Agent and MCP testing

Your tools are part of the product. Test them like it.

Zeian is designed to test the whole loop: the model's decisions, the MCP tools it calls, and the outcomes those calls produce.

01

Connect

Point Zeian at the same MCP servers your agent uses in production. No mocks, no re-implementation.

02

Exercise

Run every tool against realistic inputs, including the awkward ones your demos skip.

03

Assert

Check response schemas, argument handling, timeouts, and error paths against expectations.

A broken tool looks like a dumb agent. Zeian separates the two so you fix the layer that actually failed.

04 Trajectory evaluation

Score the path, not just the answer.

Zeian collects the full trace of each run: every tool call, every argument, every retry. Evaluation turns that trace into a verdict you can act on.

  • Did it call the right tools, in the right order?
  • Were the arguments correct and complete?
  • Did it handle errors and retries sanely?
  • Did it stay on task, or wander and loop?

Trace excerpt

# task: cancel-with-credit

01 flights.search → ok · 812ms

02 bookings.get → ok · 204ms

03 refunds.void → ok · 388ms

04 refunds.void → duplicate call, flagged

verdict: task outcome correct, trajectory unsafe

05 Regression testing

Baselines turn vibes into gates.

Promote a good run to a golden baseline. Every later run diffs against it, so a model update, prompt tweak, or MCP change that breaks behavior shows up immediately.

Baseline

Capture a known-good trajectory per task: calls, arguments, outcome, cost, latency.

Diff

Compare each new run against the baseline. Added calls, dropped calls, changed arguments, worse outcomes.

Gate

Fail the run when the diff matters. Ship the change only when the agent still does the job.

06 GitHub CI

Run your agent suite on every pull request.

Zeian is built to run inside GitHub Actions. Suites execute on each PR, results post back as a check, and a regression blocks merge just like a failing unit test.

.github/workflows/agent-evals.ymlyaml
name: agent-evals
on: [pull_request]

jobs:
  zeian:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: zeian-labs/run@v1
        with:
          suite: suites/production.yaml
          baseline: main
  • Results land as a PR check with a per-task breakdown.
  • Trajectory diffs link straight to the failing step.
  • The check fails on regression, so broken changes can't merge quietly.
07 Architecture

One pipeline from task definition to merge gate.

The shape Zeian is being built around. Each stage is explicit so a failure has an address.

01Zeian Dashboard
↓
02Agent / MCP connection
↓
03Task execution
↓
04Trace collection
↓
05Claude-powered evaluation
↓
06Results
↓
07GitHub CI

Target architecture. See docs/architecture.md for what exists today.

08 Example run

What a suite looks like when it runs.

Each task reports its outcome and its trajectory. The regression line is the part your unit tests never told you.

zeian — local run
$ zeian run suites/booking-flow.yaml --agent mcp://localhost:8123
suite booking-flow · 6 tasks
✓ search-flights pass 4 tool calls · 8.2s
✓ book-cheapest-option pass 7 tool calls · 21.4s
✗ cancel-with-credit fail refunds.void called twice
✓ rebook-same-day pass 9 tool calls · 18.9s
✓ price-check-parity pass 5 tool calls · 6.1s
~ multi-passenger skip missing fixture
result 5/6 passed · 1 regression vs baseline

Example output. Illustrative only.

09 Pricing

Coming soon.

Zeian is in early development. Pricing will be published here before any paid tier exists. The repository and this site stay the source of truth.

Start here

Build agents you can actually verify.

Zeian is in early development. The fastest way to follow the work is the repository and the docs.