IN DEVELOPMENT / AI AGENT TESTING

Make your next
agent release
less of a guess.

One broken run should make your agent better.
We’re building the tools to turn it into a test.

See how it works
Real failures. Repeatable checks. Clearer changes.
FROM “IT WORKED YESTERDAY”
runweft / compareCONCEPT PREVIEW
↳
support-agent / case-014A saved failure, turned into a fixture.
fixture
01replay "case-014.json"
02tools recorded_responses
03check required_output_fields
×
Missing expected field

The agent returned an answer without a source.

expected: response.source_url
A SMALL CHANGE. A VISIBLE DIFFERENCE.

Prompts change. Models change. Tools change.

Your checks should keep up.
01 / THE WORKFLOW

A failure is a starting point.
Not a dead end.

The workflow we’re building starts with a run you already have. No model training required.

01

Keep the failure.

Import the messages, tool calls, and results from a failed run. Redact sensitive values and save a test fixture.

CAPTURE
02

Try the change.

Run a revised prompt or model against the same recorded tool responses in a test environment.

REPLAY
03

See what moved.

Compare explicit checks, output changes, tool use, and cost estimates before making a release decision.

COMPARE
02 / THE APPROACH

Less “looks good.”
More assert().

Agent behavior can vary. Testing should make that uncertainty visible, with checks you can explain.

↳

Start with a small, useful CLI

Import a trace, define a check, and compare two configurations. A focused first version that fits the way developers work.

⤴

Replay the response, not the side effect

Use recorded tool responses for tests. Rerunning a fixture shouldn’t resend an email or repeat a real transaction.

≠

Measure what matters to your agent

Required fields. Allowed tools. Call counts. Latency. Check the properties your application actually depends on.

03 / EARLY DAYS, CLEAR DIRECTION

Built for the people
building with agents.

Runweft is an early-stage project by Humayun, focused on regression testing for AI-agent workflows. The first milestone is a local CLI and a readable comparison report.

We’re developing the concept and planning the first implementation. The interface above illustrates the planned workflow.