Keep the failure.
Import the messages, tool calls, and results from a failed run. Redact sensitive values and save a test fixture.
CAPTUREOne broken run should make your agent better.
We’re building the tools to turn it into a test.
replay "case-014.json"tools recorded_responsescheck required_output_fieldsThe agent returned an answer without a source.
expected: response.source_urlPrompts change. Models change. Tools change.
Your checks should keep up.The workflow we’re building starts with a run you already have. No model training required.
Import the messages, tool calls, and results from a failed run. Redact sensitive values and save a test fixture.
CAPTURERun a revised prompt or model against the same recorded tool responses in a test environment.
REPLAYCompare explicit checks, output changes, tool use, and cost estimates before making a release decision.
COMPAREAgent behavior can vary. Testing should make that uncertainty visible, with checks you can explain.
Import a trace, define a check, and compare two configurations. A focused first version that fits the way developers work.
Use recorded tool responses for tests. Rerunning a fixture shouldn’t resend an email or repeat a real transaction.
Required fields. Allowed tools. Call counts. Latency. Check the properties your application actually depends on.
Runweft is an early-stage project by Humayun, focused on regression testing for AI-agent workflows. The first milestone is a local CLI and a readable comparison report.
We’re developing the concept and planning the first implementation. The interface above illustrates the planned workflow.