Capture cell
Scenario runs execute in a capture cell and can be observed through strace or Tetragon-style JSONL traces.
AgentsEval
AgentsEval evaluates AI agent behavior with captured traces, deterministic scenarios, replay support, and ship/watch/fail verdicts. It is reliability evidence for specific scenarios and policies, not a broad security certification or universal benchmark.
AgentsEval is an AI agent evaluation framework that scores captured runs from syscall, file, network, intent, workspace, and final-answer evidence. It rolls multiple runs into scoped verdicts so teams can promote, watch, or reject agent versions based on observed behavior.
Captured behavior becomes a scoped verdict.
Scenario runs execute in a capture cell and can be observed through strace or Tetragon-style JSONL traces.
Safety rules inspect out-of-scope files, destructive commands, disallowed egress, privilege escalation, injection composites, and test-file edits.
Capability specs check expected commands, outputs, HTTP activity, final answers, and workspace effects.
A deterministic record/replay proxy supports regression testing without silently falling back to live providers on misses.
Agent versions can be promoted, watched, or rejected based on captured behavior rather than only prompt review.
The copied JSON files are small proof artifacts from the local scenario-library eval output. They demonstrate banding behavior, not universal agent certification.
Three failing runs roll up to band `fail` with critical safety severity.
Learn moreThree failing runs roll up to band `fail` after sensitive file and egress findings.
Learn moreDesign scenario libraries, capture boundaries, replay requirements, and promotion gates for your agent versions.
Architecture conversation
Share your deployment boundary, number of agents, work surfaces, and governance requirements. We will reply by email to arrange a focused technical discussion.
Email the Bewize teamThis opens your email application. Read our Privacy Policy.