Workflow changedStarts from GitHub on agent-workflow pull requestRun evals when prompts, tools, or guardrails change.orNightly suite starts from Cron on nightly-agent-evalsRun scheduled regressions against representative tasks.
1Scenario Designer1Creates representative happy-path and adversarial tasks.Uses1 toolproduces Scenario suitefrom Workflow spec and historical failures
2Eval Runner1Runs the suite and captures traces.Uses2 toolsproduces Raw run results and trace bundlefrom Scenario suite and workflow version
3Release Judge2Scores results and decides whether release can proceed.Uses2 toolsproduces Release decision and remediation checklistfrom Run results and thresholds
EvalForge: Agent Evaluation LabCatch behavior regressions before agent workflows reach production users.ForAI companies shipping agents that need repeatable evals, guardrail testing, and release confidence