feat(evals): add agent evaluation suite - #8409
Open
sudoKrishna wants to merge 19 commits into
Open
sudoKrishna wants to merge 19 commits into
sudoKrishna wants to merge 19 commits into
Conversation
Add a deterministic eval layer for the agent harness. Scenarios script the OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery without a provider key. - apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report - `bun run test:evals` from apps/sim runs the suite and writes the report - picked up by the normal vitest run so a regression fails CI - README documents the contract and how to add a case
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
Contributor
|
Replay the same scenarios against a real model. The model is the only thing that changes: runScenario now takes an optional completion transport and a live mode that relaxes exact assertions (ordered subsequence, minimum successes) and skips scripted-only recovery cases. - live.ts: OpenAI-compatible transport + DeepSeek factory - agent-tool-use.live.test.ts: K trials per scenario, gated on EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI - live report with pass rates, avg iterations, latency, failed checks - test:evals:live script and README knobs
…ve mode The first live DeepSeek run exposed brittle assertions, not harness bugs: the model chained the tools correctly but the checks were case-sensitive and required an internal order id. Match the retrieved value case-insensitively and let live runs accept the grounded status rather than the internal id.
|
@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel. A member of the Team first needs to authorize it. |
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor, with only executeProviderRequest mocked at the provider boundary. This covers agent-block input wiring, variable resolution from Start outputs, and executor run/error handling, which the direct loop harness cannot see. - executor-harness.ts: workflow builder + runExecutorScenario - shares the scorer (scoreExpectations) and report with the loop suite - two scenarios: Start->Agent output, and <start.message> resolution - README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the Agent block has retry enabled, and the executor replays it. The run must complete with the second response. Verifies providerCalls === 2, and fails without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the Agent block has a fallback model, and the handler serves the answer from gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails without the fallback row (checked locally: got gpt-4o, run errored).
Record a live run once, replay it forever through the real tool loop with no key. EVAL_RECORD=1 wraps the live completion and writes each model call's streamed chunks to fixtures/<scenario>.json; agent-tool-use.replay.test.ts feeds them back through createOpenAICompatStreamingToolLoopStream and scores them with the same checks. - replay.ts: recording/replay completions + fixture I/O - replay.test.ts: chunk round-trip and fixture I/O (key-free) - live test records on EVAL_RECORD=1; test:evals:record script - replay suite skips until a fixture exists; README documents the loop
4 of 6 tasks
Drive the Agent block through the executor with conversation memory on. The memory read is stubbed per conversation id, so the provider request shows what the handler assembled: prior history, then the new prompt, system prompt preserved, correct conversation id. A wrong id surfaces as missing history and fails (checked locally). - agent-context/scenarios.ts: two context scenarios - executor-harness.ts: memory seam + assembly/isolation checks - test:evals:context script; README documents the suite
5 of 6 tasks
Run the same live scenarios across a list of models and write a scenario x model matrix. models.ts resolves provider:model specs (DeepSeek, OpenAI, Groq, OpenRouter) and reads each provider's key from <PROVIDER>_API_KEY. - agent-tool-use.compare.live.test.ts: EVAL_MODELS x scenarios x trials - report.ts: buildLiveComparisonReport + JSON/Markdown matrix - report.test.ts: key-free aggregation coverage - test:evals:compare script; README documents the spec format
Pass rates alone do not say why a model lost. Aggregate the failed check names per model into the comparison report and add a Failed checks column.
Five cases that stress where models tend to fail: answering with no tool, disambiguating near-identical tools, not inventing an answer from an empty tool result, running a four-tool dependency chain, and picking settings over a near-duplicate profile tool. Scripted expectations keep them deterministic; the same cases run live.
4 of 6 tasks
Two live failures were eval design, not model failure: - empty-result-no-hallucination rejected valid 'didn't find' / 'wasn't able to find' phrasing. Broaden the grounding check. - near-duplicate-names required a userId the prompt never gave, so the model reasonably asked for it. Put the id in the prompt and the scripted call.
- long-chain-dependency: the prompt never gave a userId, so the model asked or skipped the profile step. Provide u-42 and let live runs require the three downstream calls rather than the exact four-step sequence. - near-duplicate-names: one live trial called both tools; that is over-calling, not wrong-tool selection. Drop the forbidden-tool assertion in live mode.
Substring checks measure phrasing, not correctness. judgeAnswer scores an answer against a weighted rubric with a judge model and returns structured scores; runScenario gains an optional judge that adds a judge check. The judge transport is an injectable OpenAI-compatible completion, so a recorded transcript can replay it deterministically. - judge.ts: rubric, prompt, JSON parsing/clamping, verdict - judge.test.ts: parsing/weighting/clamping (key-free) - judge.live.test.ts: grounded answer outscores an invented one (opt-in) - test:evals:judge script; README documents it
4 of 6 tasks
# Conflicts: # apps/sim/package.json
# Conflicts: # apps/sim/evals/README.md # apps/sim/package.json
# Conflicts: # apps/sim/evals/README.md # apps/sim/package.json
- Read tool feedback: the scripted model now asserts that each prior turn's tool results reached the next model call, so a loop that drops feedback fails the retrieval/planning/recovery cases. - Check tool arguments: score every executed call against the scripted arguments, so a right-name/wrong-arguments call fails. - Use absolute @/evals imports instead of relative ones, per the app rule. Verified both new checks fail under mutation (bad marker, mutated args).
Author
|
Addressed all three findings in f619479:
|
5 of 7 tasks
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a deterministic, CI-runnable evaluation layer for the agent harness, plus
opt-in live tooling. Scripted scenarios drive the real code — the OpenAI-compatible
streaming tool loop and the
DAGExecutor— and score tool selection, planning,retrieval, recovery, and context assembly. Live runs compare models and support
record/replay. No provider key is required for CI.
This is the consolidated branch: it folds the incremental stack into one PR.
What's included
Deterministic (CI, no key)
agent-tool-use/— 13 tool-loop scenarios (reliability + adversarial) and 4executor scenarios (Start→Agent, variable resolution, block retry, model fallback)
agent-context/— provider-request assembly with conversation memoryreplay.ts— record live transcripts once, replay through the real loop;replay.test.tscovers chunk round-trip and fixture I/Ojudge.ts— LLM-as-judge scoring with a weighted rubric;judge.test.tscoversparsing, weighting, clamping, and malformed responses
report.ts— JSON + Markdown reports, including the model-comparison matrixOpt-in live (
EVAL_LIVE=1, never in CI)live.ts— real model transport (any OpenAI-compatible provider)EVAL_MODELSwith pass rates, latency, tokensjudge.live.test.ts— a grounded answer outscores an invented oneHow to run
Live (needs a key):
The scripted suite is collected by the normal
bun run test, so a regressionfails CI without the dedicated commands.
Real results from live runs
no-tool-needed, near-duplicate tools, empty results, and a four-tool chain.
deepseek-chat100% /deepseek-reasoner89%, reasoner~27% more tokens — a counterintuitive result the suite surfaced.
an over-specific id, a missing input a tool required), not a model failure.
The adversarial cases are what exposed them.
Test plan
bun run test:evals→ 30/30 (1 replay suite skips with no fixtures)bun run test:evals:context→ 2/2context-isolation expectations and confirmed each failed, then reverted
bun run check:test-patternspassesbun run type-check— run in CIFollow-up