Skip to content

feat(evals): add agent evaluation suite - #8409

Open
sudoKrishna wants to merge 19 commits into
simstudioai:mainfrom
sudoKrishna:feat/agent-tool-use-evals
Open

sudoKrishna wants to merge 19 commits into
simstudioai:mainfrom
sudoKrishna:feat/agent-tool-use-evals

Conversation

@sudoKrishna

@sudoKrishna sudoKrishna commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

Adds a deterministic, CI-runnable evaluation layer for the agent harness, plus
opt-in live tooling. Scripted scenarios drive the real code — the OpenAI-compatible
streaming tool loop and the DAGExecutor — and score tool selection, planning,
retrieval, recovery, and context assembly. Live runs compare models and support
record/replay. No provider key is required for CI.

This is the consolidated branch: it folds the incremental stack into one PR.

What's included

Deterministic (CI, no key)

  • agent-tool-use/ — 13 tool-loop scenarios (reliability + adversarial) and 4
    executor scenarios (Start→Agent, variable resolution, block retry, model fallback)
  • agent-context/ — provider-request assembly with conversation memory
  • replay.ts — record live transcripts once, replay through the real loop;
    replay.test.ts covers chunk round-trip and fixture I/O
  • judge.ts — LLM-as-judge scoring with a weighted rubric; judge.test.ts covers
    parsing, weighting, clamping, and malformed responses
  • report.ts — JSON + Markdown reports, including the model-comparison matrix

Opt-in live (EVAL_LIVE=1, never in CI)

  • live.ts — real model transport (any OpenAI-compatible provider)
  • model comparison across EVAL_MODELS with pass rates, latency, tokens
  • judge.live.test.ts — a grounded answer outscores an invented one

How to run

cd apps/sim
bun run test:evals            # tool-use + executor + judge + report + replay (32 tests)
bun run test:evals:context    # context assembly

Live (needs a key):

DEEPSEEK_API_KEY=... bun run test:evals:live
EVAL_MODELS=deepseek:deepseek-chat,deepseek:deepseek-reasoner DEEPSEEK_API_KEY=... bun run test:evals:compare
DEEPSEEK_API_KEY=... bun run test:evals:judge

The scripted suite is collected by the normal bun run test, so a regression
fails CI without the dedicated commands.

Real results from live runs

  • Tool use on DeepSeek: 100% (33/33) across 11 scenarios, including
    no-tool-needed, near-duplicate tools, empty results, and a four-tool chain.
  • Model comparison: deepseek-chat 100% / deepseek-reasoner 89%, reasoner
    ~27% more tokens — a counterintuitive result the suite surfaced.
  • Every live failure so far was a test-design bug (case-sensitive phrasing,
    an over-specific id, a missing input a tool required), not a model failure.
    The adversarial cases are what exposed them.

Test plan

  • bun run test:evals → 30/30 (1 replay suite skips with no fixtures)
  • bun run test:evals:context → 2/2
  • Live runs and a key-free model-comparison report validated
  • Negative checks: broke the executor-retry, model-fallback, and
    context-isolation expectations and confirmed each failed, then reverted
  • bun run check:test-patterns passes
  • Eval modules type-check against the real loop/executor signatures
  • Full bun run type-check — run in CI

Follow-up

  • Wire the judge into scenarios where a rubric is more honest than a substring.
  • Subagent/orchestration evals (child workflow with a mocked definition loader).
  • Store per-commit reports and chart pass rate/cost over time.

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
@sudoKrishna
sudoKrishna requested a review from a team as a code owner September 29, 2026 09:51
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 29, 2026 9:51am UTC

Request Review

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds evaluation suite for the agent tool-use loop.

The evaluation suite has no identified runtime blocker, but its import paths must satisfy the repository requirement before merging.

Findings

  1. P2 Tool feedback goes unchecked ▶
  2. P2 Tool arguments are not checked ▶
  3. P2 Relative imports violate app requirement ▶

Summary

The PR adds eight deterministic scenarios that drive the production OpenAI-compatible streaming tool loop, score outcomes, and write JSON and Markdown reports.

  • The suite exercises dispatch and error accounting without a provider key.
  • Its scripted model does not inspect tool feedback, and scoring does not check dispatched arguments, limiting the regressions it can detect.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Scenario script] --> B[Scripted model turns]
  B --> C[Production streaming tool loop]
  C --> D[Stub tool results]
  D --> C
  C --> E[Scoring]
  E --> F[JSON and Markdown reports]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): add agent tool-use evaluati..."

Comment thread apps/sim/evals/agent-tool-use/harness.ts Outdated
Comment thread apps/sim/evals/agent-tool-use/harness.ts
Comment thread apps/sim/evals/agent-tool-use/harness.ts
Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the
Agent block has a fallback model, and the handler serves the answer from
gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails
without the fallback row (checked locally: got gpt-4o, run errored).
Record a live run once, replay it forever through the real tool loop with no
key. EVAL_RECORD=1 wraps the live completion and writes each model call's
streamed chunks to fixtures/<scenario>.json; agent-tool-use.replay.test.ts
feeds them back through createOpenAICompatStreamingToolLoopStream and scores
them with the same checks.

- replay.ts: recording/replay completions + fixture I/O
- replay.test.ts: chunk round-trip and fixture I/O (key-free)
- live test records on EVAL_RECORD=1; test:evals:record script
- replay suite skips until a fixture exists; README documents the loop
Drive the Agent block through the executor with conversation memory on. The
memory read is stubbed per conversation id, so the provider request shows what
the handler assembled: prior history, then the new prompt, system prompt
preserved, correct conversation id. A wrong id surfaces as missing history and
fails (checked locally).

- agent-context/scenarios.ts: two context scenarios
- executor-harness.ts: memory seam + assembly/isolation checks
- test:evals:context script; README documents the suite
Run the same live scenarios across a list of models and write a scenario x
model matrix. models.ts resolves provider:model specs (DeepSeek, OpenAI, Groq,
OpenRouter) and reads each provider's key from <PROVIDER>_API_KEY.

- agent-tool-use.compare.live.test.ts: EVAL_MODELS x scenarios x trials
- report.ts: buildLiveComparisonReport + JSON/Markdown matrix
- report.test.ts: key-free aggregation coverage
- test:evals:compare script; README documents the spec format
Pass rates alone do not say why a model lost. Aggregate the failed check
names per model into the comparison report and add a Failed checks column.
Five cases that stress where models tend to fail: answering with no tool,
disambiguating near-identical tools, not inventing an answer from an empty
tool result, running a four-tool dependency chain, and picking settings over
a near-duplicate profile tool. Scripted expectations keep them deterministic;
the same cases run live.
Two live failures were eval design, not model failure:
- empty-result-no-hallucination rejected valid 'didn't find' / 'wasn't able
  to find' phrasing. Broaden the grounding check.
- near-duplicate-names required a userId the prompt never gave, so the model
  reasonably asked for it. Put the id in the prompt and the scripted call.
- long-chain-dependency: the prompt never gave a userId, so the model asked
  or skipped the profile step. Provide u-42 and let live runs require the
  three downstream calls rather than the exact four-step sequence.
- near-duplicate-names: one live trial called both tools; that is over-calling,
  not wrong-tool selection. Drop the forbidden-tool assertion in live mode.
Substring checks measure phrasing, not correctness. judgeAnswer scores an
answer against a weighted rubric with a judge model and returns structured
scores; runScenario gains an optional judge that adds a judge check. The
judge transport is an injectable OpenAI-compatible completion, so a recorded
transcript can replay it deterministically.

- judge.ts: rubric, prompt, JSON parsing/clamping, verdict
- judge.test.ts: parsing/weighting/clamping (key-free)
- judge.live.test.ts: grounded answer outscores an invented one (opt-in)
- test:evals:judge script; README documents it
@sudoKrishna sudoKrishna mentioned this pull request Oct 1, 2026
4 of 6 tasks
@sudoKrishna sudoKrishna changed the title feat(evals): add first agent tool-use evaluation suite feat(evals): add agent evaluation suite Oct 2, 2026
- Read tool feedback: the scripted model now asserts that each prior turn's
  tool results reached the next model call, so a loop that drops feedback
  fails the retrieval/planning/recovery cases.
- Check tool arguments: score every executed call against the scripted
  arguments, so a right-name/wrong-arguments call fails.
- Use absolute @/evals imports instead of relative ones, per the app rule.

Verified both new checks fail under mutation (bad marker, mutated args).
@sudoKrishna

Copy link
Copy Markdown
Author

Addressed all three findings in f619479:

  • the scripted model now reads and asserts tool feedback, so a dropped result fails the case;
  • a new tool-arguments check compares each executed call's arguments to the expected ones;
  • eval imports are absolute now.

This branch was previously deployed

1 inactive (outdated) deployment
Preview — 86c79d78 Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant