compare_eval_runs
Compare two eval runs to decide whether a change should ship. Returns side-by-side scores and a deploy/revert recommendation. Rule: no change ships without an eval improvement.
This record as markdown: /tools/io-github-homenshum-nodebench/compare-eval-runs.md
What compare_eval_runs does on Nodebench
AI agents call compare_eval_runs to retrieve information from Nodebench without modifying anything. It is typically the context-gathering step in research, monitoring, and reporting workflows, before the agent takes action elsewhere.
Why compare_eval_runs is rated Low
The tool compares evaluation run data and returns scores with a recommendation. It does not execute, modify, or delete anything — it reads and analyzes existing eval run results to produce a comparison output. The deploy/revert recommendation is advisory only, not an action itself.
From the tool's definition Compare two eval runs to decide whether a change should ship. Returns side-by-side scores and a deploy/revert recommendation.
Attacks that exploit this kind of access
The rule that runs compare_eval_runs safely
PolicyLayer is an MCP gateway: it sits between your AI agents and Nodebench, and checks every tool call against a rule you set before the call runs. Nothing changes on the server itself. For compare_eval_runs, this is the rule to start with:
compare_eval_runs is read-only, so it stays allowed. Everything else on the server is denied unless you say otherwise.
The button opens the PolicyLayer dashboard: create your workspace, connect Nodebench, apply this rule, and every compare_eval_runs call is checked against it from then on.
Questions about compare_eval_runs
Compare two eval runs to decide whether a change should ship. Returns side-by-side scores and a deploy/revert recommendation. Rule: no change ships without an eval improvement. It is categorised as a Read tool in the Nodebench MCP Server, which means it retrieves data without modifying state.
Register the Nodebench MCP server in PolicyLayer and add a rule for compare_eval_runs: allow, deny, rate-limit, or require approval. Point your MCP client at the PolicyLayer proxy URL and the rule is enforced on every call, before it reaches Nodebench. Nothing to install.
compare_eval_runs is a Read tool with low risk. Read-only tools are generally safe to allow by default.
Yes. Add a rate_limit block to the compare_eval_runs rule in your PolicyLayer policy. For example, setting max: 10 and window: 60 limits the tool to 10 calls per minute. Rate limits are tracked per agent session and reset automatically.
Set action: deny in the PolicyLayer policy for compare_eval_runs. The AI agent will receive a policy violation error and cannot call the tool. You can also include a reason field to explain why the tool is blocked.
compare_eval_runs is provided by the Nodebench MCP server (nodebench-mcp). PolicyLayer sits as a proxy in front of this server to enforce policies before tool calls reach the server.
More on Nodebench, and thousands of servers like it.
This server
Across the catalogue