run_self_test
Run an end-to-end health check of the bridge: create a temp object, add a component, assign + measure a model, capture a screenshot, recompile a temp asset, remove the component — then clean it all up, reporting pass/fail per subsystem. Use it to confirm the install works or to catch regressions ...
This record as markdown: /tools/sbox/run-self-test.md
What run_self_test does on Sbox
AI agents invoke run_self_test to trigger actions in Sbox. What it does depends on the arguments the agent supplies, and its effects often reach beyond the immediate call: builds kicked off, notifications sent, workflows started.
Why run_self_test is rated High
While the tool is described as 'safe and self-cleaning', it executes a series of game engine operations that could have side effects depending on the engine state. Asset recompilation and object creation are executable operations that trigger external processes.
From the tool's definition Tool performs multiple operations: 'create a temp object, add a component, assign + measure a model, capture a screenshot, recompile a temp asset, remove the component — then clean it all up'.
Attacks that exploit this kind of access
The rule that runs run_self_test safely
PolicyLayer is an MCP gateway: it sits between your AI agents and Sbox, and checks every tool call against a rule you set before the call runs. Nothing changes on the server itself. For run_self_test, this is the rule to start with:
run_self_test stays usable, but rate-capped: a runaway agent can't fire it dozens of times a minute. Everything else on the server is denied unless you say otherwise.
The button opens the PolicyLayer dashboard: create your workspace, connect Sbox, apply this rule, and every run_self_test call is checked against it from then on.
Questions about run_self_test
Run an end-to-end health check of the bridge: create a temp object, add a component, assign + measure a model, capture a screenshot, recompile a temp asset, remove the component — then clean it all up, reporting pass/fail per subsystem. Use it to confirm the install works or to catch regressions before a release. Safe and self-cleaning; refuses to run in play mode. It is categorised as a Execute tool in the Sbox MCP Server, which means it can trigger actions or run processes. Use rate limits and argument validation.
Register the Sbox MCP server in PolicyLayer and add a rule for run_self_test: allow, deny, rate-limit, or require approval. Point your MCP client at the PolicyLayer proxy URL and the rule is enforced on every call, before it reaches Sbox. Nothing to install.
run_self_test is a Execute tool with high risk. Execute tools should be rate-limited and have argument validation enabled.
Yes. Add a rate_limit block to the run_self_test rule in your PolicyLayer policy. For example, setting max: 10 and window: 60 limits the tool to 10 calls per minute. Rate limits are tracked per agent session and reset automatically.
Set action: deny in the PolicyLayer policy for run_self_test. The AI agent will receive a policy violation error and cannot call the tool. You can also include a reason field to explain why the tool is blocked.
run_self_test is provided by the Sbox MCP server (sbox-mcp-server). PolicyLayer sits as a proxy in front of this server to enforce policies before tool calls reach the server.
More on Sbox, and thousands of servers like it.
This server
Across the catalogue