diagnose_crash
Diagnose a GPU/ML training crash (CUDA OOM, NCCL timeout, Xid fault, NaN loss, checkpoint corruption, etc.) and get the root cause plus an escalating fix (primary → fallback → emergency). Provide ONE of: logs (paste), file (path to a log), or dir (folder to scan for the newest *.err/*.log). Uses ...
This record as markdown: /tools/denpex-mcp/diagnose-crash.md
What diagnose_crash does on Denpex
AI agents call diagnose_crash to retrieve information from Denpex without modifying anything. It is typically the context-gathering step in research, monitoring, and reporting workflows, before the agent takes action elsewhere.
| Parameter | Type | Required | Description |
|---|---|---|---|
dir | string | — | Directory to scan for the newest *.err / *.log / *.out / *.txt. |
file | string | — | Path to a log file to read from disk. |
logs | string | — | Paste the crash log / stderr (at least 20 characters). |
gpuRate | number | — | Your effective $/GPU-hour (enables a $ cost estimate). |
jobName | string | — | Optional job or run name (saved to history). |
telemetry | object | — | Hardware telemetry captured from the failing node. Keys the engine reads: `dmesg` (output of `dmesg -T`), `nvidia_smi` (`nvidia-smi -q`), `ibstat`. Sending thes |
codeContext | object | — | Relevant source files as { path: contents }. Only send real files you read — the engine never fabricates code context, and neither should you. |
topologyMap | object | — | Rank → node placement as { nodeName: [rank, ...] }, e.g. from the launcher. Lets the engine map a failing rank to a physical node and GPU. |
runtimeHours | number | — | Hours the job ran before crashing (enables a $ cost estimate). |
Parameters from the server's own tool schema.
Why diagnose_crash is rated Low
Even though diagnose_crash only reads data, uncontrolled read access leaks sensitive information and racks up API costs: an agent caught in a retry loop can make thousands of calls a minute without anyone noticing.
Risk signalsAccepts file system path (file) · Admin/system-level operation
Attacks that exploit this kind of access
The rule that runs diagnose_crash safely
PolicyLayer is an MCP gateway: it sits between your AI agents and Denpex, and checks every tool call against a rule you set before the call runs. Nothing changes on the server itself. For diagnose_crash, this is the rule to start with:
diagnose_crash is read-only, so it stays allowed. Everything else on the server is denied unless you say otherwise.
The button opens the PolicyLayer dashboard: create your workspace, connect Denpex, apply this rule, and every diagnose_crash call is checked against it from then on.
Questions about diagnose_crash
Diagnose a GPU/ML training crash (CUDA OOM, NCCL timeout, Xid fault, NaN loss, checkpoint corruption, etc.) and get the root cause plus an escalating fix (primary → fallback → emergency). Provide ONE of: logs (paste), file (path to a log), or dir (folder to scan for the newest *.err/*.log). Uses the Denpex cloud by default, or runs fully offline with DENPEX_LOCAL=1. It is categorised as a Read tool in the Denpex MCP Server, which means it retrieves data without modifying state.
diagnose_crash accepts 9 parameters: dir, file, logs, gpuRate, jobName, telemetry, codeContext, topologyMap, runtimeHours. The full parameter table on this page comes from the server's own tool schema.
Register the Denpex MCP server in PolicyLayer and add a rule for diagnose_crash: allow, deny, rate-limit, or require approval. Point your MCP client at the PolicyLayer proxy URL and the rule is enforced on every call, before it reaches Denpex. Nothing to install.
diagnose_crash is a Read tool with low risk. Read-only tools are generally safe to allow by default.
Yes. Add a rate_limit block to the diagnose_crash rule in your PolicyLayer policy. For example, setting max: 10 and window: 60 limits the tool to 10 calls per minute. Rate limits are tracked per agent session and reset automatically.
Set action: deny in the PolicyLayer policy for diagnose_crash. The AI agent will receive a policy violation error and cannot call the tool. You can also include a reason field to explain why the tool is blocked.
diagnose_crash is provided by the Denpex MCP server (denpex-mcp). PolicyLayer sits as a proxy in front of this server to enforce policies before tool calls reach the server.
More on Denpex, and thousands of servers like it.
This server
Across the catalogue