Skip to main content

Diagnosis

One test in the run fails. The log shows expected 99, actual 0. The diagnosis names the request that produced it, the surrounding code, and what the operation changed. It answers all of that from the trace, with recorded evidence and no guessing.

The summary in a terminal​

dotnet tool install --global ProtoTest.Cli
prototest summary TestResults/run.prototrace

Every line of a failing block says one thing:

prototest summary8 notes
1ProtoTest trace 2.0 · run 29e344f9cf54431ca7d8bad3f87a1749 · 2026-09-28 09:55:33Z - 2026-09-28 09:55:33Z
22 tests · 1 failed · 1 succeeded
3
4FAILED orders match their shape (16 ms)
5Shape mismatch failed with 1 error(s):
6• [$.orderId]: Values did not match. (Expected: '7', Actual: '42')
7at artifacts/fixture-gen/Program.cs:65 (Program.<<Main)
8assert.json.shape · failed
9cause: assertion (1 mismatch)
10mismatch: $.orderId: expected 7, actual 42
  1. The document

    Trace format, run id, and the recorded time range of the run.

  2. The counts

    Every test by outcome. A partial test passed its runner outcome but something inside it failed, and the summary does not hide it.

  3. The test

    Outcome, name and duration. Every test that did not fully succeed gets a block like this one.

  4. The recorded error

    The message the run recorded, so it points at the code that failed.

  5. The source location

    The file and line of the selected failure.

  6. The selected operation

    The kind and name of the failing operation, and its status. The kind prints once when it is the name.

  7. The rule

    The diagnosis rule that matched. An assertion lists its mismatches; an operation error names the error.

  8. The mismatches

    Path, expected and actual, capped at three per block with a truncation line.

A failed run gate gets its own block at the end, with the gate's message and details. The output above is the committed MCP test fixture; a run with nothing to report prints the header, the counts and All green.

The CLI reference lists the verb's arguments, the exit codes and the other three commands.

The same document for an agent​

get_diagnosis returns the same run as one JSON document. The fields an agent works from:

FieldWhat it holds
digestVersion, traceFormatVersion, runId, traceFilewhich document this is and which archive it came from
startedAtUtc, completedAtUtc, environmentwhen the run happened and on which runtime and OS
outcomesthe per-outcome test counts
failuresevery test that did not fully succeed: its outcome, the selected failure, the rule that explains it, the mismatches, the findings and the artifacts
gates, findingsthe run's own verdicts and the evidence tests recorded
coverage or coverageAbsentReasonthe totals the embedded report published, or why there is no report

get_failure returns one test's failure entry when the agent wants the short version. A trimmed result from the same fixture:

{
"runId": "29e344f9cf54431ca7d8bad3f87a1749",
"test": {
"testId": "00002",
"name": "orders match their shape",
"outcome": "failed",
"durationMs": 16.4395
},
"failure": {
"kind": "assert.json.shape",
"name": "assert.json.shape",
"phase": "execution",
"status": "failed",
"errorType": "ProtoTest.Json.JsonShapeMismatchException",
"errorMessage": "Shape mismatch failed with 1 error(s):\r\n • [$.orderId]: Values did not match. (Expected: '7', Actual: '42')",
"sourceFile": "artifacts/fixture-gen/Program.cs",
"sourceLine": 65
},
"artifacts": [
{ "name": "00002-rest-01-expected-shape", "mediaType": "application/json", "sizeBytes": 29 },
{ "name": "00002-rest-01-response.json", "mediaType": "application/json", "sizeBytes": 30 }
]
}

The failure selector matches the viewer's. It picks the deepest failing operation. An assert.* check outranks an error. Phase spans rank last. A cancelled operation never outranks a failed one. prototest summary, the MCP tools and the viewer therefore select the same failure and tell one story.

The rules​

A non-succeeded test is explained by the first rule that matches, in this order:

RuleThe recorded evidenceThe cause the diagnosis states
assertiona failed assert.* operation with recorded checks or shape.mismatchesthe mismatch list, with path, expected and actual, and the subject the operation is about
operation errora failed operation with an error type and messagethe error plus its ancestor call
runner failurethe runner reported a failure and no operation error existsthe message and the recorded source location
findingthe runner outcome stayed green but a teardown or rollback finding attachedthe finding, and the operation it names

A failed run gate is not a per-test rule; it gets its own block in the document, as above.

Anything else is reported as unexplained, with the run and test ids and a pointer to the viewer. The diagnosis never guesses. If the evidence does not carry a cause, the document says so.

The context package​

get_diagnosis with detail=context returns what an agent needs to fix one failure. Everything the agent reads arrives around the one operation it has to fix:

ancestors (≤ 32)
│
section previews (16) ──┤
source snippet (41 lines) ─┤
state changes (≤ 10) ─────┤──▶ the selected failure ──▶ artifacts (≤ 10, metadata)
the nearest call ancestor ──┘ operation └─▶ report rows (≤ 10)
  • the selected failure operation: kind, name, phase, status, error, source file, line, function and the recorded attributes;
  • its ancestor chain to the test execution, and the nearest call ancestor (http.request, graphql.operation, grpc.call, a messaging.* operation);
  • the operation's sections (checks, diff, code, fields) with payload previews;
  • the embedded source snippet around the failing line, when the archive carries it. A file over the 512 KB embed cap, or a run with EmbedSources off, is absent with that reason;
  • the artifacts reachable from the failing operation, listed with media type and size. The MCP tool lists metadata only; reading the content is a library call (ReadContext(..., includeArtifactContent: true)), bounded at 64 KB;
  • the state changes the failing operation caused;
  • the report rows for the test: findings, run gates and the coverage row the operation touched.

Limits​

  • Deterministic and offline: one archive and one report in, one JSON document out. No model runs inside ProtoTest, no network call is made, and nothing is written. The same archive produces byte-identical JSON.
  • Coverage, gates and findings come from the JSON report a ProtoTest.Reporting sink embedded. Without a report, coverage is omitted with the reason, and gates and findings fall back to the trace's own records. Coverage is never recomputed from spans.
  • Hard caps bound every document and context:
What is cappedThe cap
mismatches in one failure25
findings, run gates and artifacts in one document10 each
an error message4,000 characters
one section preview4 KB
the source snippet41 lines, 512 KB of embedded source
the ancestor chain32
state items in the context package10
sections in the context package16
coverage rows in the context package10
  • An archive read from a stream has no file to read, so its state, sources and artifact content are absent with that reason. Open the file instead.
  • The rules read recorded evidence only. A failure whose evidence was not recorded is unexplained, and the viewer is the place to read the rest of the story.

Run prototest summary on your own trace, or ask your agent for get_diagnosis with detail=context on the newest failure. The output names the file, the line and the subject, which is enough to open the code and fix it.