Skip to main content

Benchmarks

Same short test: raw 0.33 ms, ProtoTest without trace 0.95 ms (about 3x), with trace 2.55 ms (about 8x). The trace adds about 1.6 ms and 1 MB per test here. On suites whose tests talk to a database or a browser, that difference disappears into the setup the framework replaces.

These numbers come from the in-repo harnesses, and their scope is narrow: one machine (AMD Ryzen 7 9800X3D, Windows 11, .NET 8), synthetic suites where each test records one operation, EmbedSources = false and EmbedArtifacts = false, a raw baseline that creates a client per test but never opens a broker consumer, and medians taken after warmup. TraceScaleTests.cs covers trace size. PerTestPhaseBenchmarkTests.cs covers the phase profile. OverheadBenchmarkTests.cs covers the WebApplicationFactory comparison. The OpenCSMS numbers come from that repository's own eng/run-benchmark.ps1. They are indicative, not a contract: run the harnesses on your own hardware and CI image before quoting them, and expect run-to-run medians to move by around 20%.

QuestionHarness
How big is the trace, and how does it scaletests/ProtoTest.Core.Tests/TraceScaleTests.cs
What does one test cost, phase by phasetests/ProtoTest.Core.Tests/PerTestPhaseBenchmarkTests.cs
What does the framework add over raw WebApplicationFactorytests/ProtoTest.AspNetCore.Tests/OverheadBenchmarkTests.cs
What does a per-test broker tap costtests/ProtoTest.Messaging.RabbitMq.Tests/RabbitMqTests.cs (PerTestTapLifecycle_ShouldStayWithinTheSanityBound)
TestsTrace sizeRun timeStop + exportAllocatedPeak working set growth
1000.70 MB78 ms49 ms38 MB44 MB
1,0006.95 MB389 ms59 ms377 MB23 MB

Every synthetic test records one operation, one event, one entity-state change and one observation - roughly the evidence a small API test produces. That is about 7 KB of trace and about 0.4 MB of allocation per test at this size. The harness runs with EmbedSources = false and EmbedArtifacts = false; the defaults (embedding the suite's source files and attachment bytes) produce a larger file.

A short test, written twice​

The comparison an evaluator actually cares about: one short test - POST an order, read it back, check both bodies - written twice against the same in-process application and measured a whole test at a time. The raw side uses WebApplicationFactory with System.Net.Http.Json, creates a client per test and deserializes with JsonSerializer; the ProtoTest side runs the full lifecycle and asserts with Should.HaveHttpStatus(...).Should.MatchShape(...). Same requests, same assertions, 32 warmups and 256 measured, medians:

ModePer testp95StartCallCompleteAllocated
ProtoTest, tracing on2.55 ms3.79 ms0.66 ms1.23 ms0.55 ms1,354 KB
ProtoTest, tracing off0.95 ms1.29 ms0.10 ms0.79 ms0.05 ms327 KB
Raw stack (the same test)0.33 ms0.47 ms0.01 ms0.32 ms-66 KB

Read plainly: the same test costs about 3x the raw stack without tracing and about 8x with it. The trace adds roughly 1.6 ms and 1 MB per test here, because the short test produces two recorded calls, their observations and the shape validations. A 1,000-test suite pays about 2.6 s for the framework and the trace, against 0.3 s raw - on a suite whose tests talk to a database or a browser, that difference disappears into the setup the framework is replacing.

Where the difference comes from​

A micro comparison explains the number above: a single request (GET /ping), a bare lifecycle with no request at all, and the raw equivalent - 32 warmups and 256 measured, medians:

ModePer testp95StartCallCompleteAllocated
ProtoTest, tracing on1.54 ms2.26 ms0.62 ms0.32 ms0.54 ms1,003 KB
ProtoTest, tracing off0.34 ms0.50 ms0.09 ms0.21 ms0.04 ms171 KB
ProtoTest, lifecycle only (no request)0.06 ms0.15 ms0.04 ms-0.02 ms60 KB
Raw WebApplicationFactory0.05 ms0.16 ms0.00 ms0.04 ms-17 KB

As a waterfall, each layer adds its cost on top of the last:

raw baseline 0.05 ms
lifecycle only 0.06 ms framework bookkeeping is nearly free
+ REST client 0.34 ms +0.15 ms of capture, wrapping and body reads
+ trace 1.54 ms +1.0 ms of source locations and recording

Suite startup, measured through the first completed request so both sides pay for building the in-process server: ProtoTest 22 ms with tracing, 13 ms without, raw 12 ms.

  • The lifecycle itself is cheap. Start and complete a test with no request at all: 0.06 ms and 60 KB - about the raw baseline's whole per-test work. "Tracing off" is not "framework off": Enabled = false skips the activity listener, the recorder and trace export, but the context, hook pipeline, per-test client creation and outcome release still run.
  • The REST client adds about 0.15 ms over a raw request (0.20 vs 0.05 ms). It reads and captures the response body - the raw baseline leaves it unread until asked - records the response observation and entity state, and builds the typed wrapper. A raw test that asserts a body pays part of that difference itself.
  • The trace is the rest of the cost. Tracing adds roughly 1.0 ms and 0.7 MB per test over the same suite with tracing off (start +0.5 ms of setup, complete +0.5 ms of recording and finalizing, call +0.1 ms).
  • Startup is nearly identical (22 vs 12 ms): booting the application dominates, not the framework.
  • The raw baseline creates an HttpClient per test, matching ProtoTest's per-test model; a raw suite that reuses one client will be faster than the numbers shown.
  • A real test does much more than one request. When a test starts containers, drives a browser or seeds a database, these milliseconds are noise; this table is the cost of the framework's own bookkeeping, not of a typical test.

The per-test phase profile​

One lifecycle broken into the phases the framework owns, on a bare host with no application, 64 warmups and 512 measured iterations per phase, medians. The host runs with EmbedSources = false and EmbedArtifacts = false and source locations on (the default), the same trace settings the OpenCSMS benchmark uses; the harness is tests/ProtoTest.Core.Tests/PerTestPhaseBenchmarkTests.cs.

PhaseMedianp95Allocated
DI scope, create + dispose0.0001 ms0.0002 ms0.2 KB
Context create0.0027 ms0.0034 ms6.8 KB
Context dispose (resources + scope)0.0306 ms0.0350 ms29.3 KB
Per-test clock register0.0002 ms0.0020 ms1.2 KB
Attribute + skip resolution0.0005 ms0.0007 ms0.6 KB
Attribute + skip resolution, one [Requires...]0.0017 ms0.0020 ms1.6 KB
Trace recorder, tracing off, 5 operations0.0021 ms0.0029 ms4.8 KB
Trace recorder, tracing on, 5 operations0.1077 ms0.1388 ms136.8 KB
Trace recorder, tracing on, source locations off0.0027 ms0.0031 ms8.6 KB
Lifecycle, tracing off (start + complete)0.0154 ms0.0162 ms20.6 KB
Lifecycle, tracing on0.3302 ms0.5840 ms374.8 KB
Lifecycle, tracing on, source locations off0.0200 ms0.0205 ms29.5 KB
Lifecycle, tracing on, 5 more operations0.4910 ms0.7293 ms519.8 KB
Lifecycle, tracing on, 1 declared attachment0.3782 ms0.6423 ms377.8 KB

Read plainly: the framework's own work is microseconds. DI scope creation, context creation, per-test clock registration, attribute and skip resolution and teardown all stay under 0.04 ms. Tracing costs about 0.33 ms and 375 KB per test, and source-location capture is essentially all of it: with CaptureSourceLocations = false the same lifecycle costs 0.02 ms and 30 KB.

OpenCSMS at 1,000 tests​

The OpenCSMS repository's eng/run-benchmark.ps1 runs the product itself - API, billing worker, real PostgreSQL and RabbitMQ from the environment - through one health-check test cycle per iteration (1,000 iterations, 100 warmups), against a raw WebApplicationFactory baseline, plus a 1,000-test seeded journey run. Same machine and .NET 8.

The showpiece trace is opencsms-showpiece.prototrace: the IdleFeeAfterTariffChange journey as it failed, when a reprice during an open session changed the session's billing. The fixed journey asserts the session's original tariff and runs in the OpenCSMS suite.

ModePer testStartCallComplete
ProtoTest, tracing on35.15 - 36.17 ms11.95 - 12.09 ms5.69 - 5.72 ms17.36 - 17.87 ms
ProtoTest, tracing off30.74 - 32.40 ms11.14 - 11.65 ms5.55 - 5.57 ms13.67 - 15.91 ms
Raw WebApplicationFactory4.79 - 5.00 ms0.00 ms4.79 - 5.00 ms-

Suite startup through the first completed request lands between 156 and 179 ms with tracing, 154 and 176 ms without, and 73 and 83 ms raw. The 1,000-test health-check trace is 25.8 MB (about 25 KB per test); the 1,000-journey trace is 38.8 MB, with a 49.3 to 54.8 ms journey median.

Read the table honestly: the per-test total moves inside the run-to-run band, and the tracing-off leg is the most order-sensitive, because it runs second in the harness after the first leg has created and deleted about 3,300 tap queues on the shared broker.

The run's own trace explains where the milliseconds are, and they are not the framework. Medians per operation over the 1,100 health-check tests of the recorded traces:

OperationMedianWhat it is
client.initialize for the messaging client11.09 - 11.26 msthe test's own consumer, prepared for the suite's three tap destinations
resource.release for that consumer17.00 - 17.49 msits tap queues and channels released
http.request (GET /healthz)5.65 - 5.69 msthe request itself, through the REST client
Everything elseunder 1 msHTTP client init, hooks, attributes, execution and teardown scaffold, trace included

The controlled A/B is the honest number: three tapped destinations against the same host and broker with none, the no-destination host at 0.03 ms (PerTestTapLifecycle_ShouldStayWithinTheSanityBound in the RabbitMQ suite), and the tapped prepare phase at about 11-12 ms.

So about 30 of the 36 ms is the suite's per-test broker isolation, about 6 ms is the request, and under 1 ms is ProtoTest. The raw baseline never opens a broker consumer; a test that never touches messaging still pays the tap cost, because Tap promises the destination is bound before the test acts.

  • "Tracing off" works; it is just not the cost. Enabled = false installs no activity listener, the recorder drops operations and records, and no archive is written (pinned by ProtoEvidenceBoundaryTests.DisabledTracing_ShouldNotRecordOperationsOrWriteAnArchive). The benchmark numbers move within noise because the trace is about 0.33 ms of a 36 ms cycle, and the phase profile above shows where the 0.33 ms sits: source-location capture.
  • Preparation is the floor. Each tap costs five broker round trips (channel, queue declare, two bindings, consume), and the taps are prepared concurrently, each on its own channel, so their round trips overlap: three taps' prepare phase lands at about 11-12 ms. Every destination is still attempted, so one that cannot be prepared fails only the await that names it.

The viewer at 1,000+ tests​

The viewer was measured on a generated 1,200-test trace derived from the committed demo trace: viewer/scripts/generate-scale-trace.mjs clones each demo test until the requested count, rewriting ids, names and timestamps, and writes small placeholder payloads for the artifacts. Every test keeps the demo's real operations, checks, observations, sections, events and state, so the viewer exercises its real paths. The generated file is 8.6 MB zipped (57.8 MB of spans JSON, 35.6 MB of state JSON); it is a heavier suite per test than the size table above, which counts one operation per test.

Method: a production build (npm run build), Microsoft Edge 154.0.4258.37 headless at 1440x900, no CPU or network throttling, on the same Ryzen 7 9800X3D machine. Each number is the median of five cold runs: a fresh page, the trace loaded through the viewer's own file input, then one interaction. The harness is viewer/scripts/measure-viewer.mjs; it serves the built dist over loopback, so no dev server is involved. Before and after the pass:

StepBeforeAfter
Open the trace and render the run list1,356 ms1,316 ms
Filter to "Needs attention"77 ms53 ms
Clear the filter (1,200 rows return)182 ms162 ms
Type "GraphQL" in the run search365 ms231 ms
Open a test with 169 spans167 ms151 ms
Return to the run list650 ms566 ms

Read plainly: a cold open of a suite this size takes about 1.3 seconds, and the remaining run list steps land between 50 ms and 570 ms. The metric views stay fast at this size: the spans tab, the state view and the inspector all answer in 33 to 50 ms for a 169-span test.

What the pass changed, all inside the existing views:

  • Each test computes its display name, title, group and search text once and caches it, instead of re-deriving them for every row on every render.
  • The run view is kept alive when a test opens, so returning moves the cached DOM instead of rebuilding a thousand rows.
  • Each run and rail row contains its own layout and paint (contain: layout paint), so one row's change never re-lays-out the list.

What remains, with the measurement that shows it:

  • The cold open is dominated by reading and parsing the two JSON documents and the first full layout (five long tasks, the longest about 550 ms). Removing that needs a streaming reader.
  • The run list and the outcome strip render one element per test, so returning to the run and changing the filter move about a thousand DOM rows. Removing that needs a windowed list.
  • content-visibility: auto on the rows was measured and rejected: it skipped off-screen layout on the first render but made typed filtering about twice as slow, because every row added or removed during filtering pays its bookkeeping.

Levers​

LeverEffect
trace.CaptureSourceLocations = falseDrops the code.file.path / code.line.number / code.function.name attributes; this is essentially the whole trace cost (0.32 ms -> 0.02 ms and 375 KB -> 30 KB per bare test in the phase profile)
trace.EmbedSources = falseKeeps source locations but stops embedding the files they point at
trace.EmbedArtifacts = falseDeclares attachments (name, media type, size) without reading or writing their bytes
trace.MaxArtifactBytesCaps any single artifact; an over-limit attachment becomes an error artifact (default 64 MB)
Attachment capture per integrationRequest/response/screenshot/trace capture is opt-in (CaptureAttachments(...)); leaving it off is the biggest lever

Re-running​

dotnet test tests/ProtoTest.Core.Tests --filter TraceScaleTests

The harness prints [scale] tests=... trace=... run=... stop+export=... allocated=... for each size. Its assertions are sanity bounds - a trace must stay proportional to the test count - not performance targets.

The overhead comparison re-runs with:

dotnet test tests/ProtoTest.AspNetCore.Tests --filter Category=Benchmark --logger "console;verbosity=detailed"

It prints [overhead] lines per mode - the short test written both ways, the single-request comparison, the bare lifecycle and each side's startup - and its assertions are generous sanity bounds too, not targets.

The per-test phase profile re-runs with:

dotnet test tests/ProtoTest.Core.Tests --filter PerTestPhaseBenchmarkTests --logger "console;verbosity=detailed"

The broker tap comparison re-runs with (a broker from ProtoTest__Messaging__RabbitMq__ConnectionString or a container runtime):

dotnet test tests/ProtoTest.Messaging.RabbitMq.Tests --filter PerTestTapLifecycle --logger "console;verbosity=detailed"

The viewer numbers re-run from the viewer directory:

npm run scale:trace
npm run build
npm run scale:measure -- --trace .perf/scale-1200.prototrace --tests 1200 --runs 5

The generator prints the trace's size, the harness prints every run and its median, and --profile also writes CPU profiles and renderer time (script, layout, style) per measured step. Neither script is part of the build or CI; the harness needs an installed Chromium-family browser (Edge by default) and puppeteer-core, which is a viewer dev dependency.

Parallel execution​

A suite with containers and real browsers was stable at 8 and 32 workers on a 16-core machine; at 64 workers one run showed a single failure that never reproduced, and the passing rerun overwrote the failed run's single trace file before it could be read. The concurrency page records the runs, the failure mode, and the current re-runs: Core green three times at 64 workers (380 passed each), Northstar.ProtoTest with containers and a real browser green twice at 64 workers.