Skip to main content

Inject faults on purpose

Suites rarely see a dependency fail at the wrong moment. The broker rejects a publish. A webhook answers 503 to every call. Waiting for the environment to produce that on demand does not work. The suite injects it instead.

Level 5, lesson 4About 15 minutes
By the end
  • Swap one run piece for a failing one and pin the recovery.
  • Read a target that stays down through retries to the dead letter.
  • Read a real regression from the trace the repository kept.
Before you start
  • Point the suite at a real stack (lesson 3).
  • An OpenCSMS checkout and a container runtime for the runs. The kept trace is on this site.

The scenario​

The outbox makes ending a session transactional: the ended session and its event commit together, and a publish that fails is retried from the stored row. That design is only worth trusting if a failed publish is exercised on purpose.

The same goes for the notification path. One real regression also happened, and the repository keeps its failing trace as the evidence.

The fault is a substitution​

The outbox tests replace the API's event publisher for one test. The substituted publisher fails a bounded number of matching attempts and delegates every later attempt to the product's own RabbitMQ publisher:

FailFirstAttemptsEventPublisher.cs3 notes
1public async ValueTask PublishAsync<T>(string routingKey, T message, CancellationToken cancellationToken = default)
2{
3if (_matches is not null && !_matches(routingKey, message))
4{
5await _publisher.PublishAsync(routingKey, message, cancellationToken);
6return;
7}
8
9if (Interlocked.Increment(ref _attempts) <= _failures)
10{
11Interlocked.Increment(ref _observedFailures);
12throw new InvalidOperationException(
13$"Publish attempt {Attempts} for '{routingKey}' failed (injected for this test).");
14}
15
16await _publisher.PublishAsync(routingKey, message, cancellationToken);
17}
  1. Narrow the fault to this test

    The match predicate keeps another test's pending row on the real path, so parallel tests do not spend each other's failure budget.

  2. Fail a bounded number of attempts

    The first N matching publishes throw. The count is the outage the test wants to survive.

  3. Then travel the real path

    Every attempt after the budget delegates to the product's own publisher, so a retry is a real publish, not a mock.

From the suite, registered per test with Proto.Context.Override<IEventPublisher>(publisher).

The two tests pin the recovery:

  • AFailedPublishIsRetriedByTheDispatcherAndBillsOnce fails the request's own publish. The end still commits, the event is stored as one pending row, the dispatcher retries it, and the worker stores exactly one invoice.
  • ABrokerOutageIsRiddenOutByMoreThanOneBackedOffRetry fails two attempts and keeps the event invisible to the dispatcher until both have happened. The stored row records both failed attempts, then reports itself sent, and the session is still billed once.

Run them with a container runtime:

dotnet test tests/OpenCsms.Suite --filter "FullyQualifiedName~OutboxTests"

The target that stays down​

The notification journeys point the product at per-run WireMock fakes, and one journey makes the fake answer 503 for a single entity's notifications. The worker spends its three in-process retries and dead-letters the fourth attempt:

  • the target saw four requests, every one a 503;
  • the dead letter carries the event's own facts and x-opencsms-retries: 3;
  • the invoice the target never accepted is still stored, exactly once.

The rejection is keyed to the session in the request body, because the fake is shared with the notification journeys that run in parallel beside it. A fake that rejected everything would steal their deliveries.

The billing failure path follows the same shape: a session the worker cannot bill produces a failure notification, and a target that stays down dead-letters that report with the reason the worker recorded. No invoice exists for that session.

The regression that was kept​

One fault was not injected. Billing used to read the tariff at the moment the worker processed session.ended, so an operator who repriced a tariff while a car was still plugged in silently removed the idle fee for time already spent.

The failing trace is kept at docs/static/traces/opencsms-showpiece.prototrace and served at /traces/opencsms-showpiece.prototrace. It recorded the two assertions that failed together:

FAILED OpenCsms.Suite.Journeys.IdleFeeAfterTariffChange.TheIdleFeeStillAppliesAfterAReprice (496 ms)
the idle fee the session started under still applies
Assert.That(stored.IdleFeeAmount, Is.EqualTo(10.00m))
Expected: 10m
But was: 0m
22 kWh at the tariff the session started under, the start fee and the idle fee
Assert.That(stored.Total, Is.EqualTo(20.30m))
Expected: 20.3m
But was: 13.6m

The fix copies the tariff's terms onto the session when it starts, so a later reprice never reaches an open session. The journey now asserts the same numbers and passes in container mode.

The ProtoTrace viewer on the idle-fee failure: the failed test, the assertion message and the execution entry.

The viewer on the kept trace. The failing assertion and the operation that produced it are one story, which is the point of keeping the file.

Download the trace and open it in the viewer:

opencsms-showpiece.prototrace

What injection does not cover​

  • The injected publisher fails attempts; it does not kill the broker process. A broker that dies under a live consumer is the environment's fault to produce.
  • The fakes live inside the test process. A published or topology run cannot reach them, so those journeys skip there with their reason.
  • A retry count is bounded by the product's own policy. The tests assert the count the product promises, not an unlimited recovery.

Checkpoint​

A fake rejects every notification for one invoice with 503. How many requests does the target see before the delivery is dead-lettered, and what does the dead letter carry?

Verify
Read The target that stays down below: the target saw four requests, every one a 503; the dead letter carries the event's own facts and x-opencsms-retries: 3; the invoice is still stored, exactly once. The journey NotificationTargetOutagesAreDeadLettered in the suite asserts that shape.

What you learned​

  • A test can substitute one run piece for its own failing one and keep the rest of the path real.
  • A target that stays down ends in a dead letter with a recorded retry count, and the product state stays intact.
  • The regression trace is the evidence a fault leaves when it happens for real, and the repository keeps it.

Keep exploring​