Skip to content

The Drift Log · Certifying that code does something

Property-Based Tests for Provisional Verdicts

10 October 2026 · 3 min read · 719 words · inference

A crystal artifact under orange probe beams inside a test rig

Property-based testing strengthens provisional verdicts by replacing narrow examples with invariant checks across generated inputs.

Provisional verdicts look solid until an edge case hits

A recent run of the test harness returned PROVISIONAL for a library that later crashed on a null‑pointer input. The harness had exercised the public API with a handful of hand‑crafted examples. Those examples covered the happy path but omitted the boundary where a caller passes undefined or an empty collection. The result felt flaky: the same artifact could pass one day and fail the next, depending on the exact data the harness generated.

The root cause is a narrow test surface. A provisional verdict is only as trustworthy as the properties it checks. When the test suite consists of a few concrete examples, the harness cannot distinguish “the code works for the cases we wrote” from “the code works for all valid inputs”. The ladder of verdicts described in the four‑verdict post makes this explicit: PROVISIONAL sits between “it compiles” and “it is correct”. To move up the ladder, the provisional rung must be reinforced, not replaced.

Property‑based testing fills the gaps

Property‑based testing (PBT) replaces hand‑written examples with automatically generated inputs that satisfy a declarative contract. Instead of writing:

test('sum([]) returns 0', () => {
  expect(sum([])).toBe(0);
});

you declare a property:

import { assert } from 'vitest';
import { fc } from 'fast-check';

fc.assert(
  fc.property(fc.array(fc.integer()), arr => {
    const result = sum(arr);
    // sum of an empty array is 0, otherwise result equals manual reduction
    if (arr.length === 0) {
      assert.equal(result, 0);
    } else {
      assert.equal(result, arr.reduce((a, b) => a + b, 0));
    }
  })
);

The framework generates thousands of arrays, including empty, single‑element, large, and arrays containing extreme integer values. The property encodes the invariant you care about – the relationship between input and output – rather than a single concrete case.

When such a property is added to the provisional harness, the harness no longer depends on a fixed set of examples. It exercises the artifact across a spectrum of inputs that would be impractical to write manually. The result is a higher‑signal PROVISIONAL verdict: the harness has observed the code respecting the declared contract under many more conditions.

Integrating PBT into the existing harness workflow

The harness executes the tests that are supplied to it before producing a verdict. Swapping a subset of those tests for PBT is straightforward:

  1. Identify the most brittle provisional verdicts – those that have produced false positives in the past.
  2. Write a property that captures the intended behaviour. Keep the property narrow; a single well‑defined invariant is easier to reason about than a sprawling specification.
  3. Add the property to the test harness configuration. The harness can be configured to consume a TAP or JUnit report, so a Fast‑Check suite can be dropped in without altering the harness core.
  4. Run the harness on a known‑good artifact. Verify that the verdict remains PROVISIONAL and that the new property passes consistently.
  5. Commit the property file alongside the artifact’s source. Future harness runs will automatically include the expanded coverage.

Because the harness records each verdict independently, you do not need to sum the property‑based results with existing counts. The harness will still report a single PROVISIONAL verdict, now backed by a richer exercise set.

Limits you need to accept

PBT does not eliminate the need for hand‑crafted examples. Certain behaviours – such as interaction with external services, timing constraints, or UI rendering – cannot be expressed purely as input‑output invariants. Moreover, the generator may never produce a pathological case that exists in production. In those situations you still need targeted tests.

Another practical limit is test runtime. Generating thousands of inputs can increase the harness duration. The typical compromise is to run a smaller seed set on every pull request and a larger, more exhaustive run on the nightly schedule. The harness itself can enforce a timeout, reporting INCONCLUSIVE if a property cannot be exercised within the allotted time – a signal that the property or the artifact needs refinement, not a failure of the artifact.

A concrete step for Monday

Pick the most recent provisional verdict that surprised you. Write a single property that captures the core invariant of the failing scenario. Add it to the test harness configuration and run the harness locally. If the verdict stays PROVISIONAL and the property passes, you have immediately increased confidence in that artifact. If the property fails, you have uncovered a real defect before the artifact reaches production.

Repeating this process for each flaky provisional result will steadily raise the overall signal quality of the harness without inflating the test suite with endless example cases. The ladder of verdicts becomes more informative, and your release decisions rest on a sturdier foundation.

Keep reading

Next in the log

The Strategic Master Library · written and reviewed under the house's own epistemic rules: nothing claimed that we cannot show.