Skip to main content

Testing

The API has three layers of automated verification:

LayerWhat runsWhen
Unit + staticPHPUnit (real Doctrine metadata, no database), phpcs, psalm, phpstan (incl. phpstan-doctrine)Every PR (php.yaml)
IntegrationReal repository queries executed against a local MySQL, plus schema checksLocally via composer test:integration
Functional (E2E)WebDriver/Cucumber suites from dvsa/vol-functional-testsPost-deploy per environment (cd.yaml)

This page documents how repository unit tests assert against real Doctrine, the integration layer, the two baseline mechanisms that guard the Doctrine entity metadata, the Doctrine deprecation report both suites print, and the three tests that guard output escaping in the table render pipeline.

Repository unit tests​

Location: app/api/test/module/Api/src/Domain/Repository/, base class RepositoryTestCase. They run in the normal unit suite on every PR.

These tests assert on the DQL a repository actually builds, using a real Doctrine QueryBuilder rooted in the real entity metadata. Nothing opens a database connection.

Why not a query builder double​

They used to assert against RepositoryTestCase::createMockQb(), a hand-rolled double whose mockOrderBy() was $sort . ' ' . $order — a verbatim reimplementation of ORM 2's Expr\OrderBy::add(). That double did not stub Doctrine, it forked it, so the suite went on asserting ORM 2 semantics after the ORM 3.7 upgrade and could not notice the library changing underneath it. That is how the ' ASC' sort direction bug (#1785) reached production with 887 repository tests green.

With a real builder, a mistyped field or a join to an association that does not exist fails the test where production would fail, and getDQL() is the query the repository really asked for.

The helpers​

HelperWhat it gives you
setUpRealSut($class, $mockSut = false)The repository wired to a real QueryBuilder service and the application's own query partials. Use it instead of setUpSut() for anything that builds a query.
createRealQb()A real builder rooted on the repository's own entity and alias, wired in as the one createQueryBuilder() hands back. Assert on $qb->getDQL() and $qb->getParameter(...).
createRealQbs([Entity::class => 'alias', ...])One distinct builder per entity/alias, keyed by alias, for methods that build more than one query — a shared builder would let a sub-select write into the outer query.
newRealQb()A bare real builder, nothing selected and no wiring, for repositories that build off the EntityManager rather than through createQueryBuilder().
compileDql($dql)Runs DQL through the real parser and returns the SQL, to assert a query is genuinely valid — or, for a pinned defect, that it is not.
$qb->willReturn($rows)Declares the rows the query returns. Shorthand for $qb->stubbedQuery()->shouldReceive('getResult')->andReturn($rows); pass a second argument for a different result method.

Two things stay mocked, and only two:

  • getQuery() on the builder. Repositories build and execute in a single call, so the seam has to be inside the builder rather than around it. It returns a mock of the concrete Doctrine\ORM\Query, which keeps Doctrine's real parameter types — that is how the unit lane now rejects the null hydration mode that shipped in #1787, where the old shouldReceive('getQuery->getResult') chain built an anonymous mock whose untyped __call swallowed it.
  • The EntityManager the repository holds, because it is used only to hand out repositories and to persist. The metadata EntityManager the builders and partials read from is a separate, real one, and neither path connects.

A worked example, from ReasonTest:

final class ReasonTest extends RepositoryTestCase
{
public function setUp(): void
{
$this->setUpRealSut(Repo::class, true);
}

public function testApplyListFilters(): void
{
$qb = $this->createRealQb();

$this->sut->applyListFilters($qb, ReasonList::create(['isNi' => 'Y']));

$this->assertSame(
'SELECT m FROM ' . Entity::class . ' m'
. ' WHERE m.isNi = :isNi AND m.isVisibleInInternal = :isVisibleInInternal',
$qb->getDQL(),
);
$this->assertTrue($qb->getParameter('isNi')->getValue());
}
}

What it costs​

Support\DoctrineMetadata reuses phpstan-object-manager.php — the same loader PHPStan and the integration suite build on, with serverVersion pinned so DBAL never opens a connection. It is the third consumer of that file, and the only one that needs it exactly as-is. Metadata parses once per process (~200 ms) and the DQL parser caches, so each query after that costs well under a millisecond; the 943 repository tests run in about 30 seconds.

Support\QueryPartials builds a real QueryPartialServiceManager from the application's own module/Api/config/module.config.php rather than a restatement of it, so the partials a test exercises cannot drift from the ones production wires.

What it still cannot catch​

The query is built and compiled, never executed. Anything that depends on actually running it — hydration, streaming, the schema the DQL lands on — belongs in the integration suite below.

A few repositories have no query builder to make real: DataGovUk, DataDvaNi and CompaniesHouseVsOlcsDiffs go straight to a DBAL connection with native SQL or a stored procedure, and PostcodeEnforcementArea uses findOneBy(). Those tests keep setUpSut() and expectQueryWithData().

Integration test suite​

Location: app/api/test/integration/ (namespace Dvsa\OlcsTest\Integration), own PHPUnit config app/api/phpunit-integration.xml.

cd app/api
composer test:integration

What it needs​

A running local database with the schema and test data loaded:

docker compose up -d db
npm run refresh # from the repo root; loads the Liquibase schema + testdata

If the database is unreachable the whole suite skips (it never fails a machine without Docker). Connection defaults match the compose stack and can be overridden with environment variables: VOL_TEST_DB_HOST (127.0.0.1), VOL_TEST_DB_PORT (3306), VOL_TEST_DB_USER (root), VOL_TEST_DB_PASSWORD (olcs), VOL_TEST_DB_NAME (olcs_be).

How it works​

Support/Database.php builds the minimal real object graph without booting the MVC application: an EntityManager reusing the standalone phpstan-doctrine loader (phpstan-object-manager.php — real entity metadata, custom DBAL types and DQL functions), plus the query partial / db query / repository service managers wired exactly as the production RepositoryFactory expects. IntegrationTestCase gives each test a transaction that is rolled back in tearDown(), so tests may insert whatever fixture data they need.

Repository tests fetch repositories by their production short name ($this->repo('LicenceVehicle')) and execute real DQL against the real schema. That is what separates them from the repository unit tests above, which build and compile the same queries but never run one: VOL-7445, where iterate() → toIterable() broke three CSV exports that all had green unit tests, would still pass a compiled-DQL assertion today.

Doctrine deprecation report​

Both PHPUnit configs register Support\DoctrineDeprecations as a bootstrap extension, so every run ends with the Doctrine deprecations it triggered:

5 distinct Doctrine deprecations (248 occurrences):
207x https://github.com/doctrine/orm/issues/11313
34x https://github.com/doctrine/collections/pull/472
4x https://github.com/doctrine/collections/pull/389
2x https://github.com/doctrine/orm/issues/12192
1x https://github.com/doctrine/orm/pull/12005

Each line is the link Doctrine documents that deprecation under, and how many times it was reached — which is why a schema comparison shows thousands of one and a single figure of another.

Reported, never fatal. The extension uses Doctrine's tracking mode, which records deprecations without raising an error, so failOnDeprecation is unaffected. These are advance notice of the next major, not a broken build.

The DOCTRINE_DEPRECATIONS environment variable cannot do this job. In trigger mode the library calls @trigger_error() — with the suppression operator — so PHPUnit's error handler discards it and nothing is displayed.

It is registered in both configs deliberately. CI runs bare vendor/bin/phpunit and nothing under .github/ references the integration config, so the unit suite is the only place the report is visible there; it has something to report because the repository tests build real queries against real entity metadata. The integration suite adds whatever a real connection and the schema comparison reach.

Turning the report on surfaced three first-party problems:

  • Both metadata EntityManagers pinned serverVersion to '8.0', which version_compare puts below '8.0.0', so DBAL resolved the legacy MySQLPlatform rather than MySQL80Platform. PHPStan and the schema drift comparison had been reading the wrong platform. Fixed in both loaders; the drift baseline does not move.
  • SchemaDriftTest called Table::removeForeignKey(), deprecated in favour of dropForeignKey() — 2138 of the occurrences reported. Fixed.
  • 19 entity OrderBy attributes pass 'ASC' / 'DESC' as strings where ORM 4 will require a SortDirection instance. Left alone: five sit in generated abstracts, so they are entity generator work rather than an edit to make by hand, and the nine hand-written Letter* entities should move with them.

The two metadata baselines​

ORM mapping validation (no database needed)​

test/module/Api/src/Entity/OrmMappingValidationTest.php runs Doctrine's SchemaValidator::validateMapping() over every entity. It lives in the normal unit suite, so it runs on every PR.

Its baseline (orm-mapping-validation-baseline.txt) is empty: any mapping error — a dangling inversedBy, a mismatched mappedBy, a broken association target — fails the build immediately. Fix the metadata (usually via the entity generator) rather than adding to the baseline.

# after fixing errors, to confirm the baseline stays empty:
REGENERATE_ORM_MAPPING_BASELINE=1 vendor/bin/phpunit \
test/module/Api/src/Entity/OrmMappingValidationTest.php

Schema drift (integration suite)​

test/integration/src/Schema/SchemaDriftTest.php compares the entity metadata against the real, Liquibase-migrated schema — the sync-check half of orm:validate-schema. The baseline (test/integration/schema-drift-baseline.txt) records the known, triaged disagreements; the test fails only on new drift: an entity change with no matching olcs-etl migration, or vice versa.

The comparison is faithful by default. It makes three deliberate reductions in fidelity, each named and pinned by SchemaComparisonTest so none can quietly widen:

  • Comments are silenced only where the mapping declares none of its own. The database carries comments on 2081 columns and 136 tables; the entity metadata models none, because the generator turns a column comment into the property's PHP docblock rather than an ORM\Column options entry. Comparing them would report every commented column, truthfully and uselessly. The rule is conditional rather than blanket so the blind spot shrinks by itself: a column that does declare options: ['comment' => ...] is compared like any other. (ORM 2's JoinColumn had no options, which used to be the stated reason; ORM 3's does, so that reasoning is spent.)
  • yesno / yesnonull are compared by their storage type. YesNoType::getSqlDeclaration() hardcodes the whole string tinyint(1) NOT NULL COMMENT '(DC2Type:yesno)', so it can never match the column it maps however that column is declared. (encrypted_string needed the same treatment under DBAL 3, which appended DC2Type comment hints to custom types. DBAL 4 dropped that mechanism, so the normalisation was removed — measured, it changes no statement.)
  • Foreign key constraint names are compared structurally (olcs-etl uses semantic names, Doctrine generates hashed ones).

View-backed entities (Entity\View) and tables the entity side knows nothing about (audit *_hist tables, ETL working tables, Liquibase bookkeeping) are dropped from the output entirely — neither can match by construction. Note that this covers the ALTER TABLE ... DROP FOREIGN KEY statements Doctrine emits for such tables as well as the DROP TABLE itself; missing that is how four stale document_analysis lines survived in the baseline after the table gained an entity.

Reading the baseline​

The file is grouped under two headings carrying a count. Grouping is presentational — every statement is still compared, and the headings are comments the parser ignores:

  • actionable — a real disagreement between an entity and the schema. Fix the mapping, or add a Liquibase migration in olcs-etl.
  • inexpressible from mapping — differences on implicit ManyToMany join tables that no entity change can resolve. Those tables have no entity to carry an #[ORM\Index], and ORM 3's JoinTable attribute takes only name, schema, joinColumns, inverseJoinColumns and options — no way to name an index, declare a unique constraint, or map an extra column. SchemaTool builds their indexes with no name, so DBAL falls through to _generateIdentifierName() and produces IDX_<hash>, which the Liquibase DDL will never match.

A join table's own columns are the exception: their declarations derive from the referenced entities' primary keys, so a CHANGE is fixable there as anywhere else, and a compound statement carrying both is filed as actionable so the half that can be fixed stays visible.

Normalising those generated names away, as the comparison already does for foreign key names, was considered and rejected. It trades a real difference for a quieter file, and the cost is not hypothetical: DBAL's existing rename detection had been masking a redundant index on application_tracking, which surfaced only when an unrelated attribute beside it was removed.

Shrinking the actionable section is welcome; adding to it should be a conscious, reviewed act:

REGENERATE_SCHEMA_DRIFT_BASELINE=1 vendor/bin/phpunit \
-c phpunit-integration.xml --filter SchemaDriftTest

Table output escaping​

The table render pipeline does not escape. ContentHelper::replaceContent() is a raw str_replace into <td>{{content}}</td>, so whether a column is safe depends on whether something upstream escaped it. Three tests guard that, and no one of them subsumes another — the first two catch opposite mistakes, and the third reaches code the first two structurally cannot.

TestCatchesVisible to users?
TableEscapingInvariantTesta row value reaching the output unescapedNo — it is an XSS hole
TableRenderSnapshotTestsomething that was not a row value getting escapedYes — literal &lt;b&gt; on the page
FormatterEscapingInvariantTesta formatter leaking, whatever any table doesNo — it is an XSS hole

TableRenderSnapshotTest also carries the two checks that are about the other output format, both asserted absolutely rather than against the snapshot: a table whose CSV export still contains HTML entities, and one whose CSV export contains a live formula. See CSV exports.

The first two exist three times, once per table location, all sharing one harness in olcs-common (test/Common/src/Common/Service/Table/Harness/): app/internal, app/selfserve and olcs-common's own Common/src/Common/Table/Tables. The third exists once, because formatters all live in olcs-common behind a single plugin config. All run on every PR — the apps via php.yaml, olcs-common via php-lib.yaml.

Running them​

Nothing here needs a database, a container or a network. All three are ordinary PHPUnit tests and run as part of the normal suite; these are the commands for running them alone.

# all three, for one location
cd lib/olcs-common # or app/internal, or app/selfserve
vendor/bin/phpunit --filter 'Escaping|RenderSnapshot'

# one at a time
cd lib/olcs-common
vendor/bin/phpunit --filter TableEscapingInvariantTest
vendor/bin/phpunit --filter TableRenderSnapshotTest
vendor/bin/phpunit --filter FormatterEscapingInvariantTest # olcs-common only

FormatterEscapingInvariantTest lives only in olcs-common. The other two need running in all three locations, because each has its own table directory and its own baseline — a change in olcs-common can move a digest in app/selfserve, so a green lib is not evidence on its own.

Regenerating the snapshot is the one thing that is not just "run the test":

UPDATE_TABLE_SNAPSHOTS=1 vendor/bin/phpunit --filter TableRenderSnapshotTest

CSV exports​

Escaping happens at the source here — a formatter escapes the row value it interpolates — which is right for HTML and wrong for everything else. Six controllers export a table as CSV (ResponseHelperService::tableToCsv), and a CSV is neither HTML nor a passive document. Three things follow, all handled by the renderer rather than the formatter, because the formatter has no idea which format it is feeding.

Entities. Nothing downstream decodes them, so an operator called "Smith & Sons Ltd" would reach the spreadsheet as Smith &amp; Sons Ltd. TableBuilder::renderBodyColumn() decodes when the content type is CSV.

Formulas. Excel, LibreOffice and Sheets evaluate a cell beginning =, -, +, @, tab or CR, and Excel's DDE syntax reaches outside the document — a vehicle marked =cmd|' /c calc'!A1 is code that runs on the caseworker's machine when they open the export. League\Csv\EscapeFormula prefixes such a value with an apostrophe.

Framing. Fields are quoted, internal quotes doubled, and embedded newlines kept inside the field. Not merely tidiness — it is what makes the formula handling work. Unquoted, a value of x,=cmd|... arrives as two fields, and the second starts with = having never passed through any neutralisation.

None of that is hand-written, and that is the point. TableBuilder::renderCsv() assembles the file with League\Csv\Writer — which wraps fputcsv — instead of rendering a .phtml layout and joining with implode. The quoting rules are RFC 4180 and the formula-starting characters are a published list; a local implementation of either is a maintained copy of someone else's better-tested one.

Bypassing the template also fixes a third bug that no per-field encoding could reach: render() used to pass the finished CSV back through replaceContent() with the table's variables, so a cell containing {{title}} was silently blanked.

One consequence worth knowing: EscapeFormula escapes every value starting with -, including -100.50, which a spreadsheet would otherwise read as a negative number. None of the six exports is a money table, so this costs nothing today; if one ever is, that is a decision to take deliberately rather than a default to drift into.

TableRenderSnapshotTest asserts the entity and formula properties absolutely, with no baseline: there is no legitimate reason for either to appear in an export of licence data. The formula probe is =1+1,=2+2 and the output is parsed back with a real CSV reader rather than split on commas, because splitting on commas is precisely the mistake being tested for. Its idea of "a formula" comes from EscapeFormula::FORMULA_STARTING_CHARS rather than being written out again, so the guard cannot drift from the code it guards.

The escaping contract​

Escape by provenance, not by what the value looks like:

Where the value came fromWhat to do
Row dataEscape, always, whatever it looks like
Developer-authored markupLeave raw; escape the values you interpolate into it
Another formatter's return valueLeave raw — escaping it double-escapes, and escaping its values is that formatter's job

TableBuilder::replaceContentEscapingValues() exists for the middle case and is greppable. The third case is the one that catches people out: wrapping $this->callFormatter(...) in an escape call is always wrong.

When the invariant test fails​

A table you touched now emits a row value raw. Escape the value at the point it is interpolated. The baseline (table-escaping-baseline.txt beside each test) lists tables that are known exceptions and are exempt; it should only ever shrink, and a listed table that stops leaking also fails, so the list cannot rot. Tables the synthetic probe cannot drive are reported as skipped, which means "not covered here", never "safe".

Adding a table to the baseline is not the fix, and neither is regenerating anything — the baseline only shrinks.

Every table definition in the repository renders. There is no skip list: all 233 are driven against a hostile row, and a table that stops rendering fails the build rather than quietly dropping out of the count. FormatterEscapingInvariantTest still earns its place — lva-psv-vehicles-readonly renders with no data rows at all, so its StackValue columns are exercised only there.

When a type rejects the probe​

RecursiveProbe answers any key to any depth with itself, which is what lets one value drive every table without knowing their shapes. What it cannot do is satisfy a type: number_format() wants a float, new DateTime() wants something parseable. Rather than record those tables as undrivable, both harnesses learn the constraint from the failure and retry — see RowProbeAdaptation, which they share.

Two things make this honest rather than convenient:

  • The substitution is learned, never declared. Nothing is replaced until the code actually rejects the probe, so the adaptation disappears by itself when the constraint does. A hand-maintained "these keys are numeric" map would rot the way a skip list does.
  • Substitutions that lose the payload are recorded, and then proved. A string or an array of probes still carries the marker, so the assertion is unweakened; a number or a date cannot, so each of those goes through the isolation pass below, which ends with it either escaped or rejected by its own type. Both are proofs, so neither is recorded — see Putting the payload back.

Table-level adaptation substitutes at a dotted path, not a root key: publication.pubDate can be a real date while publication.pubStatus.description stays a probe. Replacing the whole publication key instead would de-probe every sibling column to fix one, which is the difference between "this value cannot carry a payload" and "this table is no longer tested".

Putting the payload back​

A substituted value is not tested, and "not tested" is where a leak hides. Worse, the substitution lands in the row every column shares: if one column parses createdOn as a date and another interpolates it into markup, the date that satisfies the first is what the second emits — so the substitution masks the leak it should have found. The more constrained a value is, the better it is hidden.

So every payload-losing value is put back. One at a time, the marker is restored at exactly that path while the rest of the row stays as the settled render left it, and the render runs once more with no adaptation — recovering from the constraint is what the pass exists to avoid.

The outcome is binary: the value is safe, or it is not. There is no third category and no list of values that are "safe for a different reason" — that distinction is real but it is not actionable, and a reader who has just written a table cannot do anything with it.

Safe arrives two ways, and neither is weaker, so neither is recorded anywhere:

SafeMeaning
escapedThe render survived and the payload came out escaped.
constrainedThe value's own type rejected the payload before anything was written to the page.

Calling the second one safe is a complete argument, not a lenient one, and it is worth spelling out because it looks like a let-off. The type partitions every possible value: one that satisfies it is a number or a strict date, neither of which can contain <, >, & or a quote — and the output is derived from the parsed value rather than copied from the input. One that does not is rejected before a byte is written; in production that input is a 500, not a payload on a page. There is no third case, so there is nothing left to test.

Not safe also arrives two ways, and both fail the build:

Not safeMeaning
leakingThe payload reached the output raw — a real leak nothing else could see, because the ordinary run substitutes this value first.
unprovenThe render failed for a reason that is not its type rejecting it, so nothing was established either way.

unproven is the one that keeps the argument above honest. "It threw, so it cannot get out" only holds when the throw is the type rejecting the payload — an exception from anywhere else (a missing service, a formatter broken by an unrelated change) would otherwise read as proof of safety, and go on reading that way for as long as the breakage lasted. So the exception is checked against the same question the adaptation loop asks to learn a constraint, and anything it does not recognise is reported by name. There is no baseline section for it: it is empty today, and an entry means someone has to look.

Reading the output buffer even when the render throws matters for the same reason. A value echoed by one column and only later rejected by another is already on the page; the throw does not retract it.

One path at a time, never in a batch. AbstractConversationMessage::getFirstReadBy() returns early when createdBy.id equals the read's user.id, so restoring the payload to both at once would skip the branch and report "no leak" about code that never ran. Isolated, both render and come out escaped.

Attributing a failure to the right value is most of the work, and it is done in four ways, narrowest first: the failing line, the statement below it (a multi-line call puts its arguments there), a local alias of the row resolved back to its path ($licence = $data['licence'] … $licence['goodsOrPsv']['id']), and the column config the formatter was handed, read out of the exception trace. Only if all four come back empty does it widen — and only for strings, because MARKER is a string, so a widened substitution keeps its payload. Numeric and date never widen; they would silently de-probe the row.

Two consequences worth knowing:

  • The harness reads exception trace arguments, so zend.exception_ignore_args is pinned to 0 in all three phpunit.xml.dist files. With args stripped, the column config is invisible and those tables stop rendering.
  • The row is harvested from the definition and from the formatter classes it names. A key only a formatter reads would otherwise arrive as null, which makes isset() false and skips the very branch that leaks.

Keys the harvest cannot see​

That last point has a sharper edge than it reads. The harvest looks for $row['literal'], so a key read through a variable is invisible to it — and the miss is silent in the worst possible way. A key the row does not carry is absent, isset() on it is false, the branch that would have interpolated it never runs, and the table renders clean. Green, and nothing tested.

Formatter\Address was the worked example. Its field names live in a class property, not in a subscript:

protected $formats = ['FULL' => ['addressLine1', 'addressLine2', 'town', ...]];

foreach ($fields as $item) {
if (!isset($data[$item])) { continue; } // always true, so...
$parts[] = $data[$item];
}
return implode(', ', array_map(fn($p) => Escape::html($p), $parts)); // ...never ran

Nothing reads $data['addressLine1'], so the harvest never added it, so every iteration hit the continue and $parts came back empty. The formatter counted as exercised while the only line that matters had never executed.

DynamicRowKeys closes it: where a file reads the row through a variable, the quoted identifier lists in that same file are taken as candidate keys. It is a heuristic and deliberately a narrow one — only files that actually do dynamic row access contribute (19 of 149 formatters, and only a handful have a list to find), and only identifier-shaped strings in comma-separated array literals count, so a lone 'FULL' or a CSS class does not become a row key.

Guessing is justified because the costs are asymmetric. An extra key the code never reads is an unused probe and changes nothing. A missing key is a branch that silently stops being tested. And the failure direction is loud: a value the row should not have had reaching the output is reported as a leak for a human to judge, rather than passing quietly.

$column is excluded from "dynamic", because $data[$column['name']] is resolved properly elsewhere from the column config the formatter was actually handed.

One mechanism is worth knowing about, because it is not about probe data at all: shadowing. Definitions resolve by bare name against a list of directories, so two files sharing a basename cannot both be reached through one builder — the first hit wins. Two pairs exist today in app/internal: conditions.table.php (in Bus/ and in SubmissionSections/) and environmental-complaints.table.php (in the table root and in SubmissionSections/). The loser used to vanish from the report entirely, which read as full coverage; it now gets a builder of its own, with its directory ordered to win. Keys carry their directory prefix (SubmissionSections/conditions.table.php) so the two stay distinguishable.

When the snapshot test fails​

Rendered output moved. If you meant it — a column added, a label reworded — regenerate; if you did not, you have probably escaped developer markup, another formatter's output, or a value that was already escaped where it was assigned.

# from the app or lib directory that failed
UPDATE_TABLE_SNAPSHOTS=1 vendor/bin/phpunit path/to/TableRenderSnapshotTest.php

Regenerate deliberately, and read the diff. Doing it reflexively to make a red build green defeats the entire point of the test.

The snapshot stores a digest per table rather than the rendered HTML, because 185 tables of markup would be unreviewable. It renders against benign data containing an ampersand — a value of plain alphanumerics would be unchanged by escaping, so escaping it twice would be invisible. The per-render CSRF token is normalised out; anything else non-deterministic in a render will make the digest unstable and needs normalising in TableEscapingHarness::normalise().

The third test: formatters, directly​

FormatterEscapingInvariantTest (olcs-common only — formatters all live there and share one plugin config) calls every formatter directly with a hostile row.

It exists because driving formatters through a table leaves two blind spots:

  • a formatter reachable only from a table the probe cannot render never runs at all
  • TableEscapingHarness builds its row by harvesting ['literal'] subscripts from the table definition, so a key read only inside a formatter class is absent from the row and arrives as null — the formatter runs, but its leaking branch does not

Formatter\PrinterDocumentCategory is the worked example: it interpolates $row['subCategory']['category']['description'] into an <a>, but admin-printers-exceptions.table.php never names subCategory, so isset() is false, the 'Default setting' branch is taken, and the table-level test passes a formatter that leaks in production.

This harness harvests keys from the formatter's own source (and its parents), so that cannot happen here by construction. Neither test subsumes the other — the table-level one still owns inline closures, cell attributes and wrapping.

All 149 formatters are driven, none leaks, and none is skipped. Keeping it that way is the point of the baseline below.

How the probe works​

Worth understanding before you change a formatter's signature, because that is what usually makes one of these tests go red for a reason that is not escaping.

RecursiveProbe is a row stand-in that answers any key, to any depth, with itself, and stringifies to <script>xss-probe</script>. That is what lets one object drive a hundred formatters without anyone writing a fixture per table: whatever a formatter reaches for, it gets the marker, and the harness then asks whether the marker reached the output unescaped.

A probe cannot satisfy a type, though. A formatter that calls number_format() or parses a date rejects it outright. So the harness does not declare types up front — it learns them from the failure:

  1. call the formatter; if it returns, check the output for the marker
  2. if it throws, read the engine's own wording to decide what the value has to be (must be of type int|float → numeric, Failed to parse time string → date…)
  3. find which row key the failing expression touched, from the failing line first and the whole statement second
  4. substitute a value of that type into that key, and go round again

Learning rather than declaring matters: a hand-maintained "these keys are numeric" map would rot exactly the way a skip list does, whereas an adaptation disappears by itself the moment the constraint does.

Most substitutions keep the payload — a string is the marker, and an array holds probes. Numbers and dates cannot, so each of those is put back on its own and proved rather than quietly counted as covered (see Putting the payload back).

Two formatters need more than this. RecursiveProbe cannot satisfy a type at depth — getSenderName(): string returns $row['createdBy'][…]['forename'], and substituting at the root only moves the failure to the next subscript. The two conversation message formatters therefore get a hand-written row from FormatterEscapingHarness::fixtures(). A fixture is merged over the generated row, so a key the formatter starts reading tomorrow still arrives as a probe rather than missing.

The two baseline sections​

formatter-escaping-baseline.txt. Both are asserted, so neither can rot:

SectionFails when
[leaking]a formatter starts leaking, or a listed one stops (stale entry)
[skipped]a formatter becomes undrivable, or becomes drivable

Asserting [skipped] matters more than it looks. A skip is not "safe" — it is "unknown", and without this it would stay unknown long after the reason for it disappeared. If a formatter is skipped because number_format() rejects the probe, and someone later drops that call, the value starts flowing raw; comparing the set means that shows up as a failure rather than sitting unnoticed behind a stale entry.

There is deliberately no section for values a type constraint keeps off the page, and there used to be. Listing them said "safe, but for a different reason", which is true and useless: nobody reading it for the first time can act on it. Those values are safe outright — see Putting the payload back for why the argument is complete rather than lenient — so they are proved on every run and not recorded. A value that can be proved neither way fails the build by name instead.

When a formatter's constructor or signature changes​

The most common way to make this test red without touching escaping:

Symptom in the failureUsually means
unresolvable: … in the skipped listthe factory is broken, or a new dependency is missing from HarnessContainer
a formatter moved into [skipped]a new type constraint the probe cannot satisfy — add the service, or a fixture if it is nested
a formatter moved out of [skipped]good news; delete the entry
an unproven value in the failurethe formatter throws before reaching the value, for a reason unrelated to the value's type

If you add a formatter dependency, add it to HarnessContainer as the real class wherever the real thing does not need the world. A mock that answers everything with '' will let the formatter run while feeding it nothing, which looks like coverage and is not: StackHelperService was mocked that way, and StackValue, NumberStackValue, UnlicensedVehicleWeight and FeeTransactionDate all counted as exercised while formatting an empty string. Three live leaks were sitting behind it.

Adding integration tests​

Extend Dvsa\OlcsTest\Integration\IntegrationTestCase and use $this->repo() / $this->em(). Reach for this layer when the answer depends on running the query: streaming via toIterable(), hydration, stored procedures, and query paths against database views. Whether the query is well-formed — its joins, field names, predicates and ordering — is settled faster and without Docker by a repository unit test. Prefer selecting fixture rows from the seeded test data over hardcoding ids; insert your own rows where the dataset is not enough — the per-test transaction rolls them back.