The preset sweeps fan out in batches
The sweeps that make a claim about a single preset — animation, reactivity
and sanity’s loudness gate — are generated, by the same core/build.rs
glob that embeds the library, as one #[test] per batch of eight presets. A
new preset joins a batch by existing; no Rust is edited. sanity’s shape gate
and distinctness generate one test per family instead, because their claims
are about a family’s distribution and its pairwise set and do not decompose
further.
Why a batch rather than a preset. nextest runs every testcase in its own
process, so a per-preset testcase pays for that process and for the adapter,
device and pipeline set it builds inside it before it renders anything. Measured
on the reference machine through WARP, that fixed part is 38–55 % of a
testcase’s wall; a batch pays it once for the eight presets it holds, and eight
is where the return on a larger batch stops paying for the granularity it costs
(ADR-0222).
What makes that sound is that the capture primitives are pure functions of their
arguments — capture_preset reseeds every scene and resets the clock — and
core/tests/suite/batch_independence.rs holds them to it.
Three things follow that matter when you are reading a red run:
- A failure names every preset it convicted, in its message rather than in
the test’s name, and it reports the whole batch rather than stopping at the
first:
1 of the 8 presets in this batch are frozen on both readings:and then the rows.-E 'test(animation_batch_03)'re-runs the batch that held it. - The per-preset reports print per preset, one line each, attributed to the
batch’s test.
sanity’s loudness ratio does not arrive as one sorted table — sort a run’s lines to rebuild it — and the shape sweep’s flattest-preset ranking is scoped to the family whose test printed it. -P fastrenders a sample and the full run renders everything. A preset may declarerepresentative = true(see Parameter roster);build.rsbatches the representatives apart from the rest and names their batches<sweep>_rep_batch_<nn>, and thefastprofile selects on that marker. So thedevlane’s per-phase gate renders the declared sample, while a barecargo nextest run --workspace— what the plan close and CI’s coverage job run — renders the whole library plus the per-family shape and distinctness tests, which are never sampled. ADR-0081’s curation gate therefore still sees the whole library.- A test retries only by exact name, and only while a live backlog entry
diagnoses it as a flake.
.config/nextest.tomlgives such a test one more try undertest(=name), never a pattern and never profile-wide, and the test leaves the list when its entry closes. A pass on the retry still shows: nextest prints itFLAKY, and the conductor’s ledger and digest name it. A red test with no live entry is never retried. Today the list isa_preset_datagram_selects_by_name, for backlog 0219 (ADR-0261).
Individual tests (add -- --nocapture to see the printed diagnostics):
| test | kind | asserts |
|---|---|---|
reactivity | HARD | every preset moves for at least one band (bass/mid/treb/onset), driven by PCM through the real analyzer (Plan 0067 Phase 1 — see what the gates can and cannot see below); prints the per-band vector so a dead single binding — e.g. treble — is visible |
animation | HARD | every preset changes between frame N and N+k on at least one of two readings (not frozen). Since Plan 0123 Phase 1 (ADR-0136) the verdict is a disjunction over this gate’s own statistic: the silent reading below, or a driven one — the same footprint_diff between the silent capture at frame N+k and one taken against AnalysisFrame::fully_driven() at the same frame count, floored at 0.017 against the silent 0.01. A preset frozen in both still fails, and the static control is pinned failing both. Each preset’s own line, printed inside its batch’s output, names which branch carried its pass, because that weakening is real and an unprinted property is one nobody re-reads. Since the sweep fanned out across testcases (ADR-0157, ADR-0222) there is no end of a run to collect the still images in silence into one roster at; shot --presets presets --report’s anim column is where that set is read as a list now, and ADR-0136 carries the Outcome. Still no PCM: the driven capture is a synthesized frame, so the row in the stimulus table below is unchanged in kind. Scores metrics::footprint_diff since Plan 0077 Phase 1 (ADR-0091) — motion over the union of lit pixels, with bg_* stripped, so a sparse figure is measured against its own footprint rather than diluted by the frame. The rejected fifth-density emitter_squall draft passes at 0.1049 where the whole-frame statistic priced it out at 0.0057; a static control still fails on a zero numerator, and both are pinned as a standing non-vacuity test. Two things it still cannot do, both by construction rather than by tuning: a passing anim is not evidence of a watchable preset on a still family — an IFS figure is a photograph, so a slow view pan clears the floor while nothing in the figure moves (design-backlog 0066); and a rotationally symmetric figure cannot score its own spin — a figure invariant under rotation by 2*pi/k produces an identical image under that rotation, so its frame difference is zero at every resolution and no image-domain statistic lifts it (design-backlog 0009; the #[ignore]d resolution ladder from Plan 0067 Phase 1d is the recorded negative result, and it is why SIZE never moved). Such a figure must move radially, and that is an authoring constraint, not a gate defect |
sanity | HARD | every preset lights a minimum coverage and spans ≥2 quadrants (not blank, not a dot), and is not a blot. The blot check is a conjunction of two terms, and needs both to convict (ADR-0128, settled by ADR-0130, Plan 0119). Term one is tonal_flatness — the share of lit pixels inside one narrow luminance band, capped at MAX_TONAL_FLATNESS — asking does the figure have any tonal structure. Term two is metrics::assigned_boundary_density — the share of lit pixels touching an unlit 4-neighbour, perimeter over lit area, floored per system kind by boundary_floor — asking does the lit set have any interior. Which reference it reads is decided per frame by a role classifier (ADR-0200): metrics::figure_ground_ratio divides the frame’s coverage against black by its coverage against its own derived ground, and above the 1.17 cut the modal band is the frame’s figure rather than its ground, so the term reads black. Below it — every scene that draws light onto darkness, whose two references coincide — the term reads the derived ground, exactly as it did before. Read the cut for what it divides, which is narrower than figure from ground: a frame lands above it when almost none of it is near-black, so a painted canvas or a light-paper print sits on the figure side too, and its term two is read areally over a lit mask that is the whole frame. On such a frame the statistic collapses to the frame’s own border — 0.0412 at this suite’s 96x96 capture, for eight shipped presets — and says nothing about the picture. Term one is the whole of what holds them, which is the landmine the fifth caveat below prices. That choice is the whole repair: a saturated blot is its own modal band, so a term pointed unconditionally at the derived ground was handed the mass’s fringe and scored a blot as more structured the smoother its rim got, which is why the check convicted nothing at all for a plan (ADR-0161). A preset fails only when it is flat and below its family’s floor; either alone convicts legitimate content, which is why neither is a verdict. Term one alone convicted a two-ink print for being two-ink (fragment_tiledmono, held out of the set for two plans over it); term two alone would convict 62 of the 112 shipped presets, which the gate prints as a count on every run. The tonal term is Plan 0056 Phase 5: a saturated single-tone mass satisfies coverage and spread completely, which is how a run of attractor presets shipped flat. What “lit” means moved twice, and neither reading is the preset’s own backdrop. ADR-0067 strips every bg_* binding for this capture, because a bg_vignette made the frame’s corner its darkest pixel and the backdrop read as a large, well-spread figure. ADR-0126 then made the reference the frame’s own ground — the mean tone of its most populous luminance band — rather than a hardcoded black, because a scene that paints its own paper reads full coverage whatever it drew, which made three of the four statistics constants for that content and left an emptying canvas indistinguishable from a broken one. So each statistic now answers how far does this picture depart from the ground it is drawn on: coverage how much of it does, quadrant_spread and radial_shell_occupancy where, tonal_flatness whether what departs has more than one tone, and boundary_density whether it has any interior — with the one asterisk the classifier above adds, that a frame holding almost no near-black pixel is read areally instead — which is right for a blot, whose modal band really is its own figure, and is the degenerate case for a painted canvas. Most floors are measured from the library’s own distribution and printed on every run; boundary_floor’s default arm is not, and says so. It is 0.23, the midpoint of two frozen frames read under the role each is assigned — the ragged particle blot at 0.0934 against black and fragment_tiledmono at 0.3602 against its own ground — because a conjunction’s second term is judged only over the frames that failed the first, and that population is small and contains the preset the term exists to admit. Half-the-sparsest would be circular there. Both that floor and the 1.17 cut are measured at this suite’s 96x96 capture and neither travels: perimeter over area goes as ~1/L, so the same anchors at 192x192 give 0.12, and a gate reading at another size owes its own derivation (ADR-0071). The shape_collage arm (0.13) is the ordinary ceremony, on a family ADR-0123 holds under the tonemap knee so it has no additive path to the defect at all. KNOWN_FLAT is the exemption roster and it is empty — no shipped preset is excused from the blot check |
beat | HARD | a 120 BPM click track through the real DSP makes a beat-accent preset render differently on-beat vs off-beat; a zeroed beat binding does not |
distinctness | ADVISORY | prints per-family pixel + shape pairwise matrices and flags near-duplicate geometry; never asserts a similarity verdict. Covers every SystemKind, and stays covering it by construction (ADR-0234): the per-family tests in core/tests/distinctness.rs are declared through a macro that also emits an exhaustive match over the same variants, so a new system with no test there is a compile error rather than a family the report silently skips, and each test’s label is the system’s own canonical name. The report’s unit is a pairwise matrix, so each family asserts it ships two or more presets — a family with one would compare nothing and report clean. Split per family since ADR-0157 — one test each, never sampled, each asserting its own pair count against the shipped set |
frame cost (shot --report) | ADVISORY | not a test in the suite — the one row here that runs under the CLI rather than under nextest, listed so the roster says what is and is not enforced. shot --report prints a per-preset frame cost in milliseconds per frame at 1920x1080 on the report’s tier, names the adapter and the build profile it was taken on (ADR-0071), marks a reading past the 60 fps budget NFR §1 states, and never fails a run (ADR-0232). Headless, so it compares presets and predicts nothing about the app; not taken on a software adapter at all. The reading, its method and its caveats are documented with the report’s other columns in Headless capture and video |
golden | HARD (tolerance) | one frozen fixture per system, plus the EXTRA_FIXTURES escape hatch below, matches its committed baseline PNG within a mean + max-outlier tolerance |
composite | HARD (tolerance) | the post stages, one fixture each and never all at once — trails, kaleido_*, bloom_*, plus one that binds no stage and guards the composite’s arithmetic (its assertion is that no channel of that fixture reaches 255 — a claim about the fixture, not a general property of the curve; see the re-bless note below). Captured at 160x100, a size whose internal grid is not the target’s shape, so an aspect error is visible |
bloom | HARD (relative) | the bloom stage’s behaviour, beside its baseline rather than in it: halo energy rises with bloom_amount, halo extent rises with bloom_radius, the rich tier’s deeper pyramid reaches further than the floor’s, and the halo is round. Captured at 256x256 — square, and load-bearing: the roundness guard is what catches a separable kernel whose two passes step in different units, and it reads 1.001 today against 7.05 under the defect it was written for. No magic numbers: every assertion compares two captures of one fixture differing in one bound param |
reaction_diffusion | HARD | the first stateful-feedback scene: seed reproducibility, regime response |
attractor | HARD | the first compute-particle scene: seed reproducibility + beat perturbation |
line_joints | HARD (+ tolerance) | a flagged joint stops leaving a hole in the stroke (ADR-0041): against a purpose-built zigzag polyline, a vertex is not a local luminance minimum relative to the segment interiors either side of it. Threshold-free, and captured at 512x512 because the wedge it measures is a fraction of a stroke-width across. The same capture is then pinned to a committed baseline (Plan 0040), since the reported defect had no pixel guard anywhere; the relative claim runs first, even under RLX_BLESS, so the notch cannot be blessed back in. Bless with --test suite line_joints::, which cannot reach the golden roster |
attractor_trails | HARD (tolerance) | the attractor with the engine trails stage bound — the attractor’s four pipelines and the stage’s two in one command buffer, which is the densest pipeline coexistence any shipped preset produces and the thing ADR-0058’s hazard keys on. attractor.toml binds no trails and every composite_* fixture is a line scene, so nothing pinned it before Plan 0053. Captured at 160x100 (a non-square, per ADR-0037) and, like every baseline here, blessed on WARP — so it is coverage, not evidence of correctness; ADR-0058’s hardware-vs-WARP comparison is the check and this is the drift guard. Its own module, so RLX_BLESS=1 … --test suite attractor_trails:: can reach nothing else. A second, GPU-free test asserts the fixture still puts both accumulations live (trails > fade, spin non-zero), since at or below the scene’s own tail the stage is a bit-for-bit passthrough |
warp_mesh_wide | HARD (tolerance) | the converted warp_mesh chain — the bytecode VM’s mesh and a translated shader module, milk_wash_fog_tunnel.toml — captured at 160x120, a 4:3 target, against a baseline of its own. This is what the three square converted warp_mesh fixtures in golden — warp_mesh_milk, warp_mesh_shader and warp_mesh_stroke, the three carrying a [milk] table, the rostered warp_mesh.toml carrying none — structurally cannot see: at 128x128 the aspect correction is the identity in both halves — mesh::vertex_position’s x multiply and the U.aspect lanes a translated shader reads (ADR-0037) — so dropping either term moves none of those three by a byte, and 16:9 is the other shape the ADR records the confusion as invisible at. core/tests/suite/warp_mesh.rs asserts what the chain computes; this pins the picture it draws where the correction is live, which is the incidental-drift class a baseline exists for. Its own module, so RLX_BLESS=1 … --test suite warp_mesh_wide:: reaches no other baseline. What it guards is stated as a probe, not as a size: forcing self.aspect = 1.0 at the converted chain’s entry must fail this fixture, and does — on the outlier term, 75 against a tolerance of 48, while the mean stays inside its own at 0.0139 against 0.02. A warp redistributes edges rather than shifting the frame’s average, so the outlier is the term that convicts and the mean alone would have said nothing. Measured on the development machine, and the three square converted fixtures pass that same probe unchanged. The subject is the one of four candidates the probe convicts; three declare zoom, rot or warp and still read 0.0000 and 0, so declaring motion is not the same as rendering a picture the corrected space reaches. A second, GPU-free test holds the capture size off 1:1 and off 16:9, since a baseline at either would pass forever and guard nothing. Its first baseline was a person’s — it was captured, opened and judged before the test was released (Plan 0201 Phase 4b), because a first baseline is compared against nothing and no run can say whether the picture is the one the fixture meant to draw. From here it is an ordinary baseline: an absent warp_mesh_wide.png fails, as it does for every other fixture |
ink | HARD | the final tone-remap inverts tone, and ink_amount = 0 is byte-identical to an unbound frame |
geometry_extent | HARD | the in-frame geometry fraction, for the four line families only (ADR-0083): that the diagnostic is byte-identical to having it off, and that each of the two frozen over-scaled configurations measures below the shipped preset it was recovered from. Neither engine-wide nor a threshold — read the section below before using its numbers |
lit-backdrop guards (in-crate, --lib) | HARD (exact) | one per draw seam, three of them: swarm.rs’s a_lit_backdrop_survives_where_the_swarm_drew_nothing, lines/renderer.rs’s a_lit_backdrop_survives_where_the_strokes_drew_nothing, and emitter.rs’s a_lit_backdrop_survives_where_the_emitter_drew_nothing (ADR-0056). Each captures swarm_lit_backdrop.toml / lines_lit_backdrop.toml / emitter_lit_backdrop.toml three ways — lit backdrop, black backdrop, backdrop with the scene contributing nothing — and asserts that wherever the scene wrote no light the backdrop arrives intact. Bound 0 rather than a tolerance, because it reads the linear composite; see the section below. The swarm’s and the lines’ take a fourth capture at zero emitted light (Plan 0053 Phase 4), which turns the frame into a direct readout of alpha and widens the line guard’s reach from 15 channels to the whole stroke footprint |
emitter burst (in-crate, --lib) | HARD (relative) | the emitter is the first scene whose population varies, so emitter.rs’s a_spawn_rate_on_onset_bursts_and_then_idles drives emitter_onset.toml through capture_preset_over with a silent lead, a six-frame transient and a second of silence, and asserts the frame is dark before, lit after, and dark again by the end. capture_preset cannot ask this: it holds one analysis frame for every step, so it can show that a binding is live but never that the shower empties when the transient passes |
background_composite | HARD (hardware only) | RD / attractor presents alpha-blend over the bg_* backdrop; skipped on a software adapter, which mis-renders that pipeline set. The stated cause of that mis-render was identified and fixed by Plan 0053 Phase 3 — it was background-bind-layout colliding with rd-init-layout / fragment-field-uniform-layout (ADR-0058), and an explicit min_binding_size moved WARP onto the hardware numbers for the RD half (087.612 165.165 156.168 hardware, against a bare layout’s 087.543 064.538 …). Whether the skip can now be lifted is open and unmeasured: the attractor half of this test is a different layout group and nothing probed it. The module docs here and in background_composite.rs still assert the quirk as live — do not read that as evidence it is, and do not lift the gate without rendering both halves on both adapters |
transition | HARD | every switch path (cycle and select) renders intermediate blended frames as a ramp, reproducibly from the injected dt; each blend kind shows its own signature; a switch arriving mid-dissolve lands on the last index requested; a hot-reload mid-dissolve cancels cleanly; the heavy attractor ↔ reaction-diffusion pair dissolves on the freeze fallback (set RLX_TRANSITION_STRIP=<dir> to also dump filmstrips) |
easing | HARD | [smoothing] is observable: a scalar entry measures symmetric and an { attack, release } pair does not, against purpose-built near-linear fixtures (ADR-0039). Also measures the spectrum curve↔easing ordering both ways round through one renderer — every frame count in the suite is gated on segment_settled first — the shared probe’s window is 180 frames (3 s, 6 τ) because at 96 its own asymmetric arm was truncated, reading 61 where the settled answer is 69 |
preset | HARD | the expression evaluator and TOML schema: exact values, rejection without panic, zero allocation per eval, and the PARAMS ↔ set_param drift guard |
dsp / ffi / hygiene | HARD | known-signal analysis fixtures; the C ABI across the boundary; the hot-path panic pragma + exact dependency pinning |
Golden baselines pin frozen fixtures, not shipped presets. core/tests/fixtures/*.toml
is a deliberately minimal preset per SystemKind, committed alongside
core/tests/golden/*.png; the shipped presets in presets/ are guarded
behaviorally (sanity / reactivity / animation) so the preset-author lane can
tune them freely without re-blessing pixels. A new SystemKind variant fails
golden.rs to compile until its fixture exists (exhaustive match, no wildcard arm).
One narrow exception: EXTRA_FIXTURES (Plan 0063). A second fixture for a system
already in the roster, for when the rostered one structurally cannot reach the code under
test. attractor_depth.toml opened it, and the list has grown since: the rostered
attractor.toml is De Jong, and ADR-0076
gives every 2-D family an inverse depth extent of exactly 0.0 — which is precisely what
makes the perspective divide, the distance haze and the depth tint the identity there, so
no edit to that fixture could execute a line of them. The newest, warp_mesh_shader.toml
(Plan 0110), earns its place the same way one level up: it is the only fixture anywhere in
the crate that carries WGSL, and render/scenes/warp_mesh/shader.rs is built only for a
bundle that declares a shader — so no edit to the rostered warp_mesh.toml, or to the
bytecode-driven warp_mesh_milk.toml beside it, could execute a line of that file either.
The list is captured after the
roster loop, never interleaved with it, so every pre-existing baseline renders from the
device state it always did (which matters on WARP, where building GPU resources mid-run
changes what a later capture resolves to). systems_rosters_every_variant holds it to the
roster’s own two conditions plus one of its own: a stem colliding with a rostered system’s
would have the two silently overwrite each other’s baseline. The roster stays exhaustive —
ADR-0023 rests on that and this does not weaken it.
core::signal (pure, zero-dep) synthesizes the test audio; core::render::metrics
(pure) provides frame_diff, struct_diff, coverage, and quadrant_spread,
shared by the tests and the CLI report, plus the step-response pair
frames_to_settle / step_response and the segment_settled gate that says
whether either of those two is worth reading.
Built from 13c7582 at version 0.158.0. This site tracks main and is not versioned per release.