12. Runtime memory
Added 2026-07-21 and retargeted 2026-07-22 per ADR-0010.
“Lightweight” (NFR §4) caps binary size but not working set. The original §12 target — “well under
~100 MB”, to be hit primarily by compiling wgpu with only the per-OS backend — was measured and
disproved by Plan 0011 (Phase 6 landed the backend-trim; Phase 7 measured it). On the reference AMD
iGPU box, release build, the standalone sits at ~300 MB working set / 343 MB private commit — the
trim took effect (verified DX12-only, no Vulkan/GL mapped) but footprint is dominated by the DX12
driver stack’s private heap (amdxc64.dll + d3dcompiler_47 + D3D12Core …), not by wgpu’s
compiled backend code (mapped DLL code is only ~135 MB, and shared). The <100 MB absolute is not
reachable on a DX12/wgpu app; the backend-trim is retired as a memory lever (it stays as a binary-size
win under §4). See ADR-0010 for the decision and rejected alternatives.
Retargeted requirements — chosen to be enforceable by the Plan 0011 diagnostics harness
(diagnostics.log, rlx_get_metrics):
-
No session growth (the requirement that matters). Working set / private commit stays flat over a session — no monotonic growth across the §10 ≥4-hour soak. A leak is the real live-show failure; the harness is the instrument. This is the hard requirement.
-
State the cost of what we add. The GPU driver stack is a fixed, vendor-dependent floor we do not own; the actionable lever is our additions — render-pipeline / shader / resource count. A new built-in system states its working-set delta on the reference box (harness-measured), so growth is a recorded choice, not a surprise. (Footprint rose from ~200 MB to ~300 MB across Plans 0003/0010/0011, most plausibly from added pipelines — exactly this cost, previously untracked.) Now quantified (Plan 0012, reference AMD iGPU box, private commit): the fixed driver floor is ~327 MB and our entire visual system (2 scene pipelines + overlay + DSP + audio + presets) adds only ~11 MB (~3%); culling 3 dead scene pipelines saved ~2 MB, so pipeline count is a real but weak lever (~1 MB/pipeline) against a floor that dominates.
-
Soft ceiling, for regressions only: ~350 MB working set on the reference AMD iGPU box with the current built-in system set. A single-machine, vendor-dependent tripwire to catch a regression — not a portable absolute (a different GPU/driver has a different floor). Vendor spread (Intel iGPU) is a pending on-device capture —
docs/on-device-validation.md. -
Our own DSP and audio state stays <~1 MB (ring buffer ~340 ms of f32, fixed DSP buffers, a few uniform buffers) — unchanged; the target was never our allocations.
This line said “our own Rust state” and that was false by roughly an order of magnitude. A line scene preallocates its geometry buffers to the tier’s
max_segmentsso the per-frame rebuild and mirror replication never allocate, and those buffers are not DSP state. Measured element sizes areSegmentInstance44 B,ArcInstance36 B,Piece24 B, a walk point 8 B and a walk offset 4 B; at Rich’smax_segments = 60_000the parametric-curve scene alone holds 5,760,008 B acrosssegments,single_bufandpoints. It held 11,760,008 B until the four fit buffers stopped being preallocated (Plan 0149 Phase 4, closing design-backlog 0135), becausearcs,single_arcs,piecesandwalkare written only by a curve fit that no shipped preset takes.The bound this bullet can defend is the one it was measured against, so it now names DSP and audio. Scene geometry is a separate quantity with a separate ceiling, and it is a fixed allocation in the same sense the emitter pool below is: sized once from the tier, never grown on the hot path. The emitter’s object pool (Plan 0052 / ADR-0057) is the newest entry here and it stays well inside that line. It is a fixed allocation, made once at scene construction and never grown — spawning past it drops the spawn rather than reallocating, which is the whole reason it is a pool — so this is a ceiling and not an average:
floor (2 000 objects) rich (6 000) CPU pool ( Object, 44 B incl. padding)88 KB 264 KB free list ( u32)8 KB 24 KB CPU instance mirror (28 B) 56 KB 168 KB GPU instance buffer (28 B) 56 KB 168 KB total ~208 KB ~624 KB Two orders of magnitude under the ~66 MB a single post chain costs, and the reason the tier’s
emitter_objectswas sized for headroom rather than trimmed: the pool is bounded by cost of drawing the marks, not by the memory holding them. It adds one render pipeline, i.e. ~1 MB by the Plan 0012 measurement above, which is the number that actually moves. -
The linear-light composite is the largest single addition since this section was written (Plan 0045 / ADR-0046). Every intermediate upstream of the tonemap moved from the surface format to
Rgba16Float— 8 bytes a texel, not 4 — so the offscreens that were charged at the surface format doubled. The trails accumulation (PingPongField, two textures) was already float and did not move. At the floor post cap (1920x1080, 16.6 MB a full-size float texture):buffer before after trails composited 8.3 16.6 trails accumulation (x2) 33.2 33.2 kaleidoscope source 8.3 16.6 per chain, both stages live 50 66 bloom source + pyramid (only when bloom_amount > 0)— 16.6 + ~11 tonemap input (surface-sized, genuinely new) — 16.6 transition snapshot + live, while a dissolve runs 8.3 x2 16.6 x2 ink input (stays 8-bit — the tonemap hands it display-referred pixels) 8.3 8.3 Plan 0023’s dual-live dissolve holds two whole chains, so the peak is ~133 MB rather than ~100, and the worst case — dual-live, every stage on including bloom, ink on — is ~246 MB against the ~350 MB soft ceiling above, most of which is driver floor already. At the rich cap (2560x1440) the same arithmetic is ~118 MB per chain. The post cap is the relief lever if the float chain misses §1 on a floor-tier iGPU: lower it rather than re-fixing the grids, since bandwidth roughly doubled with the format and the grid policy is shared. Rich-tier frame time on the target GPU is measured and bloom is not the expensive part: windowed on the dev box’s discrete GPU,
star_lantern(the one shipped preset that bindsbloom_*) runs 164 fps at p99 8.2 ms, againstattractor_cliffordat p99 19.9 ms andattractor_leviathanat 19.0 ms — neither of which switches the stage on. The float composite plus the attractor is what puts the heaviest preset past a 60 Hz frame at Rich. Both attractor readings predate ADR-0140 and describe a configuration that no longer ships: they were taken whenRichdrew a flat 150,000 samples at any window size, and a window above the 640x360 anchor now draws up to 600,000. Plan 0128 Phase 1 measured the marginal cost of that step on the same discrete GPU as +1.454 ms p99, on an instrument that omits the present and is therefore not comparable to these figures in absolute terms. Re-taking the pair windowed is owed todocs/on-device-validation.md, not to this page. The fullscreen andFloor-pinned runs, and the whole real-iGPU side, stay withdocs/on-device-validation.md. -
Driver floor isolated (Plan 0012 Phase 2, resolved): the once-optional dev spike ran —
standalone/examples/floor.rs, a scene-less window standing up only the wgpu context — and put the hard ~327 MB private-commit floor number on the split above. It confirms ADR-0010’s diagnosis: the cost is the driver stack, not our code. Does not change ADR-0010. -
A show configuration is a different workload, and it plateaus near ~800 MB (measured 2026-08-29/30, the first full live set). 8h08m, 3,505,083 frames, zero dropped, 120.0 fps flat end to end, on the show notebook — not the reference AMD iGPU box — at Rich tier with the operator console open and 18 presets rotating from an
RLX_PRESET_DIR. Working set was 678 MB early, climbed to ~795 MB, then went flat: three consecutive half-hours at 795.4 MB, later settling 799.5 → 808.3 MB in discrete steps with flat stretches between. Handles pinned at exactly 370 across six hours, threads at 5,gpu_bytesat 15.8 MB from first sample to last. The step-then-flat shape is a resource claimed on a preset’s first pass through the rotation and then cached. This does not move the ~350 MB soft ceiling above, which is scoped to the reference box and stays a single-machine tripwire — it is a second point on the vendor/workload spread, and the first one taken at show length rather than in minutes. -
Do not call a leak on a window shorter than two flat half-hours. The same run was read mid-session as a linear leak at +0.87 MB/min, on a 20-minute window, and that reading was wrong: it was warm-up. This app’s growth is step-then-flat, so any short window through a step fits a convincing line. The discipline is to hold the measurement until two consecutive half-hours agree.
Measurement method (repeatable): PowerShell Get-Process ritmolux → WorkingSet64 vs PrivateMemorySize64,
.Modules by mapped size, and which backend loader DLLs are mapped. The private-vs-working-set split is
what proved the cost is driver heap, not our code.
Not a Plan 0001 blocker; the leak-guard folds into the §10 live-features soak, the per-system delta into each scene-adding plan.
Built from 13c7582 at version 0.158.0. This site tracks main and is not versioned per release.