---
name: rnx-perf
description: Optimize Peach engine and app performance — scroll jank, dropped frames, slow animation, idle CPU, memory growth, and "smooth normally but laggy under load" (video calls, busy machines). Covers measuring on the right thread with the right browser, proving each perf tier actually engages via its counters, a CPU-contention test harness, and the map of existing perf machinery so you improve it instead of reinventing it. For general app debugging load /rnx-debug first.
---

# Peach performance

optimize performance in Peach: the shell frame pipeline, the compositor
tiers, and the guest app. this skill exists because every real perf win in
this engine followed the same loop — measure on the right thread, read the
counters that prove machinery engaged, fix the one thing the counters
indict, remeasure — and every wasted week came from skipping a step.

> needs a connected, pinned sim (`/rnx-setup`). general "what's wrong"
> triage lives in `/rnx-debug`; this skill goes deeper on perf only.

## the four laws (each paid for by a real incident)

1. **measure the worker that painted it, in the right browser.** the tenant
   hasn't painted since the two-worker split, so never read a tenant sampler
   (the one reporting `avg 0.48ms (2069 fps)` for a visibly janky scroll was
   removed). but "the shell" is no longer the whole answer either: the
   compositor worker owns `home`, `app:one` and `app:two`, so **every guest-app
   draw call, paint boundary and raster-tier decision is in the compositor's
   render profile, not the shell's.** `perf shell stop` prints both, labelled
   by worker — read the compositor block for anything inside an app. for
   months it printed only the shell's, and a dense token-list fling
   therefore reported `raster tier: 0 promotions / 0 blits` while the
   compositor was blitting ~14 rows a frame; the giveaway was `node visits: 7`
   next to `app:one 1014 paints`. if a render-profile block describes a
   handful of nodes while an app surface is painting hundreds of times, you
   are reading the wrong worker.
   browser matters just as much: playwright's bundled headless chromium has
   NO GPU — CanvasKit falls back to software rendering and every number lies
   high. launch a dedicated GPU-backed Chrome executable with a disposable
   profile and an owning process supervisor that terminates its exact process
   tree. never automate the user's own browser profile.
2. **machinery silently not engaging is the default failure mode.** the
   raster tier compiled, gated, and did *nothing* for months — zero
   promotions on its textbook case — until per-gate rejection counters
   existed. the frame-demand gate was blind to commit storms. the vsync
   pump flapped 100 msgs/s at rest. never assume a tier works because it's
   wired; read its counter during the exact workload.
3. **check reclaimable memory and sampled CPU before blaming code.** an OOM or
   busy machine mimics every perf and infra failure. on macOS, read
   `memory_pressure -Q` and take multiple `iostat -c 3` CPU samples; on Linux,
   read `MemAvailable` from `/proc/meminfo` and sample CPU idle over time. keep
   the busiest sample. raw macOS `freemem()` and load average are not admission
   signals. attribute browser processes with `ps -eo pid,ppid,args` before
   treating one as abandoned or stopping an exact pid.
4. **never trade animation fidelity for CPU.** perpetual sub-frame animation
   must keep its authored cadence. attack per-frame cost (layers, caching,
   damage), never cadence or fidelity.

## measuring

### presentation proof must stay outside the simulator

Performance acceptance must not use `rnx record`, `liveComposite`,
`captureBitmaps`, Chrome screencast, or another in-page frame-copy path. Those
paths add work to the simulator or browser and can change the cadence being
measured. Record the visible window with the operating system's native screen
capture instead.

Use an operating-system-owned window recorder. Target the exact browser window
rather than the display when the platform supports it, start recording before
normal user input, preserve the movie, and record the operating system and
recorder version. Confirm that the recorder process, not the browser or
simulator, owns capture and encoding. If native window capture is unavailable,
report the proof as blocked instead of substituting an in-page capture path.

The visual run contains only normal user input. Do not query the performance,
debug, or test bridge while it records. Treat that video as visual evidence of
frame progression. Collect worker paint, flush, delivery, and tier counters in
a separate run with screen capture off. The two runs answer different questions
without making the simulator measure its own recording overhead.

Open the sim with `?debug=framestamp` to make that video machine-readable.
Every surface paint then draws its `Date.now()` low 20 bits as a 23-cell
barcode — `home` and `app:*` at 45% of the surface height, `shell-overlay` on
its own row at 55% so overlay paints (keyboard, native UI) are attributable
separately. Decode the documented 23-cell barcode with a compatible framestamp
decoder to report, per presented frame, which paint it shows and how late that
paint presented. If no decoder is available, keep the recording as qualitative
evidence and make no numeric presentation-latency claim from it.

This is the only way to separate a frame the compositor never produced from one
it produced and the browser never presented.

**Screen capture is not presentation-neutral, so the decoded timing expires.**
Starting a ScreenCaptureKit window capture flips the 120hz `OffscreenCanvas`
out of latest-wins presentation into sequential presentation, which falls behind
roughly 0.5ms for every 1ms of capture. A capture started mid-activity begins
near zero latency and climbs from there. A recording is therefore truthful about
content and trajectory for its whole length, and truthful about timing for
roughly its first 500ms; latency growth past that is the capture, not the
engine. Live uncaptured presentation stays latest-wins and current. Never report
a late-recording latency as an engine number, and never compare two recordings
of different lengths.

Worker counters and presentation are different claims, and a green counter never
licenses a smoothness claim. A capture can report a perfect 120hz worker cadence
for motion the user never sees move: the overscroll release shipped exactly that
way, with the compositor stepping on schedule while the presented frames showed
a clamped fling instead of the bounce spring. Counters prove a tier engaged;
only a framestamp decode of an out-of-process recording proves what presented.

When an external transform control is smooth but the simulator is not, inspect
every animation on the presented element. A separate Web Animation object does
not isolate paint work: animating `box-shadow`, `filter`, or another paint-only
property on the transform owner can move the whole trajectory to the browser's
main thread. Paint stable decoration once or give it a distinct layer.

```sh
rnx perf shell start          # arm shell frame capture (clears prior)
# reproduce — e.g. 8 fling swipes:
#   for i in {1..8}; do rnx do swipe 196 650 196 250 8 8 --no-wait; sleep 0.9; done
rnx perf shell stop           # report (add --json for scripting)
```

the report and how to read it:

- `painted frames / avg / p50 p95 p99 / jank` — work-per-painted-frame.
  budget: 8ms at 120hz, 16.7ms at 60hz; p95 > 20ms is visible jank.
- `surfaces` — per-surface layout/render/flush split. `layout` spikes on a
  scroll capture mean shell-side yoga re-layout when virtualized rows
  materialize (mid-fling commit), not steady-state cost.
- `boundaries: N records / M replays` — layer compositor health. healthy
  scroll: replays >> records, record cost near zero. records ≈ replays
  means invalidation is defeating the cache (the SkPicture attempt-1
  failure shape: `records: 1525, replays: 35`).
  the `why recorded` line under it splits those records three ways, and each
  one is a different bug: `invalidated` means something dirtied the boundary's
  content (the usual culprit, and usually a style write that did not change a
  value), `moved` means the boundary's record origin shifted so the cached
  picture no longer lines up, and `first-record` is a boundary that had no
  picture yet. a scroll dominated by `invalidated` is an invalidation bug; one
  dominated by `first-record` is a warm-up problem, so look at whether the
  idle pre-record reaches that content.
- `raster tier: promotions / blits` — blits per scrolled frame > 0 or the
  tier is not firing; `raster skips` names the first gate each candidate
  failed (rebuild / animated / cacheable). zero blits on a feed scroll is
  a bug, not a tuning matter — but check you are reading the compositor's
  block first (law 1), because the shell's is legitimately zero for an app
  workload.
  **before widening any skip bucket, read the `also fail a later gate`
  line under it.** the counters name only the FIRST failed gate, so a big
  bucket can be entirely blocked a second time and unwidenable: on one dense
  token-list fling all 8298 `animated` skips also carried a style transform
  or a non-cacheable chunk, and narrowing the animated gate moved 6007 to
  `transform` and 2350 to `cacheable` while promoting nothing. a bucket is
  worth attacking when its also-blocked share is small.
  and blits firing is not the same as the tier paying for itself: on that same
  token list the tier blits every visible row and costs 70MB of GPU
  cache for no frame-time change, while the retained row boundaries under it
  cost 2.4ms a frame. the discriminator is per-row draw density, so check
  `draw calls` in the same block before concluding a cache is earning its
  keep.
- `worst frames` — read the layout/render split of each; they are usually
  a different problem than the average.

three runs minimum; single-run p95 is noise. a fresh `perf start` between
runs clears the buffer.

### browser-site interaction matrices

for a browser-level regression, use three fresh documents per browser engine.
within each document, capture cold then warm versions of the same interaction;
do not call two independent fresh documents "cold versus warm." wait for the
home surface's first compositor paint before starting—the compositor-ready flag
only proves worker bootstrap and can precede the first visible layout.

during a measured transition, use a fixed observation window. do not poll the
shell state, test tree, or perf bridge while the animation is running; verify
the positive control once after profiling stops. query a targeted id or text on
large virtualized apps, never the full tree, because enumeration can dominate
or terminate the browser process. for a remotely populated list, require the
same scroll-container identity and content extent for a quiet window before the
gesture; "a row exists" can still be true while the app replaces the tree.
keep pointer down, moves, release, and the fixed settle window inside one
uninterrupted profile. stopping or querying between drag and release changes
velocity, cache state, and worker scheduling.

the minimum browser matrix is home drag plus snapback, a cold/warm built-in app
open, and a cold/warm app-switcher open. add one quiet heavy-app screen, one
navigation transition, and one long-list scroll when those surfaces are in
scope. use GPU-backed Chrome for Chromium claims. use installed WebKit in a
visible nonpersistent WKWebView for Safari-engine claims; Playwright WebKit is
not the installed Safari engine.

every receipt must carry the exact page and frame URLs, browser user agent,
browser executable path and content hash,
harness-source identity, capture time, engine build identity and generation,
runtime bundle receipt, sampled host capacity at run start, runtime errors, and
any partial failure. use the shared host-capacity sampler; do not record raw
`freemem()` or decide from macOS load average.
partial artifacts remain inadmissible until every requested run number for that
scenario exists with one harness, browser executable, engine, seed, and
viewport identity. browser
comparisons use the same root/target viewport tuple and backing scale. for a
wrapper or embedded surface, first prove the target has nonzero presented
geometry; a finite host viewport that responsively hides the target measures a
different surface. direct and embedded modes may intentionally have different
tuples, but browsers within one mode may not. finite but different device pixel
ratios are not comparable. if setup reloads only the child frame, replay the
parent's presentation control and wait for the new child to acknowledge it. a
selected parent toggle can outlive its child and is not proof of the child's
mode.
the profile intentionally blocks Graphite cache cleanup, so its cache snapshot
is not a terminal memory verdict. require a capture-local post-profile
compositor snapshot with before, after, and delta cleanup counters. an
over-limit profiled Context needs fresh submitted-work attempt and
successful-completion deltas; an over-limit profiled Recorder also needs a
fresh cleanup in that capture. classify Context and Recorder independently as
within-limit, locked-over-limit, or purgeable-over-limit; never compare their
aggregate, and never admit purgeable overage. a cache with zero purgeable bytes
can still hold a returned-resource queue: settled maintenance releases completed
Recorder resources, then budget enforcement evicts
least-recently-used overage before the next paint or resource insertion can
inherit the synchronous work. perpetual renderer
animation is not finite interaction demand; finite motion, readbacks, and
profiling must still block cleanup.
the host-clock `displayFrames`/`timerFrames` and shell-link
`displayPumps`/`timerPumps` counters prove whether the visible-rAF fallback
engaged. preserve failed receipts instead of retrying them out of the data set.

### confirm which build you measured, every time

Against a dev-source sim, the FIRST reload after an engine edit serves the
PREVIOUS module. The sim reports ready, the app runs, the numbers look real, and
they describe the build before yours. Two screenshots were read as a fidelity
regression this way before the cause was found.

So make every capture say which build it came from, rather than assuming the
reload took. Put a marker in something the capture already returns - a counter
name in `saveLayerReasons`, a debug event, a log line - and treat a capture with
no marker as a capture of the wrong build. Reload twice and re-read the marker.

This costs one line and it is the difference between a measurement and a guess.
The same rule covers any A/B where you edit between arms: an arm you cannot
identify from its own output is not evidence.


deeper tools when the frame report isn't enough:

```sh
rnx maestro test .maestro/scroll.yaml --profile # per-step + frame stats
rnx perf cpu --duration 5 --output /tmp/t.cpuprofile   # sampled CPU
rnx open 8089 --new --driver playwright --cdp-port 9222 # cpu-profile a driveable sim
rnx debug enable layout,render && rnx debug recent layout 40
```

`rnx perf cpu` writes one profile for the page and one for every attached
worker, plus a manifest that labels shell, compositor, tenant, worklet, and
unknown workers. exit zero requires page, shell, compositor, and tenant plus
every other discovered target, with each covering at least 99% of the requested
window. inspect start/stop offsets as well as sample counts; a late-attached
near-empty profile is incomplete. a profile of only the tenant—or a helper
that relabels the busiest worker as tenant—cannot distinguish CPU saturation
from sparse browser scheduling.

for an automated interaction, pass `--interaction-barrier` and start the cpu
command during the runner's pre-workload attachment delay, and give the runner
a post-workload delay long enough to remain open through profiler stop. the
pre-workload delay alone is not synchronization: it can end the profile while
the workload is still preparing. barrier mode pauses the workload before its
first measured capture until every discovered target is profiling, then rejects
the manifest unless the workload publishes its completion marker, every expected
capture label, and action timestamps inside the capture interval. retain its
browser executable, engine build and generation, root and target viewports, and
workload source identity in the CPU manifest. the barrier also has a bounded
page-side lease so an interrupted
profiler cannot strand the workload; lease expiry releases the runner but
invalidates the profile.

### the cadence block: frames that never happened

work-per-painted-frame cannot see a frame that was never delivered because a
worker was descheduled or its resilient clock fell back late. the report shows
low avg and zero jank while the user sees stutter.
the `cadence` section of `perf shell stop` measures the delivery side and
attributes a missing frame to its hop:

- `display clock / host rAF interval` — stretch here (gaps >1.5x median)
  means the page thread's rAF starved or was throttled, upstream of the
  shell worker entirely.
- `delivery lag` — vsync postMessage receipt latency above the run's
  best case; jitter here means the shell worker was descheduled or its
  queue was busy.
- `paint interval` — the end-to-end cadence the user perceives.
- `idle breaks` — gaps >250ms, excluded from all three: demand-gated
  quiesce is legitimate, not starvation.

the compositor `compositorHostFrameClock` keeps its historical name and reports
display/timer ticks consumed by the compositor worker's resilient clock during
the acknowledged profile window. `published` is null because renderer-main no
longer publishes compositor frames. sparse consumed ticks with low production
time identifies worker scheduling starvation. many consumed ticks with sparse
paints can be legitimate dirty-gating; correlate the frame series before
treating it as a regression. shell cadence comes from its worker frame and paint
series.

compare baseline vs contended captures of the same scripted window; also
compare **painted-frame counts**, the crudest and most robust signal.

the compositor cadence lines separate worker delivery from work submitted by
the worker:

- `engine-empty rAF` includes only sustained intervals where adjacent frames
  have no engine demand, input message, shared-slot demand epoch change,
  paint, prewarm pass, or CanvasKit flush.
- `rAF with CanvasKit flush` and `rAF after CanvasKit flush` isolate the frame
  that called `Surface.flush()` and the next delivered frame.
- `CanvasKit submission` measures intervals between final `Surface.flush()`
  completions on frames that submitted at least one surface. The JSON
  `frameSeries` preserves each phase flag and exact count for correlation.

these are engine-owned boundaries. Playwright exposes raw CDP sessions only
for Chromium, and WebKit's public Inspector timeline exposes a WebCore
`Composite` record without a RemoteLayerTree transaction id or commit event.
Do not label CanvasKit flush or compositor rAF as a RemoteLayerTree commit.

## the contention harness ("smooth alone, laggy on a zoom call")

low-load numbers do not predict feel under real-world CPU pressure. run the
same scripted capture twice, once with synthetic load, self-terminating so
it can't leak:

```sh
# 6 cores of load for 30s, self-terminating
for i in 1 2 3 4 5 6; do
  (perl -e 'my $end=time()+30; while(time()<$end){my $x=0; $x+=rand() for 1..10000}' &)
done
```

compare: painted-frame count (cadence), avg/p95 (work), layout spikes. the
diagnostic split:

- **work grows under load** (avg/p95 way up) → paint path too heavy;
  attack per-frame cost (layers, damage, raster, draw calls).
- **work flat but frames missing** (counts drop, avg barely moves) →
  delivery starvation: the host-rAF → postMessage → shell hop chain and
  worker scheduling. browsers give no thread priorities, so the only
  levers are fewer hops on the frame-critical path, less work per hop, and
  keeping motion (scroll, native anims) as close to the paint site as
  possible.
- also remember GPU contention is real (video calls encode on the GPU);
  webkit-lane numbers are Metal, chrome numbers are its GL/Metal stack.

## the perf machinery map (improve, don't reinvent)

every tier below exists and has a counter. before optimizing, identify
which tier *should* absorb your cost and prove whether it does.

| tier | what it absorbs | proof it engaged |
| --- | --- | --- |
| demand-gated frames + visibility re-marking | idle costs nothing; invisible animations quiesce | `+N skipped idle ticks`; `queryStats().needsVsync` false at rest |
| per-surface dirty scoping | one app's commit/animation repaints only its surface | idle repaints ≈ commit rate, owning surface only |
| dirty-tracked layout + tiered text measurement | color/opacity/transform commits skip yoga; repeated strings skip shaping | `layout: 0` on steady frames |
| layer compositor (`paint-boundary.ts`) | animating a child never re-records ancestors | replays >> records |
| idle boundary pre-record | the first drag replays instead of recording; cold records happen while the surface is quiet | `pre-recorded: N boundaries at idle` in the report, and the first interaction's record cost near its warm cost |
| damage rects (`damage-rect.ts`) | repaint clips to changed region (opaque-backdrop gated) | worst-frame render bounded during small updates |
| raster tier | stable rows become GPU textures; scroll skips recording | compositor block: promotions > 0 once warm, blits ≈ one per visible row on a dense feed |
| flood guard, pointer coalescing, vsync hysteresis | message-storm and pump-thrash protection | msgs/s sane during gestures and at rest |
| worker-owned resilient frame clocks | renderer-main stalls do not stop shell or compositor motion | compositor consumed display/timer frames plus shell worker frame and paint cadence |

there is no whole-screen scroll blit. one existed for a few hours in June
2026 and was reverted (`6706f3a014`): it shifted a surface RECTANGLE, so a
full-surface feed with floating translucent chrome dragged its bars along with
the content and smeared their old pixels through the feed. scroll rides the
layer, damage, and raster machinery above instead, and any future blit has to
own a scroll SUBTREE with provable backdrop correctness rather than a screen
rect. do not look for `blitBlocked`; it does not exist.

if a workload's cost lands in a tier that shows zero activity, the fix is
almost always "why didn't it engage" (a gate, a shape, an eligibility
rule), not new machinery. the area-gate fix that finally made the raster
tier fire on feeds came entirely from rejection counters.

## guest-app-side causes (fix the app, or fix the engine honestly)

- dynamic format strings (`` `${n} messages` ``) bust the shaping LRU —
  split number and label into separate `<Text>` nodes.
- `Animated.Value` on `width`/`height` forces layout per frame — use
  `transform` (scale/translate).
- per-frame React work from reanimated-style code is an engine conformance
  bug (valid worklet apps must not commit per frame) — fix Peach core,
  never add a divergent fast-path (upstream alignment overrides perf; see
  the `shouldFreezeOffscreenStackContent` incident).
- render-storm loops (poll → remount) show up as commit-rate repaints and
  RSS growth; check `debug state animations` and the timeline.

## closing the loop

done means: a named hot path + the counter or split that indicts it + the
same capture re-run showing the number moved + no new errors (`get errors`)
+ nothing torn down left running (close sims you opened, kill load loops).
for engine changes, run the affected integration tests exposed by the installed
Peach version and capture stable screenshots of surfaces you touched.
