# Profiling and performance comparisons

Use `pnpm perf:profile artifacts/profile` for a current-source run. The harness
uses installed Google Chrome and saves raw samples, JSON summaries, and Chrome
CPU profiles. Open `.cpuprofile` files in Chrome DevTools' Performance panel.
It is development-only: published library code contains no profiling hooks.

## Measured change: mesh clipping fast paths

The first diagnostic profiles identified main-thread polygon clipping and
allocation/GC as a startup bottleneck: seven gallery mesh builds plus uploads
occupied 65.7 ms, with only 0.7 ms in texture upload calls. Cells wholly outside
the silhouette now skip clipping; strictly positive interior cells emit the
same two triangles directly. Contour cells, zero-distance vertices, and map
borders retain the existing clipping/wall construction path.

A subsequent diagnostic run measured 33.4 ms for those same seven mesh builds
plus uploads (49% less synchronous wall time). These are diagnostic observations,
not six-run component benchmarks. Uninstrumented paired timings confirm the
small-gallery startup benefit in **all six pairs**, ranging from 6.6% to 13.4%.

| Workload | Ready p50 before | Ready p50 after | Synchronized redraw p50 before | After |
| --- | ---: | ---: | ---: | ---: |
| One 320px badge | 122.7 ms | 117.3 ms | 3.28 ms | 3.53 ms |
| Nine 110px badges | 165.4 ms | 148.9 ms | 9.48 ms | 9.68 ms |
| 57 110px badges | 168.7 ms | 163.1 ms | 33.07 ms | 32.72 ms |
| Nine 640px badges | 703.9 ms | 708.4 ms | 13.13 ms | 12.87 ms |

These are nearest-rank p50 values across six runs, on Chrome 154.0.8037.57,
ANGLE Metal, Apple M5 Pro. Redraw figures summarize each run's p50. No redraw
speedup is claimed: only mesh construction changed. The large workload remains
worker-bake dominated and showed no reliable readiness benefit. All outliers
are retained, including candidate startup times of 210.4 ms for 57 badges and
1186.0 ms for the large workload in the first pair. Inspect the raw paired data
rather than assuming that a single summary quantile proves a universal gain.

Evidence is saved locally in `artifacts/profile-ab/report.json`, with CPU
profiles and event/query records in the same directory. Discovery captures are
in `artifacts/profile-discovery` and `artifacts/profile-instrumented`; the former
revealed the emulated-DPR gap and predates the worker execution timer. The full
source reference is `artifacts/profile-baseline/src`.

All **108 raw RGBA captures are byte-for-byte identical** between baseline and
candidate at the verified backing dimensions. Six reference mesh cases also
preserve every vertex/index byte and the face/wall split. All 23 unit tests,
TypeScript checks, package builds, and the standalone Chrome interaction suite
pass. Captures are in `artifacts/mesh-quality-before` and
`artifacts/mesh-quality-after`.

## Repeatable A/B measurement

Save a complete source snapshot **before** editing library code:

```sh
pnpm build:worker
mkdir -p artifacts/baseline
cp -R src artifacts/baseline/
# Make the optimization, then:
pnpm perf:profile artifacts/comparison --baseline artifacts/baseline
```

The snapshot includes the generated worker bundle. Each server resolves all
library source modules from its own revision; the same fixture drives both.
Do not edit either source tree while a run is active. `PROFILE_ROUNDS` overrides
the default six rounds; use one only for harness smoke tests, not speed claims.
Run browser workloads sequentially and avoid other GPU-intensive work.

The harness warms module transforms and shader disk caches for both variants,
then alternates baseline/candidate order every round. Every measured case gets
a fresh browser context and fresh library/worker caches. Workloads are one
320px badge, nine 110px badges, 57 110px badges, and nine 640px badges. The many-
badge case repeats the nine gallery styles (including locked badges), so it is
not identical to the older benchmark's 48 added badges. Actual canvas dimensions
and map resolutions are recorded and checked, along with GPU identity.

Readiness is timed in-page from host creation through all `ready` promises;
it excludes module download, Vite transformation, and automation polling.
It means the library has baked and submitted the first image, not that a physical
display has presented it. It is not a cold network-navigation measurement.

Each uninstrumented workload records 60 frames after 15 warm-up frames:

- `submitMs`: synchronous wall time for all badge draw calls, including canvas
  copies and any implicit driver/browser waits. It is not GPU execution time.
- `rafIntervalMs`: actual animation callback cadence, including browser scheduling.
  This is measured separately from draw submission and is not inferred from it.
- A separate synchronized pass records 45 frames after five warm-ups, with draw
  time, a one-pixel GPU readback wait, and their total stored separately. This
  stress measurement is not presented as interactive FPS and does not include
  physical display presentation or guarantee completion of 2D compositing.

Reports retain every sample, nearest-rank per-run quantiles, across-run ranges, and paired
percentage changes. Negative changes mean improvement. Six rounds do not prove
universal speedups; inspect outliers and the individual paired changes. Metadata
includes Chrome version, OS/CPU, workload settings, and SHA-256 hashes of both
source trees. Different machines or browser builds require a new paired run.

## Diagnostic profiling

Diagnostic runs happen **after** timing runs. They use method wrappers, a
main-thread CDP sampling profiler, and asynchronous
[WebGL timer queries](https://registry.khronos.org/webgl/extensions/EXT_disjoint_timer_query_webgl2/).
Their numbers identify bottlenecks; the extra queries/wrappers can change GPU
batching and must not be used as the end-to-end speed comparison.

Recorded phases include renderer initialization, environment preparation,
mesh construction plus upload, texture upload, material uniforms, canvas copies,
and draw submission. Durations ending in `cpuWall` measure synchronous API wall
time, which can include waiting. Nested phases overlap: do not sum them as if
they were exclusive CPU categories. Environment-ready time is submission wall
time, not completion of environment filtering on the GPU.

A diagnostic-only worker shim records synchronous bake execution inside each
worker. A listener registered before the library's completion handler separately
records round-trip time, including queueing, transport and delivery. The shim
adds startup overhead, so production readiness comes from uninstrumented runs.
Worker jobs overlap; summing round-trip times does not give page startup time.
Main-thread CPU profiles do not represent worker CPU stacks.

GPU query results distinguish environment, relief, and shadow passes. Results
are read only after yielding; disjoint results are discarded and unsupported or
timed-out queries are reported explicitly. Query objects are deleted. These
queries cover WebGL passes, not Canvas2D copying/compositing. CDP metrics also
record heap and layout data; a single heap snapshot metric is not a leak test.
Long tasks are assigned to the phase in which they started.

## Quality protection

```sh
BADGE_BASELINE_DIR=artifacts/baseline pnpm test:quality artifacts/reference
pnpm test:quality artifacts/current artifacts/reference
pnpm test
pnpm typecheck
pnpm test:browser
pnpm build
```

The quality harness supports complete `src/` snapshots as well as the older
four-file rendering snapshot. It captures nine badge styles at actual backing
sizes of 110, 400, and 640 pixels from four poses (108 images). It checks backing
dimensions rather than trusting emulated DPR: Chrome's emulated 2× DPR did not
increase `devicePixelContentBoxSize` in the original capture harness. These
larger backing sizes exercise high-resolution rendering without claiming to
test native display scaling behavior.

Pixel thresholds, GPU/browser compatibility, blank-image checks and failures
are recorded by `scripts/quality-check.mjs`. PNGs remain available for visual
inspection. Mesh compatibility tests additionally hash complete vertex/index
buffers from the pre-optimization implementation, including contours on grid
vertices, holes, islands, map boundaries, and a dense mesh. Never regenerate
those expected values merely to silence a failed optimization test.
