Archive — history, not state. Kept for its reasoning and its evidence; its plan is closed.
Compiler optimization campaign: final performance matrix
Status: final integrated campaign seed measured and validated. External-runtime lanes are the earlier same-machine comparison run; Pit lanes below are the clean-daemon final branch measurements.
Date: 2026-07-14
Branch: compiler_optimizing
Committed base: da2a46b5901ac3f56ae9ba34b0550b832c90fa1f
Machine: Apple arm64 (Darwin 25.5.0, RELEASE_ARM64_T6050).
This matrix is the composed end-of-campaign measurement, not a collection of best historical numbers. It keeps dynamic/interpreted VM comparisons separate from native/JIT comparisons and verifies each workload result in the harness.
Interpretation groups are:
- Pit Mach, Lua, QuickJS, CRuby, and mruby are VM/interpreter comparisons.
- LuaJIT, Node/V8, Bun/JavaScriptCore, and .NET/RyuJIT are optimizing-JIT comparisons; they are the relevant aspirational field for hot AOT output, but they may use runtime feedback and allocation sinking unavailable to a cold static build.
- Pit native and
ocamloptare AOT comparisons. OCaml is compiled without a flambda flag in this matrix.
The table is deliberately about language workloads rather than identical
object representations. JavaScript/Lua/Ruby use their normal language record
objects. The statically typed ports use Dictionary/Hashtbl for the honest
generic string-keyed record rows; a C# class or OCaml record would instead be a
static offset-layout ceiling. Pit’s new dense shapes mean record_new now
sits between those categories, while the polymorphic d_field and
record_field sites remain the more useful generic-access comparisons.
Method
The comparison sources are aot_bench/cross/bench.*,
aot_bench/cross/shootout.*, and aot_bench/cross/calls.{erl,exs}. Each
micro/shootout program uses its existing calibrated harness: run the workload
once for an oracle, then for each of three inner trials double repetitions until
at least 50 ms and retain the best normalized time. This campaign executed
each complete harness three times serially and reports the median of the three
reported per-row values. Thus every table entry is a median of three
independent calibrated best-of-three values. No benchmark programs ran in
parallel.
Exact commands, each repeated three times:
lua aot_bench/cross/{bench,shootout}.lua
luajit aot_bench/cross/{bench,shootout}.lua
qjs aot_bench/cross/{bench,shootout}.js
node aot_bench/cross/{bench,shootout}.js
bun aot_bench/cross/{bench,shootout}.js
ruby aot_bench/cross/{bench,shootout}.rb
mruby aot_bench/cross/{bench,shootout}.rb
ocamlopt -O3 -o /private/tmp/pit_cross_{bench,shootout} aot_bench/cross/{bench,shootout}.ml
/private/tmp/pit_cross_{bench,shootout}
dotnet run --file aot_bench/cross/{bench,shootout}.cs -c Release
erlc -o /private/tmp aot_bench/cross/calls.erl
erl -pa /private/tmp -noshell -s calls main
elixir aot_bench/cross/calls.exs
Runtime versions: Lua 5.5.0; LuaJIT 2.1.1783773675; QuickJS 2026-06-04;
Node 26.5.0 / V8 14.6.202.34; Bun 1.3.14; CRuby 2.6.10p210; mruby 4.0.0;
OCaml 5.5.0 native (ocamlopt, no flambda flag); .NET SDK/runtime 10.0.301;
Erlang/OTP 29 with BEAM JIT; Elixir 1.20.2 on OTP 29. s7 is not installed.
The macOS /usr/bin/time -l facility could not read kern.clockrate inside
the benchmark sandbox and did not emit peak RSS. Cross-runtime peak memory is
therefore unavailable in this run. Target heap/allocation/GC evidence for Pit
comes from vs.ce’s actor statistics instead.
Cross-runtime micro suite
Milliseconds; lower is better. These are the fourteen workloads currently
implemented consistently in every bench.* source.
| bench | Pit Mach | Pit native | Lua | LuaJIT | QuickJS | Node | Bun | Ruby | mruby | OCaml | .NET |
|---|---|---|---|---|---|---|---|---|---|---|---|
| arith_int | 25.740 | 9.610 | 26.137 | 15.451 | 43.492 | 9.308 | 9.966 | 60.845 | 73.303 | 9.930 | 8.075 |
| arith_poly | 21.800 | 7.430 | 29.254 | 10.820 | 38.480 | 6.730 | 7.119 | 60.933 | 81.009 | 7.082 | 6.146 |
| loop_nested | 24.100 | 11.670 | 35.186 | 19.274 | 48.157 | 11.541 | 12.368 | 69.920 | 86.799 | 11.476 | 9.394 |
| float_math | 16.570 | 2.370 | 8.144 | 1.223 | 11.947 | 1.223 | 1.261 | 26.357 | 26.047 | 1.256 | 1.246 |
| fib | 51.860 | 13.610 | 21.968 | 2.753 | 34.553 | 4.117 | 2.609 | 38.624 | 45.684 | 2.042 | 1.118 |
| call_hot | 14.220 | 7.860 | 32.060 | 9.411 | 32.276 | 5.733 | 6.192 | 47.616 | 75.611 | 6.612 | 4.913 |
| closure | 50.880 | 26.240 | 48.272 | 11.445 | 64.442 | 7.483 | 6.942 | 83.668 | 107.010 | 8.169 | 6.322 |
| array_read | 24.180 | 23.440 | 21.644 | 13.908 | 36.291 | 8.607 | 9.297 | 48.773 | 66.777 | 10.000 | 7.384 |
| array_grow | 2.990 | 2.710 | 2.494 | 0.978 | 2.827 | 0.909 | 0.505 | 4.432 | – | 2.407 | 0.497 |
| record_new | 21.180 | 14.850 | 31.610 | 8.543 | 24.305 | 1.058 | 1.192 | 30.125 | 40.434 | 21.302 | 9.523 |
| string_concat | 0.420 | 0.350 | 11.063 | 8.928 | 0.505 | 0.039 | 0.044 | 10.605 | 34.944 | 9.133 | 27.430 |
| gc_churn | 7.300 | 6.150 | 17.001 | 4.969 | 10.124 | 0.388 | 0.538 | 8.300 | 12.732 | 4.683 | 3.140 |
| d_field | 22.770 | 11.820 | 13.529 | 3.220 | 15.302 | 4.195 | 2.694 | 32.265 | 46.522 | 17.630 | 7.916 |
| record_field | 115.200 | 52.410 | 78.791 | 21.943 | 79.315 | 24.392 | 14.874 | 166.201 | 220.711 | 97.210 | 42.949 |
mruby 4.0.0 rejects the 300,000-element array_grow workload with
array size too big; this is a missing result, not a zero or a timeout.
Cross-runtime composed shootouts
| bench | Pit Mach | Pit native | Lua | LuaJIT | QuickJS | Node | Bun | Ruby | mruby | OCaml | .NET |
|---|---|---|---|---|---|---|---|---|---|---|---|
| nbody | 258.930 | 150.790 | 71.858 | 3.060 | 75.449 | 1.834 | 2.242 | 145.548 | 242.881 | 1.583 | 1.599 |
| binarytrees | 78.790 | 50.060 | 64.896 | 20.188 | 43.206 | 2.794 | 3.283 | 34.702 | 31.799 | 1.192 | 2.820 |
| spectralnorm | 167.770 | 44.550 | 181.180 | 7.424 | 281.940 | 7.005 | 3.891 | 422.066 | 379.857 | 10.539 | 3.491 |
| fannkuch | 227.960 | 111.350 | 140.500 | 17.265 | 191.065 | 12.043 | 9.818 | 447.382 | 400.094 | 13.638 | 10.649 |
| mandelbrot | 89.150 | 9.210 | 51.872 | 41.649 | 61.044 | 2.219 | 2.934 | 140.832 | 172.403 | 3.873 | 2.900 |
All runtime results matched the shared workload oracles.
BEAM call comparison
These are seven-sample medians inside each run, then the median of three
complete process runs. dynamic_call0 measures indirect selection among
eight zero-argument immutable BEAM closures. It is useful call-dispatch
evidence but is not semantically identical to Pit’s mutable captured-cell
closure fixture.
| bench | Erlang/OTP 29 | Elixir 1.20 |
|---|---|---|
| call_hot | 8.253 | 8.173 |
| dynamic_call0 | 26.051 | 25.970 |
| fib | 2.194 | 2.259 |
| tco_self | 4.072 | 3.666 |
Pit-only extended matrix
The Pit harness performs two warmups followed by seven timed samples per lane
and reports the median. The final serialized 68-row aggregate completed with
check=Y for every row. The selected rows below are the clean-daemon medians;
allocation and collection columns are shown where the final summary captured
them. Unlike the earlier exploratory matrix, nbody now statically binds its
math/radians.sqrt dependency and executes in both lanes.
| bench | category | Mach ms | native ms | allocation bytes Mach/native | GC Mach/native |
|---|---|---|---|---|---|
| arith_int | compute | 25.74 | 9.61 | about 1 KiB | 0 / 0 |
| arith_poly | compute | 21.80 | 7.43 | about 1 KiB | 0 / 0 |
| pow2_divrem | compute | 44.00 | 15.13 | about 1 KiB | 0 / 0 |
| loop_nested | compute | 24.10 | 11.67 | about 1 KiB | 0 / 0 |
| float_math | compute | 16.57 | 2.37 | about 1 KiB | 0 / 0 |
| intrinsic_ops | builtin | 24.92 | 13.88 | about 1 KiB | 0 / 0 |
| call0_plain | dynamic call | 45.35 | 40.65 | about 1 KiB | 0 / 0 |
| call_hot | known call | 14.22 | 7.86 | about 1 KiB | 0 / 0 |
| closure | dynamic closure | 50.88 | 26.24 | about 2 KiB | 0 / 0 |
| fib | recursion | 51.86 | 13.61 | about 1 KiB | 0 / 0 |
| tco_self | tail recursion | 44.14 | 4.98 | about 1 KiB | 0 / 0 |
| switch_chain | dispatch | 59.82 | 15.30 | about 1 KiB | 0 / 0 |
| array_read | memory | 24.18 | 23.44 | about 17 KiB | 0 / 0 |
| arrfor_sum | higher order | 43.12 | 24.97 | about 134 KiB | 0 / 0 |
| record_field | generic record read | 115.20 | 52.41 | about 2 KiB | 0 / 0 |
| pgo_record_field, unprofiled | dynamic shaped read | 38.19 | 25.83 | about 2 KiB | 0 / 0 |
| record_template | shaped construction | 34.17 | 26.84 | – | – |
| record_template_suspend | shaped/suspending | 20.59 | 16.12 | – | – |
| record_nullable_shape | shape fallback | 3.96 | 4.80 | – | – |
| mach_guard_fusion | guard control | 11.33 | 11.28 | – | – |
| array_grow | allocation | 2.99 | 2.71 | 8,389,464 / 8,389,608 | 3 / 2 |
| record_new | allocation | 21.18 | 14.85 | 16,817,096 / 16,817,264 | 20 / 11 |
| string_concat | allocation | 0.42 | 0.35 | 492,816 / 492,912 | 0 / 1 |
| gc_churn | allocation | 7.30 | 6.15 | 9,625,416 / 9,625,600 | 10 / 7 |
| rec_alloc | recursion/allocation | 6.66 | 4.72 | 4,800,744 / 4,800,744 | 4 / 1 |
| record_grow_keys | structural mutation | 0.12 | 0.23 | about 65 KiB | 0 / 0 |
| native_call_generic | C leaf | 14.82 | 8.81 | about 1 KiB | 0 / 0 |
| native_call | static-export C leaf | 11.66 | 8.39 | about 1 KiB | 0 / 0 |
| native_nested_generic | nested C leaf | 20.52 | 14.74 | about 1 KiB | 0 / 0 |
| native_nested | nested static-export C leaf | 17.51 | 13.49 | about 1 KiB | 0 / 0 |
| d_loop | decomposition | 3.00 | 1.58 | about 1 KiB | 0 / 0 |
| d_rec1 | decomposition | 2.44 | 4.45 | about 1 KiB | 0 / 0 |
| d_rec3 | decomposition | 2.91 | 4.99 | about 1 KiB | 0 / 0 |
| d_field | decomposition | 22.77 | 11.82 | about 2 KiB | 0 / 0 |
| d_arr3 | decomposition | 15.34 | 11.07 | – | – |
| licm_reset | correctness/perf canary | 9.24 | 1.53 | – | – |
| native_stone_forward | stone canary | 13.50 | 10.84 | – | – |
| stone_field | memory | 7.70 | 3.89 | about 1 KiB | 0 / 0 |
| mandelbrot | shootout | 89.15 | 9.21 | about 1 KiB | 0 / 0 |
| fannkuch | shootout | 227.96 | 111.35 | about 1 KiB | 0 / 0 |
| spectralnorm | shootout | 167.77 | 44.55 | about 25 KiB | 0 / 0 |
| binarytrees | shootout | 78.79 | 50.06 | about 42.2 MiB | about 27 GC |
| nbody | shootout | 258.93 | 150.79 | – | – |
The nursery remained disabled throughout this final matrix: all nursery allocation and collection counters were zero.
Current readout
- Pure integer loops now land where intended: Pit native is essentially tied
with OCaml/.NET on
arith_int,arith_poly, andloop_nested; Pit Mach is competitive with or faster than Lua on those rows. - Known/inlined calls are also healthy:
call_hotnative is close to OCaml and faster than LuaJIT. Honest dynamic activation is not: the polymorphic zero-argumentcall0_plaincosts 40.65 ms and the captured/dynamicclosurerow is 3.2x OCaml and 2.3x LuaJIT. The retained Mach outer-frame cache does help the closure row, but it does not change native activation. The BEAM indirect-zero-arg result around 26 ms is a useful directional comparison, although BEAM’s fixture has immutable closures and does less accumulator arithmetic than Pit’s fixture. - Boxed array traffic is a clear remaining native weakness.
array_readis no faster native than Mach, 2.4x OCaml, and 1.7x LuaJIT. This agrees with the rejected exact boxed-array proof: check removal did not remove the dominant representation traffic. - Dense shapes materially fixed record allocation density:
record_newallocates 16.8 MB rather than the prior roughly 43.2 MB and is faster than the ordinary Lua/QuickJS/Ruby lanes. It is still allocation-bound and far behind JS JIT object allocation. Honest polymorphicrecord_fieldremains 2.4x LuaJIT but now beats OCaml’sHashtblanalog. - The shootouts identify different ceilings. Native
mandelbrotis 4.5x faster than LuaJIT and within 2.2x of OCaml, so scalar numeric lowering is working.spectralnormis 6.0x LuaJIT and 4.2x OCaml, dominated by calls and boxed array access.fannkuchis 18.5% faster than interpreted Lua, but remains 6.4x LuaJIT and 8.2x OCaml because of dense array/index mutation.binarytreesis slower than Lua and spends about 43 MB/27 collections; object allocation and collector behavior dominate it.nbodyis an especially poor result (150.79 ms native versus 1.58 ms OCaml), showing that repeated array/record access and call boundaries still overwhelm otherwise healthy scalar arithmetic.
Movement from the contained campaign baseline
compiler-optimizing-handoff.md records the fresh contained-campaign baseline
at 81685c37. It predates several mechanisms already present at this arc’s
da2a46b5 starting point, so this table measures the full contained compiler
journey rather than claiming every delta for the linked/shape/PGO arc alone.
| benchmark | baseline Mach/native | current Mach/native | Mach change | native change |
|---|---|---|---|---|
| mandelbrot | 108.30 / 28.01 | 89.15 / 9.21 | -17.7% | -67.1% |
| fannkuch | 297.85 / 241.13 | 227.96 / 111.35 | -23.5% | -53.8% |
| spectralnorm | 305.97 / 172.10 | 167.77 / 44.55 | -45.2% | -74.1% |
| binarytrees | 97.66 / 60.71 | 78.79 / 50.06 | -19.3% | -17.5% |
| dynamic closure | 58.07 / 30.43 | 50.88 / 26.24 | -12.4% | -13.8% |
| Fibonacci | 54.24 / 28.79 | 51.86 / 13.61 | -4.4% | -52.7% |
| self TCO | 89.13 / 30.30 | 44.14 / 4.98 | -50.5% | -83.6% |
The final same-artifact canary revises the earlier optimistic shape result.
Dense/default binarytrees measured 74.79/47.25 ms and 43.25 MB, versus
83.63/44.86 ms and 53.96 MB for generic-template and 80.37/51.66 ms with shapes
off; a generic repeat was 80.34/44.73 ms. Dense shapes are therefore a clear
memory and Mach win and beat no-shape native, but they carry a repeatable
approximately 5% native regression versus the generic-template representation
on this one allocation-heavy workload. They remain retained because the shared
representation cuts each common three-field record from 144 to 56 bytes,
removes roughly 20% of binarytrees allocation traffic, and improves the
default path on both lanes relative to no shapes. The native regression is an
explicit follow-up target, not hidden as noise.
PGO production result and artifact identity
The final program-scoped profile run returned 3,200,000 in every mode:
| mode | Mach ms | native ms |
|---|---|---|
| ordinary | 53.875 | 30.147 |
| collect | 56.283 | – |
| auto/consume | 45.954 | 22.895 |
The generation-1 portable profile is 7,088 bytes. Its semantic hash is
d28969b9ced728b4e4753a5c998de137131819c715560fe0c05d50042db1f91f and
its content hash is
48c924e48b69c3e487cac489e4fc988a4ec3a2dcd4a8db8ce4bde1e2409af9a9.
Consumption uses only record-load sites with at least 1,024 observations and
90% valid top-four shape coverage.
Final strict artifacts contain 65 validated mcode modules. Exact sizes are
74,842,189 bytes for boot.qop, 67,509,488 bytes for boot/firmware, and
1,573,945 bytes for boot/boot (143,925,622 bytes combined). Compared with the
immediate pre-audit campaign artifacts, the combined result is 1,095,242 bytes
smaller; the engine is 126,372 bytes larger because it contains the strict
portable validator/hydrator.
The final SHA-256 identities are
a58afba087c65d9d04b31c10ade9c16400fc3766b3a3f8a8295749462eeb7d57
for boot.qop,
797db28fc9b1bec61db89da36a5656049db9da3389ce72a6f0dc61c12b175e89
for boot/firmware, and
6fff594fcb6f130c9c2ab5db5c1f595081320ec56d1c7f6f11ea1ebb4bf17174
for boot/boot.
Integrated validation and memory hygiene
make seedpassed; Meson passed 7/7; focused compiler tests passed 145/145.- The isolated VM suite passed 1,086/1,086 and the warmed full suite passed
1,921/1,921. One cold full attempt timed out only while charging realization
to
vm_suite; its isolated and warmed reruns passed. - Deterministic fuzz passed 3,733/3,733 checks over 500 programs with seed
20260713. The 68-row aggregate returned
check=Yfor every row. - Final validation found and fixed two general wrong-code hazards: LICM no
longer treats branch-local/conditional writes as dominating replacements,
and QBE now records helper-written
is_array,is_func, andis_recordpredicate destinations as frame writes.
A clean node mapped 64 MiB: 4 MiB of actor heap, 1 MiB of node constants, and
about 20 MiB of shared Mach code; the shop was about 2 MiB live in an 8 MiB
block and the builder about 59 KiB live in a 256 KiB block. Running the entire
aggregate and test fleet without restarting grew the node to 704 MiB mapped,
278 MiB of actor heaps, 37 MiB of node constants, and a 242/256 MiB builder
block. That is compiler/test-daemon retention, not target benchmark memory: it
slowed binarytrees to 216/136 ms and record_new to 70/90 ms while leaving
call_hot near 15.5/6.67 ms. Restarting restored 78/48 and 23/15 ms. Final
performance claims use only the restarted clean-daemon measurements.
Source: plans/archive/perf-2026-07/perf-campaign-final-matrix.md