Packed-Q4 projection scoreboard
Resident inputs; min of three run medians in each of two process sessions. The headline uses the slower session. Metal packs 128 GEMVs or 16 GEMMs in one command buffer, with a uniform 250 ms untimed soak, then three warmups and five samples per run. CUDA uses cuEvent-timed bench_kernel, five warmups and 30 samples per run. Opposite variant ordering in the second session. Host uptime/load is preserved in every source log.
Weights include F16 block scales: 18 bytes per 32 weights (0.5625 bytes/weight). Floors use decimal 500 GB/s for M5 Max and 273 GB/s for GB10. Repeated resident inputs may benefit from cache; reported effective bandwidth is not a measurement of physical DRAM bandwidth. The weight-only floor is a lower bound for prefill, which also has substantial arithmetic and shared-memory work.
| Backend | M × N × K | Selected kernel | Iron ms | Floor ms | Ratio | Effective GB/s | Speedup session 1 / 2 |
|---|---|---|---|---|---|---|---|
| metal | 1 × 1024 × 5120 | coalesced_2row | 0.010280 | 0.005898 | 1.74× | 286.9 | 1.51× / 1.44× |
| metal | 1 × 5120 × 6144 | coalesced_2row | 0.037197 | 0.035389 | 1.05× | 475.7 | 1.97× / 1.97× |
| metal | 1 × 5120 × 17408 | coalesced_2row | 0.104623 | 0.100270 | 1.04× | 479.2 | 1.81× / 1.83× |
| metal | 1 × 6144 × 5120 | coalesced_2row | 0.038336 | 0.035389 | 1.08× | 461.6 | 1.89× / 1.98× |
| metal | 1 × 10240 × 5120 | coalesced_2row | 0.058321 | 0.058982 | 0.99× | 505.7 | 1.99× / 1.95× |
| metal | 1 × 12288 × 5120 | coalesced_2row | 0.076078 | 0.070779 | 1.07× | 465.2 | 1.77× / 1.78× |
| metal | 1 × 17408 × 5120 | coalesced_2row | 0.102744 | 0.100270 | 1.02× | 488.0 | 1.88× / 1.82× |
| metal | 128 × 1024 × 5120 | packed | 0.125758 | 0.005898 | 21.32× | 23.5 | 1.85× / 1.85× |
| metal | 128 × 5120 × 6144 | bm32 | 0.289078 | 0.035389 | 8.17× | 61.2 | 2.93× / 3.00× |
| metal | 128 × 5120 × 17408 | bm32 | 0.920432 | 0.100270 | 9.18× | 54.5 | 2.69× / 2.69× |
| metal | 128 × 6144 × 5120 | bm32 | 0.319297 | 0.035389 | 9.02× | 55.4 | 2.82× / 2.78× |
| metal | 128 × 10240 × 5120 | bm32 | 0.482247 | 0.058982 | 8.18× | 61.2 | 2.98× / 3.09× |
| metal | 128 × 12288 × 5120 | bm32 | 0.599068 | 0.070779 | 8.46× | 59.1 | 2.92× / 2.97× |
| metal | 128 × 17408 × 5120 | bm32 | 0.850148 | 0.100270 | 8.48× | 59.0 | 2.88× / 2.91× |
| metal | 512 × 1024 × 5120 | packed | 0.270737 | 0.005898 | 45.90× | 10.9 | 2.72× / 2.71× |
| metal | 512 × 5120 × 6144 | bm32 | 1.180750 | 0.035389 | 33.36× | 15.0 | 2.90× / 2.91× |
| metal | 512 × 5120 × 17408 | bm32 | 3.679971 | 0.100270 | 36.70× | 13.6 | 2.67× / 2.68× |
| metal | 512 × 6144 × 5120 | bm32 | 1.182729 | 0.035389 | 33.42× | 15.0 | 2.98× / 3.00× |
| metal | 512 × 10240 × 5120 | bm32 | 1.967768 | 0.058982 | 33.36× | 15.0 | 2.93× / 2.95× |
| metal | 512 × 12288 × 5120 | bm32 | 2.345464 | 0.070779 | 33.14× | 15.1 | 2.98× / 3.02× |
| metal | 512 × 17408 × 5120 | bm32 | 3.334711 | 0.100270 | 33.26× | 15.0 | 2.95× / 2.95× |
| metal | 1024 × 1024 × 5120 | packed | 0.508617 | 0.005898 | 86.23× | 5.8 | 2.50× / 2.52× |
| metal | 1024 × 5120 × 6144 | bm32 | 2.436708 | 0.035389 | 68.85× | 7.3 | 2.88× / 2.82× |
| metal | 1024 × 5120 × 17408 | bm32 | 8.115315 | 0.100270 | 80.93× | 6.2 | 2.54× / 2.52× |
| metal | 1024 × 6144 × 5120 | bm32 | 2.425771 | 0.035389 | 68.55× | 7.3 | 2.92× / 2.90× |
| metal | 1024 × 10240 × 5120 | bm32 | 4.047036 | 0.058982 | 68.61× | 7.3 | 2.85× / 2.87× |
| metal | 1024 × 12288 × 5120 | bm32 | 4.885531 | 0.070779 | 69.03× | 7.2 | 2.83× / 2.87× |
| metal | 1024 × 17408 × 5120 | bm32 | 6.981328 | 0.100270 | 69.63× | 7.2 | 2.83× / 2.86× |
| cuda | 1 × 1024 × 5120 | coalesced_x4+vector | 0.016768 | 0.010803 | 1.55× | 175.9 | 2.95× / 2.96× |
| cuda | 1 × 5120 × 6144 | coalesced_x4+vector | 0.075968 | 0.064816 | 1.17× | 232.9 | 3.59× / 3.68× |
| cuda | 1 × 5120 × 17408 | coalesced_x4+vector | 0.211264 | 0.183645 | 1.15× | 237.3 | 3.63× / 3.63× |
| cuda | 1 × 6144 × 5120 | coalesced_x4+vector | 0.074304 | 0.064816 | 1.15× | 238.1 | 3.64× / 3.64× |
| cuda | 1 × 10240 × 5120 | coalesced_x4+vector | 0.124096 | 0.108026 | 1.15× | 237.6 | 3.67× / 3.62× |
| cuda | 1 × 12288 × 5120 | coalesced_x4+vector | 0.146848 | 0.129632 | 1.13× | 241.0 | 3.66× / 3.65× |
| cuda | 1 × 17408 × 5120 | coalesced_x4+vector | 0.213280 | 0.183645 | 1.16× | 235.1 | 3.57× / 3.57× |
| cuda | 128 × 1024 × 5120 | bm32 | 0.153888 | 0.010803 | 14.25× | 19.2 | 11.02× / 11.03× |
| cuda | 128 × 5120 × 6144 | bm32 | 0.663552 | 0.064816 | 10.24× | 26.7 | 13.14× / 11.82× |
| cuda | 128 × 5120 × 17408 | bm32 | 1.909952 | 0.183645 | 10.40× | 26.2 | 11.56× / 11.69× |
| cuda | 128 × 6144 × 5120 | bm32 | 0.579648 | 0.064816 | 8.94× | 30.5 | 11.94× / 11.83× |
| cuda | 128 × 10240 × 5120 | bm32 | 1.078304 | 0.108026 | 9.98× | 27.3 | 10.88× / 10.80× |
| cuda | 128 × 12288 × 5120 | ld40 | 1.260576 | 0.129632 | 9.72× | 28.1 | 11.07× / 10.93× |
| cuda | 128 × 17408 × 5120 | ld40 | 1.855776 | 0.183645 | 10.11× | 27.0 | 10.56× / 10.56× |
| cuda | 512 × 1024 × 5120 | bm32 | 0.430592 | 0.010803 | 39.86× | 6.8 | 11.56× / 11.58× |
| cuda | 512 × 5120 × 6144 | bm32 | 2.287904 | 0.064816 | 35.30× | 7.7 | 11.96× / 12.10× |
| cuda | 512 × 5120 × 17408 | bm32 | 6.611136 | 0.183645 | 36.00× | 7.6 | 11.78× / 11.85× |
| cuda | 512 × 6144 × 5120 | bm32 | 2.306048 | 0.064816 | 35.58× | 7.7 | 11.82× / 11.86× |
| cuda | 512 × 10240 × 5120 | bm32 | 4.201888 | 0.108026 | 38.90× | 7.0 | 10.74× / 10.74× |
| cuda | 512 × 12288 × 5120 | ld40 | 5.092128 | 0.129632 | 39.28× | 6.9 | 10.69× / 10.65× |
| cuda | 512 × 17408 × 5120 | ld40 | 7.071872 | 0.183645 | 38.51× | 7.1 | 10.79× / 10.80× |
| cuda | 1024 × 1024 × 5120 | bm32 | 0.794752 | 0.010803 | 73.57× | 3.7 | 12.36× / 12.33× |
| cuda | 1024 × 5120 × 6144 | bm32 | 4.578400 | 0.064816 | 70.64× | 3.9 | 11.78× / 11.80× |
| cuda | 1024 × 5120 × 17408 | bm32 | 13.428288 | 0.183645 | 73.12× | 3.7 | 11.32× / 11.39× |
| cuda | 1024 × 6144 × 5120 | bm32 | 4.600864 | 0.064816 | 70.98× | 3.8 | 11.75× / 11.74× |
| cuda | 1024 × 10240 × 5120 | bm32 | 8.356288 | 0.108026 | 77.35× | 3.5 | 10.72× / 10.80× |
| cuda | 1024 × 12288 × 5120 | ld40 | 10.143008 | 0.129632 | 78.24× | 3.5 | 10.57× / 10.62× |
| cuda | 1024 × 17408 × 5120 | ld40 | 13.948032 | 0.183645 | 75.95× | 3.6 | 10.90× / 10.89× |
Exploratory history: the original short Metal packing protocol had unstable clock ordering on small shapes; all its raw logs are retained. The final uniform soak and packing protocol was applied to every variant in both confirmation sessions. An initial BM32 prototype had an incorrect cooperative tile setup and failed the oracle; its failure log is retained. Only corrected, passing BM32 results are used above. BM32 did not consistently beat the packed 64-row tile at Metal N=1024, so that shape retains the 64-row tile.
One-variable steps: half shared staging; packed-word/scale hoisting; M tile 64→32; shared leading dimension 32→40; GEMV rows per group; two rows per SIMD group; grouped activation loads; CUDA vectorization flag. Activation grouping alone and the vectorization flag alone were both flat on CUDA; grouping followed by enabling vectorization produced the retained gain. Unselected experiments and all session timings remain in the raw logs.
The initial live trace used M=128. The unchanged E2E campaign submits whole M=512/1024 projection batches at those prompt lengths; the original shape-limited selector fell back there. Both larger batches were subsequently measured in two opposite-order sessions and received separate full-output production oracles before selector admission. Unsupported batch sizes still return None.
Integration contract
Section titled “Integration contract”The Q4 change is based on dev (dc856bf4cc808a69594af2f8dd2d13940f82e478). The performance measurements and Butter integration use the same Q4 source on Iron 5b3fb4922de057af71193da1e2e50f5d4cc5ca4e, which supplies PR #299’s CUDA WMMA lowering and corrected chain runtime. This Q4 change does not edit the attention kernels, dense NAX split-K kernel, CUDA generator, vectorization/unroll passes, or chain dispatch. CUDA performance requires that prerequisite; software lowering on older dev is only a correctness fallback.
q4_projection_for(F32, Some(10 or BACKEND_CUDA), M, N, K) returns kernel, grid, TPG and rows-per-threadgroup. Supported M values are 1/128/512/1024 and the seven weight shapes above. Weights are signed Q4 with four U32 words plus one F16 scale per 32 values. X and output are F32. M=1 uses existing two-row-per-SIMD-group GEMV on M5, and grouped four-element activation loads with CUDA vectorized I/O on GB10. Prefill stages X and dequantized W into F16 shared memory and accumulates in F32. It requires finite F16-representable values and caller acceptance of rounding from F32. No scratch or extra host synchronization is needed.
The production oracle uses independently generated packed weights with row period 17, per-block F16 scales, non-dyadic activations, and seven signed row factors. It checks every output against an independent F64 sum; the row periods differ from every matrix tile. There are also ragged M/N edge oracles for each new kernel. Butter’s live-model validation compares every selected projection against the existing path on the same inputs, checks finiteness and cosine >=0.999, and reports maximum absolute error.
Reproduction
Section titled “Reproduction”Build cargo build -p wh-iron-std --example q4_projections on macOS, or add --features cuda on Linux. Always pass --backend metal or --backend cuda to the executable. Mac GPU commands must use flock -w 3600 /tmp/iron-gpu.lock on the shared host. Q4_M, Q4_N, Q4_PREFILL=baseline|packed|bm32|ld40, Q4_GEMV=x4, Q4_VECTOR=1, and Q4_RPT select a cell or variant. --sweep compares GEMV geometry; --reverse reverses its ordering. The historical half-staging-only prototype was discarded after the packed-load variant won; its source and original timings are archived with Butter’s campaign artifacts.
The tiny CUDA decode shape remains around 1.55 times its weight-bandwidth floor after the hunt. RPT=1/2/4/8 were essentially tied across two sessions; RPT=16 lost. No extra launch or reduction was selected. Metal’s larger GEMVs are close to the approximate bandwidth floor. Prefill remains above a weight-only floor because matrix arithmetic, dequantization and shared-memory staging are also required.
Legacy oracle diagnostic
Section titled “Legacy oracle diagnostic”The existing gemm_q4_mpp.rs test module is gated by cfg(test) and is absent from the standalone Iron inventory. An isolated diagnostic enabled it, aligned its declared dispatch mode, and used dtype-rounded operands with an F64 reference. On CUDA at 5b3fb492, its BF16 cases then passed, while F16 tile/edge cases still reported 0.03125/0.0625 versus their unchanged 0.03 gate. The diagnostic patch and failure logs are preserved in Butter’s campaign artifacts. The legacy kernel and its tests are unchanged by this Q4 patch. The new selector admits only F32 outputs and passes its independent production and ragged-edge oracles; this legacy F16-output discrepancy remains a follow-up.
Q4 end-to-end A/B
Section titled “Q4 end-to-end A/B”Whole eager forward plus final synchronization; checkpoint load, cache allocation/restoration, sampling and diagnostics excluded. Three warmups, then median of five runs in each of two sessions. Attention selectors remain enabled; only BUTTER_IRON_Q4_SELECTORS=0/1 changes. QWEN35_MARLIN=0 and the same checkpoint, packing, capacities and prompt tokens are held fixed. Prefill uses whole 128/512/1024-row batches in both arms. Every raw log records host uptime/load and the test-binary SHA256.
| Backend | Workload | Length | Old → Q4 tok/s, session 1 | Old → Q4 tok/s, session 2 | Speedup, sessions 1 / 2 |
|---|---|---|---|---|---|
| metal | prefill | 128 | 74.368 → 98.627 | 52.787 → 86.103 | 1.33× / 1.63× |
| metal | prefill | 512 | 101.200 → 119.321 | 97.516 → 121.885 | 1.18× / 1.25× |
| metal | prefill | 1024 | 101.453 → 145.805 | 98.214 → 146.438 | 1.44× / 1.49× |
| metal | decode | 129 | 2.192 → 2.752 | 1.516 → 2.967 | 1.26× / 1.96× |
| metal | decode | 513 | 2.196 → 2.769 | 1.408 → 2.897 | 1.26× / 2.06× |
| metal | decode | 1025 | 1.277 → 2.665 | 1.292 → 2.867 | 2.09× / 2.22× |
| cuda | prefill | 128 | 22.294 → 217.299 | 22.293 → 218.581 | 9.75× / 9.80× |
| cuda | prefill | 512 | 23.913 → 232.595 | 23.735 → 232.602 | 9.73× / 9.80× |
| cuda | prefill | 1024 | 23.971 → 228.754 | 23.955 → 229.381 | 9.54× / 9.58× |
| cuda | decode | 129 | 1.175 → 12.420 | 1.174 → 12.445 | 10.57× / 10.60× |
| cuda | decode | 513 | 1.173 → 12.394 | 1.173 → 12.430 | 10.56× / 10.60× |
| cuda | decode | 1025 | 1.173 → 12.363 | 1.173 → 12.387 | 10.54× / 10.56× |
Across the three context sizes, live validation checked 2,400 selected projections per GPU (4,800 total). Every output was finite. Minimum cosine against the existing projection on identical inputs was 0.999999945 on Metal and 0.999999944 on CUDA; maximum absolute differences were 0.0965271 and 0.131073 respectively. These are F32 projection outputs using the requested GEMM cosine gate of 0.999; the attention precision gates are unchanged.
Metal prefill512 attribution (median of five after three warmups) shows the complete projection bucket at 2700.792→1003.953 ms GPU time and 3687.859→2015.808 ms wall time. Dispatch count remains 449. The attention bucket remains about 8 ms GPU time (16 dispatches), and other work retains about 1.46–1.48 seconds of wall time (1011 dispatches). This accounts for a smaller end-to-end improvement than the resident kernel speedup. Shared-host variation is visible in both complete E2E sessions; both orders and all samples are retained.
The detailed raw timings, diagnostic patch, live comparisons and per-dispatch attribution are preserved in Butter’s benchmarks/iron-q4/ campaign, alongside the fixed-head Phase 1 report. The E2E comparisons use the same kernel source as this change over the fixed 5b3fb492 prerequisite.
