Measured on your GPU · 20 / 20 verified on both APIs

CUDA ↔ WebGPU

RTX 5080 · CUDA 13.3 · NVIDIA driver 616.64 · 2026-09-11 UTC
Ten identical kernel sources, two sizes, real GPU execution. Median microseconds per operation. Small text shows the observed sample range.

WebGPU / CUDA above 1 means CUDA graph replay was faster. CUDA graph timing and WebGPU compute-pass timing both batch GPU-resident operations. Ordinary CUDA launches expose submission gaps. WebGPU wall time also includes encoding and completion waiting. Tiny differences can be noise.

KernelWebGPU GPU µsCUDA graph GPU µsWebGPU / CUDACUDA ordinary GPU µsWebGPU wall µs
saxpysmall1.631.63–1.791.491.49–1.801.09×
6.663.31
saxpy_vec4small1.441.41–1.601.401.27–1.721.03×
6.462.96
matmul_naivesmall19.9719.33–20.356.626.62–6.963.01×
10.2524.87
matmul_tiledsmall3.203.20–3.332.882.88–3.011.11×
7.005.52
matmul_registersmall5.895.76–6.533.063.06–3.181.92×
8.0011.11
reduce_sumsmall3.683.68–3.713.553.39–3.591.04×
18.637.78
convolutionsmall2.052.05–2.111.811.67–1.981.13×
6.343.10
histogramsmall4.104.10–4.293.823.82–4.061.07×
13.907.10
transposesmall1.761.73–1.951.211.21–1.381.46×
6.733.07
particlessmall1.441.41–1.541.741.74–1.820.83×
6.743.31
saxpylarge16.2616.16–16.2916.1616.06–16.331.01×
18.8817.51
saxpy_vec4large10.029.95–10.0510.079.97–10.320.99×
12.6711.25
matmul_naivelarge231.42230.53–232.06107.18106.95–108.532.16×
109.64235.83
matmul_tiledlarge64.2663.62–64.7763.0862.21–63.341.02×
65.2066.86
matmul_registerlarge46.2145.95–47.4943.0842.92–44.101.07×
45.8249.00
reduce_sumlarge20.6120.48–20.7718.7118.64–18.861.10×
25.3625.57
convolutionlarge21.0920.90–21.2516.3916.34–16.551.29×
18.9922.23
histogramlarge24.7024.64–24.9919.8619.73–19.951.24×
23.3627.85
transposelarge13.4113.31–13.609.169.07–9.201.46×
11.1615.04
particleslarge4.514.51–4.744.934.72–5.030.91×
7.686.22

Nine samples; 2,048 repeats per sample, 512 for matrices. Full reduction chains and histogram clearing are included. Particle compute excludes rendering. Inputs reset outside each sample; buffers are reused inside the sample. Float32 reference checks pass; integer histograms match exactly. This measures these kernels, not optimized CUDA libraries or cold transfers.

Full methodology and reproduction commands · Combined raw results · Open live particle app