Measured on your GPU · 20 / 20 verified on both APIs
CUDA ↔ WebGPU
RTX 5080 · CUDA 13.3 · NVIDIA driver 616.64 · 2026-09-11 UTC
Ten identical kernel sources, two sizes, real GPU execution. Median microseconds per operation. Small text shows the observed sample range.
WebGPU / CUDA above 1 means CUDA graph replay was faster. CUDA graph timing and WebGPU compute-pass timing both batch GPU-resident operations. Ordinary CUDA launches expose submission gaps. WebGPU wall time also includes encoding and completion waiting. Tiny differences can be noise.
Nine samples; 2,048 repeats per sample, 512 for matrices. Full reduction chains and histogram clearing are included. Particle compute excludes rendering. Inputs reset outside each sample; buffers are reused inside the sample. Float32 reference checks pass; integer histograms match exactly. This measures these kernels, not optimized CUDA libraries or cold transfers.
Full methodology and reproduction commands · Combined raw results · Open live particle app