CW CUDA → WEB SHADER
THE SHOWCASE EXPLORER

CUDA ideas.
Running in your browser.

Explore live GPU experiments, inspect the CUDA behind them, and compare the results with native execution on an RTX 5080.

Illustration of a collapsing blue SPH water columnEXPERIMENTAL CUDA PORT
3D / SPH WATER BSD-3-Clause kernels

Chrono water dam-break

Run Project Chrono's original SPH kernels on 16,731 fluid particles. Inspect all 15 CUDA-to-WGSL passes, with neighbours rebuilt every step. Currently slower than real time.

Original CUDA physicsNative comparisonsPressure colours
Native CUDA quadtree partitions and 1024 points rendered directly from GPU buffersOFFICIAL NVIDIA KERNEL
2D / RECURSIVE SPATIAL PARTITION BSD-3-Clause kernel

NVIDIA recursive quadtree

The original CUDA kernel splits points into quadrants and recursively launches its children. Explore the complete GPU-built tree, verified against native CUDA.

1,024 points · 193 nodesExact native outputGPU recursive launches
256 coloured Bezier curves rendered from actual GPU-allocated verticesOFFICIAL NVIDIA KERNELS
2D / DYNAMIC TESSELLATION BSD-3-Clause kernels

NVIDIA Bezier curves

The original CUDA parent chooses curve detail, allocates vertices and launches child kernels. GPU-generated queues drive the children; zoom into their output in the sandbox.

256 curves · 3,958 verticesNative CUDA comparedGPU allocation + child launches
Actual GPU path trace of glass, metal and diffuse spheresCUDA PATH TRACER
3D / MONTE CARLO RENDERING Public-domain CUDA source

Glass, metal and 488 spheres

Roger Allen’s final Ray Tracing in One Weekend scene. Original CUDA device functions compile into five WGSL passes, with persistent objects, material dispatch and seeded random rays.

1200 × 800 pixels10 samples per pixelNative CUDA compared
Five offset texture planesOFFICIAL NVIDIA KERNEL
46 / LAYERED TEXTURE BSD-3-Clause kernel

NVIDIA layered texture

Sample five full-resolution scalar texture layers with the original CUDA kernel. Compare every output with native CUDA and inspect the generated WGSL.

1,310,720 native matchesFive 512 × 512 layers
A cube with three visible texture facesOFFICIAL NVIDIA KERNEL
45 / CUBEMAP TEXTURE BSD-3-Clause kernel

NVIDIA cubemap texture

Sample all six faces through 3D directions with the original CUDA kernel. Inspect face orientation and native-verified edge filtering.

24,576 native matchesSix texture faces
Walsh transform butterfly connections and alternating signal blocksOFFICIAL NVIDIA KERNELS
44 / WALSH CONVOLUTION BSD-3-Clause kernels

NVIDIA complete Walsh convolution

Transform 8.4 million values, modulate with the original convolution kernel, then transform back. Inspect the complete native-verified output.

8,388,608 valuesThree full transforms
Branches of a binomial tree beside a grid of option valuesOFFICIAL NVIDIA KERNEL
43 / BINOMIAL TREE BSD-3-Clause kernel

NVIDIA binomial options

Price 1,024 options through 2,048 time steps with the original kernel. Inspect each result in the heatmap and compare CUDA with its generated WGSL.

1,024 native comparisonsOriginal input data
Alternating comparator connections in a sorting networkOFFICIAL NVIDIA KERNELS
42 / SORTING NETWORK BSD-3-Clause kernels

NVIDIA odd-even merge sort

Sort a million key/value pairs using the original shared-memory stage and 155 in-place global merge passes. Compare the input and final ordering.

1,048,576 pairs30 native-verified cases
Unordered bars becoming a sorted sequenceOFFICIAL NVIDIA KERNELS
41 / SORTING NETWORK BSD-3-Clause kernels

NVIDIA complete bitonic sort

Sort a million key/value pairs through the original shared-memory and in-place global merge stages. Compare the original and sorted keys side by side.

1,048,576 pairs30 native-verified cases
Multiple coordinate planes projected into a point cubeOFFICIAL NVIDIA KERNEL
40 / HIGH-DIMENSIONAL SEQUENCE BSD-3-Clause kernel

NVIDIA Sobol projections

Compute ten million values across 100 dimensions. Explore a 3D projection of 100,000 vectors and choose other coordinate planes in the launch settings.

100 dimensionsNative bit-for-bit match
A cube filled with evenly distributed pointsOFFICIAL NVIDIA KERNEL
39 / 3D POINT DISTRIBUTION BSD-3-Clause kernel

NVIDIA quasirandom cube

Generate over a million three-dimensional points with NVIDIA’s Niederreiter sequence. Orbit the complete GPU-generated distribution and compare CUDA with WGSL.

1,048,576 pointsNative bit-for-bit match
Image pixels grouped into compressed four-by-four blocksOFFICIAL NVIDIA KERNELS
38 / TEXTURE COMPRESSION BSD-3-Clause kernel

NVIDIA DXT compression

Compress the original teapot image into DXT1 blocks, then decode the GPU result for display. Inspect the complete compressor beside its generated WGSL.

16,384 blocksNative byte-for-byte match
Separate transform stagesOFFICIAL NVIDIA KERNELS
36 / FFT IMAGE FILTER BSD-3-Clause kernels

NVIDIA custom FFT convolution

Original real/complex preprocessing and frequency multiplication, with four million input values and all GPU stages exposed.

2000 × 2000 inputNative CUDA verified
Combined frequency processingOFFICIAL NVIDIA KERNELS
37 / FFT IMAGE FILTER BSD-3-Clause kernels

NVIDIA fused FFT convolution

Combines spectrum conversion, multiplication and inverse preparation in NVIDIA’s original fused processing kernel.

2000 × 2000 inputNative CUDA verified
An image grid transformed into a frequency spectrum and filtered outputOFFICIAL NVIDIA KERNELS
35 / FFT IMAGE FILTER BSD-3-Clause kernels

NVIDIA FFT convolution

Filter four million input values through texture padding, real FFTs and frequency multiplication. Original CUDA kernels with native-verified output.

2000 × 2000 input2048 × 2048 FFT
Coloured arrows revealing motion across an imageOFFICIAL NVIDIA KERNELS
34 / MOTION ESTIMATION BSD-3-Clause kernels

NVIDIA optical flow

Reveal motion between two camera frames. Six original CUDA kernels build an image pyramid, warp the image and solve the displacement field.

614,400 native-exact values5 pyramid levels7,500 solver iterations
Two camera images combining into layers of disparityOFFICIAL NVIDIA KERNEL
33 / STEREO VISION BSD-3-Clause kernel

NVIDIA stereo disparity

Recover image shifts from a stereo camera pair. Original CUDA block matching uses packed colour textures, shared memory and SIMD byte differences.

341,120 native-exact pixels17 disparity candidatesDepth-related image
Green particle streams curling through a fluid fieldOFFICIAL NVIDIA KERNELS
32 / INTERACTIVE FLUID BSD-3-Clause kernels

NVIDIA fluid flow

Drag to stir 262,144 particles. Original CUDA advection, forces and diffusion run with GPU FFT projection in a 512 × 512 fluid field.

64 native steps comparedDrag to stirGPU FFT solver
A softly lit smoke cloud with layered shadowsOFFICIAL NVIDIA KERNELS
31 / VOLUMETRIC PARTICLES BSD-3-Clause kernels

NVIDIA smoke particles

Orbit a cloud of 16,384 particles. Original CUDA integration and depth calculation feed GPU sorting and a 32-slice smoke renderer with volumetric shadows.

64 native steps comparedVolumetric shadowsGPU feedback
A three-dimensional volume with green and warm-coloured transfer layersOFFICIAL NVIDIA KERNELS
30 / VOLUME RAY MARCHING BSD-3-Clause kernels

NVIDIA pre-integrated volume

See two colour and opacity tables reveal the original Bucky volume. Integration, layered tables and ray marching run from unchanged CUDA functions.

Two 1024² transfer layers6 native image comparisonsCUDA / WGSL comparison
Spheres falling and colliding above a simulation gridOFFICIAL NVIDIA KERNELS
29 / PARTICLE PHYSICS BSD-3-Clause kernels

NVIDIA particle collisions

Watch 1,024 spheres fall, collide and settle. The original integration and collision code runs with GPU sorting and shared render positions.

64 native steps comparedGPU spatial gridAnimated 3D spheres
Waves crossing a three-dimensional ocean surfaceOFFICIAL NVIDIA KERNELS
28 / OCEAN SIMULATION BSD-3-Clause kernels

NVIDIA FFT ocean

Watch a wind-driven wave spectrum become an animated 3D surface. Spectrum, inverse FFT, heights and slopes run on the GPU.

256 × 256 ocean130,050 trianglesNative cuFFT compared
Eight image blocks processed together in a shared-memory tileOFFICIAL NVIDIA KERNELS
27 / SHARED IMAGE TRANSFORM BSD-3-Clause kernels

NVIDIA optimized DCT

Run NVIDIA’s optimized floating-point transform. Threads share a 32 × 16 image tile and process rows and columns before reconstructing the teapot.

512 × 512 imageNative CUDA matchedShared-memory helpers
An image tile transformed into frequency coefficients and reconstructedOFFICIAL NVIDIA KERNELS
26 / IMAGE TRANSFORM BSD-3-Clause kernels

NVIDIA DCT image

Transform the teapot into 8 × 8 frequency blocks, quantize the coefficients and reconstruct the image. Compare the original CUDA kernels with generated WGSL.

512 × 512 imageNative CUDA matchedGPU intermediate copies
Noisy pixels filtered into a smoother portraitOFFICIAL NVIDIA KERNEL
23 / IMAGE DENOISING BSD-3-Clause kernel

NVIDIA KNN denoising

Smooth image noise using colour-distance weights in a local window. Run NVIDIA’s original noisy portrait and compare CUDA with WGSL.

320 × 408 portraitNative CUDA checkedEditable filter settings
Noisy pixels filtered into a smoother portraitOFFICIAL NVIDIA KERNEL
24 / IMAGE DENOISING BSD-3-Clause kernel

NVIDIA non-local means

Compare neighbourhood patches to preserve image detail while reducing noise. Run NVIDIA’s original noisy portrait and compare CUDA with WGSL.

320 × 408 portraitNative CUDA checkedEditable filter settings
Noisy pixels filtered into a smoother portraitOFFICIAL NVIDIA KERNEL
25 / IMAGE DENOISING BSD-3-Clause kernel

NVIDIA shared non-local means

Reuse patch weights across a shared-memory tile using the original optimized kernel. Run NVIDIA’s original noisy portrait and compare CUDA with WGSL.

320 × 408 portraitNative CUDA checkedEditable filter settings
Sobel edges computed through a shared image tileOFFICIAL NVIDIA KERNEL
22 / SHARED IMAGE TILE BSD-3-Clause kernel

NVIDIA Sobel shared tile

Run the shared-memory version of Sobel edge detection. CUDA lanes load a pixel tile together and write four packed output pixels per record.

1024² grayscale image4 native image comparisonsOriginal shared kernel
A teapot traced by the edges of a Sobel filterOFFICIAL NVIDIA KERNELS
21 / EDGE DETECTION BSD-3-Clause kernels

NVIDIA Sobel edges

Find edges in the original teapot image. The unchanged CUDA kernel reads a 3 × 3 neighbourhood and writes each pixel’s edge intensity.

1024² grayscale image8 native image comparisonsEditable edge scale
A sliding filter window softening rows and columnsOFFICIAL NVIDIA KERNELS
20 / COLOUR BLUR BSD-3-Clause kernels

NVIDIA box filter

Blur the original teapot image with a sliding window. The CUDA row and column passes share the intermediate image on the GPU.

1024² colour image8 native image comparisonsEditable radius
Faceted molecular shell reconstructed from a volumeOFFICIAL NVIDIA KERNELS
19 / VOLUME TO MESH BSD-3-Clause kernels

NVIDIA Bucky mesh

Extract a triangle surface from NVIDIA’s original Bucky volume. Shared vertex pointers and surface normals stay on the GPU.

11,126 triangles3 native mesh comparisonsOriginal 32³ volume
Triangulated implicit surface inside a voxel latticeOFFICIAL NVIDIA KERNELS
18 / 3D MESH BSD-3-Clause kernels

NVIDIA marching cubes

Build a mesh from the original implicit field. Classification, scans, compaction and triangle generation run on the GPU.

2,080 trianglesNative CUDA verifiedOrbit the GPU mesh
Layered volume with a highlighted cross-sectionOFFICIAL NVIDIA KERNEL
17 / 3D VOLUME FILTER BSD-3-Clause kernel

NVIDIA volume filter

Filter the original Bucky volume and explore its slices. Compare voxel and upstream texture coordinates using the unchanged CUDA kernel.

32³ voxels16 native volume comparisonsInteractive Z slices
Colour bands surrounding the Mandelbrot cardioidOFFICIAL NVIDIA KERNEL
16 / FRACTAL RENDERING BSD-3-Clause kernel

NVIDIA Mandelbrot

Render Mandelbrot and Julia sets with NVIDIA’s original float kernel. Edit the viewport and colours, accumulate frames and compare CUDA with generated WGSL.

Six native frame comparisonsMandelbrot + JuliaEditable colours
Noisy image smoothed while keeping its angular edgesOFFICIAL NVIDIA KERNEL
15 / EDGE-PRESERVING FILTER BSD-3-Clause kernel

NVIDIA bilateral filter

Smooth the original photograph while preserving colour boundaries. Run NVIDIA’s filter, edit its colour-distance setting and compare CUDA with generated WGSL.

Native CUDA compared640 × 480 photographEditable smoothing
A bright pixel spreading through a circular blur neighbourhoodOFFICIAL NVIDIA KERNEL
14 / IMAGE POST-PROCESSING BSD-3-Clause kernel

NVIDIA post-process glow

Blur the original teapot image and amplify its highlights with NVIDIA’s CUDA kernel. Inspect RGBA texture sampling, shared-memory tiles and the generated WGSL.

Five native image comparisons512 × 512 colour imageEditable highlights
Pixel steps becoming a smooth interpolated curveOFFICIAL NVIDIA KERNELS
13 / IMAGE INTERPOLATION BSD-3-Clause kernels

NVIDIA bicubic filtering

Compare nearest, bilinear, bicubic, fast bicubic and Catmull–Rom sampling on the teapot image. Original CUDA filters, editable zoom and pan, and packed colour output.

Five filter modes15 native image comparisonsCUDA / WGSL comparison
Horizontal and vertical filter passes across an image gridOFFICIAL NVIDIA KERNELS
12 / TEXTURE CONVOLUTION BSD-3-Clause kernels

NVIDIA texture convolution

Filter a teapot image with the original row and column kernels. Edit the 17 coefficients and compare both generated shaders; the intermediate image stays on the GPU.

Two original kernelsNative CUDA compared512 × 512 image
Illustration of a mint and blue wireframe sine waveOFFICIAL NVIDIA SAMPLE
01 / GRAPHICS INTEROP BSD-3-Clause kernel

NVIDIA sine wave

The original simpleGL vertex kernel, unchanged. Orbit a moving surface of GPU-generated points.

12 native CUDA checks12 WebGPU checks
Three-dimensional gravitational particle clusterOFFICIAL NVIDIA KERNEL
03 / 3D GRAVITY BSD-3-Clause kernel

NVIDIA N-body

The original tiled gravitational force and integration kernels. Watch 512 bodies evolve in 3D with position feedback entirely on the GPU.

Native CUDA verifiedNVIDIA WebGPU verifiedCUDA / WGSL comparison
Three-dimensional lattice of field samplesOFFICIAL NVIDIA KERNEL
04 / VOLUME SIMULATION BSD-3-Clause kernel

NVIDIA FDTD3d

A tiled 3D finite-difference stencil with configurable coefficients. Explore 8,192 interior field samples, colored directly from the GPU output.

Native CUDA verifiedNVIDIA WebGPU verified3D scalar field
An input buffer written to a surface and sampled into an imageOFFICIAL NVIDIA KERNELS
11 / SURFACE AND TEXTURE BSD-3-Clause kernels

NVIDIA surface write

Write the teapot image into a GPU surface, then sample it with the original rotation kernel. Both passes share the image entirely on the GPU.

Two original kernelsNative CUDA comparedGPU surface writes
A teapot image rotated inside a texture gridOFFICIAL NVIDIA KERNEL
10 / 2D TEXTURE SAMPLING BSD-3-Clause kernel

NVIDIA texture rotation

Rotate the original teapot image with NVIDIA’s unchanged texture kernel. Edit the angle and compare nearest or linear sampling in the sandbox.

512 × 512 imageNative CUDA comparedFloat GPU texture
Coloured volume viewed through a ray-marching gridOFFICIAL NVIDIA KERNEL
09 / VOLUME RAY MARCHING BSD-3-Clause kernel

NVIDIA volume renderer

Ray-march the original Bucky volume using NVIDIA’s unchanged device code. Edit the camera, density and colour transfer table, then compare CUDA with generated WGSL.

Full ray-marching kernelNative image comparison32³ volume
A slice through a three-dimensional volumeOFFICIAL NVIDIA KERNEL
08 / 3D TEXTURE SAMPLING BSD-3-Clause kernel

NVIDIA 3D texture slice

Sample the original Bucky volume with a real 3D GPU texture. Choose a depth and compare point or linear filtering with wrapped coordinates.

32³ volumeNative CUDA verified3D GPU sampler
A colour image blurred with a recursive Gaussian filterOFFICIAL NVIDIA KERNELS
07 / RECURSIVE IMAGE FILTER BSD-3-Clause kernels

NVIDIA recursive Gaussian

Filter a colour image with the original recursive kernel and tiled transpose. Four GPU passes produce an RGBA preview with editable coefficients.

Four GPU passesNative CUDA verifiedRGBA image preview
A signal decomposed into coarse and detailed wavelet coefficientsOFFICIAL NVIDIA KERNEL
06 / SIGNAL TRANSFORM BSD-3-Clause kernel

NVIDIA Haar wavelet

Decompose 4,096 signal values into coarse and detailed coefficients. Two original-kernel launches exchange coefficients directly on the GPU.

Full transformNative CUDA verifiedTwo GPU passes
Row and column image filtering stagesOFFICIAL NVIDIA KERNELS
05 / IMAGE FILTERING BSD-3-Clause kernels

NVIDIA separable convolution

The original row and column kernels run in sequence. Keep the intermediate image on the GPU and compare the generated shader for each pass.

Two GPU passesNative CUDA verifiedEditable coefficients
Illustration of gold and cyan particles in an orbital fieldPROJECT SHOWCASE
02 / PARTICLE SIMULATION MIT

Orbital field

A GPU-resident particle simulation and live kernel workbench. Edit CUDA, inspect generated WGSL, and test the actual GPU.

Editable particle countShared render buffer
SOURCE → SHADER → GPU

Real kernel code. Visible results.

The browser translates a supported subset of CUDA C into WGSL. Native comparisons use NVCC on an NVIDIA GPU. Each showcase explains its source, validation, and any host-side adaptation.