GPU compute POC: compute shader → Texture2DRD → canvas shader #84

Closed
claude wants to merge 12 commits from feature/gpu-compute-poc into master
Collaborator

Summary

  • Adds a standalone proof-of-concept scene in tools/gpu_poc/ for the GPU compute ball rendering pipeline
  • ball_sim.glsl: compute shader animates 2000 balls with Lissajous motion, writes pos/radius/alive + color into a 2×N RGBA32F texture via imageStore
  • ball_render.gdshader: canvas_item shader reads ball data via texelFetch(ball_data, ivec2(0/1, INSTANCE_ID), 0), overrides VERTEX and COLOR per instance, clips to circle in fragment
  • gpu_poc.gd: full RenderingDevice setup — compute pipeline, Texture2DRD bridge, params SSBO, MultiMesh with identity transforms, per-frame buffer_update + dispatch + barrier(); gracefully no-ops when _rd == null (headless/no GPU)
  • gpu_poc.tscn: minimal standalone Node2D scene to run the POC

Test plan

  • Open tools/gpu_poc/gpu_poc.tscn in the Godot editor and run — should show 2000 animated colored circles moving in Lissajous patterns
  • Bump BALL_COUNT to 10k, 50k, 100k to benchmark rendering throughput vs current GDScript multimesh approach
## Summary - Adds a standalone proof-of-concept scene in `tools/gpu_poc/` for the GPU compute ball rendering pipeline - `ball_sim.glsl`: compute shader animates 2000 balls with Lissajous motion, writes pos/radius/alive + color into a 2×N RGBA32F texture via `imageStore` - `ball_render.gdshader`: canvas_item shader reads ball data via `texelFetch(ball_data, ivec2(0/1, INSTANCE_ID), 0)`, overrides `VERTEX` and `COLOR` per instance, clips to circle in fragment - `gpu_poc.gd`: full RenderingDevice setup — compute pipeline, `Texture2DRD` bridge, params SSBO, MultiMesh with identity transforms, per-frame `buffer_update` + dispatch + `barrier()`; gracefully no-ops when `_rd == null` (headless/no GPU) - `gpu_poc.tscn`: minimal standalone Node2D scene to run the POC ## Test plan - [ ] Open `tools/gpu_poc/gpu_poc.tscn` in the Godot editor and run — should show 2000 animated colored circles moving in Lissajous patterns - [ ] Bump `BALL_COUNT` to 10k, 50k, 100k to benchmark rendering throughput vs current GDScript multimesh approach
Proof-of-concept for rendering large ball counts via GPU pipeline:
- ball_sim.glsl: compute shader animates 2000 balls with Lissajous
  motion, writes pos/radius/alive and color into a 2×N RGBA32F texture
  via imageStore
- ball_render.gdshader: canvas_item shader reads ball_data texture via
  texelFetch(ball_data, ivec2(0/1, INSTANCE_ID), 0), overrides VERTEX
  position and COLOR in the vertex stage, clips to circle in fragment
- gpu_poc.gd: RenderingDevice setup (compute pipeline, Texture2DRD
  bridge, params SSBO, MultiMesh with identity transforms), per-frame
  buffer_update + compute dispatch + barrier(); gracefully skips setup
  when _rd == null (headless)
- gpu_poc.tscn: minimal standalone scene

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Remove #[compute] pragma from ball_sim.glsl — only valid when loaded
  via Godot's import system (RDShaderFile), not raw FileAccess text reads
- Replace early return in ball_render.gdshader vertex() with if/else —
  canvas_item vertex functions do not support return statements

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- canvas_item shaders write COLOR directly; ALBEDO/ALPHA are spatial-only
- _rd.barrier() is deprecated in Godot 4.6, RenderingDevice adds barriers
  automatically between compute and rendering passes

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replaces the Lissajous animation POC with a real simulation:

Compute shader (ball_sim.glsl):
- Reads/writes a flat ball state SSBO (pos, vel, radius, alive per ball)
- Each frame: checks explosion events, integrates velocity, bounces off
  arena walls, writes updated state back + render texture
- Colors balls by velocity direction so motion is visible

GDScript (gpu_poc.gd):
- Initializes BALL_COUNT (10000) balls at random positions/velocities
- Ball state SSBO, explosion events SSBO (max 16 per frame), params SSBO
- Timer fires a random explosion every 1.5s; events are uploaded and
  cleared each frame so the SSBO slot is reused
- BALL_COUNT const at the top to tune for benchmarking

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace 2×BALL_COUNT texture (exceeded device max ~8192) with a
TEX_ROW_SIZE×H texture where H = 2*ceil(BALL_COUNT/TEX_ROW_SIZE):

  ball idx → col = idx % 4096,  block = idx / 4096
  row block*2+0 : pos.x, pos.y, radius, alive
  row block*2+1 : color rgba

At 10k balls: 4096×6. At 500k balls: 4096×246.
Updated imageStore in ball_sim.glsl, texelFetch in ball_render.gdshader,
and texture_create dimensions in gpu_poc.gd.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Compute shader (ball_sim.glsl):
- Two new SSBOs: RockData (center.x, center.y, radius per rock) and
  RockParams (count)
- Rock bounce: push-out overlap correction + velocity reflection off
  surface normal; only applies when dot(v, normal) < 0 to avoid sticking
- Color by ball type (inferred from radius match to 4 type radii) using
  the game's actual healthy_color values: white / light-blue / purple /
  salmon for standard / small / big / chain

GDScript (gpu_poc.gd):
- BALL_TYPES const with radius, speed range, and spawn weight per type;
  weighted-random assignment at init
- 8 rocks placed in a ring, sizes cycling 1x/1.75x/2.5x base radius 64
- Rock SSBOs initialized once and never updated (static obstacles)
- _draw_rocks(): fan-triangulated MeshInstance2D circles at rock positions
  so they're visible; z_index=-1 so balls render on top

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
CPU arms random balls; GPU tracks fuse and reports explosions back:

Compute shader (ball_sim.glsl):
- Ball state extended to 7 floats (adds fuse_time at index 6)
- New DetonationOut SSBO at binding 6: int count (atomic) + vec4 events[]
- When fuse_time expires: atomicAdd claims a slot, writes blast center +
  radius (br*2+40), marks ball dead
- Armed balls (fuse_time > 0 at frame start) render red

GDScript (gpu_poc.gd):
- BALL_STRIDE 6 → 7; ball init sets fuse_time = 0
- _arm_random_ball(): patches only the 4-byte fuse_time field of a random
  ball via buffer_update at the exact SSBO byte offset — no full re-upload
- _process(): reads DetonationOut via buffer_get_data (GPU stall / readback
  cost measurement), parses count + events, converts to _pending_explosions
  for the next frame's explosion SSBO upload, then resets count to 0 before
  the next dispatch
- Explosion events SSBO is now driven entirely by detonation readback
  (no more CPU-side random position explosions)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add use_custom_data=true to the MultiMesh. Each instance stores its ball
index in INSTANCE_CUSTOM.r so the vertex shader can look up the correct
texture row independently of instance slot order.

CPU maintains two parallel PackedInt32Arrays (_instance_to_ball,
_ball_to_instance). When a ball is armed, it is swapped into the last
available normal slot (_armed_start-1 and decrements _armed_start).
Armed instances occupy [_armed_start, BALL_COUNT), which renders on top
because MultiMesh uses painter's order (later = in front).

_arm_random_ball now picks ball_idx from _instance_to_ball[randi()%_armed_start]
so it always picks from the normal (non-armed) region, then does the swap
with 2 set_instance_custom_data calls and updates both index arrays.

Vertex shader: replaces INSTANCE_ID lookup with int(round(INSTANCE_CUSTOM.r)).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fuse countdown disabled — armed balls accumulate indefinitely.
Each armed ball checks against the CPU-uploaded armed list for overlap;
if any two armed balls touch, both detonate via atomicAdd.

Compute shader (ball_sim.glsl):
- New ArmedList SSBO at binding 7: float[0]=count, float[1..N]=ball indices
- After physics, each armed ball iterates the armed list, skips self and
  dead balls, checks circle-circle overlap; on hit calls detonate() which
  atomicAdds a det slot and writes (px,py,blast_r,ball_idx) as vec4
- detonate() extracted to a helper function for clarity
- Fuse countdown block present but commented out in spirit (replaced by
  intersection check as the sole detonation trigger)

GDScript (gpu_poc.gd):
- MAX_ARMED = 256; _armed_list_buffer created at startup
- _upload_armed_list(): builds [count, idx0..N] from _instance_to_ball
  [_armed_start..BALL_COUNT) and uploads to SSBO each frame (tiny buffer)
- _arm_random_ball(): sets fuse_time=1.0 (armed marker, no countdown)
- det event .w now carries ball_idx; _remove_from_armed_region() swaps
  the detonated ball back into the normal region and increments _armed_start
- Armed list buffer freed in _exit_tree()

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Owner

Successful test

Successful test
cammymoop closed this pull request 2026-03-06 10:28:46 -06:00

Pull request closed

Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cammymoop/semi-vibed-game!84
No description provided.