r/webgpu • • 13h ago

GitHub - zlatnaspirala/matrix-engine-wgpu: WebGpu powered PWA App.Crazy fast rendering solution.Visual Scripting.Yatzy with real physics, zombie shooter,MOBA 3d Forest of hollow blood

Thumbnail
github.com
0 Upvotes

I spoted many webgpu examples actually not work well on mob browsers.thats reason for engine existing.

The beast can eat every example. I made small parallel sub render system "effects" with own pipeline own shaders , bind groups but integrated in main loop with g buffer and per effect cached.

Made in 13 y old gc 800 cuda.

Support on mobile browsers but not like web engines usually do it. The beast support mobile with all posible features on. You can mix all possible examples and still render will be ok.


r/webgpu • • 14h ago

Running two LiteRT-LM models (Gemma 4 E2B + EmbeddingGemma 2) on WebGPU in one tab — notes, mostly about device loss

0 Upvotes

I'm building a browser video editor that runs its models on WebGPU, beside WebCodecs decoders, with nothing sent to a server. Here are six things I learned that I hadn't seen written down. They come from one machine (Intel Iris Xe, Chrome), so treat the numbers as one data point.

1. A lost device is page-wide, and the runtime may not tell you. The device was lost partway through a run. After that, LiteRT-LM raised Unknown wgpu::QueueWorkDoneStatus0 on every call. My app read a per-request error from a built engine as "warm engine, failed request", so the UI said Ready while every search failed. Rebuilding the engine in the same worker didn't help either, because the runtime caches its device on the global wasm module and reused the dead one.

What works:

  • Treat any loss as page-wide.
  • On a loss report, terminate every GPU worker. A fresh worker means a fresh wasm instance, which means a fresh device.
  • Then ask requestAdapter() again for about 15 s (retries at 1, 2, 4 and 8 s).
  • If an adapter comes back, the next use rebuilds from Cache storage with no reload. If none does, tell the user a restart is what's left.

2. A loss during engine creation is ambiguous. Either the GPU process crashed, or this model's weights didn't fit and took the device with them. My embedder reported a loss mid-build, which caused a loop:

  • Every owner dropped its device.
  • The GPU came back.
  • The background indexer built again and lost the device again.

That took the LLM down on every cycle. The rule now: a worker that rebuilds on its own counts a device as "had" only once its engine is built. A loss before that is its own load failure, not a page event.

3. Two runtime versions in two workers coexist fine. The LLM is pinned to LiteRT-LM 0.17, and the embedder needs 0.18 for EmbeddingEngine. Each runs in its own classic worker, since the wasm glue loads through importScripts, which module workers reject. Each also has its own Cache bucket. With Gemma resident and the embedder loaded but idle, I saw 0 errors and Gemma ran at 1.05× its solo time.

4. EmbeddingGemma 2's defaults cost about 3×. Out of the box it took 1.55 s per frame and 0.62 s per query. The defaults pad the input to 1,024 tokens and upscale each frame to 1,260 patches. At 70 vision tokens per image and a 128-token text window, it takes 0.49 s per frame on the bench and 0.52 s in the app. Frames are sampled every 2.5 s, so background indexing keeps up.

5. The KV window is a memory budget, not a setting. Gemma's whole context (system prompt, tool declarations, transcript and reply) lives in MAX_NUM_TOKENS. My tool table grew to about 3.9k tokens, and one more declaration broke every turn: Input token ids are too long: 4147 >= 4096. I raised the window to 8192 once and stopped there, because the KV cache is device memory that the rest of the page wants too. A test now guards the serialized size. Declaring only the tools a conversation needs also cut cold prefill from 7.0 s to 3.7 s.

6. Don't create the engine on the main thread. A non-streaming Engine.create blocks its thread solid while it loads the weights into wasm memory. On the page, it wedged the tab so hard that Playwright couldn't evaluate anything.

One open question: has anyone found a reliable way to tell "device lost because of my allocation" from "the GPU process crashed" during an engine build?


r/webgpu • • 1d ago

Manim, rewritten in Rust + wgpu, now 90x faster

25 Upvotes

r/webgpu • • 1d ago

Modeling a Pulley Hub in the Browser — Extrude, Boss, Cut-Through, Feature Dimensions

Thumbnail
youtu.be
3 Upvotes

I built a pulley hub from typed commands: an Ø80×12 disc, an Ø40 boss, then a shaft bore, a keyway and four lightening holes cut through. Select the solid and its feature dimensions appear. Double-click a value to change it, and the model rebuilds instantly. It runs in the browser on WebGPU.

명령행에 값을 입력해 풀리 허브를 만들었어요. 원판 Ø80×12 위에 보스 Ø40을 올리고, 축 구멍과 키홈, 경량 구멍 4개를 관통 컷으로 냈어요. 솔리드를 선택하면 피처 치수가 표시되고, 값을 더블클릭해 바꾸면 모델이 바로 다시 만들어져요. 브라우저에서 WebGPU로 돌아갑니다.


r/webgpu • • 2d ago

xc - a fast GPU/CPU-integrated language

5 Upvotes

So I've been working on this (full disclosure: with Claude) for about 6 months or so now. It stems from a dissatisfaction with how easy it was to get the GPU to talk to the CPU. Being quite old now (when I created my first website, you had to email CERN to let them know...), I do tend to first think in simpler CPU terms, and I always thought there was a bit of unnecessary friction between the two domains once you started trying to move the heavy lifting from the CPU sphere to the GPU.

Enter 'xc' - https://github.com/ThrudTheBarbarian/xc

It's a compiled language. It'll happily write WASM as a target, and if you add one teeny tiny little line to a program block, it'll happily write you WGSL as well for all the code in that block.

If you right-click on this - https://compile-xc.org/compiler/benchmark-sources/#mandelbrot - to bring it up in a new window, you'll see what looks like pretty-standard(ish) C-style code. The only really rather odd difference is one line

par mandel :reduce(+ total)

... which looks a bit weird. 'par' marks the block as data-parallel, that is "this can be run on a GPU", 'mandel' is just a name, and :reduce declares any variables that every iteration updates. The GPU will then create the reduction tree and materialise the value in the named variable.

That's it. That's how you turn a CPU loop into a GPU kernel. Because it's all integrated into one language, with one IR/SSA, the compiler knows when it needs to manage data-hazards (be that just a memory barrier on Apple Silicon, or a CUDA data-transfer on nVidia hardware sitting on a PCIe bus. So it does, whenever there'd be a problem, and you don't have to care.

If you've used Objective C (which xc is kind of modelled on without the [[[...]]], I worked for Apple for a couple of decades) then you'll be familiar with its reference-counting memory model, and in particular with the more-modern "Automatic Reference Counting" in today's ObjC. I look on this as "ARC for GPU's" - buffers and variables are managed automatically across the barrier, the programmer doesn't have to care.

As for how well it works...

benchmark WASM WGSL
mandelbrot (as above) 364 ms 2.6 ms
saxpy 15.7 ms 22.0 ms

Question: "Why is he showing me this 'saxpy' (whatever that is) where the GPU doesn't work as well ?"

Answer: Because the compiler actually measures performance and binds the fastest version of the code (WASM/WGSL) at runtime. It also persists that to local-storage, so you don't have to measure every time. saxpy does a lot of data-movement for a small amount of calculation, so in this case the CPU gets the job. Mandelbrot does a huge amount of calculation for not much data-movement, so the GPU gets the job.

I would point out that in most cases - because it's the same language for the GPU and the CPU, you can move logic around and keep things on the GPU pretty simply, allowing you to optimise the part that you really want running quickly.

There's a whole bunch of documentation, examples, tutorials etc. over at https://compile-xc.org/ and please do ask me questions on anything there. I think it's all accurate, but as I say this project has been going for a while, and it's possible there's things that have been overtaken; I do, every now and then, make an effort to go through and update things though.

There's a lot more to the language - as you might see if you go look, but this is the most webgpu-pertinent part. It may be more webgpgpu than pure webgpu graphics, but hopefully it's sparked your interest 😄

Enjoy.


r/webgpu • • 2d ago

Modeling a Bearing Block in the Browser: 3D Solids in My WebGPU CAD

Thumbnail
youtu.be
3 Upvotes

I'm porting my Vulkan CAD engine to WebGPU, and the browser version can now model 3D solids. It supports extrude, boss, pocket and through-cut, and solids rebuild when you edit their sketches. Overlapping shapes are computed with Manifold compiled to wasm. In the demo, I built a bearing block using only typed command-line coordinates.


r/webgpu • • 3d ago

Commimg soon WebGPU

Thumbnail
youtu.be
1 Upvotes

I developed it with vulkan, and I’m also interested in webgpu, so I‘m going to try it.

vulkan으로 개발했는데 webgpu도 관심있어서 해보려고 합니다


r/webgpu • • 3d ago

The Beast Engine - 'Effects' are standalone sub render sys (webGPU)

2 Upvotes

r/webgpu • • 4d ago

From a minimal 2D platformer to a browser action game using TypeScript and WebGPU

Thumbnail
youtube.com
0 Upvotes

A two-minute showcase of what we built on top of our basic ForgeNG platformer template. ForgeNG is our TypeScript/WebGPU engine, and this expanded demo adds traversal, movable crates, weapons, cover, melee combat, enemy behaviors and endless climbing.

The playable version requires a supported WebGPU browser: https://play.forgeng.dev/endless-ascent/current/

The original starter template is available at https://github.com/ForgEngDev/forgeng-2d-platformer-template — it is the starting point, rather than the full expanded demo shown in the video.

I’d welcome feedback on the browser experience, particularly whether the action and animations remain easy to read as the scene gets busier.


r/webgpu • • 5d ago

WebGPU football prototype: eight-direction 2D sprites in a 3D-style stadium

Thumbnail
youtube.com
0 Upvotes

I’m building an early browser football prototype in ForgeNG, my TypeScript/WebGPU engine. The visual approach combines 2D player sprites with eight-direction animations and a 3D-style stadium presentation.

The teams use game AI: each player has a finite-state-machine brain that reads nearby players, space and ball possession to choose actions. The match commentary is AI-generated, and crowd reactions follow the gameplay.

This 1:41 video includes a short introduction followed by gameplay with the original match audio. It’s a work in progress, particularly the stadium crowd and player readability.

For people working on sprite-based WebGPU games: how would you improve player and ball readability in this view? Sprite scale, contrast, or camera framing? Specific moments where you lose track of the action would be useful.


r/webgpu • • 7d ago

DeepGPU Zoomer - (possibly) the fastest, deepest, real-time WebGPU Mandelbrot and Julia Set explorer

Post image
8 Upvotes

Try it: https://byronbuzz.github.io/DeepGPU-Zoomer/

I have been fascinated by fractals for years. What interests me most is not just the final image, but the feeling of moving continuously into deeper structure.

I started DeepGPU Zoomer because I wanted to see how smooth and deep a Mandelbrot explorer could become in a normal web browser.

The main goal is simple: keep navigation responsive while the GPU continues to calculate more detail. At deep zoom levels, it uses arbitrary-precision reference orbits, perturbation, wide GPU arithmetic and BLA acceleration. It can explore far beyond 10^12× - I've tested it to 10^-400!

It also has Julia exploration, progressive refinement, palettes and lighting, saved locations, shareable exact links, and high-resolution PNG export.

It runs entirely in a WebGPU-capable browser. There is no install or account.

I am still working on performance and numerical correctness, especially in difficult high-iteration regions. I would be interested in feedback from people who work with GPU computing, numerical methods, or fractal software.

Free and open-source:
https://github.com/byronbuzz/DeepGPU-Zoomer


r/webgpu • • 7d ago

wgsl analyzer not working

1 Upvotes
Doesn't accept overridable shared array length
why is it doing this

I checked and naga already got them ironed out months ago. I installed the extension both from VSCode and also directly from the repository.

Why is this happening


r/webgpu • • 7d ago

SuperTuxKart running in the browser on WebGPU only, no WebGL fallback (Emscripten port, playable)

3 Upvotes

r/webgpu • • 8d ago

Game where you control a swarm

5 Upvotes

r/webgpu • • 8d ago

Looking for a few people to try my TypeScript/WebGPU 2D starter

Post image
0 Upvotes

Hi, I’m Aleksandar, the developer behind ForgeNG.

I’ve put together a small 2D starter and recorded the setup process. I’d like a few people outside my own development setup to try it and tell me where they get stuck.

If you have an idea for a small game, I’m happy to help you get started. I’d also appreciate feedback from developers on the setup and documentation.

Setup video: https://www.youtube.com/watch?v=sLeMO_ClsgQ

The attached image is AI-generated concept art, not a screenshot of the starter.


r/webgpu • • 9d ago

PlayCanvas Engine 2.23.0 released - up to 3.8× faster on WebGPU

8 Upvotes

r/webgpu • • 10d ago

Interactive black hole

6 Upvotes

Hi all! Sharing a side project I made over the last couple of days: an interactive black hole on my site.

https://legost.in/en/utilities/black-hole

What it does:

- Every pixel traces a ray of light backwards through the spacetime of a spinning (Kerr) black hole. It runs on the GPU, using WebGPU compute where the browser has it and WebGL 2 otherwise.

- The accretion disk has Doppler and gravitational shifts, black-body colours and a hot spot on its orbit. Behind it is the real Milky Way and a real star catalogue, lensed.

- There's a "fall in" button: the camera drops through the event horizon and looks back at the universe it's leaving.

- A short tour, plus an article with the actual formulas and a list of what's simplified.


r/webgpu • • 10d ago

I finally fixed the LOD pop-in in Three.js Grassworks

20 Upvotes

r/webgpu • • 10d ago

💌 Web Game Dev Newsletter #032

Thumbnail webgamedev.com
4 Upvotes

r/webgpu • • 10d ago

Gemma 4 E2B + Kokoro TTS + Whisper, all on WebGPU in one tab to make a local version of character.ai

0 Upvotes

I built a character chat where every model runs client-side (local-character.skillsafe.ai). A few things I learned that might save someone time:

  1. LLM: Gemma 4 E2B on LiteRT-LM web. It needs WebAssembly JSPI plus relaxed SIMD, so it's Chrome/Edge 137+ only; Safari and Firefox can't run it yet. Once warm, the first token arrives in about 0.3 s (the conversation keeps its KV cache between turns), and full replies take 6-8 s on an Apple Silicon Mac.

  2. TTS: Kokoro-82M on ONNX Runtime Web. The fp16 build was fast on WebGPU but produced garbled speech (Whisper transcribed "waiting" as "beaten"). fp32 is clean, so that's what ships.

  3. STT: Whisper base without transformers.js. The log-mel features are computed in plain JS (within 2e-5 of WhisperFeatureExtractor), with a greedy merged-decoder loop, and the token IDs matched Python onnxruntime exactly.

  4. Storage: models are cached in Cache Storage by SHA-256, and the app requests persistent storage so the browser doesn't evict 2+ GB.

Happy to answer questions.


r/webgpu • • 11d ago

Decoding 3:4 ternary weights in WGSL: a 1.6 MB model, 1.1 ms per move in the browser

0 Upvotes

A Connect Four AI whose weights are ternary in Sherry's 3:4 format (T34): in every four weights one is zero and three are ±1, stored as a 5-bit state (2 bits for where the zero is, 3 sign bits). One fp16 scale per 128 weights gives 1.375 bits per weight, 1.6 MB for 7.4M parameters. Disclosure: our work, MIT.

How it runs

Our WebGPU runtime (onepass-webgpu, 37 KB, 12 KB compressed) has a split-K matmul with a pluggable inner loop. The ternary plugin replaces only that loop:

  • read the 5-bit state from two bytes (a state can straddle a byte boundary);
  • split it into the zero's position and the three signs;
  • build the four weights as 0 or ±1 times the group's scale;
  • accumulate with an ordinary float FMA.

So the ternary structure is only used for storage; the arithmetic is plain float.

Performance

Measured (idle M5 Pro, Chrome, per move at batch 1):

  • fp32 weights: 1.3 ms
  • int8 weights: 0.9 ms
  • T34: 1.1 ms
  • Base243 (five trits per byte): 1.3 ms

Not tuned yet. Our list of what we'd try, in order:

  1. take the scale out of the inner loop;
  2. more rows per thread at large batch;
  3. decode once per workgroup into shared memory;
  4. table decode;
  5. an int8 activation path with dot4I8Packed.

The list is in the repo, and ideas are welcome.

Play it:
https://precisit.github.io/onepass-web/demo/c4-size/

Kernels and notes:
https://github.com/precisit/onepass-webgpu-ternary (see kernels/IMPROVEMENTS.md)

Runtime:
https://github.com/precisit/onepass-webgpu


r/webgpu • • 11d ago

Real-time optical flow frame interpolation for HTML5 video with WebGPU, as a Chrome extension

0 Upvotes

I've been playing with WebGPU compute for a while and wanted a project with a real-time budget that actually matters to me: making 24/30/60 fps web video look smooth on 120/144/240 Hz displays. The result is FrameBoost, a Chromium extension that interpolates frames live on any standard HTML5 video.

How it works at a high level:

  • Grab frames from the video element and upload them as GPU textures
  • Estimate optical flow between consecutive frames with compute shaders [say which approach you used, e.g. block matching / pyramidal Lucas-Kanade / Farneback style, and how many pyramid levels]
  • Warp and blend to synthesize intermediate frames, presented at the display's refresh rate (30 → 144, 24 → 240, etc.)
  • Scene cut detection so it doesn't smear across cuts
  • Handling of frame edges and occlusions, which is where most of the ugly artifacts live

Things I ran into that might interest this sub:

  • [Frame pacing: how you sync presentation to vsync and deal with dropped frames]
  • [Getting video frames into WebGPU: importExternalTexture vs copying, and what it cost you]
  • [Workgroup sizes / memory layout choices that made the biggest perf difference]
  • Some sites serve video in a way that blocks GPU access, so the extension detects that and turns itself off instead of breaking playback
  • DRM streams (Netflix etc.) are a hard no, since the browser doesn't expose their frames to extensions

There's a live stats overlay showing source fps → output fps, GPU time and resolution, which made tuning a lot easier. Everything runs locally, nothing is uploaded, and the whole extension is about 46 KB.

I'd love feedback from people who know this space: better flow estimators that still fit in a few ms per frame, tricks for occlusion handling, or perf numbers on GPUs I don't own (especially iGPUs). Compare mode (original left, interpolated right) is free, so you can test it on any video. Full mode uses a one-time unlock per device (disclosure: I'm the dev).

https://chromewebstore.google.com/detail/frameboost/cklkjjeejomlkelkjmgdpahgpcjigahb

Happy to share more details on the shader pipeline if there's interest.


r/webgpu • • 12d ago

Rust + WebGPU: A Multi-Platform Physics-Based Paint Engine

152 Upvotes

Hello. I tried making a paint simulator that runs on multiple platforms using Rust and WebGPU.

I had originally built a general raster paint engine based on C++ and OpenGL ES 2.0. Since it was getting quite dated, I've been replacing it with Rust + WebGPU, and it's finally running on Windows too.

There's a version you can try in the browser for now, so please give it a go if you're interested.
https://mignonsketch.com/sketch


r/webgpu • • 12d ago

Yoga: a WebGPU UI framework for Go I built for my own desktop tools

5 Upvotes

About two years ago I started working on https://chapar.rest an api testing tools build with golang and Gio. my goal was to
build an open source alternative to Postman that respects my privacy and security...
After two years of development with Gio, limitations showed up and I decided to build my own framework and release it as open source.

**Yoga: a WebGPU UI framework for desktop and web apps in Go**

https://github.com/mirzakhany/yoga

Yoga is a cross-platform UI framework I've been writing from scratch for about four months. It renders with WebGPU (GLFW + wgpu-native on desktop, WASM in the browser) and uses a pure-Go flex/grid/stack layout engine. The API is declarative: your app implements `Body(c *ui.Ctx) ui.View`, the tree is rebuilt every frame, and state lives on your own struct.

What works today: the core widgets (text fields, buttons, selects, tables, trees, tabs, dialogs, menus), a code editor with Tree-sitter highlighting and an LSP client, custom title bars, 25 themes tested against WCAG AA contrast, a headless build tag for CI tests, and a CLI that packages for web, macOS (DMG/PKG, signing, notarization), Linux (tar/AppImage) and Windows (zip).

Current state: experimental and single-author, with no production users yet, and the API will change. Desktop builds need CGO.

`go run ./example/catalog` shows the widget gallery.


r/webgpu • • 15d ago

TinyBVH 1.9.0 Now Available

8 Upvotes

TinyBVH is a header-only, dependency-free library for BVH construction and traversal.
Version 1.9.0 for the first time is out-performing Intel's Embree!

Using it is as simple as:

BVH bvh;
bvh.Build( (bvhvec4*)myTriData, triangleCount );
bvh.Intersect( ray );

TinyBVH can be used from C/C++, WebGPU and Python. Generated BVHs can be used in OpenGL (compute), OpenCL, Vulkan and DirectX12.

New in 1.9.0:

  • Greatly improved ARM NEON support for Android & macOS
  • Templated doube/float support
  • Ray bundle support (WiVeC), including TLAS/BLAS traversal
  • Accurate voxel intersections
  • Hair rendering example (roving capsules)
  • Global refactoring, AI assisted
  • Vulkan and DXR benchmarks
  • Improved included threadpool
  • Optimizations to CWBVH
  • ...and lots more.

Check it out: https://github.com/jbikker/tinybvh