About this project
OpenDLSS-NR is a Vulkan-based reimplementation of NVIDIA's DLSS 5 Neural Rendering network that produces bit-exact results against the original. It implements the same 71-block Swin/ViT architecture (as found in build 310.8.0) using FP8 (E4M3) activations with FP16 accumulation on GPU tensor cores. Unlike upscalers, this is a generative neural rendering network: it takes a frame at the same resolution as output and regenerates detail from injected noise while adjusting tone, structure, and style.
The project also includes an independent WebGPU port (`ports/browser-webgpu/`) that runs the same network in a browser without tensor cores or FP8, achieving bit-exact intermediate results across all 75 block boundaries.
Host code is written in C++20, with GLSL kernels for the reference route (cooperative-matrix FP8 GEMMs) and Python-generated PTX kernels for the optimized fast route (using `mma.sync` E4M3, `cp.async` rings, barrier-free chaining, and split-K GEMMs). A demo application built on a patched version of Google's Filament renderer showcases the network with several glTF scenes.
Performance on an RTX 4070 SUPER: 2.8 ms at 768×768, 7.8 ms at 1920×1080, 12.6 ms at 2560×1440, and 29.3 ms at 3840×2160. The WebGPU port runs at 72 ms for 512×512.
Requirements include Windows, an NVIDIA Ada (or newer) GPU with specific Vulkan extensions, Visual Studio 2022+, Python 3, and Node.js + npm. The repository does not contain NVIDIA weights or software — users must supply their own model directory. The project is MIT licensed and not affiliated with or endorsed by NVIDIA.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.