TL;DR
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
OpenDLSS is a GitHub project that reimplements NVIDIA’s DLSS 5 neural rendering network using Vulkan and also provides an independent browser-based WebGPU version. Its author reports byte-for-byte agreement with reference captures at 75 network boundaries, but the project requires users to supply model weights and has demanding hardware and driver requirements.
A GitHub project called OpenDLSS presents a Vulkan reimplementation of NVIDIA’s DLSS 5 neural rendering network, with a separate version designed to run in browsers through WebGPU. The project’s author says its outputs match reference captures byte for byte at all 75 block boundaries; the code does not include the model weights, which users must provide separately.
The project describes the network as a 71-block shifted-window transformer with a global vision transformer component, arranged across six pooling levels. Its documentation says the model uses E4M3 FP8 activations with FP16 accumulation and has 141 MiB of weights. OpenDLSS is intended to reproduce the network’s computation, not to provide NVIDIA’s model files or a general-purpose replacement for its graphics software.
According to the repository, the network takes a rendered frame along with noise lanes, a reprojected prior output and five conditioning values. It returns four channels per pixel: an RGB residual and a temporal-blend logit. The project describes the method as generative neural rendering; it operates at the input frame’s resolution and is not an upscaler. Its scope excludes DLSS-SR, which the repository identifies as a different network.
The author reports performance figures measured on an RTX 4070 SUPER: minimum frame times over 40 frames were 2.8 milliseconds at 768-by-768, 7.8 ms at 1920-by-1080, 12.6 ms at 2560-by-1440 and 29.3 ms at 3840-by-2160. The repository says the GPU alternates between two clock states under sustained load, making medians a few percent higher. These are the project’s own measurements, not independent benchmark results.
What Reimplementation Makes Possible
OpenDLSS could give graphics developers and researchers a way to examine and run a reconstruction of a specialized neural rendering pipeline outside NVIDIA’s supplied implementation. The claimed agreement across intermediate blocks is relevant because it offers a more detailed comparison point than matching only a final rendered image. Whether that claim holds across other inputs and environments has not been independently established in the supplied material.
The browser port also offers a separate route to experiment with the network’s computation without tensor cores or FP8. The repository reports 72 ms at 512-by-512 for that version, compared with 2.7 ms for its Vulkan implementation at the same resolution. The comparison illustrates a substantial performance gap in the author’s tests, while making the WebGPU port a potential reference or experimentation tool rather than a performance substitute for the optimized GPU path.
There are practical limits. The Vulkan build requires a Windows system, an NVIDIA Ada-generation or newer GPU, and drivers exposing several specified Vulkan and NVIDIA extensions. The project also requires users to supply compatible model weights. Those requirements narrow who can run it and mean the repository alone is not a ready-to-use, self-contained DLSS 5 package.
NVIDIA DLSS 5 neural rendering GPU
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Two Implementations Differ
The repository divides its work between a Vulkan reference route and a faster path built with generated PTX kernels. The project says the Vulkan shaders implement the network’s operations directly, while the PTX route uses NVIDIA GPU features including FP8 matrix instructions and asynchronous data movement. Both are part of the same implementation effort; the supplied report does not describe an independent audit of either route.
The browser version, in ports/browser-webgpu/, is described as an independent implementation that matches the same captures. It does not use tensor cores or FP8 and does not fuse or chain operations between blocks. The project’s documentation says the demo integrates the network with the Filament renderer and supports a temporal feedback path, while the command-line tool runs single frames without history.
For comparison, NVIDIA’s DLSS 5 project page is cited by the repository as describing the model. The OpenDLSS project is a reimplementation of that network, not an announcement from NVIDIA and not a claim that other DLSS components are included. In particular, the repository explicitly says DLSS-SR is not implemented.
“The intermediates match too, not just the final image: all 75 block boundaries, byte for byte.”
— OpenDLSS GitHub repository
As an affiliate, we earn on qualifying purchases.
What the Parity Claim Establishes
The byte-for-byte result is a claim made by the project. The supplied material does not include an independent reproduction, external review, or details about how broadly the reference captures cover inputs, model files and hardware configurations. It is therefore unclear whether the reported parity extends beyond the project’s tested fixtures.
The repository says users must provide model weights, but the supplied information does not establish their source, availability or licensing terms. It also does not specify when the project was published. The reported performance figures are tied to one GPU and the author’s test method, so results on other cards or drivers remain unknown.
As an affiliate, we earn on qualifying purchases.
Running and Testing OpenDLSS
The repository provides build, benchmark, profiling and parity commands, as well as instructions for preparing the model directory and building a Filament demo. The immediate next step for readers seeking to validate the claims is to review the project documentation and, if they have the required hardware, drivers and weights, run its fixtures and benchmarks on their own systems.
Further confirmation would require independent testing of the intermediate outputs and performance, along with clearer information about compatible weight files and their terms of use. The supplied source gives no announcement of an NVIDIA response, an external evaluation or a planned release milestone, so those developments remain unreported.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is OpenDLSS?
OpenDLSS is a GitHub project that reimplements NVIDIA’s DLSS 5 neural rendering network in Vulkan and provides a separate browser-based WebGPU version. The repository says it does not implement DLSS-SR.
Does OpenDLSS include the model weights?
No. The project documentation says users must supply the model weights in a directory with the expected layout. The source material does not confirm where compatible weights can be obtained or the terms governing their use.
What hardware does the Vulkan version require?
The repository lists Windows, an NVIDIA Ada-generation or newer GPU, and a driver supporting several named Vulkan and NVIDIA extensions. Those requirements may exclude many systems.
Is the reported bit-exactness independently verified?
The repository author reports byte-for-byte matches at 75 block boundaries. The supplied source material contains no independent verification, so the result should be treated as the project’s claim.
Is this an upscaler?
No. The repository says the network takes and returns frames at the same resolution. It describes the project as neural re-rendering, and says DLSS-SR is a separate network not implemented here.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
