Pimp My IDE / garage dispatch
Back to garage
October 1, 2026 | graphics / compatibility / release proof

Bit-exact is a strong claim with a hard edge.

OpenDLSS-NR reconstructs a neural-rendering graph and tests exact bytes at internal boundaries. That proves something rare. It does not bundle the model, reference captures, renderer contract, or a portable fast path.

Keep graph parity, model custody, frame history, and target hardware on separate test cards. Then decide which card the product needs.

Exactness needs a named boundary.

OpenDLSS-NR says its Vulkan implementation matches DLSS-NR build 310.8.0 at all 75 comparable block and encoder-transition boundaries. Its parity checker treats a signed-zero change as failure, requires every boundary to have a reference or an omission reason, and compares the production schedule with a barrier-heavy schedule.[1]

That is a precise contract. The repository also says the CPU reference does not cover the expert feed-forward network, the 512-channel split block, or the global transformer. Those sections rely on two GPU routes matching native captures at block boundaries. The evidence remains strong, but the diagnostic depth changes inside the graph.

"Bit-exact" answers a byte question. It does not answer every integration question.

The missing files are part of the result.

The repository supplies code, documentation, build scripts, and an MIT license. It does not supply weights or parity fixtures. The model loader expects a manifest with stage hashes and tensor offsets. The parity tool expects recorded references that account for the whole declared check. Without those files, a fresh clone cannot reproduce the headline claim.[2]

The license boundary is explicit. The project license covers the repository. Its notice says the repository contains no NVIDIA software, weights, or headers. It grants no rights to model data. That distinction belongs beside the build command, not in a footnote after adoption.

A frame is more than one graph pass.

The command-line tool runs single frames without history. The demo adds the temporal loop. It builds a display proxy, records motion, reprojects the previous output, samples history, uses the network's blend logit, and writes the next history image. The frame document says this surrounding pipeline is as important to matching output as the network itself.[3]

This is the integration trap. A team can reproduce the graph and still ship different pixels because motion, padding, history publication, renderer synchronization, or display conversion differs. Test the graph to debug arithmetic. Test moving scenes to judge the product.

The fast route is hardware-specific.

The native build requires Windows, an NVIDIA Ada or newer GPU, and named Vulkan extensions. One fast route launches CUDA kernels inside a Vulkan command buffer. Khronos describes VK_NV_cuda_kernel_launch as an NVIDIA extension for loading CUDA fat binaries and launching CUDA kernels from Vulkan.[4]

The repository's WebGPU port matters because it separates graph identity from one fast implementation. The authors report the same reference bytes through a slower route without tensor cores or FP8. That is portability evidence for the specification. It is not performance parity or proof that another renderer has the same frame contract.

Interactive makeover / compatibility rack

Choose the claim, then couple the evidence

Traditional purpose replaced: one generic compatibility checkbox. Better version: a three-position acceptance selector and four ordered evidence couplers generate one review card without pretending that the tests ran.

Set the acceptance target

The selector changes the required test. The couplers choose which evidence sections the card will request. They do not attach model files, fixtures, captures, or results.

Compatibility target
Review-card sections
4 review-card sections openGraph parity
Acceptance line / requested evidence

Neural rendering compatibility rack

The target is selected. The source section is open.

Pin the implementation revision, reference build, graph scope, and exact checker.

All four couplers means the review-card structure is ready. It does not mean the model was obtained lawfully, the fixtures exist, parity passed, the renderer matched, or the target hardware met its frame budget.

One port / four contracts

Do not flatten compatibility into yes or no.

01 / GRAPH

Arithmetic identity

Pin publication points, rounding, boundary coverage, checker behavior, and the exact reference build.

02 / PAYLOAD

Model custody

Keep model hashes, manifests, provenance, distribution terms, and permitted use beside the code revision.

03 / FRAME

Temporal behavior

Test proxy conversion, motion, history, blend, synchronization, and display output on moving scenes.

04 / HOST

Shipping fit

Verify renderer hooks, driver extensions, target hardware, frame time, failure behavior, and fallback.

Sources read

Source log and evidence boundary
  1. OpenDLSS-NR numerical exactness contract at commit 9d08f41, read October 1, 2026. It defines bit-exact verdicts, complete boundary-fixture requirements, production-versus-instrumented schedule checks, current coverage, and the sections without a CPU reference.
  2. OpenDLSS-NR repository and README at commit 9d08f41, read and cloned October 1, 2026. The README documents the graph, build requirements, model-directory contract, missing weights and fixtures, reported performance, WebGPU port, and license boundary. We also read the repository's MIT license and NOTICE.
  3. OpenDLSS-NR frame pipeline document at commit 9d08f41, read October 1, 2026. It documents proxy conversion, motion vectors, history sampling, temporal composition, renderer command-buffer integration, frame timing, and resolution behavior.
  4. Khronos Vulkan reference for VK_NV_cuda_kernel_launch, read October 1, 2026. It describes the NVIDIA extension used to load CUDA fat binaries, obtain CUDA functions, and launch CUDA kernels from Vulkan command buffers.
  5. Hacker News discussion, item 49906100, read October 1, 2026. We resolved the exact item through the Hacker News API. The discussion raised useful questions about missing weights, performance, inputs, and source authorship. We treated those comments as questions, not evidence for technical claims.

Artifact check: We cloned commit 9d08f4184bbcb9d858e2fb7a7834ec0837a9d2f1 and ran the repository's test_fast_divmod.py helper with divisors 1 through 4. It reported exact results for every unsigned value below 224. This is a narrow smoke check of one generated PTX arithmetic helper. We did not build or run the renderer because this host is Linux, has no designated compatible NVIDIA test GPU, and has none of the required model or fixture files. We did not reproduce the project's parity or timing claims.