A developer going by maanHimself has published OpenDLSS-NR, a from-scratch Vulkan reimplementation of NVIDIA's DLSS 5 Neural Rendering network that produces bit-exact results against the original proprietary implementation. The project matches not just the final rendered image but all 75 internal block boundaries, byte for byte, across the 71-block shifted-window transformer / ViT architecture that NVIDIA ships as a black box inside its driver stack. The network in question is not the upscaler most gamers associate with DLSS. It is what NVIDIA calls a 'generative neural rendering' system: a U-net of Swin transformer blocks with a global Vision Transformer at the bottom, taking a rendered frame at native resolution and re-rendering it with injected noise, adjusting tone, structure, and detail. The model uses FP8 (E4M3) activations with FP16 accumulation across 141 MiB of weights and 241 dispatches per frame. On an RTX 4070 SUPER, the full network runs in 2.8 ms at 768×768 and 29.3 ms at 4K. The project ships two independent implementations. The primary route uses GLSL cooperative-matrix kernels and hand-generated PTX for NVIDIA tensor cores, with techniques including cp.async ring buffers, barrier-free counter chaining, and split-K GEMMs. The second is a WebGPU port that runs the identical model in a browser — no tensor cores, no FP8 hardware, no kernel fusion — achieving 72 ms at 512×512 against the native implementation's 2.7 ms. The bit-exactness holds across both, proving the specification is what matters, not the hardware path. Critically, the repository contains no NVIDIA weights, headers, or software. Users must supply their own model directory in a specified format. The project explicitly disclaims any NVIDIA affiliation. The entire codebase is MIT-licensed. This is a clean-room reimplementation: the developer worked from NVIDIA's published technical report and the observable behavior of the network to reconstruct every kernel, every scheduling decision, every numerical detail. The engineering rigor is unusual for an open-source project. A formal parity testing system compares against recorded captures of the original at every block boundary. A verification mode does kernel-by-kernel bisection of block 0. Every tuning switch — and there are over a dozen, controlling fusion, chaining, PTX vs GLSL fallback paths — is constrained to produce byte-identical output. The fixture validation system refuses to run if declared checks lack references, if references name nothing in the graph, or if comparable boundaries have neither a reference nor an explicit reason for omission. The strategic significance is straightforward. NVIDIA's DLSS stack is a proprietary moat — it runs only on NVIDIA hardware, distributed only through NVIDIA's driver, with no public specification of the network internals beyond a marketing-grade technical report. A bit-exact open reimplementation, especially one that runs on commodity WebGPU, converts a vendor lock-in mechanism into a documented, portable specification. It doesn't liberate the weights, but it liberates the architecture and the inference path. This is the kind of project that matters more as precedent than as product. If a single developer can reconstruct a 71-block neural renderer to byte-level parity using only published information and observable behavior, the defensibility of proprietary inference stacks as competitive moats becomes a question of weight secrecy alone — and weight secrecy is a thinner wall than architectural opacity.