DLSS 5: RTX 50 Series Sees Minor Gains from NVFP4 Hybrid Mod
Modders have successfully implemented a hybrid FP8/NVFP4 version of NVIDIA's DLSS 5 on RTX 50 Series GPUs. However, the performance improvements observed are minimal, suggesting more complex optimizations are needed.

Enthusiast modders are pushing the boundaries of NVIDIA's latest AI upscaling technology, DLSS 5, with a new experimental hybrid implementation for GeForce RTX 50 Series GPUs. This latest effort, part of the OptiScaler-DLSSNR-PreSR-Multipass project, aims to leverage the newer NVFP4 format alongside the existing FP8, but early results show only marginal performance boosts. Version 0.7.1 of the project introduced the hybrid model, and subsequent updates have further refined it, with version 0.7.2 now recommending the combined approach for NVIDIA's upcoming Blackwell architecture GPUs.
The motivation behind this experiment is to reduce the significant computational demands of DLSS 5. NVIDIA's fifth-generation Tensor Cores, found in Blackwell GPUs, offer native acceleration for the NVFP4 format, which promises reduced memory usage compared to FP8. Theoretically, shifting parts of the DLSS 5 neural rendering process to NVFP4 should yield a noticeable performance increase. However, real-world testing indicates that the gains are exceedingly small. According to the project's release notes, the hybrid model only managed to reduce the Neural Rendering pass time at 4K resolution by approximately 1–2% when compared to the pure FP8 implementation. The developer behind the project characterized the improvement on Blackwell GPUs as "VERY minor." Crucially, this reduction in rendering time does not directly translate to a similar uplift in overall game frame rates.
Challenges in Precision Optimization
The observed minimal improvement highlights the complexities of efficiently utilizing low-precision formats like NVFP4 for AI inference. Simply converting an existing FP8 model to NVFP4 is insufficient. Achieving significant performance gains requires sophisticated quantization techniques and a deep understanding of the neural network's architecture to optimize it for the target hardware. The project's developer notes that unsupported model shapes within the hybrid implementation still necessitate a fallback to FP8, further limiting potential benefits. While the potential of NVFP4 is acknowledged by NVIDIA for its memory efficiency and Blackwell Tensor Core acceleration, realizing that potential in practice proves challenging.
Despite the minor gains from this specific mod, NVIDIA itself is actively working on improving DLSS 5's performance. The company has stated that the technology is already significantly faster than its initial GTC 2026 demonstration, and further optimizations are planned. Official support for the RTX 40 Series GPUs is also in development, which will broaden the accessibility of DLSS 5. As DLSS 5 is NVIDIA's most computationally intensive upscaling model to date, reducing its performance cost is a key objective for the company. This community-driven experiment, while not delivering a major breakthrough, offers valuable insights into the practical challenges of optimizing AI models for next-generation hardware and low-precision formats.
In a specific test case involving Baldur's Gate 3 at 4K, the hybrid model achieved approximately 55.16 rendered frames per second, a slight improvement over the 54.57 FPS recorded with the pure FP8 version. However, the developer cautioned that such a small difference might not be consistently repeatable across different sessions or hardware configurations. This underscores the need for rigorous testing and validation when assessing the impact of such optimizations. The evolution of DLSS 5 continues, with both community efforts and NVIDIA's internal development contributing to its progress.
