As AI deployments shift from model training to large-scale production inference, organizations are increasingly focused on improving AI infrastructure efficiency and reducing operational costs. DDN, Nebul, and NVIDIA have expanded their collaboration to advance AI inference economics through high-performance KV Cache acceleration, combining AI infrastructure, accelerated computing, and intelligent data architecture to improve GPU utilization, reduce latency, and lower the cost per token generated.
The collaboration combines Nebul's AI inference platform, DDN's Infinia data intelligence architecture, and NVIDIA accelerated computing technologies to address one of the biggest challenges facing enterprise AI deployments—making AI inference economically sustainable at scale. Rather than focusing solely on model performance, the initiative emphasizes infrastructure efficiency by maximizing GPU utilization, accelerating token generation, and reducing operational costs.
As part of the ongoing proof-of-concept, DDN and Nebul are validating next-generation KV Cache acceleration capabilities designed for NVIDIA DSX-based AI factory deployments. The technology improves inference efficiency through distributed KV Cache services, GPU-native data movement, intelligent data orchestration, and high-performance storage architecture.
By reducing bottlenecks associated with data movement during inference, the solution aims to improve time-to-first-token, increase throughput, and maximize the return on expensive GPU infrastructure investments.
According to the companies, AI adoption has entered a new phase where business metrics such as GPU utilization, cost-per-token, tokens-per-watt, and time-to-production have become key indicators of AI success. As organizations deploy larger foundation models and agentic AI workloads into production, efficient data infrastructure is emerging as a critical component for delivering scalable and profitable AI operations.
"The AI conversation has fundamentally changed," said Alex Bouzari, CEO and Co-Founder at DDN.
"For years, the industry focused on acquiring GPUs. Today, the question is how efficiently those GPUs generate value. Inference has become the economic engine of AI, and reducing the cost of every token produced is now one of the most important challenges facing the industry."
"For years, the industry focused on building larger models. Today, the challenge is making those models economically viable in production," said Arnold Juffer, CEO at Nebul.
"Every organization is looking for ways to generate more value from its AI infrastructure investments. Through our collaboration with DDN and NVIDIA, we are demonstrating how KV Cache optimization and high-performance data architectures can improve inference efficiency, reduce latency, and help unlock the next phase of AI adoption."
"AI infrastructure is increasingly defined by efficiency at scale," said Rod Evans, Vice President of Cloud Infrastructure at NVIDIA.
"As organizations deploy larger models and agentic AI workloads into production, technologies that improve GPU utilization, reduce latency, and accelerate token generation become critical. DDN continues to be an important collaborator in advancing the data and infrastructure capabilities needed to support the next generation of AI factories."
Beyond technology validation, the collaboration has expanded to include joint work on benchmarking methodologies, scalability testing, and future technical publications. The initiative aligns with the growing industry focus on AI factories, where infrastructure performance directly impacts operational efficiency and return on AI investments.
As enterprises continue scaling AI workloads, advancements in KV Cache acceleration and intelligent data infrastructure are expected to play an increasingly important role in enabling cost-effective, high-performance AI inference across production environments.
DDN is the world's leading AI and data intelligence company, powering the world's most demanding AI workloads by keeping GPUs fed, efficient, and productive—at massive scale—so organizations can train, checkpoint, and infer faster with less footprint and power while achieving tremendous ROI from their AI investments. From hyperscalers and next-gen cloud builders to enterprises, governments, and research institutions, DDN delivers proven data intelligence at exabyte scale across millions of GPUs—so customers can deploy AI with confidence, accelerate time-to-value, and realize outsized returns. Discover more at ddn.com.
Nebul offers European values on privacy and sovereignty combined with the technology, convenience and scale of hyperscale into a genuine European Sovereign-AI Cloud. Using the full range of NVIDIA technologies, new capabilities like Agentic AI and AI Coding can be enabled more easily and quickly without sharing sensitive corporate data with third parties you do not control. Nebul AI Cloud provides access to the latest AI technologies to empower European organizations to harness the power of AI.