WEKA has introduced NeuralMesh 6, the latest version of its AI data and memory infrastructure platform designed to support production-scale AI training, inference, and accelerated compute workloads. The release combines native multi-tenancy, AI-optimized data mobility, unified object and file storage, Kubernetes-native operations, and observability into a single software platform, helping organizations simplify AI infrastructure while improving performance and efficiency.
As enterprises shift from AI model training to large-scale production inference, infrastructure requirements are evolving rapidly. NeuralMesh 6 is designed to address growing demands for long-context reasoning, agentic AI, retrieval-augmented generation (RAG), and GPU-intensive workloads by reducing memory bottlenecks and improving AI infrastructure utilization.
WEKA says NeuralMesh 6 is its most significant software release to date, delivering a unified platform that eliminates the need for organizations to assemble AI infrastructure from multiple vendors.
The platform combines capabilities including native hyperscale multi-tenancy, a full S3 protocol implementation on NVMe storage, metadata-first intelligent replication, always-on data reduction, Kubernetes-native operations, and centralized observability within a single software stack.
According to the company, these capabilities are designed to simplify AI infrastructure deployment while improving scalability for enterprise AI environments.
WEKA highlights that the AI industry is increasingly prioritizing production inference over model training, creating new infrastructure challenges related to memory management, metadata processing, and storage performance.
NeuralMesh 6 is built to support these workloads by improving inference efficiency for AI cloud providers, enterprises, government agencies, and organizations deploying large-scale AI applications.
The company also shared production benchmark results on Oracle Cloud Infrastructure (OCI) using its Augmented Memory Grid technology. By extending GPU memory through accelerated persistent KV cache access to NVMe storage managed by NeuralMesh, production deployments demonstrated:
"As agentic AI workloads push context windows and GPU utilization to new limits, WEKA and Oracle Cloud Infrastructure are helping customers scale inference more efficiently without simply adding more GPUs," Pablo Selem, senior director, software development, Oracle Cloud Infrastructure. "WEKA's NeuralMesh platform with Augmented Memory Grid on OCI helps remove memory bottlenecks, delivering substantially more throughput and concurrent users from the same GPU footprint. For customers, that means higher ROI on infrastructure investments and a clearer path to cost-efficient AI at scale."
NeuralMesh 6 introduces several capabilities aimed at supporting modern AI infrastructure.
The platform adds native multi-tenancy, combining physical hardware isolation with logical network isolation. It also supports both Composable Clusters for dedicated hardware resources and Virtual Multi-Tenancy, enabling organizations to scale to thousands of isolated tenants within a single deployment.
WEKA also introduces a native S3 implementation that allows the same data blocks to be accessed simultaneously through both S3 and POSIX protocols, eliminating duplicate datasets across AI training, fine-tuning, and inference workflows while increasing object storage performance.
To support distributed AI environments, NeuralMesh 6 includes metadata-first intelligent replication, allowing organizations to browse remote datasets immediately while transferring only the data required during execution. This approach reduces WAN traffic and enables workloads to move more easily across cloud regions and data centers based on GPU availability.
The release also adds:
"WhiteFiber is building a distributed GPU platform across multiple data centers connected by high-speed dark fiber. At our scale, data mobility isn't a nice-to-have; it's foundational," said Sam Tabar, CEO at WhiteFiber. "NeuralMesh's intelligent replication makes data mobility real at scale: we can make datasets visible across sites and pull exactly the data each job needs to the next GPU allocation as it becomes available. That shifts replication from a back-end protection function to a core part of how our distributed AI infrastructure needs to operate, improving workload mobility, capacity efficiency, and the resiliency our customers depend on. WEKA's data and memory infrastructure provides the foundation to scale our footprint without compromise. We're excited to keep building on that together."
"The infrastructure operators running production AI today have been forced to assemble platforms from vendors that were never designed to work together. Separate stacks for file and object, manual data movement between them, multi-tenancy bolted on after the fact. NeuralMesh 6 delivers what they've actually needed all along: a single platform that handles the high-performance file layer and the high-capacity object layer on the same blocks, with native multi-tenancy, intelligent data mobility, and always-on data efficiency built in from the start. This is what production inference infrastructure looks like when it's designed for the workload, not retrofitted for it," said Ajay Singh, Chief Product Officer at WEKA.
WEKA announced that NeuralMesh 6 will become generally available during the second half of 2026. Existing customers will be able to upgrade through standard upgrade channels at no additional cost.
Alongside the software release, the company also introduced the third generation of WEKApod Nitro, WEKApod Prime, and WEKApod Prime Max, AI storage appliances engineered specifically to optimize the NeuralMesh platform.
WEKA is the AI data and memory infrastructure company transforming the economics of agentic AI. Its NeuralMesh™ platform unifies high-performance data storage with extended GPU memory, giving enterprises, AI cloud providers, and AI builders a single foundation for training, inference, and agentic workloads. With Augmented Memory Grid, NeuralMesh extends GPU memory capacity by 1000x, accelerates time to first token by up to 20x, and delivers 10x more concurrent users from the same GPU footprint, proven in production benchmarks. Trusted by 30% of the Fortune 50, WEKA enables organizations to scale AI faster, optimize GPU utilization, and reduce the cost of every token served. Learn more at www.weka.io or connect with us on LinkedIn and X.