Lightbits Labs, inventor of NVMe over TCP and the Inferra KV cache acceleration engine for AI inference, today announced the appointment of former Infineon executive Ramesh Chettuvetty as Senior Vice President of Product and Business for AI Solutions. Chettuvetty will drive the business forward and build the playbook for Lightbits' AI solutions in the next generation of data centers and NeoClouds, ensuring that the company's technology solves hardware efficiency and performance challenges faced by organizations running GPU-heavy workloads.
Lightbits appoints former Infineon executive Ramesh Chettuvetty as SVP of Product and Business for AI Solutions.
Inferra KV cache engine eliminates GPU stalls in long-context AI inference by enabling disaggregated KV cache architectures.
Inferra enables context windows to scale from 32K to 1M tokens in production systems today, with path to 10M tokens.
Improves Time to First Token (TTFT) and reduces Inter-Token Latency (ITL) while providing per-agent QoS.
Lightbits backed by Cisco Investments, Dell Technologies Capital, Intel Capital, Lenovo, and Micron.
Lightbits is inventor of NVMe over TCP storage protocol.
“Lightbits is currently working with NeoCloud providers evaluating high-performance scaleout software-defined architectures for long context, high-density GPU environments,” said Eran Kirzner, Co-Founder and CEO of Lightbits Labs. “Ramesh brings deep business and technology expertise across memory systems and AI infrastructure that will help accelerate Inferra's adoption as a high-performance backend for modern inference environments.”
“I am excited to join Lightbits at this pivotal time,” said Ramesh Chettuvetty. “Inferra directly addresses one of the most critical bottlenecks in AI inference — scalable KV cache management. Combining Lightbits' proven software-defined storage leadership with next-generation AI inference infrastructure creates a compelling opportunity.”
Inferra effectively breaks the memory wall by enabling practically infinite KV cache capacity through smart tiering and pre-fetching algorithms, allowing context windows to scale from 32K to 1M tokens in production systems today, and to expand to 10M tokens in the next generation of production systems without scaling HBM. Inferra enables NeoCloud operators to accelerate Time to First Token (TTFT) and reduce Inter-Token Latency (ITL), provide per-agent Quality of Service with SLAs, scale context windows with a path to 10M tokens, improve GPU utilization, and reduce dependence on costly HBM capacity expansion.
Lightbits Labs (Lightbits) is the inventor of the NVMe over TCP storage protocol, which is natively built into its industry-leading block storage, and the first KV cache prefetch engine acceleration for AI. Lightbits data storage solutions are engineered to deliver unmatched high performance and maximum hardware efficiency for LLM inference, real-time analytics, and transactional workloads at scale. Lightbits is backed by enterprise technology leaders Cisco Investments, Dell Technologies Capital, Intel Capital, Lenovo, and Micron.