As enterprises accelerate the adoption of agentic AI, the focus is shifting from simply running AI models to building secure, governed environments where autonomous agents can execute real business workflows. Organizations increasingly require infrastructure that combines AI inference, policy enforcement, auditing, and enterprise-grade governance to support production AI deployments.
Timed with AMD Advancing AI 2026, Embedded LLM has announced the launch of TokenVisor Spaces, a new infrastructure solution designed for AMD-powered AI clouds and enterprise environments. Built for the AMD Instinct AI & HPC Software Ecosystem and validated with Red Hat OpenShift, TokenVisor Spaces enables organizations to deploy governed agentic AI services while combining CPU-based execution with GPU-accelerated inference.
Unlike traditional AI workloads that primarily generate model outputs, agentic AI systems require environments capable of executing code, managing files, interacting with enterprise applications, running browser automation, and maintaining workflow state over extended periods.
TokenVisor Spaces addresses these requirements by providing persistent agent workspaces running on AMD EPYC processors while leveraging AMD Instinct GPUs for AI inference. The platform combines execution environments with governance controls, policy enforcement, usage metering, and audit capabilities, enabling organizations to deploy production-ready AI agents on infrastructure they control.
The platform is designed to help AI cloud providers move beyond selling GPU compute alone by offering governed AI agent services.
For enterprise customers, TokenVisor Spaces allows AI agents to interact with existing business applications inside isolated containers while maintaining security, compliance, and operational oversight. This architecture enables organizations to deploy long-running AI agents capable of executing business processes without compromising governance.
"The next AI cloud product is not a GPU hour. It is a governed agent service," said Ghee Leng Ooi, CEO of Embedded LLM. "TokenVisor Spaces turns AMD-powered infrastructure into a place where agents can actually work: execute tools, persist state, follow policy, leave traces, and connect back to high-performance inference. Agentic AI runs where enterprise software runs, and that makes the CPU plus GPU platform the center of the next AI cloud."
Validated on Red Hat OpenShift, the platform provides isolated execution environments while TokenVisor manages access policies, usage controls, routing, budgets, rate limits, metering, and auditability.
"Agentic AI requires a balanced compute platform that brings CPU-based execution and GPU-accelerated inference together," said Dan McNamara, Senior Vice President and General Manager, Compute and Enterprise AI, AMD. "By combining AMD EPYC processors and AMD Instinct accelerators, TokenVisor Spaces shows how ecosystem software can help cloud providers and enterprises deploy open, governed agent services on infrastructure they control."
The platform enables organizations to deploy AI agents capable of long-running operations, secure code execution, human approval workflows, replayable event histories, and governed model access—all essential capabilities for enterprise-scale AI deployments.
Embedded LLM also highlighted advances in KV-cache reuse, an increasingly important capability for long-context, multi-turn AI agents.
The company reports validating storage-backed KV-cache reuse on AMD Instinct MI355X, achieving 3.31× lower warm-turn median latency and 2.23× faster total wall-clock performance compared with a matched HBM prefix-cache baseline during production-style synthetic agentic replay testing.
To further enhance enterprise AI infrastructure, Embedded LLM is collaborating with VAST Data and Tensormesh on platform-scale KV-cache reuse, persistent agent state, trace capture, replay, evaluation, and reinforcement learning infrastructure.
"KV-cache reuse is what makes long-running, multi-turn agents economically viable," said Kuntai Du, Chief Scientist and Co-founder at Tensormesh. "By adopting LMCache, Embedded LLM brings that infrastructure to more of the AI cloud market, giving agents persistent, reusable context so they run faster, are more cost-effective, save energy and stay auditable at scale."
"Agentic AI requires infrastructure that can efficiently manage context, data, and state across long-running AI workflows," said Anat Heilper, Director of AI Architecture at VAST Data. "Our collaboration with Embedded LLM helps bring the VAST AI Operating System to AMD-powered AI environments, giving customers the persistent data foundation they need to scale production AI with greater performance and efficiency."
TokenVisor Spaces is now available for partner deployments, enterprise evaluations, proof-of-concept projects, and AI cloud implementations, supporting organizations seeking governed AI infrastructure for large-scale production environments.
Embedded LLM is an agentic inference infrastructure company in the AMD Instinct AI & HPC Software Ecosystem and a Red Hat ecosystem partner. The company helps AI clouds and enterprises turn GPU fleets into production AI services through vLLM-based serving, TokenVisor for governed model APIs and monetization, TokenVisor Spaces for stateful agent execution, JamAI Base for traceable AI operations, and operator-level RL/post-training infrastructure.