Workato has announced its contribution to the MLPerf Agentic Inference Benchmark, helping establish an industry-standard framework for evaluating AI agents in real-world enterprise environments. Working alongside AMD, Intel, and NVIDIA within the MLPerf Inference Agentic Taskforce, Workato contributed 500 enterprise workflow trajectories based on complex business processes commonly orchestrated across large organizations.
Workato announced its contribution to the MLPerf Agentic Inference Benchmark, the first industry-standard benchmark specifically designed to evaluate AI agents executing complex enterprise workflows rather than isolated prompts.
As a member of the MLPerf Inference Agentic Taskforce, Workato collaborated with AMD, Intel, and NVIDIA to help create benchmarks that more accurately represent how enterprise AI systems operate in production environments.
The benchmark measures how AI agents complete multi-turn business processes involving planning, reasoning, tool usage, follow-up actions, and workflow completion instead of evaluating single question-and-answer responses.
Workato contributed 500 synthetic workflow trajectories modeled on real-world enterprise business operations.
These workflows represent common customer service and operational scenarios where AI agents retrieve orders, track shipments, verify policies, escalate support cases, and resolve account issues by interacting with multiple business systems, databases, and knowledge repositories.
Although synthetic, the trajectories reflect the production environments Workato supports for enterprise customers, providing realistic evaluation scenarios for AI infrastructure providers.
“As AI moves from the edge of the business to its core, the orchestration that directs and governs agent work becomes as critical as the models themselves,” said Adam Seligman, Chief Technology Officer and General Manager of the AI Lab at Workato. “By grounding this benchmark in the kind of complex business processes Workato orchestrates in production every day, we're helping give model providers an independent standard for whether their AI agents can run reliably and at scale."
Unlike traditional AI benchmarks that focus on language model accuracy, the MLPerf Agentic Inference Benchmark evaluates complete agent execution across multi-step enterprise tasks.
The benchmark assesses how AI systems interpret requests, invoke external tools, process intermediate results, ask follow-up questions, and continue executing workflows until business objectives are completed.
Workato's enterprise expertise complements the infrastructure knowledge contributed by AMD, Intel, and NVIDIA, helping ensure the benchmark reflects production-grade enterprise AI deployments rather than isolated laboratory scenarios.
The contribution also highlights Workato's ongoing AI research efforts through its research laboratories in San Francisco and Singapore, where the company focuses on autonomous enterprise agents, reinforcement learning, synthetic evaluation, and customer-specific model optimization.
By contributing to MLPerf, Workato is extending its research into an open industry benchmark that allows organizations to evaluate AI serving systems using reproducible and standardized enterprise workloads.
The MLPerf Agentic Inference Benchmark will be made available through the MLPerf Endpoints framework, including reference implementations, datasets, and accuracy thresholds published by MLCommons, supporting broader adoption of standardized AI performance evaluation across the enterprise technology ecosystem.
Workato is the leading Control and Execution Platform for Enterprise AI — the neutral platform enterprises trust to put AI to work across their business. Workato unifies data, applications, and processes into a single platform so AI can reliably orchestrate business processes in production at enterprise scale. Built on more than a decade of running mission-critical processes for over half the Fortune 500 — including Nasdaq, Amazon, Cisco, Vodafone, Atlassian, and Lucid Motors — Workato turns over 14,000 enterprise systems AI needs to act on into one governed execution layer. For more information, visit workato.com.
MLCommons® is an open engineering consortium that produces MLPerf®, a widely adopted suite of industry-standard benchmarks for measuring machine learning performance. MLPerf benchmarks are developed collaboratively by member organizations spanning industry and academia to provide fair, reproducible, and representative measurements of AI systems.