New research from NTT Research's Physics of Artificial Intelligence (PAI) Lab, in collaboration with Harvard University's Center for Brain Science, challenges the assumption that simply adding more AI agents automatically improves enterprise AI performance. The research shows that multi-agent AI systems perform best within an optimal operating range, with collective accuracy peaking at approximately 16 agents in complex tasks before declining.
Multi-agent AI systems perform best within an optimal operating range; performance declines beyond a certain point due to communication challenges and competing viewpoints.
In the Flag Game experiment, collective accuracy peaked with approximately 16 AI agents before declining.
Diverse AI models outperform homogeneous teams, suggesting diversity improves collective intelligence.
Human guidance—clear instructions and communication strategies—had a greater impact on collective performance than simply increasing the number of AI agents.
Organizational design matters as much as scale in building effective AI agent teams.
The research was presented at the AI4Good Workshop at ICML.
As organizations rapidly deploy AI agents across customer service, software development, cybersecurity, scientific research and business operations, one question is becoming increasingly important: How should organizations structure AI agent teams to achieve the best results?
New research from Dr. Hidenori Tanaka and Elizabeth Pavlova from NTT Research's Physics of Artificial Intelligence (PAI) Lab, in collaboration with Harvard University's Center for Brain Science, challenges the assumption that simply adding more AI agents automatically improves enterprise AI performance. Instead, the research shows that multi-agent AI systems perform best within an optimal operating range. Beyond that range, additional AI agents can surface competing interpretations of the same evidence, causing groups to split into camps rather than converge, depending on the task.
The research demonstrates that multi-agent AI systems reach an optimal operating range. As AI agents collaborate, collective performance initially improves. Beyond that point, however, additional AI agents can make it more difficult for groups to communicate effectively, reach consensus and solve complex problems.
According to the research, in one experiment called the Flag Game, collective accuracy peaked with approximately 16 AI agents before declining. In this game, each agent sees only a small, randomly assigned piece of a hidden flag, and the group must communicate to figure out which country's flag it is. Like having too many cooks in the kitchen, simply adding more AI agents does not guarantee better results. Beyond an optimal operating range, communication becomes more difficult, competing viewpoints emerge and collective performance can decline.
In the flag game experiment, each AI agent receives only part of the available information and must communicate with other agents to identify the correct answer. The experiment illustrates how AI agents combine distributed knowledge, build consensus and solve complex problems—and how communication challenges emerge as AI organizations grow. Rather than studying individual AI models in isolation, the research applies principles from physics, mathematics and machine learning to understand how populations of AI agents organize knowledge, communicate and develop collective intelligence.
"As organizations begin deploying hundreds or even thousands of AI agents, one of the most important questions becomes how collective intelligence emerges from their interactions," said Dr. Hidenori Tanaka, Group Leader of the Physics of Artificial Intelligence (PAI) Lab at NTT Research and the Physics of Intelligence Program at Harvard University's Center for Brain Science. "Our research shows that simply adding more AI agents does not necessarily improve performance — just as hiring more people does not automatically make a company more effective. Communication becomes harder, and groups can split into competing camps. Organizations also need to consider how AI agents communicate, how they are structured and how humans design and guide these systems to achieve the best outcomes."
Key findings include:
Multi-agent AI systems perform best within an optimal operating range. Adding more AI agents does not necessarily improve collective performance and may reduce accuracy for certain tasks.
Organizational design matters as much as scale. How AI agents communicate, exchange information and collaborate significantly influences collective performance.
Diverse AI models outperform homogeneous teams. Researchers found that teams combining AI models with complementary strengths produced better results than teams composed of a single AI model, suggesting that diversity can improve collective intelligence.
Human direction shapes effective AI organizations. Better guidance from humans improves outcomes. Clear instructions and communication strategies had a greater impact on collective performance than simply increasing the number of AI agents.
The findings suggest organizations should focus less on deploying the largest possible AI workforce and more on designing effective AI organizations that balance scale, communication, organizational structure and model diversity. Success depends not simply on the number of AI agents, but on how they collaborate, how they are guided by humans and how they complement one another.
About NTT Research's Physics of Artificial Intelligence (PAI) Lab
NTT Research's Physics of Artificial Intelligence (PAI) Lab studies the fundamental principles of intelligence to better understand how intelligent systems learn, reason, communicate and collaborate. Through research spanning the physics of AI, neuroscience and AI, AI interpretability and multi-agent AI systems, the PAI Lab develops the scientific foundation for the next generation of trustworthy, scalable and collaborative artificial intelligence.
About NTT Research
NTT Research is the Silicon Valley research arm of NTT, one of the world's largest technology and business solutions providers. Founded in 2019, NTT Research invents the future of foundational science while accelerating its real-world impact across NTT's global ecosystem. From its headquarters in Sunnyvale, California, NTT Research brings together world-class scientists across four research pillars: the Physics & Informatics (PHI) Lab, the Cryptography & Information Security (CIS) Lab, the Medical & Health Informatics (MEI) Lab, and the Physics of Artificial Intelligence (PAI) Lab.