TwelveLabs, a video understanding company specializing in AI-powered video intelligence, has announced the launch of Rodeo, its first application-layer product designed to bring AI agents directly into video production workflows. The platform enables creators to search, edit, and assemble video footage using natural language, allowing production teams to streamline content creation and reduce time spent manually reviewing video archives.
The launch marks an important expansion for TwelveLabs as the company moves beyond infrastructure and foundation models to deliver creator-focused applications powered by its video intelligence technology.
According to TwelveLabs, Rodeo introduces a new workflow model for video production by allowing creators to describe the content they need while the platform automatically searches and assembles relevant footage from large video libraries.
The company stated that Rodeo’s AI-powered contextual understanding enables the platform to identify not only what appears in a video, but also why the footage is contextually important within a scene or narrative.
This allows creators to quickly locate relevant clips without manually reviewing hours of archived footage.
TwelveLabs said the platform is designed to help creative professionals spend more time producing content and less time managing file searches and footage organization.
"Video is inherently a creative medium, so we wanted to deliver all of the foundation model power and innovation directly into creative workflows without any technical barriers."
Rodeo is positioned as a production assistant for creators working with large-scale video archives and complex content workflows.
The company said the platform is intended for:
According to TwelveLabs, Rodeo functions as an AI assistant capable of understanding an organization’s entire video library before surfacing the most relevant content within the workflow.
The company also noted that creators can repurpose existing footage more efficiently and accelerate story assembly through natural language-driven video search and editing capabilities.
Rodeo is powered by TwelveLabs’ video understanding models, including Marengo 3.0 and Pegasus 1.5.
According to TwelveLabs, Marengo 3.0 treats video as a dynamic, multimodal system capable of interpreting:
The company stated that the model enables video to be searched, navigated, and understood at scale rather than processed purely as isolated visual frames.
TwelveLabs also highlighted Pegasus 1.5, its long-form video understanding model designed to analyze videos up to one hour in length while maintaining low latency and high contextual accuracy.
The model generates descriptive text outputs by analyzing both visual and audio information within video content.
Together, Marengo 3.0 and Pegasus 1.5 provide the foundation for Rodeo’s ability to understand and organize video content using conversational prompts and contextual interpretation.
TwelveLabs stated that recent advancements in AI infrastructure and MCP servers have enabled AI agents to move beyond experimentation into real production environments.
According to the company, Rodeo bypasses traditional technical implementation requirements, allowing creative professionals to work directly with advanced AI-powered video intelligence tools without complex setup or infrastructure management.
"We've shown what is possible for the enterprise and businesses across a wide range of verticals at the model and infrastructure layer through Marengo and Pegasus. Now we're opening another door," said Jae Lee, CEO and founder of TwelveLabs.
"Video is inherently a creative medium, so we wanted to deliver all of the foundation model power and innovation directly into creative workflows without any technical barriers. We've accomplished this with Rodeo, empowering creators to move from raw footage to finished stories faster than ever before."
As AI-powered video understanding technology continues advancing, tools such as Rodeo are expected to play an increasingly important role in streamlining content production, media workflows, and creative collaboration across industries.
About TwelveLabs
TwelveLabs is the world's most powerful video intelligence platform, enabling machines to see, hear, and reason about video like humans do. From semantic search to automated summaries and multimodal embeddings, TwelveLabs empowers developers, enterprises, and creatives to unlock the full potential of video data across industries including media, advertising, government, security, and automotive.