Versos AI launches NVIDIA NeMo agents for video training data curation
Versos AI, a Canadian startup positioning itself as infrastructure for premium video licensing, has launched an agentic curation capability built on NVIDIA NeMo that allows AI teams to describe a required training dataset in natural language and have the system locate, grade and assemble matching footage automatically. The company will demonstrate the feature at IBC2026 in Amsterdam from 11 to 14 September.
The release addresses a recognised friction point in AI model development. Assembling a bespoke video dataset typically requires teams to translate a detailed content specification into multiple database searches, manually inspect results, and cross-check individual clips against technical, linguistic and licensing criteria before packaging the output for delivery. Versos says its new workflow compresses that process into a single conversational prompt.
How the technology stack fits together
The agentic layer sits on top of Versos' existing Video Library Intelligence Platform, which indexes footage at scene and frame level. Three components underpin the new workflow. NVIDIA's CUDA Toolkit accelerates the inference that turns raw footage into structured metadata. NVIDIA Nemotron Ultra runs the agents responsible for interpreting buyer requests and directing discovery and grading of candidate clips. LangChain provides the orchestration scaffolding, coordinating the multi-step sequence from natural language input through search, metadata evaluation and dataset assembly.
Versos says the architecture is deliberately model-agnostic: Nemotron Ultra handles agent execution today, but the same workflow can be pointed at alternative open-weight or frontier models where a customer's cost, latency or deployment constraints demand a different choice.
Chris Keevill, co-founder and chief executive of Versos AI, said the aim is parity of experience between AI builders and the users of the products they create. "The people building AI deserve to benefit from it just like everyone else," he said, "and our agentic interface makes that happen by streamlining the process for curating training data."
Rights provenance is presented as a core design constraint rather than an afterthought. The platform restricts discovery to footage that has already been prepared for controlled AI licensing and carries documented ownership and permissions metadata, which Versos says shortens the path for AI teams to obtain datasets with cleared rights.
Market context and competitive positioning
The market for licensed AI training data has expanded sharply as model developers face growing pressure from rights holders and regulators to demonstrate provenance for the content used in training. Several high-profile lawsuits against foundation model developers in the United States and Europe have accelerated demand for datasets where permissions are documented upfront.
Versos competes in a space that includes a mix of established stock-media platforms extending into AI licensing, specialist data curation providers, and in-house data operations teams at larger AI labs. The natural-language interface lowers the technical bar for procurement teams without video-database expertise, which may widen Versos' addressable market beyond pure engineering buyers.
The company did not disclose pricing, the size of its indexed library, the number of content-owner partners, or any named AI lab customers in its release. Revenue model and scale remain opaque, making independent assessment of the platform's commercial traction difficult at this stage.
From a regulatory standpoint, the EU AI Act's forthcoming transparency obligations for general-purpose AI model training are likely to increase scrutiny of dataset provenance documentation across the industry. Versos' rights-centric architecture is positioned to align with those requirements, though the company has not cited specific compliance certifications.