AI Video · 29 Sep 2026 · 17:55 CEST
How Condé Nast built multimodal video discovery with Amazon Bedrock

Publisher preview · OZZZER analysis pending editorial review.
PUBLISHER ARTICLE PREVIEW
From the original article
Condé Nast’s editorial teams had no fast way to do multimodal video discovery. They were spending an average of 250 minutes per content discovery task, manually scrubbing through a library of more than 140,000 videos. They relied on titles and descriptions to find relevant clips. In a media environment where speed-to-market directly determines revenue capture, this process created measurable operational drag across brands such as Vogue, GQ, Vanity Fair, and Wired.
The core problem was structural: Existing search tools can’t look inside video content. Teams depended on institutional knowledge to locate assets, creating single points of failure when specific individuals were unavailable. Meanwhile, underutilized content sat in the archive undiscoverable because no keyword in a title or description connected it to the queries editors were actually running.
To solve this, Condé Nast partnered with the AWS Generative AI Innovation Center (GenAIIC) to build an AI-powered multimodal video discovery solution. Built on Amazon Bedrock and Amazon OpenSearch Service, the solution runs intent-based semantic search across video transcripts, visual elements, and audio. The team selected the TwelveLabs Marengo embedding model for its native ability to jointly encode visual, audio, and transcript signals.
Marengo powers all five capabilities described in the following section. This reduced discovery time from 250 minutes to under 2 minutes per task.
In this post, we describe the architecture, explain why we chose specific technology choices, and share the business outcomes the solution delivered.
When the team scoped the problem, two constraints shaped the solution design.
First, keyword search was fundamentally insufficient. Editorial teams don’t search for “yoga_tutorial_march_2024.mp4.” Instead, they search for “beginner yoga content with calming backgrounds” or “behind-the-scenes fashion week moments.” The search layer needed to understand intent, not match strings. This pointed directly to vector embeddings that capture semantic meaning across modalities (visual, audio, and transcript).
Human-authored metadata alone could not provide that level of understanding.
Second, the video library was large enough (over 140,000 videos) that any solution needed to separate the expensive, compute-heavy work of generating embeddings
Source
AWS AI · 29 Sep 2026 · 17:55 CEST
Open the original at AWS AI ↗