Skip to main content Scroll Top

Industrial Manufacturer · Sales Enablement & Training

PROBLEM

A manufacturing subsidiary within a large holding company had accumulated an extensive internal library of product demonstrations, machine walk-throughs, and training videos. The content captured years of useful product knowledge, but finding a specific moment inside that library was difficult.

Experienced salespeople often remembered that a video demonstrated a particular machine capability, material, or operating sequence, but locating the right video and scrubbing to the relevant moment during a customer conversation was rarely practical.

New salespeople faced the opposite problem. The same archive contained valuable training material, but much of it was long-form and organized by video rather than by question, feature, or topic. Learning about one machine capability could require searching through multiple recordings and manually reviewing lengthy sections of footage.

The company already had the knowledge it needed. The problem was that much of that knowledge was effectively locked inside video and difficult to retrieve at the moment it was needed.

OBSTACLES

  • Traditional keyword and transcript search could identify what was said, but many important machine demonstrations were primarily visual. An operator might change a fixture, feed a specific material, or perform a safety procedure without explicitly describing every action aloud.
  • Search results needed to identify the relevant moment within a video rather than simply return a filename. Finding the correct ten-second sequence inside a 20-minute demonstration was the actual retrieval problem.
  • Field use required fast retrieval. A search process that required several queries or manual video review would not be practical during a live sales conversation.
  • Training questions needed to remain grounded in the company’s own products and demonstrations rather than relying on generic AI knowledge about industrial machinery.
  • The underlying workflow combined multimodal video indexing, semantic retrieval, timestamps, transcripts, visual context, and AI reasoning, but the user experience still needed to remain simple enough for both experienced salespeople and new hires.

OUTCOME

We developed a shared AI video intelligence layer that transformed the company’s existing archive into a searchable resource for both sales enablement and employee training.

Using TwelveLabs’ multimodal video search, product demonstrations and training videos were indexed across their visual and audio content. Sales representatives could search in natural language using requests such as “show the press handling corrugated steel” or “find the safety interlock demonstration.”

Rather than returning only a full-length video, the application surfaced relevant moments with timestamps, surrounding context, and video previews. Users could jump directly to the portion of the demonstration that matched their question instead of manually searching through multiple recordings.

This formed the first use case, a sales-focused retrieval tool that made the archive practical to use during live customer conversations. A salesperson who remembered seeing a particular capability no longer needed to remember which file contained it or prepare the clip in advance.

The same indexed library also supported an interactive training experience. New employees could watch a training video and ask questions about what they were seeing, such as “What is happening at this point?”, “What does the operator adjust before startup?”, or “What safety step is being demonstrated here?”

Relevant video segments, transcripts, visual context, and metadata were passed into a multimodal AI reasoning layer using Gemini 2.5 Flash, allowing responses to remain grounded in the company’s actual training material. Broader questions could also retrieve related moments from elsewhere in the indexed archive rather than limiting the response to the video currently being viewed.

Videos were indexed before retrieval rather than analyzed from scratch for every request, keeping search responsive enough for practical field use. Application-level caching and stored metadata further reduced unnecessary processing for frequently accessed content.

Despite the complexity underneath, the interface remained centered around familiar actions: search, watch, jump to a relevant moment, and ask a question.

RESULTS

  • Years of product demonstrations and training footage became searchable through natural-language queries instead of relying on filenames, folders, or individual employee memory.
  • Timestamp-level retrieval allowed sales teams to locate specific machine demonstrations quickly enough to use the archive during customer conversations rather than only during meeting preparation.
  • Multimodal search improved retrieval for visually demonstrated actions and machine capabilities that could not be identified reliably through transcript search alone.
  • Interactive training allowed new hires to investigate unfamiliar features, terminology, and operating procedures while viewing the relevant source material instead of relying entirely on passive video playback.
  • AI-generated explanations remained tied to retrieved company video content, reducing reliance on unsupported general-purpose answers when discussing specific products and procedures.
  • The shared video intelligence layer created a reusable foundation for additional use cases such as technical support, marketing-content discovery, internal knowledge retrieval, and product documentation.