Skip to main navigation Skip to search Skip to main content

Localizing Events in Videos with Multimodal Queries

  • Gengyuan Zhang
  • , Mang Ling Ada Fok
  • , Jialu Ma
  • , Yan Xia
  • , Daniel Cremers
  • , Philip Torr
  • , Volker Tresp
  • , Jindong Gu
  • Ludwig-Maximilians-Universität München
  • Munich Center for Machine Learning
  • Technical University of Munich
  • University of Oxford

Research output: Contribution to journalConference articlepeer-review

2 Scopus citations

Abstract

Localizing events in videos based on semantic queries is a pivotal task in video understanding research and user-oriented applications like video search. Yet, current research predominantly relies on natural language queries (NLQs), overlooking the potential of using multimodal queries (MQs) that incorporate images to flexibly represent semantic queries, particularly when it is difficult to express non-verbal or unfamiliar concepts in words. To bridge this gap, we introduce ICQ, a new benchmark designed for localizing events in videos with MQs, alongside an evaluation dataset ICQ-Highlight. To adapt and reevaluate existing video localization models for this new task, we propose 3 Multimodal Query Adaptation methods and a novel Surrogate Fine-Tuning strategy, serving as strong baseline methods. ICQ systematically benchmarks 12 state-of-the-art backbone models, spanning from specialized video localization models to Video Large Language Models. Our extensive experiments highlight the high potential of using MQs in real-world applications.

Original languageEnglish
Pages (from-to)3339-3351
Number of pages13
JournalProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOIs
StatePublished - 2025
Event2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, United States
Duration: 11 Jun 202515 Jun 2025

Keywords

  • multimodal large language model
  • temporal grounding
  • video grounding
  • video understanding

Fingerprint

Dive into the research topics of 'Localizing Events in Videos with Multimodal Queries'. Together they form a unique fingerprint.

Cite this