Zijie Xin (辛梓杰)

email: xinzijie@ruc.edu.cn

I am a second-year Ph.D. student in the AI & Media Computing Lab at the Renmin University of China, advised by Prof. Xirong Li.

I obtained my Bachelor’s degree with honors in the Top-notch Program (a class of 15 elite students selected from 400+) from Sichuan University in 2024, under the supervision of Prof. Qijun Zhao. I’ve interned at Tencent and KuaiShou.

My research primarily revolves around video understanding, VideoLLM/OmniLLM, and cross-modal retrieval, complemented by a broad curiosity in generative model, RAG, RL.

News

Mar 16, 2026	I was awarded the 2025 Outstanding Innovative Talents Cultivation Funded Programs of RUC.
Feb 21, 2026	Our one paper on Speech-Aware Video Representation for Video-Text Retrieval has been accepted to CVPR 2026! 🎉
Jul 7, 2025	I joined Tencent as a research internship on video understanding.
Jul 5, 2025	Our one paper on Ad-hoc Video Search has been accepted to ACMMM 2025! 🎉
Jun 26, 2025	Our two papers on Music Grounding by Short Video and Sketch Animation have been accepted to ICCV 2025! I’m proud to be the first author of MGSV. 🎉
Apr 20, 2025	We released SeekWorld, an open-source project exploring o3-like visual clue-tracking reasoning for geolocation. We open-sourced the SeekWorld dataset and an RL-trained model, SeekWorld-7B! 🌍
Mar 21, 2025	Our one paper on Text-based Person Search has been accepted to ICME 2025! Congratulations to Yuchuan! 🎉
Sep 7, 2024	I officially started my PhD at RUC under the supervision of Professor Xirong Li in the AIMC Lab. 👨‍🎓
Jun 28, 2024	I successfully graduated with my bachelor’s degree from SCU and been honored as an Outstanding Graduate of Sichuan University! 👨‍🎓
Feb 27, 2024	Our one paper on Multi-Grained Teaching Strategy for Efficient Text-to-Video Retrieval has been accepted to CVPR 2024! 🎉
Nov 29, 2023	I joined Kuaishou as a research internship on video-music retrieval.
Oct 8, 2023	After heading to Beijing and joining the AI & Media Computing Lab, I unofficially started my PhD adventure!
Apr 20, 2023	I joined GeWu-Lab as a short-term intern.

Selected Publications (full list)

* Equal Contribution | † Corresponding Author

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

Zijie Xin, Jie Yang†, Ruixiang Zhao, Tianyi Wang, Fengyun Rao, Jing LYU, and Xirong Li†

preprint arXiv:2605.20035, 2026

Abs Paper Code Website

Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window level. However, processing these dense non-textual tokens throughout the LLM incurs substantial computational overhead. Although training-free token selection can reduce this cost, existing methods either focus on visual-only inputs or prune om-LLM tokens only before the LLM with fixed per-modality ratios, failing to capture how cross-modal token importance evolves across layers. To address this limitation, we first analyze the layer-wise token dependency of om-LLMs. We find that visual and audio dependencies follow a block-wise pattern and gradually weaken with depth, indicating that many late-layer non-textual tokens become redundant after cross-modal fusion. Motivated by this observation, we propose SEATS, a training-free, stage-adaptive token selection method for efficient om-LLM inference. Before the LLM, SEATS removes spatiotemporal redundancy via attention-weighted diversity selection. Inside the LLM, it progressively prunes tokens across blocks and dynamically allocates the retention budget from temporal windows to modalities using query relevance scores. In late layers, it removes all remaining non-textual tokens once cross-modal fusion is complete. Experiments on Qwen2.5-Omni and Qwen3-Omni demonstrate that SEATS effectively improves inference efficiency. Retaining only 10% of visual and audio tokens, it achieves a 9.3\times FLOPs reduction and a 4.8\times prefill speedup while preserving 96.3% of the original performance.
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Ruixiang Zhao, Jie Yang†, Zijie Xin, Tianyi Wang, Fengyun Rao, Jing LYU, and Xirong Li†

preprint arXiv:2605.18577, 2026

Abs Paper Code Data Website

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchmarks fall short in three key aspects: they rely primarily on visual signals, adopt polling or fixed-timestamp protocols instead of true proactive evaluation, and cover only a limited range of tasks, preventing reliable assessment and differentiation of omni-proactive streaming models. We present OmniPro, the first benchmark to jointly evaluate omni-modal perception, proactive responding, and diverse video understanding tasks. It comprises 2,700 human-verified samples spanning 9 sub-tasks and 3 cognitive levels, covering 6 basic video understanding capabilities. Notably, 84% of samples require audio signals (speech or non-speech), and each sample is annotated with modality-isolation labels to enable fine-grained multimodal analysis. We further introduce a dual-mode evaluation protocol: \textitProbe mode assesses content understanding by querying the model before and after each ground-truth trigger, while \textitOnline mode evaluates full proactive ability by requiring models to autonomously decide when to respond in streaming input. Evaluating 11 representative models reveals three key findings: (1) audio provides consistent gains but with highly variable utilization across models, (2) performance degrades significantly over time, indicating limited long-horizon robustness, and (3) non-speech audio perception remains the weakest dimension.

ICCV

Music Grounding by Short Video

Zijie Xin, Minquan Wang, Jingyu Liu, Ye Ma, Quan Chen, Peng Jiang, and Xirong Li†

In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

Abs Paper Code Data Poster Website

Adding proper background music helps complete a short video to be shared. Previous work tackles the task by video-to-music retrieval (V2MR), aiming to find the most suitable music track from a collection to match the content of a given query video. In practice, however, music tracks are typically much longer than the query video, necessitating (manual) trimming of the retrieved music to a shorter segment that matches the video duration. In order to bridge the gap between the practical need for music moment localization and V2MR, we propose a new task termed Music Grounding by Short Video (MGSV). To tackle the new task, we introduce a new benchmark, MGSV-EC, which comprises a diverse set of 53k short videos associated with 35k different music moments from 4k unique music tracks. Furthermore, we develop a new baseline method, MaDe, which performs both video-to-music matching and music moment detection within a unified end-to-end deep network. Extensive experiments on MGSV-EC not only highlight the challenging nature of MGSV but also set MaDe as a strong baseline.
ACMMM

Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search

Fan Hu, Zijie Xin, and Xirong Li†

In Proceedings of the 33rd ACM international conference on Multimedia (ACMMM), 2025

Abs Paper Code

Ad-hoc Video Search (AVS) involves using a textual query to search for multiple relevant videos in a large collection of unlabeled short videos. The main challenge of AVS is the visual diversity of relevant videos. A simple query such as "Find shots of a man and a woman dancing together indoors" can span a multitude of environments, from brightly lit halls and shadowy bars to dance scenes in black-and-white animations. It is therefore essential to retrieve relevant videos as comprehensively as possible. Current solutions for the AVS task primarily fuse multiple features into one or more common spaces, yet overlook the need for diverse spaces. To fully exploit the expressive capability of individual features, we propose LPD, short for Learning Partially Decorrelated common spaces. LPD incorporates two key innovations: feature-specific common space construction and the de-correlation loss. Specifically, LPD learns a separate common space for each video and text feature, and employs de-correlation loss to diversify the ordering of negative samples across different spaces. To enhance the consistency of multi-space convergence, we designed an entropy-based fair multi-space triplet ranking loss. Extensive experiments on the TRECVID AVS benchmarks (2016-2023) justify the effectiveness of LPD. Moreover, diversity visualizations of LPD’s spaces highlight its ability to enhance result diversity.
CVPR

Holistic Features are almost Sufficient for Text-to-Video Retrieval

Kaibin Tian*, Ruixiang Zhao*, Zijie Xin, Bangxiang Lan, and Xirong Li†

In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

Abs Paper Code Poster

For text-to-video retrieval (T2VR), which aims to retrieve unlabeled videos by ad-hoc textual queries, CLIP-based methods currently lead the way. Compared to CLIP4Clip which is efficient and compact, state-of-the-art models tend to compute video-text similarity through fine-grained cross-modal feature interaction and matching, putting their scalability for large-scale T2VR applications into doubt. We propose TeachCLIP, enabling a CLIP4Clip based student network to learn from more advanced yet computationally intensive models. In order to create a learning channel to convey fine-grained cross-modal knowledge from a heavy model to the student, we add to CLIP4Clip a simple Attentional frame-Feature Aggregation (AFA) block, which by design adds no extra storage / computation overhead at the retrieval stage. Frame-text relevance scores calculated by the teacher network are used as soft labels to supervise the attentive weights produced by AFA. Extensive experiments on multiple public datasets justify the viability of the proposed method. TeachCLIP has the same efficiency and compactness as CLIP4Clip, yet has near-SOTA effectiveness.