Zijie Xin (辛梓杰)

email: xinzijie@ruc.edu.cn

I am a second-year Ph.D. student in the AI & Media Computing Lab at the Renmin University of China, advised by Prof. Xirong Li.

I obtained my Bachelor’s degree with honors in the Top-notch Program (a class of 15 elite students selected from 400+) from Sichuan University in 2024, under the supervision of Prof. Qijun Zhao. I’ve interned at Tencent and KuaiShou.

My research primarily revolves around multi-media learning, video understanding, cross-modal retrieval, and open-set recognition, complemented by a broad curiosity in generative model, RAG, RL, and LLM.

News

Jul 7, 2025	I joined Tencent as a research internship on video understanding.
Jul 5, 2025	Our one paper on Ad-hoc Video Search has been accepted to ACMMM 2025! 🎉
Jun 26, 2025	Our two papers on Music Grounding by Short Video and Sketch Animation have been accepted to ICCV 2025! I’m proud to be the first author of MGSV. 🎉
Mar 21, 2025	Our one paper on Text-based Person Search has been accepted to ICME 2025! Congratulations to Yuchuan! 🎉
Sep 7, 2024	I officially started my PhD at RUC under the supervision of Professor Xirong Li in the AIMC Lab. 👨‍🎓
Jun 28, 2024	I successfully graduated with my bachelor’s degree from SCU and been honored as an Outstanding Graduate of Sichuan University! 👨‍🎓
Feb 27, 2024	Our one paper on a Multi-Grained Teaching Strategy for Efficient Text-to-Video Retrieval has been accepted to CVPR 2024! 🎉
Nov 29, 2023	I joined Kuaishou as a research internship on video-music retrieval.
Oct 8, 2023	After heading to Beijing and joining the AI & Media Computing Lab, I unofficially started my PhD adventure!
Apr 20, 2023	I joined GeWu-Lab as a short-term intern.

Publications

* Equal Contribution | † Corresponding Author

ICCV

Music Grounding by Short Video

Zijie Xin, Minquan Wang, Jingyu Liu, Ye Ma, Quan Chen, Peng Jiang, and Xirong Li†

In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

Abs Paper Code Data Poster Website

Adding proper background music helps complete a short video to be shared. Previous work tackles the task by video-to-music retrieval (V2MR), aiming to find the most suitable music track from a collection to match the content of a given query video. In practice, however, music tracks are typically much longer than the query video, necessitating (manual) trimming of the retrieved music to a shorter segment that matches the video duration. In order to bridge the gap between the practical need for music moment localization and V2MR, we propose a new task termed Music Grounding by Short Video (MGSV). To tackle the new task, we introduce a new benchmark, MGSV-EC, which comprises a diverse set of 53k short videos associated with 35k different music moments from 4k unique music tracks. Furthermore, we develop a new baseline method, MaDe, which performs both video-to-music matching and music moment detection within a unified end-to-end deep network. Extensive experiments on MGSV-EC not only highlight the challenging nature of MGSV but also set MaDe as a strong baseline.
CVPR

Holistic Features are almost Sufficient for Text-to-Video Retrieval

Kaibin Tian*, Ruixiang Zhao*, Zijie Xin, Bangxiang Lan, and Xirong Li†

In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

Abs Paper Code Poster

For text-to-video retrieval (T2VR), which aims to retrieve unlabeled videos by ad-hoc textual queries, CLIP-based methods currently lead the way. Compared to CLIP4Clip which is efficient and compact, state-of-the-art models tend to compute video-text similarity through fine-grained cross-modal feature interaction and matching, putting their scalability for large-scale T2VR applications into doubt. We propose TeachCLIP, enabling a CLIP4Clip based student network to learn from more advanced yet computationally intensive models. In order to create a learning channel to convey fine-grained cross-modal knowledge from a heavy model to the student, we add to CLIP4Clip a simple Attentional frame-Feature Aggregation (AFA) block, which by design adds no extra storage / computation overhead at the retrieval stage. Frame-text relevance scores calculated by the teacher network are used as soft labels to supervise the attentive weights produced by AFA. Extensive experiments on multiple public datasets justify the viability of the proposed method. TeachCLIP has the same efficiency and compactness as CLIP4Clip, yet has near-SOTA effectiveness.
ACMMM

Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search

Fan Hu, Zijie Xin, and Xirong Li†

In Proceedings of the 33rd ACM international conference on Multimedia (ACMMM), 2025

Abs Paper Code

Ad-hoc Video Search (AVS) involves using a textual query to search for multiple relevant videos in a large collection of unlabeled short videos. The main challenge of AVS is the visual diversity of relevant videos. A simple query such as "Find shots of a man and a woman dancing together indoors" can span a multitude of environments, from brightly lit halls and shadowy bars to dance scenes in black-and-white animations. It is therefore essential to retrieve relevant videos as comprehensively as possible. Current solutions for the AVS task primarily fuse multiple features into one or more common spaces, yet overlook the need for diverse spaces. To fully exploit the expressive capability of individual features, we propose LPD, short for Learning Partially Decorrelated common spaces. LPD incorporates two key innovations: feature-specific common space construction and the de-correlation loss. Specifically, LPD learns a separate common space for each video and text feature, and employs de-correlation loss to diversify the ordering of negative samples across different spaces. To enhance the consistency of multi-space convergence, we designed an entropy-based fair multi-space triplet ranking loss. Extensive experiments on the TRECVID AVS benchmarks (2016-2023) justify the effectiveness of LPD. Moreover, diversity visualizations of LPD’s spaces highlight its ability to enhance result diversity.
ICCV

Multi-Object Sketch Animation by Scene Decomposition and Motion Planning

Jingyu Liu, Zijie Xin, Yuhan Fu, Ruixiang Zhao, Bangxiang Lan, and Xirong Li†

In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

Abs Paper Code Poster Website

Sketch animation, which brings static sketches to life by generating dynamic video sequences, has found widespread applications in GIF design, cartoon production, and daily entertainment. While current sketch animation methods perform well in single-object sketch animation, they struggle in multi-object scenarios. By analyzing their failures, we summarize two challenges of transitioning from single-object to multi-object sketch animation: object-aware motion modeling and complex motion optimization. For multi-object sketch animation, we propose MoSketch based on iterative optimization through Score Distillation Sampling (SDS), without any other data for training. We propose four modules: LLM-based scene decomposition, LLM-based motion planning, motion refinement network and compositional SDS, to tackle the two challenges in a divide-and-conquer strategy. Extensive qualitative and quantitative experiments demonstrate the superiority of our method over existing sketch animation approaches. MoSketch takes a pioneering step towards multi-object sketch animation, opening new avenues for future research and applications.
ICME

DAPL: Integration of Positive and Negative Descriptions in Text-Based Person Search

Yuchuan Deng, Zhanpeng Hu, Zijie Xin, Chuang Deng, and Qijun Zhao†

In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), 2025

Abs Paper

Text-based person search (TBPS) aims to retrieve specific images of individuals from large datasets using textual descriptions. Existing TBPS methods focus primarily on identifying explicit positive attributes, often neglecting the critical role of negative descriptions. This oversight can lead to false positives, where images that should be excluded based on negative descriptions are incorrectly included, due to partial alignment with the positive criteria. To address this limitation, we propose the Dual Attribute Prompt Learning (DAPL) framework, which incorporates both positive and negative descriptions to improve the interpretative accuracy of vision-language models in TBPS tasks. DAPL combines Dual Image-Attribute Contrastive (DIAC) learning with Sensitive Image-Attribute Matching (SIAM) learning to enhance the detection of previously unseen attributes. Furthermore, to achieve a balance between coarse and finegrained alignment of visual and textual embeddings, we introduce the Dynamic Token-wise Similarity (DTS) loss. This loss function refines the representation of both matching and non-matching descriptions at the token level, providing more precise and adaptable similarity assessments, and ultimately improving the accuracy of the matching process. Empirical results demonstrate that DAPL outperforms state-of-the-art methods, enhancing both precision and robustness in TBPS tasks.