This is your work, valued
PhD student at Zhejiang University
Awesome-Video-Agent. A collection of awesome think with videos papers.
TVG-R1. [EMNLP 2025 Industry] Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
PAD. [ICLR 2025] Pad: Personalized alignment of llms at decoding-time
DiffPO. [ACL 2025]
OptimSyn. A framework for influence-guided synthetic data generation through rubric optimization and reinforcement learning.
science-taste-skills. Two-layer Codex-native science taste skills: a persona layer plus a reusable gate.
Awesome-Agent-Environments. Awesome Agent Environments
Qwen-CodePercept. [CVPR2026] CodePercept: Code-Grounded Visual STEM Perception for MLLM
review-prompts. AI review prompts
Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Ego-R1. [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
Siamese-Diffusion. [CVPR 2025] Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
LGGPT. [IJCV 2025] Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
SPECTRUM. [ACM MM 2025 Oral] Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification
Awesome-Generative-Models-for-OCR. [arXiv 25] OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
unary-feedback. Python
SynCamMaster. [ICLR'25] SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
ReCamMaster. [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
BiasAlert. A Plug-and-play Tool for Social Bias Detection in LLMs. [EMNLP' 24]
Elysium. [ECCV 2024] Elysium: Exploring Object-level Perception in Videos via MLLM
MT-R1-Zero. [EMNLP'25] Code for paper "MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning"
VideoMind. 🧠 VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning (ICLR 2026)
Video-R1. Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]
nano-mdm. Tiny re-implementation of MDM in style of LLaDA and nano-gpt speedrun
Awesome-Personalized-LLMs. The latest progress of Personalized Large Language Models (LLMs).
FairMT-bench. Python
samurai. Official repository of "SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory"
LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)