This is your work, valued
SWE-Vision. Python
★ 172jan. Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
★ 44kEgoMemo. Official Repo of Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
★ 10Liger-Kernel. Efficient Triton Kernels for LLM Training
★ 6.5kViaRL. ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
★ 11Qwen-Agent. Agent framework and applications built upon Qwen>=3.0, featuring Function Calling, MCP, Code Interpreter, RAG, Chrome extension, etc.
★ 17kThinkingWithVideos. The official code of "Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning"
★ 102LongVT. [CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
★ 257Video-o3. [ICML 2026] Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
★ 134Video-CoM. Video-CoM: Interactive Video Reasoning via Chain of Manipulations
★ 22WorldVQA. Python
★ 120EasyR1. EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kFlashPortrait. [CVPR2026]We present FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving, infinite-length videos while achieving up to 6$\times$ acceleration in inference speed.
★ 480mllm-npu. mllm-npu: training multimodal large language models on Ascend NPUs
★ 95AISystem. AISystem 主要是指AI系统,包括AI芯片、AI编译器、AI推理和训练框架等AI全栈底层技术
★ 17kVLM-R1. Solve Visual Understanding with Reinforced VLMs
★ 6kLogic-RL. Reproduce R1 Zero on Logic Puzzle
★ 2.4kStableAnimator. [CVPR2025] We present StableAnimator, the first end-to-end ID-preserving video diffusion framework, which synthesizes high-quality videos without any post-processing, conditioned on a reference image and a sequence of poses.
★ 1.4kllm.c. LLM training in simple, raw C/CUDA
★ 31kMovieChat. [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
★ 706Awesome_Long_Form_Video_Understanding. Awesome papers & datasets specifically focused on long-term videos.
★ 381llm_interview_note. 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
★ 15kMotionFollower. [ICCV2025] MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
★ 245Open-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29k