This is your work, valued
Ph.D. candidate. Used to be young, sometimes simple, always naive.
LLaRA. [ICLR'25] LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
★ 228naive-scrcpy-client. A naive client of Scrcpy in Python.
★ 77crossway_diffusion. [ICRA'24] Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
★ 72open_x_pytorch_dataloader. An unofficial pytorch dataloader for Open X-Embodiment Datasets https://github.com/google-deepmind/open_x_embodiment
★ 25qtCyberDIP. CyberDIP driver for windows in C++ 11. 按C++ 11标准编写的CyberDIP在Windows环境下的配套驱动。
★ 14naive-android-ssr. Naive Android Screen Stream Reader, this project decodes screenrecord stream from an Android device to OpenCV. Written in Python
★ 13RemoteWifiSpyTank. Hacking the Wifi Spy Tank YD-211S
★ 12elo-rainbow. [NeurIPS'22] Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?
★ 4elo-sac. [NeurIPS'22] Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?
★ 4naive-rl. Naive implementations of CSE525 HWs
★ 2fallguys-env. Python
★ 2motionEstimation. Code for <A Point and Line Features Based Method for Disturbed Surface Motion Estimation>
★ 2sim-evals. A simulation evaluation platform for DROID
★ 1simple-raspi-car. A simple smart car project based on raspberry pi
★ 1pybullet-gym. Open-source implementations of OpenAI Gym MuJoCo environments for use with the OpenAI Gym Reinforcement Learning Research Platform.
★ 1generative-models. Generative Models by Stability AI
★ 1DAWN. HTML
★ 1pyCyberCar. A driver for Raspberry PI 3B+ Car.
★ 1visualTriangulation. A triangulation demo with detection and visualization
★ 1KanColleRecorder. 一个坎口垒玩家的自言自语 / A record for my journey of KanColle.
★ 1FlowWAM. Official repo of "FlowWAM: Optical Flow as a Unified Action Representation for World Action Models"
★ 39REGEN. Official Codebase for "World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays"
★ 14academic-research-skills-codex. Codex-native Academic Research Skills suite for human-in-the-loop academic research workflows
★ 7.4kpaperwise. Deep-reading pipeline for research papers — LLM-powered reports, vector KB, and knowledge graph.
★ 87tauri-terminal.
★ 1FluxVLA. An all-in-one VLA engineering platform for embodied AI — from data to real-robot deployment.
★ 578zvec. A lightweight, lightning-fast, in-process vector database
★ 15kFOFPred. Python
★ 39TCO. Learning 3D Reconstruction with Priors in Test Time
★ 5claude-howto. A visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
★ 41kOpenWorldLib. Unified Codebase for Advanced World Models.
★ 849ai-agents-for-beginners. 18 Lessons to Get Started Building AI Agents
★ 71klearn-claude-code. Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
★ 73kKohakuRAG. Simple Hierarchical RAG Framework
★ 91UltraDexGrasp. [ICRA 2026] UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data
★ 84filebrowser. 📂 Web File Browser
★ 7.6kvidi. The official repo for "Vidi: Large Multimodal Models for Video Understanding and Editing"
★ 645claude-code-my-workflow. A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols.
★ 2awesome-vla-wam. A Curated List of Vision-Language-Action (VLA) and World Action Models (WAM) Research and Beyond
★ 881sd-webui-all-in-one. 支持安装、下载、启动和管理多种 AI WebUI / 训练工具的一体化项目,提供 Installer、整合包下载器、Launcher、CLI 与 Colab / Kaggle Notebook。
★ 239qwen-tts-webui. 基于 Qwen3 TTS 的 WebUI
★ 315claude-code-skills. Professional Claude Code skills marketplace featuring production-ready skills for enhanced development workflows.
★ 1.3kani2xcur. 一个功能强大的命令行工具,用于在 Windows 和 Linux 平台上发现、转换、安装和管理鼠标指针主题。它支持双向转换,可将 Linux 光标主题 (XCursor) 转为 Windows 格式 (.cur/.ani),亦可将 Windows 主题转为 Linux 格式,并提供安装、应用和卸载鼠标主题的全套管理功能。
★ 31psh2bat. 将 PowerShell 脚本转换为 Bat 脚本
★ 1system_prompts_leaks. Extracted system prompts from Anthropic - Claude Fable 5, Opus 5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-5.6-Sol, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.
★ 61ksceniris. A fast procedural scene generation framework
★ 28sam-3d-objects. SAM 3D Objects
★ 7.2kDiffSynth-Studio. Enjoy the magic of Diffusion models!
★ 13kDepth-Anything-3. Depth Anything 3
★ 6kLLaVA_1005_new. Python
★ 2LBMamba. [TMLR] LBMamba: Locally Bi-directional Mamba
★ 8Dita. ICCV2025
★ 171InternVLA-M1. InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
★ 418VisCoP. [ECCV 2026] VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
★ 12Ouroboros. Official Repository for Ouroboros - ICCV 2025
★ 26UrbanSparse. Python
★ 3PixCell. A generative foundation model for digital histopathology images
★ 34index-tts. An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
★ 22kwhole_body_tracking. BeyondMimic Motion Tracking Code
★ 45HDM. Home Made Diffusion Models
★ 197GMR. [ICRA 2026] GMR: General Motion Retargeting. Retarget human motions into diverse humanoid robots in real time on CPU. Retargeter for TWIST.
★ 2.5kdinov3. Reference PyTorch implementation and models for DINOv3
★ 11kASAP. [RSS 2025] "ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills"
★ 2.1kexif-photo-blog. Photo blog, reporting 🤓 EXIF camera details (aperture, shutter speed, ISO) for each image.
★ 1.7kRynnVLA-001. [ICRA 2026] RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
★ 303OmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kPBHC. Official Implementation of "KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills"
★ 1kRynnVLA-002. RynnVLA-002: A Unified Vision-Language-Action and World Model
★ 1.1kskland-daily-attendance-shell. 纯 Shell 实现的森空岛每日签到
★ 22Franca. Official code of Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
★ 278LTX-Video. Official repository for LTX-Video
★ 11kaugmax. Efficiently Composable Data Augmentation on the GPU with Jax
★ 42Lightwheel-simready-asset. Open-source 3D digital assets for simulation and robot training. Free for non-commercial use.
★ 171NER_With_LLMs.
★ 1aloha_sim. A collection of tabletop tasks in Mujoco
★ 316sim-evals. A simulation evaluation platform for DROID
★ 235json_repair. Repair malformed JSON from LLMs, APIs, logs, and user input in Python.
★ 5.1kLLaVA-VLA. LLaVA-VLA: A Simple Yet Powerful Vision-Language-Action Model [ICRA 2026]
★ 205Hunyuan3D-2.1. From Images to High-Fidelity 3D Assets with Production-Ready PBR Material
★ 3.8kmolmo. Code for the Molmo Vision-Language Model
★ 923vjepa2. PyTorch code and models for VJEPA2 self-supervised learning from video.
★ 4.4kdroid_policy_learning. DROID Policy Learning and Evaluation
★ 289CSR_Adaptive_Rep. Official Code for Paper: Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
★ 141P4L. Python
★ 1Qwen-VL-Series-Finetune. An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1.9kBunnyVisionPro. Bimanual Dexterous Teleoperation with Real-Time Retargeting using VisionPro
★ 355VisionProTeleop. VisionOS App + Python Library to stream hand tracking data from Vision Pro, video/audio stream to Vision Pro.
★ 800CSR. Official Repository for CSR - ICML 2025 Oral
★ 20OneTwoVLA. Official implementation of "OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning"
★ 236awesome-embodied-vla-va-vln. A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
★ 3.4kimmich. High performance self-hosted photo and video management solution.
★ 109kLangToMo. [WIP] Code for LangToMo
★ 21Awesome-VLA-Robotics. A comprehensive list of excellent research papers, models, datasets, and other resources on Vision-Language-Action (VLA) models in robotics.
★ 488LLaRA. [ICLR'25] LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
★ 228mvu. 🤖 [ICLR'25] Multimodal Video Understanding Framework (MVU)
★ 58MM-EUREKA. MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
★ 771PALM-E. Implementation of "PaLM-E: An Embodied Multimodal Language Model"
★ 339embodied_gaussians. Python
★ 226fast-constrained-sampling. Code for 'Fast constrained sampling in pre-trained diffusion models'
★ 13RoboVerse. RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
★ 1.8kRAFT. Python
★ 4.1ktouch-spot. Based on E2E-spot
★ 1ReCamMaster. [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
★ 1.8kPhysTwin. [ICCV 2025] PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos
★ 435Qwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
★ 4.1kawesome-rl-for-legged-locomotion. A curated list of awesome material on legged robot locomotion using reinforcement learning (RL) and sim-to-real techniques.
★ 200few-shot-scanpath. Python
★ 16EgoExo4ADL. [ECCV 2026] From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision
★ 12ZoomLDM. CVPR 2025: 'ZoomLDM: Latent Diffusion Model for multi-scale image generation'
★ 34TopoCellGen. Python
★ 28Spark-TTS. Spark-TTS Inference Code
★ 11kEmbodied-AI-Paper-TopConf. [Actively Maintained🔥] A list of Embodied AI papers accepted by top conferences (ICLR, NeurIPS, ICML, RSS, CoRL, ICRA, IROS, CVPR, ICCV, ECCV).
★ 729language-table. Suite of human-collected datasets and a multi-task continuous control benchmark for open vocabulary visuolinguomotor learning.
★ 363ERQA. Embodied Reasoning Question Answer (ERQA) Benchmark
★ 282Awesome-LLM-Post-training. Awesome Reasoning LLM Tutorial/Survey/Guide
★ 2.5kSpace-awareVLM. Python
★ 14Emma-X. Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
★ 84fca. Repository of the paper Few-Class Arena in ICLR 2025🚀
★ 6gpu-fryer. Where GPUs get cooked 👩🍳🔥
★ 401Magma. [CVPR 2025] Magma: A Foundation Model for Multimodal AI Agents
★ 1.9kAgiBot-World. [IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
★ 3.1kVPN. Pose driven attention mechanism
★ 44ear_challenge. ViFi-CLIP Baseline for Elderly Action Recognition (EAR) Challenge 2025
★ 1VAPO. Python
★ 5Ultralight-Digital-Human. 一个超轻量级、可以在移动端实时运行的数字人模型
★ 2.6kLAPA. [ICLR 2025] LAPA: Latent Action Pretraining from Videos
★ 562TRACE. [ICLR 2025] TRACE: Temporal Grounding Video LLM via Casual Event Modeling
★ 157MM-Eureka-V0. MM-Eureka V0 also called R1-Multimodal-Journey, Latest version is in MM-Eureka
★ 325openpi. Python
★ 13kRoboVLMs. Python
★ 474YOLO-World. [CVPR 2024] Real-Time Open-Vocabulary Object Detection
★ 6.5kTopoDiffusionNet. This repository contains the implementation for our work "TopoDiffusionNet: A Topology-aware Diffusion Model", accepted to ICLR 2025.
★ 28cat-catch. 猫抓 浏览器资源嗅探扩展 / cat-catch Browser Resource Sniffing Extension
★ 21kDeepSeek-R1.
★ 92kGo-with-the-Flow. The official implementation of CVPR'25 Oral paper "Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise"
★ 1.1kgowiththeflowpaper.github.io. Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise
★ 42LatentCRF. Code for LatentCRF
★ 9LLAVIDAL. This is the offical repository of LLAVIDAL
★ 25NapCatQQ. Modern protocol-side framework based on NTQQ
★ 10kDaVinciResolve-API-Docs. Ruby
★ 71Auralis. A Fast TTS Engine
★ 626A3VLM. [CoRL2024] Official repo of `A3VLM: Actionable Articulation-Aware Vision Language Model`
★ 122AdaCache. Code for our ICCV 2025 paper "Adaptive Caching for Faster Video Generation with Diffusion Transformers"
★ 172DoRA_ICLR24. This repo contains the official implementation of ICLR 2024 paper "Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video""
★ 96Vision-Pro-Head-Hand-Tracking-Demo. Starter project for the Apple Vision Pro that tracks each hand joint your head. Tested on visionOS 1.0 and 1.1
★ 38IMUPoser. Code for IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and Earbuds
★ 2SPA. [ICLR 2025] SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
★ 177finetune-Qwen2-VL. Python
★ 395locvlm. Unofficial Implementation of "Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs"
★ 6dakpn. Python
★ 2theia. Theia: Distilling Diverse Vision Foundation Models for Robot Learning
★ 277TopoSemiSeg. The official implementation of "Semi-supervised Segmentation of Histopathology Images with Noise-Aware Topological Consistency".
★ 14rp. This is a python library. Install with "python3 -m pip install rp" then run with "python3 -m rp" or just "rp". Requires python≥3.5
★ 13sapiens. High-resolution models for human tasks.
★ 5.4krobocasa. RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
★ 1.6kLS-sample-quality. Assessing Sample Quality via the Latent Space of Generative Models (ECCV 2024)
★ 6JEAN.
★ 2Swin-MAE. Pytorch implementation of Swin MAE https://arxiv.org/abs/2212.13805
★ 107RoboFlamingo. Code for RoboFlamingo
★ 437MAGICK. Giant open source dataset of 140,000 captioned RGBA images!
★ 11botorch. Bayesian optimization in PyTorch
★ 3.6kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kgradient-checkpointing. Make huge neural nets fit in memory
★ 2.8kLIBERO. Benchmarking Knowledge Transfer in Lifelong Robot Learning
★ 2.1kcalvin. CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
★ 963anygrasp_sdk. Python
★ 952awesome-world-models-manipulation. Awesome world models for manipulation
★ 55mixture-of-experts. PyTorch Re-Implementation of "The Sparsely-Gated Mixture-of-Experts Layer" by Noam Shazeer et al. https://arxiv.org/abs/1701.06538
★ 1.2ktheia. Fork of official theia github.com/bdaiinstitute/theia
★ 1luckyrobots. Python SDK for LuckyEngine robotics simulation
★ 278CitationMap. A simple pip-installable Python tool to generate your HTML citation world map from your Google Scholar ID.
★ 722RADIO. Official repository for "AM-RADIO: Reduce All Domains Into One"
★ 1.9ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20khelping_hand_for_egocentric_videos. Implementation of paper 'Helping Hands: An Object-Aware Ego-Centric Video Recognition Model'
★ 33gaussian-splatting. Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
★ 23kInternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kVILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8kminbpe. Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.
★ 11kroboagent. Repository to train and evaluate RoboAgent
★ 376openvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
★ 6.7kgenima. Official Code Repo for GENIMA
★ 77Latte. [TMLR 2025] Latte: Latent Diffusion Transformer for Video Generation.
★ 1.9kCoPT. [ECCV 2024] Official Implementation of CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
★ 10NewtonRewards.
★ 16PAAB.
★ 1UnSAM. [NeurIPS 2024] Code release for "Segment Anything without Supervision"
★ 503SimplerEnv. Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo) in simulation under common setups (e.g., Google Robot, WidowX+Bridge) (CoRL 2024)
★ 1.1ktoolkit. Shell
★ 1.3kweb2code. Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
★ 103SS-cVAE. Code Repository for Semi-Supervised Contrastive VAE for Disentanglement of Digital Pathology Images
★ 5Fibottention. Official Repository of "Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads"
★ 17cambrian. Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2kVista. [NeurIPS 2024] A Generalizable World Model for Autonomous Driving
★ 889LVNet. [Main Conference @ EACL'26] [Workshop @ NeurIPS'24] 🎞️ LVNet.
★ 44PianoMotion10M. Code release for PianoMotion10M
★ 117finetune-detr. Fine-tune Facebook's DETR (DEtection TRansformer) on Colaboratory.
★ 154matryoshka-mm. Matryoshka Multimodal Models
★ 123nonebot_plugin_game_collection. nonebot 游戏合集
★ 52ml-visuals. 🎨 ML Visuals contains figures and templates which you can reuse and customize to improve your scientific writing.
★ 17klerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kVideo-LLaVA. 【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
★ 3.5kViP-LLaVA. [CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
★ 339RoLD. [MMM 2025 Best Paper] RoLD: Robot Latent Diffusion for Multi-Task Policy Modeling
★ 24llama3. The official Meta Llama 3 GitHub site
★ 29kllama-cookbook. Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
★ 19kARTrack. PyTorch implementation of paper "ARTrack" and "ARTrackV2"
★ 317Hotshot-XL. ✨ Hotshot-XL: State-of-the-art AI text-to-GIF model trained to work alongside Stable Diffusion XL
★ 1.1kAwesome-Video-Diffusion-Models. [CSUR] A Survey on Video Diffusion Models
★ 2.3kVAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7klifelong-memory. Code for LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
★ 33VLM_survey. Collection of AWESOME vision-language models for vision tasks
★ 3.1kComp4D. "Comp4D: Compositional 4D Scene Generation", Dejia Xu*, Hanwen Liang*, Neel P. Bhatt, Hezhen Hu, Hanxue Liang, Konstantinos N. Plataniotis, and Zhangyang Wang
★ 78Grounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kllm-colosseum. Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM
★ 1.5kopen_x_pytorch_dataloader. An unofficial pytorch dataloader for Open X-Embodiment Datasets https://github.com/google-deepmind/open_x_embodiment
★ 25LLMs-from-scratch. Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
★ 100kStARformer. [ECCV2022] [T-PAMI] StARformer: Transformer with State-Action-Reward Representations.
★ 97sugarl. Code for NeurIPS 2023 paper "Active Vision Reinforcement Learning with Limited Visual Observability"
★ 56VIMABench. Official Task Suite Implementation of ICML'23 Paper "VIMA: General Robot Manipulation with Multimodal Prompts"
★ 327LangRepo. Code for our ACL 2025 paper "Language Repository for Long Video Understanding"
★ 36