This is your work, valued
segment-anything-2-real-time. Run Segment Anything Model 2 on a live video stream
★ 593Tracking. 主要实现了基于Sort的MOT的Tracking模块
★ 18RM-2021. Robomaster-2021-XDU-IRobot-哨兵视觉代码
★ 8qwen-mcp-tool. TypeScript
★ 7IRobot-SYSU-CICR. C
★ 2gy920.github.io. CSS
★ 1segment-anything-2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 1GordenSuperPPTSkills. AI PPT赛道终结者,史上最最最强 PPT Skill!!! 使用GPT生成豪华的图片格式PPT,然后转换为完全可编辑的PPTX文件。
★ 1.7kopenworker. Python
★ 11kJoyAI-Image. JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
★ 2.2krobocasa. RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
★ 1.6khssd. Code repository for the Habitat Synthetic Scenes Dataset (HSSD) paper.
★ 120giga-world-1. A Roadmap to Build World Models for Robot Policy Evaluation
★ 866OmTrackVLA. Open & Reproducible Research for Tracking VLAs
★ 265ABot-World. Infinite Interactive World Rollout on a Single Desktop GPU
★ 1.3k3DGen-Playground. 3D generation made easy!
★ 458BEHAVIOR-1K. BEHAVIOR-1K: a platform for accelerating Embodied AI research. Join our Discord for support: https://discord.gg/bccR5vGFEx
★ 1.6kEmbodiedGen. Towards a Generative 3D World Engine for Embodied Intelligence
★ 592infinigen. Infinite Photorealistic Worlds using Procedural Generation
★ 7.2kOmniNavBench. [RSS 2026] Official code & data for "OmniNavBench: Beyond Isolation — A Unified Benchmark for General-Purpose Navigation"
★ 87lingbot-vla-v2. From Foundation to Application
★ 686MTU3D. Python
★ 266RLinf. RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
★ 4.4kArtiFixer. Python
★ 581daily_stock_analysis. LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
★ 60k40-questions. Questions that I ask myself at the end of each year and each decade.
★ 1.6kopen-design. 🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
★ 83kgs_playground. Jupyter Notebook
★ 453pi. AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
★ 81kcodexpro. Use ChatGPT Developer Mode as a local coding agent for your repo through MCP.
★ 1.5kFC-Vision. [RA-L'26] Online Visibility-Aware Replanning for Occlusion-Free Aerial Scanning
★ 13openpi. Python
★ 13kOmniNav. 【ICLR 2026】 Official implementation of [OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation]
★ 201FlyCo. Foundation-Model-Empowered and Prompt-Driven System for Open-World Aerial 3D Structure Scanning
★ 62VLM3. Official implementation of paper "VLM³: Vision Language Models Are Native 3D Learners".
★ 403rerun. Visualize, query, and stream to train on multimodal robotics data.
★ 11kTripoSplat. TripoSplat converts a single 2D image into high-quality and variable number of 3D Gaussians, developed by TripoAI.
★ 1.1ktau-0-wm. Python
★ 301VQASynth. Compose multimodal datasets 🎹
★ 584TriSplat. TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction
★ 347Qwen-VLA. The official repository of Qwen-VLA
★ 725Awesome-Memory-for-Robotics. Research reading list on memory for robotics
★ 156UrbanVerse. Scaling Urban Simulation - Infinite Physically-Plausible Urban Simulation = IsaacSim(Physically-Accurate Assets × Real-World City-Tour Layouts)
★ 61open-eqa. OpenEQA Embodied Question Answering in the Era of Foundation Models
★ 365AwareVLN. [CVPR 2026] AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
★ 71InfinityStar. [NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
★ 774le-wm. Official code base for LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
★ 4.2koh-my-codex. OmX - Oh My codeX: Your codex is not alone. Add hooks, agent teams, HUDs, and so much more.
★ 32kTideGS. Python
★ 157video-3d-reconstruction-gsplat. a streamlined pipeline for converting videos into 3D Gaussian Splatting models
★ 46vggt-omega. [CVPR 2026 Oral] VGGT Omega
★ 3.8kedge-fm-x. 边缘端大模型高效推理引擎
★ 31Warp-as-History. Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
★ 225realtime-vla-flash. Python
★ 93H-OmniStereo. H-OmniStereo: Zero-Shot Omnidirectional Stereo Matching with Heading-Aligned Normal Priors
★ 50splat-transform. CLI tool and library for 3D Gaussian splat processing and conversion
★ 1.3kAnySplat. [SIGGRAPH Asia 2025 (ACM TOG)] AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
★ 901Awesome-WAM. A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.
★ 1.2kvlash. Real-Time VLAs via Future-state-aware Asynchronous Inference.
★ 455hermes-desktop. Desktop Companion for Hermes Agent
★ 14ktaste-skill. Taste-Skill - gives your AI good taste. stops the AI from generating boring, generic slop
★ 70kstarVLA. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
★ 3.3kdexbotic. Dexbotic: Open-Source Vision-Language-Action Toolbox
★ 1.3kUMI-3D. UMI-3D SLAM and Data Processing Pipeline: https://umi-3d.github.io/
★ 260Awesome-VLA-Post-Training. A collection of vision-language-action model post-training methods.
★ 231VLA-Adapter. VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
★ 2.3kMotus. Official code of Motus: A Unified Latent Action World Model
★ 1.2kMixture-of-Transformers. Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models. TMLR 2025.
★ 281Supervisor-Skills. 将博导十年科研经验炼化为可直接调用的 AI 技能。从 Idea 构思到论文投稿,你的 AI 科研副导师。
★ 4.7kcraft-agents-oss. TypeScript
★ 7kInfraTech. 分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等
★ 3.3kNavGSim. C++
★ 21UrbanVideo-Bench.code. [ACL'25 Oral] Code for the paper "UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces"
★ 31HY-Embodied. HY-Embodied: Embodied Foundation Models for Real-World Agents
★ 841MemMachine. Universal memory layer for AI Agents. It provides scalable, extensible, and interoperable memory storage and retrieval to streamline AI agent state management for next-generation autonomous systems.
★ 3.3kHy3-preview. Hy3 preview (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency
★ 456efficientsam3. EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking.
★ 646cc-haha. Local-first cross-platform desktop workspace for Claude Code / agents: multi-agent, Git worktrees, code diffs, skill marketplace, multi-model, Computer Use, task-aware desktop pets, with WeChat, Feishu, DingTalk, Telegram, WhatsApp and H5 access.
★ 14kvla_foundry. Python
★ 420claude-desktop-debian. Claude Desktop for Linux
★ 5.3kAutoTeam. ChatGPT Team 账号自动轮转管理 - Codex 额度监控、自动换号、邮箱注册、CPA/Sub2API 认证同步
★ 1.3kandroid-reverse-engineering-skill. Claude Code skill to support Android app's reverse engineering
★ 6.6kAwesome-3D-Gaussian-Splatting-in-Robotics.
★ 118TokenGS. [CVPR'26] TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens
★ 241lingbot-map. A feed-forward 3D foundation model for reconstructing scenes from streaming data
★ 16khabitat-gs. [ECCV 2026] Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting
★ 258ABot-Claw. Python
★ 200ABot-Navigation. Python
★ 205Efficient-VLAs-Survey. 🔥This is a curated list of "A survey on Efficient Vision-Language Action Models" research. We will continue to maintain and update the repository, so follow us to keep up with the latest developments!!!
★ 171NavDP. [ICRA 2026] NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
★ 728paseo. Orchestrate multiple coding agents from desktop and mobile
★ 12kStepFlow-Duck. cpa+duck结合体
★ 180free-proxy-list. 🚀 Free HTTP, SOCKS4, & SOCKS5 proxy list * Updated every 5 minutes * and rotating proxy API (100+ countries)
★ 6.3kAeroVLA. [Accepted to ECCV 2026] Official repository for AeroVLA (formerly known as AerialVLA): A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
★ 122AgentVLN.
★ 180cap-x. A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
★ 681DuckMail. 优雅易用的临时邮箱客户端,支持 API 无鉴权调用与接入私有域名,一键获取专属的临时邮箱。
★ 833cli. The official Lark/飞书 CLI tool, maintained by the larksuite team — built for humans and AI Agents. Covers core business domains including Messenger, Docs, Base, Sheets, Calendar, Mail, Tasks, Meetings, and more, with 200+ commands and 20+ AI Agent Skills.
★ 16kcolleague-skill. 将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶
★ 21kdeer-flow. An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
★ 78kCiteScan. Scan the Hallucination Citation of Academic papers. Convert second-hand citation to official version
★ 233Fast-FoundationStereoPhysics. A standlone demo from ManiDreams: Fast-FoundationStereo + SAM2 + Newton, zero-shot real-time simulation-based world model
★ 141MetaGPT. 🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming
★ 70kcodex-autoresearch. Codex Autoresearch Skill — A self-directed iterative system for Codex that continuously cycles through: modify, verify, retain or discard, and repeat indefinitely. Inspired by Karpathy’s autoresearch concept.
★ 2kCarlaAir. CarlaAir: Fly Drones Inside a CARLA World!! A Unified Infrastructure for Air-Ground Embodied Intelligence
★ 1.1kAwesome-PhD-CV. Curated academic CV templates and guidelines for PhD students, researchers, and faculty job applicants.
★ 1.2kMemRoPE. [ECCV 2026] Official implementation of "MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens"
★ 46Attention-Residuals.
★ 3.4kdocker-easyconnect. 使深信服(Sangfor)开发的非自由的 VPN 软件 EasyConnect 和 aTrust 运行在 docker 或 podman 中,并作为网关和/或提供 socks5、http 代理服务
★ 5.4kUnrealClaude. Claude Code CLI integration for Unreal Engine 5.7 - Get AI coding assistance with built-in UE5.7 documentation context directly in the editor.
★ 873superpowers. An agentic skills framework & software development methodology that works.
★ 264ksekai-codebase. [NeurIPS 2025] Sekai: A Video Dataset towards World Exploration
★ 302metaurban. [ICLR 2025 Spotlight] MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility
★ 250hierarchical-3d-gaussians. Official implementation of the SIGGRAPH 2024 paper "A Hierarchical 3D Gaussian Representation for Real-Time Rendering of Very Large Datasets"
★ 1.5kOpenFly-Platform. C++
★ 350Research-Paper-Writing-Skills. Skill package for ML/CV/NLP paper writing, curated and adapted from Prof. Peng Sida's open notes for Codex, Claude Code, and Gemini.
★ 5.7kros-mcp-server. Connect AI models like Claude & GPT with robots using MCP and ROS.
★ 1.4kremodex. Remote Control for Codex.
★ 3.3kCodexDesktop-Rebuild. Codex Desktop App - Cross-platform Rebuild
★ 2.6kCLI-Anything. "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
★ 46kRoboClaw. RoboClaw is an Embodied AI Assistant.
★ 5243dgrut. Ray tracing and hybrid rasterization of Gaussian particles
★ 2.3kFriendlySplat. FriendlySplat is a user-friendly, open-source Gaussian Splatting toolkit, integrating SOTA features into a unified platform for training, pruning, meshing and segmentation.
★ 154codex-desktop-linux. Unofficial ChatGPT desktop app for Linux (formerly the Codex app), built locally from OpenAI’s official macOS app. Includes Chat, Work, and Codex. Packages for Debian/Ubuntu (.deb), Fedora/openSUSE (.rpm), Arch (pacman), Nix/NixOS, and AppImage, with Wayland and X11 support.
★ 3.2kCodex-App-Linux. Codex App (macOS) ported to Linux
★ 55OnFly. [IROS'26] OnFly: Onboard Zero-Shot Aerial Vision-Language Navigation toward Safety and Efficiency
★ 89C2-Explorer. [IROS'26] A Flexible and Contiguous Decentralized Multi-UAV Exploration Framework
★ 87DiffPhysDrone. Published on Nature Machine Intelligence! The first real robot(quadrotor) based on differentiable physics training.
★ 589Utonia. [ICML'26] Official repository of Utonia: Toward One Encoder for All Point Clouds
★ 713CrazySim2Real. A Sim2Real Benchmarking Framework for Crazyflie Drones
★ 21cc-connect. Bridge local AI coding agents (Claude Code, Cursor, Gemini CLI, Codex) to messaging platforms (Feishu/Lark, DingTalk, Slack, Telegram, Discord, LINE, WeChat Work). Chat with your AI dev assistant from anywhere — no public IP required for most platforms.
★ 15kam-planner. Whole-Body Integrated Motion Planning for Aerial Manipulators
★ 48robocode. Agents for robot physical reasoning
★ 7figures4papers. My Python scripts to make high-quality figures for publications in top AI conferences and journals.
★ 2.9kagno. Build, run, and manage agent platforms.
★ 42kSpatialTree. CVPR 2026 (Highlight); Spatial Intelligence; MLLMs
★ 48worldexplorer. [SIGGRAPH Asia 2025] WorldExplorer: Towards Generating Fully Navigable 3D Scenes
★ 190scenesmith. Code for "SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes", ICML 2026 Spotlight
★ 538awesome-claude-skills. A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows
★ 71kHKUST-ELEC5660-Introduction-to-Aerial-Robotics. Repo for HKUST ELEC5660 Course Notes & Lab Tutorial & Project Docker
★ 185nanobot. Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
★ 46kAI-Research-SKILLs. Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
★ 11kvideo2tasks. Video2Tasks: Split multi-task robot videos into single-task segments with auto-generated instruction labels for VLA (pi0, OpenVLA) training
★ 82SparseVideoNav. Sparse Video Generation Model for Embodied Navigation conditioned on loose language guidance, 100% real world verification
★ 113trl. Train transformer language models with reinforcement learning.
★ 19kdreamzero. Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
★ 2.5kawesome-multi-agent-papers. A compilation of the best multi-agent papers
★ 1.6kawesome-ai-research-writing. Elevate your AI research writing, no more tedious polishing ✨
★ 32kTrackVLA. [CoRL 2025] Repository relating to "TrackVLA: Embodied Visual Tracking in the Wild"
★ 420ros-foxglove-bridge. Foxglove WebSocket bridge for ROS 1
★ 286SceneSplat_Benchmark. [NeurIPS 2025] Implementation of "SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting"
★ 56SAGE-3D_Official. This is the official repository of the paper "Towards Physically Executable 3D Gaussian for Embodied Navigation".
★ 194lingbot-va. [RSS 2026] Causal video-action world model for generalist robot control
★ 1.7klingbot-world. Advancing Open-source World Models
★ 4.3kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kRealSee3D. RealSee3D: A multi-view RGB-D dataset combining real-world captures and procedurally generated scenes, with extensible annotations for diverse 3D vision research.
★ 282ACT. Action Chunking Transformer implementation for low cost robot
★ 437dive-into-llms. 《动手学大模型Dive into LLMs》系列编程实践教程
★ 47koasis. 🏝️ OASIS: Open Agent Social Interaction Simulations with One Million Agents.
★ 5kRoboBrain2.5. RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. 🎉🎉🎉
★ 1.1kAether. [ICCV 2025 & ICCV 2025 RIWM Outstanding Paper] Aether: Geometric-Aware Unified World Modeling
★ 6043D-Mem. [CVPR 2025] Source codes for the paper "3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning"
★ 270ShapeR. Code for the ShapeR research paper
★ 864ProjectAirSim. Project AirSim is Microsoft's evolution of AirSim, an advanced simulation platform for building, training, and testing autonomous systems in high-fidelity virtual environments
★ 757Awesome-LLMs-for-Video-Understanding. 🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
★ 3.3kOverleaf-Workshop. Open Overleaf/ShareLaTex projects in vscode, with full collaboration support.
★ 1.6kAcontext. Agent Skills as a Memory Layer
★ 3.6khello-agents. 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
★ 70kN3D-VLM. Official code for paper: N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
★ 117spirit-v1.5. Spirit-v1.5: A Robotic Foundation Model by Spirit AI
★ 631TensorRT-Edge-LLM. High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
★ 488LLaVA-3D. [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D World
★ 388CronusVLA. [AAAI26 oral] CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
★ 112MemoryVLA. [ICLR 2026] Code of "MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation"
★ 309navitrace_evaluation. Jupyter Notebook
★ 35Ai-Review. Large model-assisted paper review
★ 589SG-Reg. [T-RO 2025] SG-Reg: Generalizable and Efficient Scene Graph Registration
★ 138Rex-Omni. [CVPR2026] Detect Anything via Next Point Prediction
★ 1.5kEgoX. Code for "EgoX: Egocentric Video Generation from a Single Exocentric Video"
★ 741Odin-Nav-Stack. An open-source navigation stack based on Odin1.
★ 298file-transfer-go. Go/React开发的端到端webrtc的文件传输/文字传输/桌面共享,安全,隐私,数据不经过服务器。
★ 5.1kfms-fsdp. 🚀 Efficiently (pre)training foundation models with native PyTorch features, including FSDP for training and SDPA implementation of Flash attention v2.
★ 288MiMo-Embodied. MiMo-Embodied
★ 399UniPred. Code for UniPred:Unifying Deep Predicate Invention with Foundation Models
★ 7SparseVLMs. [ICML'25][TPAMI'26] Official implementation of paper "SparseVLM" and "SparseVLM+".
★ 268vla-cache. [NeurIPS 2025] VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
★ 95Awesome-Embodied-AI-Job. Lumina Robotics Talent Call | Lumina社区具身智能招贤榜 | A list for Embodied AI / Robotics Jobs (PhD, RA, intern, etc
★ 1.5kDynam3D. Official implementation of "Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation" (NeurIPS'25 Oral)
★ 90GMemory. SAS
★ 263Agent-Memory-Paper-List. The paper list of "Memory in the Age of AI Agents: A Survey"
★ 2.3kAwesome-VLA.
★ 619TRELLIS.2. Native and Compact Structured Latents for 3D Generation
★ 9.7kVisPruner. [ICCV 2025] Official code for paper: Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
★ 84Avoid-MPC. Mapless Collision-Free Flight via MPC using Dual KD-Trees in Cluttered Environments(IROS 2025)
★ 130UnityGaussianSplatting. Toy Gaussian Splatting visualization in Unity
★ 3.4kInteriorGS. InteriorGS: 3D Gaussian Splatting Dataset of Semantically Labeled Indoor Scenes
★ 284Flexible-Radio-Mapping. Open-source code for the RSS 2025 paper "FERMI: Flexible Radio Mapping with a Hybrid Propagation Model and Scalable Autonomous Data Collection".
★ 17VLM-FO1. VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
★ 330ACMMM25-UAV_ON. Python
★ 92chat-gpt. ChatGPT conversation saving bookmark
★ 376nwm. Official code for the CVPR 2025 paper "Navigation World Models".
★ 661MuseBot. supports Telegram, Discord, Slack, Lark(飞书),钉钉, 企业微信, QQ, 微信, compatible with various LLMs including OpenAI, Gemini, DeepSeek, Doubao, and OpenRouter. It offers intelligent conversation, image generation, video creation, and more. Works seamlessly in both private chats and group settings.
★ 1.6kUAV-Flow. Python
★ 157TravelUAV. Python
★ 270habitat-matterport-3dresearch.
★ 746Awesome-Memory-for-Agents. A Collection of Papers about Memory for Language Agents
★ 621ZHO-nano-banana-Creation. 我的 nano-banana 创意玩法大合集! 持续更新中!
★ 3.7k