This is your work, valued
Teaching robots to see and act 🤖 | Assistant Professor @ SUAT | CV · Robotics · Embodied AI
Bayesian-Crowd-Counting. Official Implement of ICCV 2019 oral paper Bayesian Loss for Crowd Count Estimation with Point Supervision
★ 343awesome-video-action-models. 🤖 A curated list of Video Action Models (VAMs) — papers using video generation models to produce executable robot actions. Covers UniPi, UVA, mimic-video, Motus, Cosmos Policy, DreamZero, and more.
★ 10A-Prompts. The official Implement of paper "Remind of the Past: Incremental Learning with Analogical Prompts"
★ 5group_sparcity. using L2,1 norm base one tensorflow
★ 3tensorflow_face. An academic oriented face recognition framework base on tensorflow and tensorflow.slim
★ 3Awsome-Multimodal-In-Context-Learning. A curated list of multimodal in context learning.
★ 2Awesome-Robotics-Lifelong-Learning. Paper Collecttions of Robotics Lifelong Learning
★ 1obsidian-fast-note-sync. Can be privately deployed, focusing on providing Obsidian users with a seamless, distraction-free note synchronization plugin with real-time sync across multiple platforms, supporting Mac, Windows, Android, iOS, and offering multilingual support.可私有化部署,专注为 Obsidian 用户提供无打扰、丝般顺滑、多端实时同步的多平台笔记同步插件。
★ 2.7kXRZero-G0. XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios
★ 67RoboOrchard. Python
★ 36hiddify-app. Multi-platform auto-proxy client, supporting Sing-box, X-ray, TUIC, Hysteria, Reality, Trojan, SSH etc. It’s an open-source, secure and ad-free.
★ 32kAiScientist. Python
★ 141DiT4DiT. This is the official code repo for DiT4DiT, a Vision-Action-Model (VAM) framework that combines video generation model with flow-matching-based action prediction for generalizable robotic manipulation.
★ 417freemocap. Free Motion Capture for Everyone 💀✨
★ 9.9kdeer-flow. An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
★ 78klast30days-skill. AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary
★ 56khermes-agent. The agent that grows with you
★ 223kskills. Skills for Real Engineers. Straight from my .agents directory.
★ 198kclaude-howto. A visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
★ 41ksherlock. Hunt down social media accounts by username across social networks
★ 87kfff. The fastest and the most accurate file search SDK for AI agents, Neovim, Rust, C, Python, Bun and NodeJS
★ 9.9kawesome-autoresearch. awesome autoresearch list
★ 667learn-claude-code. Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
★ 73kPointWorld. PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
★ 470LATENT. Official implementation of Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data
★ 673autoresearch. AI agents running research on single-GPU nanochat training automatically
★ 93kAuto-claude-code-research-in-sleep. ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
★ 14kawesome-video-action-models. 🤖 A curated list of Video Action Models (VAMs) — papers using video generation models to produce executable robot actions. Covers UniPi, UVA, mimic-video, Motus, Cosmos Policy, DreamZero, and more.
★ 15kai0. Code for kai0, including training, inference and data collection.
★ 406OpenReal2Sim. A toolbox for real-to-sim reconstruction and robotic simulation
★ 237mimicdroid-robocasa. MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
★ 52FastTD3. Python
★ 456XLeRobot. XLeRobot: Practical Dual-Arm Mobile Home Robot for $660
★ 5.4kAwesome-Imitation-Learning. A curated list of awesome imitation learning resources and publications
★ 608xr_teleoperate. This repository implements teleoperation of the Unitree humanoid robot using XR Devices.
★ 1.6kawesome-implicit-representations. A curated list of resources on implicit neural representations.
★ 2.6kAwesome-Video-Robotic-Papers. This repository compiles a list of papers related to the application of video technology in the field of robotics! Star⭐ the repo and follow me if you like what you see🤩.
★ 193BridgeVLA. ✨✨【NeurIPS 2025】Official implementation of BridgeVLA
★ 194SimpleVLA-RL. [ICLR 2026] SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
★ 1.8kDexGraspVLA. [AAAI'26 Oral] DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
★ 556Hybrid-VLA. HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
★ 352RynnVLA-002. RynnVLA-002: A Unified Vision-Language-Action and World Model
★ 1.1kOpenVLA. 多模态具身智能大模型 OpenVLA 的复现以及在 LIBERO 数据集上的微调改进
★ 4973D-VLA. [ICML 2024] 3D-VLA: A 3D Vision-Language-Action Generative World Model
★ 630diffusion_policy. [RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion
★ 4.4kSpatialVLA. 🔥 SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.
★ 711My-awesome-PINN-papers.
★ 41awesome-pinn. A curated list of awesome Physics Informed Neural Network, projects and communities.
★ 54SimplerEnv-OpenVLA. Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo, and OpenVLA) in simulation under common setups (e.g., Google Robot, WidowX+Bridge)
★ 271openvla-oft. Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
★ 1.3kopenvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
★ 6.7kCogACT. A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
★ 429torchdiffeq. Differentiable ODE solvers with full GPU support and O(1)-memory backpropagation.
★ 6.5ksimple_neural_ode. Simple implementation of neural ODE in autograd. For learning only
★ 7QueST. Official code for "QueST: Self-Supervised Skill Abstractions for Continuous Control" [NeurIPS 2024]
★ 114awesome-AI4EDA. This repo awesome-AI4EDA contains the source for the webpage: https://ai4eda.github.io, which is a curated paper list of awesome AI for EDA.
★ 211Awesome-LLM4EDA.
★ 293Awesome-VLA-RL. This repository summarizes recent advances in the VLA + RL paradigm and provides a taxonomic classification of relevant works.
★ 427awesome-embodied-vla-va-vln. A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
★ 3.4kEmbodied-AI-Paper-TopConf. [Actively Maintained🔥] A list of Embodied AI papers accepted by top conferences (ICLR, NeurIPS, ICML, RSS, CoRL, ICRA, IROS, CVPR, ICCV, ECCV).
★ 729Awesome-MCoT. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
★ 1kuqlm. UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection
★ 1.2kcontinuous-thought-machines. Continuous Thought Machines, because thought takes time and reasoning is a process.
★ 2kmem0. Universal memory layer for AI Agents
★ 62kmindshub. Make AI do actual work. Swap the model anytime — keep everything you've built.
★ 40kgraphiti. Build Real-Time Knowledge Graphs for AI Agents
★ 29khumanoid-gym. Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer https://arxiv.org/abs/2404.05695
★ 2.1kAgenticRAG-Survey. Agentic-RAG explores advanced Retrieval-Augmented Generation systems enhanced with AI LLM agents.
★ 1.7kagents-course. This repository contains the Hugging Face Agents Course.
★ 31kEmbodied-AI-Guide. [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
★ 15kAwesome-Robotics-Manipulation. A comprehensive list of papers about Robot Manipulation, including papers, codes, and related websites.
★ 1.1kawesome-robotic-manipulation. A curated list of robotic manipulation papers
★ 5awesome-humanoid-learning. Humanoid Robots Resources
★ 935Awesome-VLA.
★ 619awesome-vision-language-action-model. Latest Advances on Vison-Language-Action Models.
★ 142awesome-humanoid-robot-learning. A Paper List for Humanoid Robot Learning.
★ 2.6kxiaozhi-esp32. An MCP-based chatbot | 一个基于MCP的聊天机器人
★ 29kLLMEvaluation. A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the adoption of best practices in LLM assessment, and critically assess the effectiveness of these evaluation methods.
★ 196gorilla. Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)
★ 13kLightRAG. [EMNLP2025] "LightRAG: Simple and Fast Retrieval-Augmented Generation"
★ 38kself-llm. 《开源大模型食用指南》针对中国宝宝量身打造的基于Linux环境快速微调(全参数/Lora)、部署国内外开源大模型(LLM)/多模态大模型(MLLM)教程
★ 32kevals. Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
★ 19kllm-action. 本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
★ 25kllama-cookbook. Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
★ 19kAwesome-Chinese-LLM. 整理开源的中文大语言模型,以规模较小、可私有化部署、训练成本较低的模型为主,包括底座模型,垂直领域微调及应用,数据集与教程等。
★ 23kAwesome-LLM. Awesome-LLM: a curated list of Large Language Model
★ 27kcomposio. Composio powers 1000+ toolkits, tool search, context management, authentication, and a sandboxed workbench to help you build AI agents that turn intent into action.
★ 29k1Panel. 🔥 1Panel is a modern, open-source VPS control panel — and the only one with native AI agent support. Run Ollama models, deploy OpenClaw agents, and manage your entire server stack from one clean web interface.
★ 36kawesome-LLM-resources. 🧑🚀 全世界最好的LLM资料总结(多模态生成、Agent、辅助编程、AI审稿、数据处理、模型训练、模型推理、o1 模型、MCP、小语言模型、视觉语言模型) | Summary of the world's best LLM resources.
★ 8.8kRAG_Techniques. This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
★ 29kvanna. 🤖 Chat with your SQL database 📊. Accurate Text-to-SQL Generation via LLMs using Agentic Retrieval 🔄.
★ 24kRagaAI-Catalyst. Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced analytics with timeline and execution graph view
★ 16kpandas-ai. Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.
★ 24kawesome-llm-apps. 100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
★ 129kagno. Build, run, and manage agent platforms.
★ 42kLLMs-from-scratch. Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
★ 100kllm-course. Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
★ 81kcs-self-learning. 计算机自学指南
★ 75kKnowledgeEditingPapers. Must-read Papers on Knowledge Editing for Large Language Models.
★ 1.2kPRAG. Code for Parametric RAG, SIGIR 2025 Full Paper
★ 233topoloss. Induce brain-like topographic structure in your neural networks
★ 74ComprehendEdit. Python
★ 5awesome_lists. Awesome Lists for Tenure-Track Assistant Professors and PhD students. (助理教授/博士生生存指南)
★ 1.6kacad-homepage.github.io. AcadHomepage: A Modern and Responsive Academic Personal Homepage
★ 2.9kR1-V. Witness the aha moment of VLM with less than $3.
★ 4.1kopenai-structured-outputs-samples. Sample apps to help developers get started with Structured Outputs
★ 681victoreke.com. My personal portfolio website built with Next.js, Sanity and Tailwind CSS.
★ 437EnzymeCAGE. A geometric foundation model for enzyme retrieval with evolutionary insights.
★ 99EnzymeFlow. Official repository of EnzymeFlow
★ 103GENzyme. Official repository of GENzyme
★ 66bRAG-langchain. Everything you need to know to build your own RAG application
★ 4.1kPDFMathTranslate. [EMNLP 2025 Demo] PDF scientific paper translation with preserved formats - 基于 AI 完整保留排版的 PDF 文档全文双语翻译,支持 Google/DeepL/Ollama/OpenAI 等服务,提供 CLI/GUI/MCP/Docker/Zotero
★ 36kReactZyme. Official repository of ReactZyme
★ 45lingua. Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
★ 4.8kcircuit_training. Python
★ 1.7kProteusAI. ProteusAI is a library for the machine learning driven engineering of proteins. The library enables workflows from protein structure prediction, prediction of mutational effects to protein ligand interactions powered by artificial intelligence.
★ 87Emu3. Next-Token Prediction is All You Need
★ 2.4kllm-mutate. Python
★ 15PixWizard. [ICLR2025] A versatile image-to-image visual assistant, designed for image generation, manipulation, and translation based on free-from user instructions.
★ 211Awsome-Multimodal-In-Context-Learning. A curated list of multimodal in context learning.
★ 2MIC. MMICL, a state-of-the-art VLM with the in context learning ability from ICL, PKU
★ 361Mu-Protein. Jupyter Notebook
★ 23computer-science. 🎓 Path to a free self-taught education in Computer Science!
★ 207kkotaemon. An open-source RAG-based tool for chatting with your documents.
★ 26kChannelViT. Channel Vision Transformers: An Image Is Worth C x 16 x 16 Words
★ 74warm-up-kit. Warm-up Kit for Erasing the Invisible Competition @ NeurIPS 2024
★ 8Awesome-Phenotypic-Drug-Discovery. PDD: Awesome Phenotypic Drug Discovery
★ 41GOT-OCR2.0. Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
★ 8.2kAwesome-Bio-Foundation-Models. A collection of awesome bio-foundation models, including protein, RNA, DNA, gene, single-cell, and so on.
★ 317awesome-described-object-detection. A curated list of papers and resources related to Described Object Detection, Open-Vocabulary/Open-World Object Detection and Referring Expression Comprehension. Updated frequently and pull requests welcomed.
★ 359ML-YouTube-Courses. 📺 Discover the latest machine learning / AI courses on YouTube.
★ 17kAwesome-Video-Diffusion. A curated list of recent diffusion models for video generation, editing, and various other applications.
★ 5.7kMemoRAG. Empowering RAG with a memory-based data interface for all-purpose applications!
★ 2.3kknowledge_graph. Convert any text to a graph of knowledge. This can be used for Graph Augmented Generation or Knowledge Graph based QnA
★ 3.6kAwesome-Graph-LLM. A collection of AWESOME things about Graph-Related LLMs.
★ 2.4kKG-LLM-Papers. [Paper List] Papers integrating knowledge graphs (KGs) and large language models (LLMs)
★ 2.2knano-graphrag. A simple, easy-to-hack GraphRAG implementation
★ 4kSciAgentsDiscovery. Python
★ 629paper-qa. High accuracy RAG for answering questions from scientific documents with citations
★ 9kai-for-grant-writing. A curated list of resources for using LLMs to develop more competitive grant applications.
★ 4.2kprompts.chat. f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
★ 167kComfyUI. The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
★ 123kfirecrawl. The API to search, scrape, and interact with the web at scale. 🔥
★ 159kDeep-Live-Cam. real time face swap and one-click video deepfake with only a single image
★ 95kVane. Vane is an AI-powered answering engine.
★ 36kanything-llm. Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
★ 64kAwesome-Model-Merging-Methods-Theories-Applications. Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities. ACM Computing Surveys, 2026.
★ 771relik. Retrieve, Read and LinK: Fast and Accurate Entity Linking and Relation Extraction on an Academic Budget (ACL 2024)
★ 513ADAS. [ICLR 2025] Automated Design of Agentic Systems
★ 1.6kAutoSurvey. Python
★ 471Awesome-LLM-3D. Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resources
★ 2.2kMinerU. Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
★ 76kdata-juicer. Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6.8kOpenResearcher. OpenResearcher, an advanced Scientific Research Assistant
★ 507AI-Scientist. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑🔬
★ 14k3DeeCellTracker. A python algorithm for tracking cells in 3D deforming organs
★ 70Segment-Any-Anomaly. Official implementation of "Segment Any Anomaly without Training via Hybrid Prompt Regularization (SAA+)".
★ 843Tranception. Official repository for the paper "Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval"
★ 170LitSearch. [EMNLP 2024] A Retrieval Benchmark for Scientific Literature Search
★ 109LLM101n. LLM101n: Let's build a Storyteller
★ 38kNOS. Protein Design with Guided Discrete Diffusion
★ 143esm. Jupyter Notebook
★ 2.9kmagiclens. [ICML'24 Oral] "MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions"
★ 211EVE. EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 376hyperui. Free Tailwind CSS v4 components for your next project, designed to enhance your web development with the latest features and styles 🚀
★ 12kultimate-react-course. Starter files, final projects, and FAQ for my Ultimate React course
★ 4.5kreact-complete-guide-course-resources. React - The Complete Guide Course Resources (Code, Attachments, Slides)
★ 3.9kdify. Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
★ 151kphotoshot. An open-source AI avatar generator web app - https://photoshot.app
★ 3.9kragflow. RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
★ 87kAutoWebGLM. An LLM-based Web Navigating Agent (KDD'24)
★ 930VAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kml-aim. This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects.
★ 1.4kLlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kwalk-jump. Official repository for discrete Walk-Jump Sampling (dWJS)
★ 57ssms_event_cameras. [CVPR'24 Spotlight] The official implementation of "State Space Models for Event Cameras"
★ 136ProteinInvBench. The official implementation of the NeurIPS'23 paper ProteinInvBench: Benchmarking Protein Design on Diverse Tasks, Models, and Metrics
★ 203esm. Evolutionary Scale Modeling (esm): Pretrained language models for proteins
★ 4.2kprotein-sequence-models. Python
★ 259ByProt. Python
★ 218Machine-learning-for-proteins. Listing of papers about machine learning for proteins.
★ 1.7kProteinDT. A Text-guided Protein Design Framework, Nat Mach Intell 2025 (https://www.nature.com/articles/s42256-025-01011-z)
★ 107papers_for_protein_design_using_DL. List of papers about Proteins Design using Deep Learning
★ 2kELLA. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
★ 1.3kGPU-Benchmarks-on-LLM-Inference. Multiple NVIDIA GPUs or Apple Silicon for Large Language Model Inference?
★ 1.9kawesome-large-multimodal-agents.
★ 497all-seeing. [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of the Open World"
★ 506bgpt. Beyond Language Models: Byte Models are Digital World Simulators
★ 333PseCo. (CVPR 2024) Point, Segment and Count: A Generalized Framework for Object Counting
★ 126Osprey. [CVPR2024] The code for "Osprey: Pixel Understanding with Visual Instruction Tuning"
★ 843LLM-in-Vision. Recent LLM-based CV and related works. Welcome to comment/contribute!
★ 871Awesome-LLM-Robotics. A comprehensive list of papers using large language/multi-modal models for Robotics/RL, including papers, codes, and related websites
★ 4.4kScientific-LLM-Survey. Scientific Large Language Models: A Survey on Biological & Chemical Domains
★ 359symbolicai. A neurosymbolic perspective on LLMs
★ 1.7kMoE-LLaVA. 【TMM 2025🔥】 Mixture-of-Experts for Large Vision-Language Models
★ 2.3kMobileVLM. Strong and Open Vision Language Assistant for Mobile Devices
★ 1.4kSynthCLIP. Code base of SynthCLIP: CLIP training with purely synthetic text-image pairs from LLMs and TTIs.
★ 104COMM. Pytorch code for paper From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
★ 211Awesome-Mamba. Awesome list of papers that extend Mamba to various applications.
★ 142EVA. EVA Series: Visual Representation Fantasies from BAAI
★ 2.7kjepa. PyTorch code and models for V-JEPA self-supervised learning from video.
★ 4.1kLWM. Large World Model -- Modeling Text and Video with Millions Context
★ 7.4kprismatic-vlms. A flexible and efficient codebase for training visually-conditioned language models (VLMs)
★ 1kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kawesome-ssm-ml. Reading list for research topics in state-space models
★ 368Awesome-state-space-models. Collection of papers on state-space models
★ 620mamba. Mamba SSM architecture
★ 19k