This is your work, valued
PhD student at USC and student researcher at Google DeepMind.
3dcodebench. Benchmarking Agentic Procedural 3D Modeling Via Code
★ 70AcroFOD. (ECCV2022) The official PyTorch implementation of the "AcroFOD: An Adaptive Method for Cross-domain Few-shot Object Detection".
★ 68AsyFOD. (CVPR2023) The PyTorch implementation of the "AsyFOD: An Asymmetric Adaptation Paradigm for Few-Shot Domain Adaptive Object Detection".
★ 44opentopos. OpenTopos — unified, code-driven 3D object generation: static + articulated + texture (incl. UV atlas), multi-agent + Blender
★ 7LAM. The official implementation of LAM: Language Articulated Object Modelers
★ 33dcode_toolkit. Toolkit for 3D Code Data Curation
★ 1Kimi-K3. Open Frontier Intelligence
★ 7.5kArticraft. A simple AI agent that turns prompts into articulated 3D objects.
★ 20manim. Animation engine for explanatory math videos
★ 89kvideos. Code for the manim-generated scenes used in 3blue1brown videos
★ 11kData-Engine-OmniX. UE5-based Data Engine used in OmniX
★ 100triton. Development repository for the Triton language and compiler
★ 20kGLM-5. GLM-5: From Vibe Coding to Agentic Engineering
★ 6.8kSWE-IF. SWE-IF: Aligning Code Evaluation with Human Preference (ICML 2026)
★ 4EasyR1. EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1kartiverse. Python
★ 23LongLive-RAG. Official Implementation of LongLive-RAG: A general retrieval-augmented framework for long video generation.
★ 103cosmos-framework. Our inference and training framework to run on the Cosmos Models
★ 424OpenHands. 🙌 OpenHands: AI-Driven Development
★ 83k3dcodebench. Benchmarking Agentic Procedural 3D Modeling Via Code
★ 70articraft. Superseded by https://github.com/articraftresearch/Articraft. This repository is kept for reference.
★ 1.4kECC. The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
★ 236kkilocode. Kilo is the all-in-one agentic engineering platform. Build, ship, and iterate faster with the most popular open source coding agent.
★ 27kautoresearch. AI agents running research on single-GPU nanochat training automatically
★ 92kVibeVoice. Open-Source Frontier Voice AI
★ 52khumanlm. HumanLM: Simulating Users with State Alignment Beats Response Imitation
★ 86VRAG. Multimodal Retrieval-augmented Generation Framework Built by Tongyi Lab, Alibaba Group.
★ 972HY3D-Bench. Python
★ 348IsaacLab. Unified framework for robot learning built on NVIDIA Isaac Sim
★ 7.8kgemma. Gemma open-weight LLM library, from Google DeepMind
★ 5.6klingbot-world. Advancing Open-source World Models
★ 4.3kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kskills. Public repository for Agent Skills
★ 165koriginal_performance_takehome. Anthropic's original performance take-home, now open for you to try!
★ 4.1kShapeR. Code for the ShapeR research paper
★ 863VIGA. VIGA: Vision-as-Inverse-Graphics Agent
★ 1.3kunity-mcp. Unity MCP acts as a bridge between AI assistants and your Unity Editor. Give your LLM tools to manage assets, control scenes, edit scripts, and automate tasks within Unity.
★ 13kQwen3-VL-Embedding. Python
★ 1.3kterminal-bench. A benchmark for LLMs on complicated tasks in the terminal
★ 2.5kskillsbench. SkillsBench evaluates how well skills work and how effective agents are at using them.
★ 1.6kblender-mcp. Open-source MCP to use Blender with any LLM
★ 25kharbor. Framework for evaluating and improving agents
★ 3.7kopen-thoughts. Fully open data curation for reasoning models
★ 2.3kUltraRAG. A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
★ 5.7kVisRAG. Parsing-free RAG supported by VLMs
★ 975Awesome-Latent-Space. A paper list of Awesome Latent Space.
★ 952Awesome-RAG. 😎 Awesome list of Retrieval-Augmented Generation (RAG) applications in Generative AI.
★ 1.3kAwesome-RAG-Vision. Awesome-RAG-Vision: a curated list of advanced retrieval augmented generation (RAG) for Computer Vision
★ 339Pancap. [NeurIPS 2025] Panoptic Captioning: An Equivalence Bridge for Image and Text
★ 38ragflow. RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
★ 86kthreefiner. An interface for text-guided mesh refinement.
★ 188TRELLIS. Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
★ 13kfullpart. A part-based 3D generation framework & the largest and most comprehensively annotated 3D part dataset.
★ 144PartNeXt. [Neurips DB 2025] PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
★ 101PartCrafter. [NeurIPS 2025] PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
★ 2.5kBesiegeField. Official repo of BesiegeField, an interactive and real-time environment for machine construction and simulation (arXiv:2510.14980).
★ 62nanochat. The best ChatGPT that $100 can buy.
★ 57klynx. Lynx: Towards High-Fidelity Personalized Video Generation
★ 336Hunyuan3D-Part. Hunyuan 3D Part Segmentation and Generation Pipeline
★ 519ll3m. LL3M writes Python code that generates 3D assets in Blender.
★ 546urdf-visualizer. A VSCode extension for visualizing the URDF file and xacro file.
★ 220MeshCoder. Jupyter Notebook
★ 500LeetCUDA. LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
★ 12kMUSAR.
★ 30gpt-oss. gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
★ 20kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kcopart. CoPart (ICCV 2025): A part-based 3D generation framework & the first large-scale part-level 3D dataset.
★ 207context7. Context7 Platform -- Up-to-date code documentation for LLMs and AI code editors
★ 60kflux. Official inference repo for FLUX.1 models
★ 26kMSRVTT-Personalization. Benchmark dataset and code of MSRVTT-Personalization
★ 52Phantom-Data. Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset
★ 118Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kVACE. [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
★ 3.9kPhantom. Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
★ 1.5kAwesome-Video-Diffusion. A curated list of recent diffusion models for video generation, editing, and various other applications.
★ 5.7kMAGI-1. MAGI-1: Autoregressive Video Generation at Scale
★ 3.8kpython-sdk. The official Python SDK for Model Context Protocol servers and clients
★ 24kAny2Caption. This is the project for 'Any2Caption', Interpreting Any Condition to Caption for Controllable Video Generation
★ 49SkyReels-A2. SkyReels-A2: Compose anything in video diffusion transformers
★ 713codex. Lightweight coding agent that runs in your terminal
★ 103kArtFormer. [CVPR 2025] ArtFormer: Controllable Generation of Diverse 3D Articulated Objects
★ 43X-Dyna. [CVPR 2025 Highlight] X-Dyna: Expressive Dynamic Human Image Animation
★ 269SWELancer-Benchmark. This repo contains the dataset and code for the paper "SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?"
★ 1.4klangchain. The agent engineering platform.
★ 143kHunyuan3D-2. High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models.
★ 14kcosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kAgiBot-World. [IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
★ 3.1kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30karticulate-anything. [ICLR 2025] Official implementation of Articulate-Anything
★ 206generative-models. Generative Models by Stability AI
★ 27kHOISwap. [NeurIPS 2024] HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
★ 26stretch_ai. Python
★ 230HunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12krobust-rearrangement. From Imitation to Refinement -- Residual RL for Precise Assembly
★ 248furniture-bench. FurnitureBench: Real-World Furniture Assembly Benchmark (RSS 2023)
★ 236IKEA-Manuals-at-Work. IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
★ 65openvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
★ 6.7kdexart-release. DexArt: Benchmarking Generalizable Dexterous Manipulation with Articulated Objects, CVPR 2023
★ 147scene-language. (CVPR 2025 Highlight) The Scene Language: Representing Scenes with Programs, Words, and Embeddings
★ 263LLaMA-Mesh. Unifying 3D Mesh Generation with Language Models
★ 1.2kscreenpipe. YC (S26) | Record your screen 24/7 and plug into your agents. Local, private, secure. Connect to OpenClaw, Hermes agent and 100+ apps
★ 21kShapeAssembly. Public code release for our Siggraph Asia 2020 paper: ShapeAssembly: Learning to Generate Programs for 3D Shape Structure Synthesis
★ 57Moore-AnimateAnyone. Character Animation (AnimateAnyone, Face Reenactment)
★ 3.5kHolodeck. CVPR 2024: Language Guided Generation of 3D Embodied AI Environments.
★ 562ComfyUI. The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
★ 123kxformers. Hackable and optimized Transformers building blocks, supporting a composable construction.
★ 11kLGM. [ECCV 2024 Oral] LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation.
★ 2.1kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kCogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kHPT. Heterogeneous Pre-trained Transformer (HPT) as Scalable Policy Learner.
★ 543cage. [CVPR 2024] Official Implementation of the paper "CAGE: Controllable Articulation GEneration"
★ 82real2code. Python
★ 117Microsoft-Activation-Scripts. Open-source Windows and Office activator featuring HWID, Ohook, TSforge, and Online KMS activation methods, along with advanced troubleshooting.
★ 185kAwesome-Articulated-Object-Understanding. A curated list of resources for articulated objects understanding.
★ 130urdf-viz. visualize URDF/XACRO file, URDF Viewer works on Windows/MacOS/Linux
★ 586phobos. An add-on for Blender allowing to create URDF, SDF and SMURF robot models in a WYSIWYG environment.
★ 898kubric. A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
★ 2.8kurdformer. code release for URDFormer
★ 202axolotl. Go ahead and axolotl questions
★ 12kMetaworld. Collections of robotics environments geared towards benchmarking multi-task and meta reinforcement learning
★ 1.9kmujoco. Multi-Joint dynamics with Contact. A general purpose physics simulator.
★ 14kmujoco-py. MuJoCo is a physics engine for detailed, efficient rigid body simulations with contacts. mujoco-py allows using MuJoCo from Python 3.
★ 3.1kPanoGen. Code and Data for Paper: PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation
★ 83viper_rl. Using advances in generative modeling to learn reward functions from unlabeled videos.
★ 143atari-py. A packaged and slightly-modified version of https://github.com/bbitmaster/ale_python_interface
★ 392DrawingSpinUp. (SIGGRAPH Asia 2024) This is the official PyTorch implementation of SIGGRAPH Asia 2024 paper: DrawingSpinUp: 3D Animation from Single Character Drawings
★ 656DepthCrafter. [CVPR 2025 Highlight] DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
★ 1.6klerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kAwesome-Robotics-3D. A curated list of 3D Vision papers relating to Robotics domain in the era of large models i.e. LLMs/VLMs, inspired by awesome-computer-vision, including papers, codes, and related websites
★ 820dl-engineer-guidebook. 深度学习工程师生存指南
★ 921ModernRobotics. Modern Robotics: Mechanics, Planning, and Control Code Library --- The primary purpose of the provided software is to be easy to read and educational, reinforcing the concepts in the book. The code is optimized neither for efficiency nor robustness.
★ 2.9kpuppeteer. Code for "Hierarchical World Models as Visual Whole-Body Humanoid Controllers"
★ 213easy-rl. 强化学习中文教程(蘑菇书🍄),在线阅读地址:https://datawhalechina.github.io/easy-rl/
★ 14kawesome-reinforcement-learning-zh. 中文整理的强化学习资料(Reinforcement Learning)
★ 2.2ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kLLM101n. LLM101n: Let's build a Storyteller
★ 38kbehavior-vision-suite.github.io. CSS
★ 170infinigen. Infinite Photorealistic Worlds using Procedural Generation
★ 7.2kMoE-Adapters4CL. Code for paper "MoE-Adapters" CVPR2024 and "MoE-Adapters++" TPAMI2025
★ 278lightplane. Lightplane implements a highly memory-efficient differentiable radiance field renderer, and a module for unprojecting features from images to 3D grids.
★ 286DriveAGI. [CVPR 2024 Highlight] GenAD: Generalized Predictive Model for Autonomous Driving
★ 802DreamView. (ECCV 2024) Official implementation of Paper ''DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation''
★ 40grok-1. Grok open release
★ 52kOpen-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kawesome-tips.
★ 4.7kLWM. Large World Model -- Modeling Text and Video with Millions Context
★ 7.4kobjaverse_filter. naive filter of objaverse
★ 152OLMo. Modeling, training, eval, and inference code for OLMo
★ 6.6kawesome-3d-diffusion. A collection of papers on diffusion models for 3D generation.
★ 1.3kstable-dreamfusion. Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
★ 8.8kPointLLM. [ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds
★ 1kmamba. Mamba SSM architecture
★ 19kmamba.py. A simple and efficient Mamba implementation in pure PyTorch and MLX.
★ 1.5kWonderJourney. Python
★ 770GPTs. leaked prompts of GPTs
★ 32kApple-Silicon-Guide. Apple Silicon Guide. Learn all about the A17 Pro, A16 Bionic, R1, M1-series, M2-series, and M3-series chips. Along with all the Devices, Operating Systems, Tools, Gaming, and Software that Apple Silicon powers.
★ 1.9kconsistencydecoder. Consistency Distilled Diff VAE
★ 2.2kMixCon3D. [CVPR 2024] The official implementation of paper "Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training"
★ 35Wonder3D. Single Image to 3D using Cross-Domain Diffusion for 3D Generation
★ 5.4kAwesome-Anything. General AI methods for Anything: AnyObject, AnyGeneration, AnyModel, AnyTask, AnyX
★ 1.9kawesome-instruction-learning. Papers and Datasets on Instruction Tuning and Following. ✨✨✨
★ 512ml-visuals. 🎨 ML Visuals contains figures and templates which you can reuse and customize to improve your scientific writing.
★ 17kopencsapp.github.io. Open CS Application | 开源CS申请
★ 2.3kUni3D. [ICLR'24 Spotlight] Uni3D: 3D Visual Representation from BAAI
★ 678DiT-3D. 🔥🔥🔥Official Codebase of "DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation"
★ 321PointCLIP_V2. [ICCV 2023] PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning
★ 291PointCloudDatasets. 3D point cloud datasets in HDF5 format, containing uniformly sampled 2048 points per shape.
★ 519MVTN. pytorch implementation of the ICCV'21 paper "MVTN: Multi-View Transformation Network for 3D Shape Recognition"
★ 107objaverse-xl. 🪐 Objaverse-XL is a Universe of 10M+ 3D Objects. Contains API Scripts for Downloading and Processing!
★ 1.3kobjaverse. Python
★ 2objaverse-rendering. 📷 Scripts for rendering Objaverse
★ 265pytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
★ 102kdgl. Python package built to ease deep learning on graph, on top of existing DL frameworks.
★ 14ktuning_playbook. A playbook for systematically maximizing the performance of deep learning models.
★ 30kMETER. METER: A Multimodal End-to-end TransformER Framework
★ 377academicpages.github.io. Github Pages template based upon HTML and Markdown for personal, portfolio-based websites.
★ 17k3D-LLM. Code for 3D-LLM: Injecting the 3D World into Large Language Models
★ 1.2kllm-attacks. Universal and Transferable Attacks on Aligned Language Models
★ 4.8kReSTE. Official implementation of Rectified Straight Through Estimator (ReSTE).
★ 3RandBox. [ICCV 2023] PyTorch implementation of RandBox
★ 58point-e. Point cloud diffusion for 3D model synthesis
★ 6.9kSemantic-SAM. [ECCV 2024] Official implementation of the paper "Semantic-SAM: Segment and Recognize Anything at Any Granularity"
★ 2.9kaccelerate. 🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
★ 9.8kxla. Enabling PyTorch on XLA Devices (e.g. Google TPU)
★ 2.8kawesome-point-cloud-analysis-2023. A list of papers and datasets about point cloud analysis (processing) since 2017. Update every day!
★ 1.6kSummer2027-Internships. Summer 2026 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.
★ 46kgrounded-segment-any-parts. Grounded Segment Anything: From Objects to Parts
★ 416VLPart. [ICCV2023] VLPart: Going Denser with Open-Vocabulary Part Segmentation
★ 395MinkowskiEngine. Minkowski Engine is an auto-diff neural network library for high-dimensional sparse tensors
★ 3kopen_flamingo. An open-source framework for training large multimodal models.
★ 4.1kOpenShape_code. official code of “OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding”
★ 319Cap3D. [NeurIPS 2023] Scalable 3D Captioning with Pretrained Models
★ 278ULIP. Python
★ 611stanford-shapenet-renderer. Scripts for batch rendering models using Blender. Tested with models from stanfords shapenet library.
★ 542OmniObject3D. [ CVPR 2023 Award Candidate ] OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation
★ 531jax. Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
★ 36kshap-e. Generate 3D objects conditioned on text or images
★ 12kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kopen_clip. An open source implementation of CLIP.
★ 14kflax. Flax is a neural network library for JAX that is designed for flexibility.
★ 7.3kGroundingDINO. [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
★ 10kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kx-clip. A concise but complete implementation of CLIP with various experimental improvements from recent papers
★ 724awesome-computer-vision. A curated list of awesome computer vision resources
★ 23kImage2Paragraph. [Image 2 Text Para] Transform Image into Unique Paragraph with ChatGPT, BLIP2, OFA, GRIT, Segment Anything, ControlNet.
★ 822