This is your work, valued
๐ฑ ๐ cloneofsimo@gmail.com
lora. Using Low-rank adaptation to quickly fine-tune diffusion models.
โ 7.5kminDiffusion. Self-contained, minimalistic implementation of diffusion models with Pytorch.
โ 1.2kpaint-with-words-sd. Implementation of Paint-with-words with Stable Diffusion : method from eDiff-I that let you generate image from text-labeled segmentation map.
โ 646minRF. Minimal implementation of scalable rectified flow transformers, based on SD3's approach
โ 641minSDXL. Huggingface-compatible SDXL Unet implementation that is readily hackable
โ 439vqgan-training. Train VAE like a boss
โ 313d3pm. Minimal Implementation of a D3PM in pytorch
โ 310consistency_models. Unofficial Implementation of Consistency Models in pytorch
โ 259min-max-gpt. Minimal (400 LOC) implementation Maximum (multi-node, FSDP) GPT training
โ 132realformer-pytorch. Implementation of RealFormer using pytorch
โ 101magicmix. Unofficial Implementation of MagicMix
โ 98scaling-guide. WIP
โ 96min-fsdp. Python
โ 93ezmup. Simple implementation of muP, based on Spectral Condition for Feature Learning. The implementation is SGD only, dont use it for Adam
โ 88t2i-adapter-diffusers. Python
โ 86sdxl_inversions. Jupyter Notebook
โ 85promptplusplus. Jupyter Notebook
โ 73ptx-tutorial-by-aislop. PTX-Tutorial Written Purely By AIs (Deep Research of Openai and Claude 3.7)
โ 66sd-various-ideas. Jupyter Notebook
โ 56karras-power-ema-tutorial. Python
โ 53insightful-nn-papers. These papers will provide unique insightful concepts that will broaden your perspective on neural networks and deep learning
โ 48clipping-CLIP-to-GAN. Python
โ 42imagenet.int8. Python
โ 40fim-llama-deepspeed. Python
โ 33zeroshampoo. Python
โ 33repa-rf. Python
โ 32infinite-fractal-stream. Jupyter Notebook
โ 30minSAE. Python
โ 30auto_llm_codebase_analysis. Python
โ 27min-max-in-dit. Python
โ 27minVJEPA. Python
โ 25project_RF. Python
โ 24minDinoV2. Python
โ 24efae. Python
โ 24inversion_edits. Jupyter Notebook
โ 21planning-with-diffusion-tutorial. Jupyter Notebook
โ 18zeroshot-storytelling. Github repository for Zero Shot Visual Storytelling
โ 15blocked-decorr. Python
โ 13ptar. C++
โ 13poly2SOP. Transformer takes a polynomial, expresses it as sum of powers.
โ 11minMomentMatching. Python
โ 11smallest_working_performer. Python
โ 10minmoe. Python
โ 6n-body-dynamic-cuda. Cuda
โ 6smallest_working_gpt. gpt that is even smaller
โ 6rectified-flow. Jupyter Notebook
โ 5torchcu. Python
โ 5reverse_eng_deepspeed_study. DeepSpeed Study, focused on reverse engineering and enhancing documentation
โ 5imgdataset_process. Python
โ 4lora_dreambooth_replicate. Jupyter Notebook
โ 4SemanticSegmentationTrainerTemplate. Python
โ 1cv. Simple CV (pdf, Latex)
โ 1netfilter. C
โ 1unn-lstm-torch. Python
โ 1Super-Simple-LSTM-Template. Python
โ 1send-arp. C++
โ 1Freshman_2. Lecture notes from my Freshman 2nd semester
โ 1quack. A Quirky Assortment of CuTe Kernels
โ 1.1kfast.cu. Fastest kernels written from scratch
โ 587infinite-kanvas. An infinite canvas image editor using fal.ai
โ 292lavender-data. Load & manage evolving datasets efficiently
โ 22nano-aha-moment. Single File, Single GPU, From Scratch, Efficient, Full Parameter Tuning library for "RL for LLMs"
โ 626NativeSparseAttention. research impl of Native Sparse Attention (2502.11089)
โ 62llmdifftracker. Lightweight package that tracks and summarizes code changes using LLMs (Large Language Models)
โ 32video-starter-kit. Enable AI models for video production in the browser
โ 2.4kdiffusion-speedrun. Focused on fast experimentation and simplicity
โ 77Cosmos-Tokenizer. A suite of image and video neural tokenizers
โ 1.7klingua. Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
โ 4.8kmochi. The best OSS video generation models, created by Genmo
โ 3.7kmodded-nanogpt. NanoGPT (124M) in 90 seconds
โ 5.6kai-robotics. AI Robotics tutorials for hobbyists
โ 108SV3D-fine-tune. Fine-tuning code for SV3D
โ 115Motion-LoRA. Learning Motion from Low-Rank Adaptation
โ 46lancedb. Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.
โ 11kllm.c. LLM training in simple, raw C/CUDA
โ 31kefficient_cross_entropy. Python
โ 124SpeeD. SpeeD: A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training
โ 188edm2. EDM2 and Autoguidance -- Official PyTorch implementation
โ 848llama3. The official Meta Llama 3 GitHub site
โ 29kDEADiff. [CVPR 2024] Official implementation of "DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations"
โ 280Score-Entropy-Discrete-Diffusion. [ICML 2024 Best Paper] Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution (https://arxiv.org/abs/2310.16834)
โ 740sglang. SGLang is a high-performance serving framework for large language models and multimodal models.
โ 31ktexify. Math OCR model that outputs LaTeX and markdown
โ 1.1kmermaid. Generation of diagrams like flowcharts or sequence diagrams from text in a similar manner as markdown
โ 89kresource-stream. GPU programming related news and material links
โ 2.2kmin-max-gpt. Minimal (400 LOC) implementation Maximum (multi-node, FSDP) GPT training
โ 132OLMo. Modeling, training, eval, and inference code for OLMo
โ 6.6kImprovedTokenMerge. Jupyter Notebook
โ 49StanfordQuadruped. Python
โ 1.8kchakra-ui. Chakra UI is a component system for building SaaS products with speed โก๏ธ
โ 41klmql. A language for constraint-guided and efficient LLM programming.
โ 4.2kmamba. Mamba SSM architecture
โ 19kgpt-fast. Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.
โ 6.2kAwesome-Video-Diffusion. A curated list of recent diffusion models for video generation, editing, and various other applications.
โ 5.7koptimistix. Nonlinear optimisation (root-finding, least squares, ...) in JAX+Equinox. https://docs.kidger.site/optimistix/
โ 606Hotshot-XL. โจ Hotshot-XL: State-of-the-art AI text-to-GIF model trained to work alongside Stable Diffusion XL
โ 1.1kHPSv2. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
โ 677lm-evaluation-harness. A framework for few-shot evaluation of language models.
โ 13kBayesian-Red-Teaming. About Official PyTorch implementation of "Query-Efficient Black-Box Red Teaming via Bayesian Optimization" (ACL'23)
โ 15streaming. A Data Streaming Library for Efficient Neural Network Training
โ 1.5kNeMo-Framework-Launcher. Provides end-to-end model development pipelines for LLMs and Multimodal models that can be launched on-prem or cloud-native.
โ 521facechain. FaceChain is a deep-learning toolchain for generating your Digital-Twin.
โ 9.5knougat. Implementation of Nougat Neural Optical Understanding for Academic Documents
โ 10kLVDM. LVDM: Latent Video Diffusion Models for High-Fidelity Long Video Generation
โ 503Qwen. The official repo of Qwen (้ไนๅ้ฎ) chat & pretrained large language model proposed by Alibaba Cloud.
โ 22kcog-sdxl-webui. Stable Diffusion XL training and inference as a cog model
โ 13MOSS-RLHF. Secrets of RLHF in Large Language Models Part I: PPO
โ 1.4kSeedSelect. Code for our papers : "Generating images of rare concepts using pre-trained diffusion models" (AAAI 24) and "Norm-guided latent space exploration for text-to-image generation" (Neurips 23)
โ 87cog-sdxl. Stable Diffusion XL training and inference as a cog model
โ 234ToolLearningPapers.
โ 923awesome-information-geometry. About A collection of AWESOME things about information geometry Topics
โ 195LISA. Project Page for "LISA: Reasoning Segmentation via Large Language Model"
โ 2.7kGPU-Puzzles. Solve puzzles. Learn CUDA.
โ 12kAwesome-LLM-Robotics. A comprehensive list of papers using large language/multi-modal models for Robotics/RL, including papers, codes, and related websites
โ 4.4kaisys2023.
โ 102Awesome_Quadrupedal_Robots. Awesome Quadrupedal Robots
โ 1.1kawesome-test-time-adaptation. Collection of awesome test-time (domain/batch/instance) adaptation methods
โ 1.3kawesome-llm-security. A curation of awesome tools, documents and projects about LLM Security.
โ 1.7kMorphoSymm. Tools for exploiting Morphological Symmetries in robotics
โ 108choreographer. [ICLR 2023] Choreographer: a world-model-based agent that discovers and learns unsupervised skills in latent imagination, and it's able to efficiently coordinate and adapt the skills to solve downstream tasks.
โ 42sd-webui-3d-editor. A custom extension for sd-webui that with 3D modeling features (add/edit basic elements, load your custom model, modify scene and so on), then send screenshot to txt2img or img2img as your ControlNet's reference image, basing on ThreeJS editor
โ 143viper_rl. Using advances in generative modeling to learn reward functions from unlabeled videos.
โ 143consistency_models_cifar10. Consistency models trained on CIFAR-10, in JAX.
โ 153audiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
โ 24kFlagAI. FlagAI (Fast LArge-scale General AI models) is a fast, easy-to-use and extensible toolkit for large-scale model.
โ 3.9kefficient-ai-study.
โ 90Wuerstchen. Official implementation of Wรผrstchen: Efficient Pretraining of Text-to-Image Models
โ 555Minari. A standard format for offline reinforcement learning datasets, with popular reference datasets and related utilities
โ 1.3kMDT. Masked Diffusion Transformer is the SOTA for image synthesis. (ICCV 2023)
โ 596google-research. Google Research
โ 38ktrident. A performance library for machine learning applications.
โ 183CelebBasis. Official Implementation of 'Inserting Anybody in Diffusion Models via Celeb Basis'
โ 255prolificdreamer. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation (NeurIPS 2023 Spotlight)
โ 1.6kToolBench. [ICLR'24 spotlight] An open platform for training, serving, and evaluating large language model for tool learning.
โ 5.7kProFusion. Code for Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach
โ 468KORani. Python
โ 108Asyrp_official. official repo for Asyrp : Diffusion Models already have a Semantic Latent Space (ICLR2023)
โ 289fastcomposer. [IJCV] FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention
โ 715ImageBind. ImageBind One Embedding Space to Bind Them All
โ 9.1kfairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
โ 32kopen_llama. OpenLLaMA, a permissively licensed open source reproduction of Meta AIโs LLaMA 7B trained on the RedPajama dataset
โ 7.5knumba. NumPy aware dynamic Python compiler using LLVM
โ 11kGraphit. Official Pytorch implementation of "Graphit: A Unified Framework for Diverse Image Editing Tasks"
โ 200IF. Python
โ 7.8klora. Using Low-rank adaptation to quickly fine-tune diffusion models.
โ 7.5kStableLM. StableLM: Stability AI Language Models
โ 16kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
โ 13kget-started-with-JAX. The purpose of this repo is to make it easy to get started with JAX, Flax, and Haiku. It contains my "Machine Learning with JAX" series of tutorials (YouTube videos and Jupyter Notebooks) as well as the content I found useful while learning about the JAX ecosystem.
โ 783mujoco_mpc. Real-time behaviour synthesis with MuJoCo, using Predictive Control
โ 1.7ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
โ 55kdreamerv3-torch. Implementation of Dreamer v3 in pytorch.
โ 887Kandinsky-2. Kandinsky 2 โ multilingual text2image latent diffusion model
โ 2.8kdreamerv3. Mastering Diverse Domains through World Models
โ 3.6kIsaacGymEnvs. Isaac Gym Reinforcement Learning Environments
โ 3kcleanrl. High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
โ 10keai-vc. The repository for the largest and most comprehensive empirical study of visual foundation models for Embodied AI (EAI).
โ 509AlpacaDataCleaned. Alpaca dataset from Stanford, cleaned and curated
โ 1.6k3DFuse. Official implementation of "Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation"
โ 732EVA. EVA Series: Visual Representation Fantasies from BAAI
โ 2.7kdiffusers-play. Repository with which to explore k-diffusion and diffusers, and within which changes to said packages may be tested.
โ 55champagne. An official codebase for paper ":champagne: CHAMPAGNE: Learning Real-world Conversation from Large-Scale Web Videos (ICCV 23)"
โ 52EVAL. EVAL(Elastic Versatile Agent with Langchain) will execute all your requests. Just like an eval method!
โ 866rl-book. Source codes for the book "Reinforcement Learning: Theory and Python Implementation"
โ 1kcloob-latent-diffusion. CLOOB Conditioned Latent Diffusion training and inference code
โ 113unidiffuser. Code and models for the paper "One Transformer Fits All Distributions in Multi-Modal Diffusion"
โ 1.5kELITE. ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation (ICCV 2023, Oral)
โ 541pytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
โ 102kprismer. The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".
โ 1.3kGenLoco. Python
โ 297composer. Official implementation of "Composer: Creative and Controllable Image Synthesis with Composable Conditions"
โ 1.6kcog. Containers for machine learning
โ 9.4kllama. Inference code for Llama models
โ 60kreduce_reuse_recycle. ICML 2023: Reduce, Reuse, Recycle: Composing Energy-Based Diffusion Models with MCMC
โ 151T2I-Adapter. T2I-Adapter
โ 3.8kscribble-diffusion. Turn your rough sketch into a refined image using AI
โ 3kUniversal-Guided-Diffusion. Jupyter Notebook
โ 511ControlNet. Let us control diffusion models!
โ 34krl. A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.
โ 3.5kINT8-Flash-Attention-FMHA-Quantization. Cuda
โ 165plug-and-play. Official Pytorch Implementation for โPlug-and-Play Diffusion Features for Text-Driven Image-to-Image Translationโ (CVPR 2023)
โ 1khard-prompts-made-easy. Python
โ 648X-Decoder. [CVPR 2023] Official Implementation of X-Decoder for generalized decoding for pixel, image and language
โ 1.3kpix2pix-zero. Zero-shot Image-to-Image Translation [SIGGRAPH 2023]
โ 1.1klora-training. LoRA training model packaged with Cog
โ 114lora-inference. LoRA inference model packaged with Cog
โ 74Attend-and-Excite. Official Implementation for "Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models" (SIGGRAPH 2023)
โ 770stylegan-t. [ICML'23] StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis
โ 1.2kdadaptation. D-Adaptation for SGD, Adam and AdaGrad
โ 531tuning_playbook. A playbook for systematically maximizing the performance of deep learning models.
โ 30kinstruct-pix2pix. Python
โ 6.9kprojUNN. Fast training of unitary deep network layers from low-rank updates
โ 30DeepFashion2. DeepFashion2 Dataset https://arxiv.org/pdf/1901.07973.pdf
โ 2.6kimg2dataset. Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
โ 4.4ksd-leap-booster. Fast finetuning using a booster model that puts the initial state to a local minimum
โ 113sd-webui-additional-networks. Python
โ 1.8kGLM-130B. GLM-130B: An Open Bilingual Pre-Trained Model (ICLR 2023)
โ 7.7ksimple-DVMPC-implemetation. Python
โ 24polyglot. Polyglot: Large Language Models of Well-balanced Competence in Multi-languages
โ 487pix2latent. Code for: Transforming and Projecting Images into Class-conditional Generative Networks
โ 195Gradient-Free-Textual-Inversion. Gradient-Free Textual Inversion for Personalized Text-to-Image Generation
โ 44mediapipe. Cross-platform, customizable ML solutions for live and streaming media.
โ 36kDiffusionDisentanglement. Official implementation of the paper "Uncovering the Disentanglement Capability in Text-to-Image Diffusion Models
โ 174portfolio-ideas. A curation of awesome portfolio website ideas for developers and designers to draw inspiration from. Raise a pull request to add more. ๐
โ 6.2kAwesome-Dataset-Distillation. A curated list of awesome papers on dataset distillation and related applications.
โ 2knestjs-oauth2. NestJS module for customizable, multiple OAuth2 clients
โ 5bbangjo.kr. MDX
โ 3buffalo. blog landing page(threejs + Typescript + blender)
โ 16Dreambooth. Fine-tuning of diffusion models
โ 97lora. Using Low-rank adaptation to quickly fine-tune diffusion models.
โ 2multimodal. TorchMultimodal is a PyTorch library for training state-of-the-art multimodal multi-task models at scale.
โ 1.7kgradient-descent-the-ultimate-optimizer. Code for our NeurIPS 2022 paper
โ 370pbrt-v4. Source code to pbrt, the ray tracer described in the forthcoming 4th edition of the "Physically Based Rendering: From Theory to Implementation" book.
โ 3.6kstochastic_process_2022W.
โ 8Piece-of-Cake. Make your investment life easier.
โ 3react-infinite-scroller. โฌ Infinite scroll component for React in ES6
โ 3.3klovely-tensors. Tensors, for human consumption
โ 1.4kstable-diffusion-webui. Stable Diffusion web UI
โ 164kgatsby-advanced-starter. A high performance skeleton starter for GatsbyJS with an advanced feature set.
โ 1.6kflash-attention. Fast and memory-efficient exact attention
โ 25kcuda-samples. Samples for CUDA Developers which demonstrates features in CUDA Toolkit
โ 9.4kLoRA. Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"
โ 14kpaint-with-words-sd. Implementation of Paint-with-words with Stable Diffusion : method from eDiff-I that let you generate image from text-labeled segmentation map.
โ 646AITemplate. AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.
โ 4.7kText2LIVE. Official Pytorch Implementation for "Text2LIVE: Text-Driven Layered Image and Video Editing" (ECCV 2022 Oral)
โ 886paint-with-words-sd. Unofficial Implementation of Paint-with-words, method from eDiffi that let you generate image from text-labeled segmentation map.
โ 4GODEL. Large-scale pretrained models for goal-directed dialog
โ 880torchrec. Pytorch domain library for recommendation systems
โ 2.6kmagicmix. Unofficial Implementation of MagicMix
โ 98bitsandbytes. Accessible large language models via k-bit quantization for PyTorch.
โ 8.4kequiformer-pytorch. Implementation of the Equiformer, SE3/E3 equivariant attention network that reaches new SOTA, and adopted for use by EquiFold for protein folding
โ 290stable-diffusion-pytorch. Yet another PyTorch implementation of Stable Diffusion (probably easy to read)
โ 592gimi. Cloud-based malicious URL detection system.
โ 3NATTEN. Fast Multi-dimensional Sparse Attention
โ 779Neighborhood-Attention-Transformer. Neighborhood Attention Transformer, arxiv 2022 / CVPR 2023. Dilated Neighborhood Attention Transformer, arxiv 2022
โ 1.2kPoisson_flow. Code for NeurIPS 2022 Paper, "Poisson Flow Generative Models" (PFGM)
โ 874webgl_study. JavaScript
โ 5RecBole. A unified, comprehensive and efficient recommendation library
โ 4.5kwhisper. Robust Speech Recognition via Large-Scale Weak Supervision
โ 106knerf-factory. An awesome PyTorch NeRF library
โ 1.3ksurrealdb. A scalable, distributed, collaborative, document-graph database, for the realtime web
โ 33kDreambooth-Stable-Diffusion. Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
โ 7.7kCrossAttentionControl. Unofficial implementation of "Prompt-to-Prompt Image Editing with Cross Attention Control" with Stable Diffusion
โ 1.3ksd-various-ideas. Jupyter Notebook
โ 55nestjs-slack-listener. The NestJS module for building the listeners and handlers of Slack events and interactivities.
โ 26shortest-even-cycle. Python
โ 9RWKV-LM. RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
โ 15kLIVE-Layerwise-Image-Vectorization. [CVPR 2022 Oral] Towards Layer-wise Image Vectorization
โ 619sygil-webui. Stable Diffusion web UI
โ 7.9kstable-diffusion.
โ 1.7kGFPGAN. GFPGAN aims at developing Practical Algorithms for Real-world Face Restoration.
โ 38kauto-animate. A zero-config, drop-in animation utility that adds smooth transitions to your web app. You can use it with React, Vue, or any other JavaScript application.
โ 14kwandb. The AI developer platform. Use Weights & Biases to train and fine-tune models, and manage models from experimentation to production.
โ 11kthree.js. JavaScript 3D Library.
โ 114kPipeDream. Parallel Training to PipeDream
โ 4stable-diffusion. A latent text-to-image diffusion model
โ 73kkaolin. A PyTorch Library for Accelerating 3D Deep Learning Research
โ 5.1kfluentui-emoji. A collection of familiar, friendly, and modern emoji from Microsoft
โ 10k