This is your work, valued
aphantasia. CLIP + FFT/DWT/RGB = text to image/video
★ 789stylegan2ada. StyleGAN2-ada for practice
★ 181stylegan2. StyleGAN2 for practice
★ 171SDfu. Stable Diffusers for studies
★ 131stargan2. StarGAN2 for practice
★ 95SD. Stable Diffusion for studies
★ 14assembly. speculative experimental setup for exploring agentic narratives
★ 7FoleyCrafter. FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds. AI拟音大师,给你的无声视频添加生动而且同步的音效 😝
★ 6stylegan3. StyleGAN3 for practice
★ 6disspress. Dissociated Press
★ 4dx11-vvvv. DirectX11 Rendering within vvvv
★ 1lucent. Lucid library adapted for PyTorch
★ 1langchain-graphrag. GraphRAG / From Local to Global: A Graph RAG Approach to Query-Focused Summarization
★ 1723DCellForge. AI-powered interactive 3D model generation, inspection, and presentation studio.
★ 2.6kMiroFish. A Simple and Universal Swarm Intelligence Engine, Predicting Anything. 简洁通用的群体智能引擎,预测万物
★ 70kEmergence-World. Emergence World: A world designed to reveal what no benchmark can: emergent intelligence.
★ 553superpowers. An agentic skills framework & software development methodology that works.
★ 264kBiRefNet. [CAAI AIR'24] Bilateral Reference for High-Resolution Dichotomous Image Segmentation
★ 4ksoul.md. The best way to build a personality for your agent. Let Claude Code / OpenClaw ingest your data & build your AI soul.
★ 631cursed_browser. True AI-Native Browser — a VLM reads the HTML and hallucinates the page.
★ 273claude-code-best-practice. from vibe coding to agentic engineering - practice makes claude perfect
★ 64krelsim. 🍑 relsim: Relational Visual Similarity | pip install relsim 🌍 (CVPR 2026)
★ 87Paper2Video. Automatic Video Generation from Scientific Papers
★ 2.3kCode2Video. [ICML 2026] Video generation via code
★ 1.9kagenticSeek. Fully Local Manus AI. No APIs, No $200 monthly bills. Enjoy an autonomous agent that thinks, browses the web, and code for the sole cost of electricity.
★ 27kHYPIR. Official implementation of HYPIR: Harnessing Diffusion-Yielded Score Priors for Image Restoration (SIGGRAPH 2025)
★ 1.2kFlashVSR. [CVPR 2026] Towards Real-Time Diffusion-Based Streaming Video Super-Resolution — An efficient one-step diffusion framework for streaming VSR with locality-constrained sparse attention and a tiny conditional decoder.
★ 1.7ksoulx-livekit-avatar. LiveKit Talking Head Avatar with open AI S2S
★ 17Shadowbroker. Open-source intelligence for the global theater. Track everything from the corporate/private jets of the wealthy, and spy satellites, to seismic events in one unified interface. Hook an AI agent up to have it parse through data and find previously unseen correlations. The knowledge is available to all but rarely aggregated in the open, until now.
★ 10kExiv. Modular and extensible open source GenAI engine
★ 12ouroboros. Active mirror of https://github.com/razzant/ouroboros — open issues and PRs there.
★ 900infinite-ai-web. HTML
★ 10ComfyUI-Copilot. An AI-powered custom node for ComfyUI designed to enhance workflow automation and provide intelligent assistance
★ 5.4kautogen. A programming framework for agentic AI
★ 60kdify. Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
★ 151kWan2GP. A fast AI Video Generator for the GPU Poor. Supports Wan 2.1/2.2, LTX-2, Qwen Image, Hunyuan Video, LTX Video and Flux.
★ 7.1kmgrep. A calm, CLI-native way to semantically grep everything, like code, images, pdfs and more.
★ 4.3kuv. An extremely fast Python package and project manager, written in Rust.
★ 88kclaude-context-local. Code search MCP for Claude Code. Make entire codebase the context for any coding agent. Embeddings are created and stored locally. No API cost.
★ 8smolagents. 🤗 smolagents: a barebones library for agents that think in code.
★ 29kn8n. Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
★ 199kworkshops. One repository for some of the workshops I'm giving
★ 163dremio-local-agentic-ai-workshop. Python
★ 1graphiti. Build Real-Time Knowledge Graphs for AI Agents
★ 29kphoenix. AI Observability & Evaluation
★ 11kSelf-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5kLLM-Engineering-Essentials. Materials for the LLM Engineering Essentials course
★ 829servers. Model Context Protocol Servers
★ 89kunsloth. Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek, GLM and other models.
★ 69kpublic-talk-slides.
★ 3btnpy. Jupyter Notebook
★ 6sliderspace. SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
★ 124ComfyUI-TangoFlux. ComfyUI Custom Nodes for "TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching". This generates high-quality 44.1kHz audio up to 30 seconds using just a text prompt.
★ 107yt-dlp. A feature-rich command-line audio/video downloader
★ 181kagent.exe. TypeScript
★ 3.5kmochi. The best OSS video generation models, created by Genmo
★ 3.7kAllegro. Allegro is a powerful text-to-video model that generates high-quality videos up to 6 seconds at 15 FPS and 720p resolution from simple text input.
★ 1.1kPyramid-Flow. [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
★ 3.2kdeveloper. the first library to let you embed a developer agent in your own app!
★ 12kE2B. Open-source, secure environment with real-world tools for enterprise-grade agents.
★ 13kSyncthingWindowsSetup. Syncthing Windows Setup
★ 4kPatchFusion. [CVPR 2024] An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth Estimation
★ 1kComfyUI-GGUF. GGUF Quantization support for native ComfyUI models
★ 3.9kai-digest. A CLI tool to aggregate your codebase into a single Markdown file for use with Claude Projects or custom ChatGPTs.
★ 679ControlNetPlus. ControlNet++: All-in-one ControlNet for image generations and editing!
★ 2.1kaura-sr. AuraSR: GAN-based Super-Resolution for real-world
★ 526Fooocus. Focus on prompting and generating
★ 52kLGM. [ECCV 2024 Oral] LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation.
★ 2.1kComfyUI-to-Python-Extension. A powerful tool that translates ComfyUI workflows into executable Python code.
★ 2.4kFreeU. FreeU: Free Lunch in Diffusion U-Net (CVPR2024 Oral)
★ 1.9kMarkovJunior. Probabilistic language based on pattern matching and constraint propagation, 153 examples
★ 8.1kinsightface. State-of-the-art 2D and 3D Face Analysis Project
★ 29kPHDNet-Painterly-Image-Harmonization. [AAAI 2023] Painterly image harmonization in both spatial domain and frequency domain.
★ 55AnimateDiff. Official implementation of AnimateDiff.
★ 12kCoDeF. [CVPR'24 Highlight] Official PyTorch implementation of CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
★ 4.8kDDPM_inversion. Official pytorch implementation of the paper: "An Edit Friendly DDPM Noise Space: Inversion and Manipulations". CVPR 2024.
★ 368GigaGAN. JavaScript
★ 412ConceptLab. Official Implementation for "ConceptLab: Creative Generation using Diffusion Prior Constraints"
★ 255style2paints. sketch + style = paints :art: (TOG2018/SIGGRAPH2018ASIA)
★ 18knerfstudio. A collaboration friendly studio for NeRFs
★ 12kinstant-ngp. Instant neural graphics primitives: lightning fast NeRF and more
★ 18kComfyUI. The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
★ 123kainodes-engine. Python
★ 279IF. Python
★ 7.8kImageReward. [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1.7kKandinsky-2. Kandinsky 2 — multilingual text2image latent diffusion model
★ 2.8ktomesd. Speed up Stable Diffusion with this one simple trick!
★ 1.4kPython-SpoutGL. Python wrapper for SpoutGL
★ 36Text2Video-Zero. [ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
★ 4.2klatentblending. Create butter-smooth transitions between prompts, powered by stable diffusion
★ 365choreography. 'I didn’t want to imitate anybody. Any movement I knew, I didn’t want to use.' – Pina Bausch
★ 45diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kphycv. PhyCV: The First Physics-inspired Computer Vision Library
★ 527sd-leap-booster. Fast finetuning using a booster model that puts the initial state to a local minimum
★ 113custom-diffusion. Custom Diffusion: Multi-Concept Customization of Text-to-Image Diffusion (CVPR 2023)
★ 2kmuse-maskgit-pytorch. Implementation of Muse: Text-to-Image Generation via Masked Generative Transformers, in Pytorch
★ 918stable-diffusion. A latent text-to-image diffusion model
★ 67voltaML. ⚡VoltaML is a lightweight library to convert and run your ML/DL deep learning models in high performance inference runtimes like TensorRT, TorchScript, ONNX and TVM.
★ 1.2kpirounet. Conditional dance generator
★ 25Mubert-Text-to-Music. A simple notebook demonstrating prompt-based music generation via Mubert API
★ 2.7kEvoGen-Prompt-Evolution. Jupyter Notebook
★ 124Cold-Diffusion-Models. Official implementation of Cold-Diffusion for different transformations in pytorch.
★ 1.1kGFPGAN. GFPGAN aims at developing Practical Algorithms for Real-world Face Restoration.
★ 38kInvokeAI. Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest AI-driven technologies. The solution offers an industry leading WebUI, and serves as the foundation for multiple commercial products.
★ 28kstable-diffusion. A latent text-to-image diffusion model
★ 73kSymphonyNet. Symphony Generation with Permutation Invariant Language Model
★ 256basic-pitch. A lightweight yet powerful audio-to-MIDI converter with pitch bend detection
★ 5.4keg3d. Python
★ 3.3keqgan-sa. [CVPR 2022] Improving GAN Equilibrium by Raising Spatial Awareness
★ 137MaGNET. Official repository for MaGNET, ICLR 2022
★ 24aesthetic-predictor. A linear estimator on top of clip to predict the aesthetic quality of pictures
★ 728CogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kASSET. ASSET: Autoregressive Semantic Scene Editing with Transformers at High Resolutions (SIGGRAPH 2022 - Journal Track)
★ 112wombot. Unofficial API and discord bot for wombo.art
★ 34lucid_stylegan3_datasets_models. artistic stylegan3 datasets and models we created during our ongoing stylegan3 trip.
★ 36PyTorch-Deep-Dream. Minimal PyTorch implementation of Deep Dream
★ 147insgen. [NeurIPS 2021] Data-Efficient Instance Generation from Instance Discrimination
★ 104latent-diffusion. High-Resolution Image Synthesis with Latent Diffusion Models
★ 60PyTorch-OptimalStyleTransfer. An unofficial PyTorch implementation of paper "A Closed-form Solution to Universal Style Transfer - ICCV 2019"
★ 8v-diffusion-pytorch. v objective diffusion inference code for PyTorch.
★ 719lucid-sonic-dreams. Python
★ 774fourier_feature_nets. Supplemental learning materials for "Fourier Feature Networks and Neural Volume Rendering"
★ 179ic_gan. Official repository for the paper "Instance-Conditioned GAN" by Arantxa Casanova, Marlene Careil, Jakob Verbeek, Michał Drożdżal, Adriana Romero-Soriano.
★ 535PaddleFormers. PaddleFormers is an easy-to-use library of pre-trained large language model zoo based on PaddlePaddle.
★ 13kThemeTransformer. The official implementation of Theme Transformer. A Theme-based music generation. IEEE TMM
★ 126DeceiveD. [NeurIPS 2021] Deceive D: Adaptive Pseudo Augmentation for GAN Training with Limited Data
★ 234awesome-pretrained-stylegan3. A collection of pretrained models for StyleGAN3
★ 298shadertoy-render. Render a ShaderToy script directly to a video file.
★ 252MoCoGAN-HD. [ICLR 2021 Spotlight] A Good Image Generator Is What You Need for High-Resolution Video Synthesis
★ 246torchextractor. Feature extraction made simple with torchextractor
★ 101clipit. CLIP + VQGAN / PixelDraw
★ 285pixray. neural image generation
★ 402NansAreNumbers. An esoteric data type built entirely of NaNs.
★ 74articulated-animation. Code for Motion Representations for Articulated Animation paper
★ 1.3kknowyourdata. A tool to help researchers and product teams understand datasets with the goal of improving data quality, and mitigating fairness and bias issues.
★ 295rife-ncnn-vulkan. RIFE, Real-Time Intermediate Flow Estimation for Video Frame Interpolation implemented with ncnn library
★ 1.1kbuilding-a-poor-mans-supercomputer. I've built a 4x V100 box for less than $5,500.
★ 175StarGANv2-VC. StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion
★ 522MiXLab. Yet another multi-purpose Colab Notebook
★ 253style-transfer-pytorch. Neural style transfer in PyTorch.
★ 491copilot-clone. VSCode extension for code suggestion
★ 1.8kannotated_deep_learning_paper_implementations. 🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), gans(cyclegan, stylegan2, ...), 🎮 reinforcement learning (ppo, dqn), capsnet, distillation, ... 🧠
★ 67kMelSpecVAE. Variational Autoencoder in the mel-spectrogram domain for one-shot audio synthesis
★ 146stegcloak. Hide secrets with invisible characters in plain text securely using passwords 🧙🏻♂️⭐
★ 3.9kmotion-models.
★ 33silero-models. Silero Models: pre-trained text-to-speech models made embarrassingly simple
★ 6kFlare. Flaretic new repository, creative coding framework in c# and c++
★ 2torch-dreams. Flexible Feature visualization on PyTorch, for research and art :mag_right: :computer: :brain: :art:
★ 246gansformer. Generative Adversarial Transformers
★ 1.3kDISTS. IQA: Deep Image Structure and Texture Similarity Metric
★ 486mediapy. This Python library makes it easy to display images and videos in a notebook.
★ 450PlotNeuralNet. Latex code for making neural networks diagrams
★ 25kwit. WIT (Wikipedia-based Image Text) Dataset is a large multimodal multilingual dataset comprising 37M+ image-text sets with 11M+ unique images across 100+ languages.
★ 1.1ktfpyth. Putting TensorFlow back in PyTorch, back in TensorFlow (differentiable TensorFlow PyTorch adapters).
★ 646lucent. Lucid library adapted for PyTorch
★ 664CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kstylegan2-ada-pytorch. StyleGAN2-ADA - Official PyTorch implementation
★ 4.5kavatars4all. Live real-time avatars from your webcam in the browser. No dedicated hardware or software installation needed. A pure Google Colab wrapper for live First-order-motion-model, aka Avatarify in the browser. And other Colabs providing an accessible interface for using FOMM, Wav2Lip and Liquid-warping-GAN with your own media and a rich GUI.
★ 372minGPT. A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training
★ 25kNeuralBeatbox_ML_Examples. Jupyter Notebook
★ 14choreo-graph. Learning latent graph representations of the dancing body with GNNs
★ 27SariGAN. [NeurIPS'20] Learning Semantic-aware Normalization for Generative Adversarial Networks
★ 52Dain-App. Source code for Dain-App
★ 1.1kfew-shot-gan. Few-shot adaptation of GANs.
★ 127tn2-wg. Tacotron2 + Waveglow Russian
★ 43Audio-transcriptor-russian-. [Russian] This script will split audio file on silence, transcript it with google recognition and save it in LJSpeech-1.1 dataset manner.
★ 9vvvv-texturefx-workshops. 🖼️ Materials for compositing workshops at NODE20
★ 3VL.RunwayML. C#
★ 14colmap. COLMAP - Structure-from-Motion and Multi-View Stereo
★ 12krewriting. Rewriting a Deep Generative Model, ECCV 2020 (oral). Interactive tool to directly edit the rules of a GAN to synthesize scenes with objects added, removed, or altered. Change StyleGANv2 to make extravagant eyebrows, or horses wearing hats.
★ 535stylegan2-surgery. StyleGAN2 fork with scripts and convenience modifications for creative media synthesis
★ 139gan-vis. Visualization of GAN training process
★ 92SketchSynthVideo-Simple. A video version of SketchSynth.
★ 54Sequencer. An algorithm that detects one-dimensional sequences in complex datasets
★ 127data-efficient-gans. [NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
★ 1.3kScholarWithCode. A simple chrome extension to present the number of available code implementions (via Papers With Code) for articles listed on Google Scholar.
★ 88fourier-feature-networks. Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
★ 1.4kwebpages-to-ebook. Create an EPUB from a list of URLs. Standing on the shoulders of Wget, Readability and Pandoc.
★ 206AGD. [ICML 2020] "AutoGAN-Distiller: Searching to Compress Generative Adversarial Networks" by Yonggan Fu, Wuyang Chen, Haotao Wang, Haoran Li, Yingyan (Celine) Lin, and Zhangyang (Atlas) Wang.
★ 105Open4D. 4D visualization of dynamic events from unconstrained multi-view videos
★ 41obs-spout2-plugin. A Plugin for OBS Studio to enable Spout2 (https://github.com/leadedge/Spout2) input / output
★ 1kadjustable-real-time-style-transfer. Python
★ 14Pyno. Python-based visual programming
★ 172instagram-3d-photo. A Chrome extension that adds a 3d photo effect to instagram pages.
★ 672RigNet. Code for SIGGRAPH 2020 paper "RigNet: Neural Rigging for Articulated Characters"
★ 1.5kawesome-pretrained-stylegan. A collection of pre-trained StyleGAN models to download
★ 333awesome-pretrained-stylegan2. A collection of pre-trained StyleGAN 2 models to download
★ 1.3kgamma-course. vvvv gamma course, based on processing-course repo
★ 5avatarify-python. Avatars for Zoom, Skype and other video-conferencing apps.
★ 17kVL.Elementa. Collection of UI widgets for easy UI prototyping
★ 303d-photo-inpainting. [CVPR 2020] 3D Photography using Context-aware Layered Depth Inpainting
★ 7.1kGANalyze. The authors' official implementation of GANalyze, a framework for studying cognitive properties such as memorability, aesthetics, and emotional valence using genenerative models.
★ 131ganspace. Discovering Interpretable GAN Controls [NeurIPS 2020]
★ 1.8kPIFu. This repository contains the code for the paper "PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization"
★ 1.8kface-parsing.PyTorch. Using modified BiSeNet for face parsing in PyTorch
★ 2.6kpytorch-optimizer. torch-optimizer -- collection of optimizers for Pytorch
★ 3.2kBachGAN. Python
★ 62ru_transformers. Python
★ 764ruGPT2. Russian GPT2 model
★ 62Rotate-and-Render. Code for Rotate-and-Render: Unsupervised Photorealistic Face Rotation from Single-View Images (CVPR 2020)
★ 494synthesizing_human_like_sketches. Code for the WACV20 paper "Synthesizing human-like sketches from natural images using a conditional convolutional decoder"
★ 44LipGAN. This repository contains the codes for LipGAN. LipGAN was published as a part of the paper titled "Towards Automatic Face-to-Face Translation".
★ 616SinGAN. Official pytorch implementation of the paper: "SinGAN: Learning a Generative Model from a Single Natural Image"
★ 3.3kcolab-tricks. Tricks for Colab power users
★ 1683d-ken-burns. an implementation of 3D Ken Burns Effect from a Single Image using PyTorch
★ 1.6ktext-to-text-transfer-transformer. Code for the paper "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer"
★ 6.5ktransformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kcompress-fasttext. Tools for shrinking fastText models (in gensim format)
★ 187plink-plonk. Chrome Extension: A minimal auditory debugger for web pages
★ 47listen-to-transformer. An app to make it easier to explore and curate output from a Music Transformer
★ 118GANLatentDiscovery. The authors official implementation of Unsupervised Discovery of Interpretable Directions in the GAN Latent Space
★ 428pytorch_Realtime_Multi-Person_Pose_Estimation. Python
★ 1.4k