This is your work, valued
Machine Learning and AI Lead
awesome-feature-engineering. A curated list of resources dedicated to Feature Engineering Techniques for Machine Learning
★ 598retrivex. Explainability toolkit for retrieval models. Explain prediction of vector search models (embeddings similarity models, siamese encoders, bi-encoders, dense retrieval models). Debug your vector search models for RAG or agentic AI system.
★ 15awesome-nlp. :book: A curated list of resources dedicated to Natural Language Processing (NLP)
★ 5awesome-deep-vision. A curated list of deep learning resources for computer vision
★ 1awesome-rnn. Recurrent Neural Network - A curated list of resources dedicated to RNN
★ 1bdm-tool. Simple lightweight dataset versioning utility based purely on the file system and symbolic links.
★ 1awesome-computer-vision. A curated list of awesome computer vision resources
★ 1ir-examples. Examples of problem-solving in Information Retrieval, RAG, and Agentic Systems
★ 1MSformer. A novel molecular representation framework via meta structures
★ 18Search-R1. Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL
★ 5.2kopenfold. Trainable, memory-efficient, and GPU-friendly PyTorch reproduction of AlphaFold 2
★ 3.4klerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kir-examples. Examples of problem-solving in Information Retrieval, RAG, and Agentic Systems
★ 1beir. A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
★ 2.3kretrivex. Explainability toolkit for retrieval models. Explain prediction of vector search models (embeddings similarity models, siamese encoders, bi-encoders, dense retrieval models). Debug your vector search models for RAG or agentic AI system.
★ 15minihack. MiniHack the Planet: A Sandbox for Open-Ended Reinforcement Learning Research
★ 43teemi. teemi: A Python package for reproducible and FAIR microbial strain construction. Simulate the entire dbtl-cycle, generate genetic parts, design libraries, and track samples. Open-source Python platform for workflow flexibility and automated tasks, accelerating metabolic engineering. Try teemi with our Google Colab notebooks!
★ 44lm-polygraph. Python
★ 498nanochat. The best ChatGPT that $100 can buy.
★ 57kbdm-tool. Simple lightweight dataset versioning utility based purely on the file system and symbolic links.
★ 1RFdiffusion. Code for running RFdiffusion
★ 3kgraphframes-rs. GraphFrames but in DataFusion
★ 27chispa. PySpark test helper methods with beautiful error messages
★ 771incubator-graphar. An open source, standard data file format for graph data storage and retrieval.
★ 370graphframes. GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs
★ 1.2kCTranslate2. Fast inference engine for Transformer models
★ 4.6kLLMs-Planning. An extensible benchmark for evaluating large language models on planning
★ 470SKURG. Python
★ 20LyCORIS. Lora beYond Conventional methods, Other Rank adaptation Implementations for Stable diffusion.
★ 2.5kt-few. Code for T-Few from "Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning"
★ 460langchain. The agent engineering platform.
★ 143kmlc-llm. Universal LLM Deployment Engine with ML Compilation
★ 23kmedmcqa. A large-scale (194k), Multiple-Choice Question Answering (MCQA) dataset designed to address realworld medical entrance exam questions.
★ 287GENRE. Autoregressive Entity Retrieval
★ 801ThoughtSource. A central, open resource for data and tools related to chain-of-thought reasoning in large language models. Developed @ Samwald research group: https://samwald.info/
★ 1kmedical-reasoning. Medical reasoning using large language models
★ 94knnlm. Python
★ 331quart. An async Python micro framework for building web applications.
★ 3.7kalfred. ALFRED - A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
★ 526hivemind. Decentralized deep learning in PyTorch. Built to train models on thousands of volunteers across the world.
★ 2.5kpygaggle. a gaggle of deep neural architectures for text ranking and question answering, designed for Pyserini
★ 354pyserini. Pyserini is a Python toolkit for reproducible information retrieval research with sparse and dense representations.
★ 2.1kcomposer. Supercharge Your Model Training
★ 5.5kivy. Convert Machine Learning Code Between Frameworks
★ 14kir_datasets. Provides a common interface to many IR ranking datasets.
★ 390AR2. Python
★ 71haystack. Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search, and conversational systems.
★ 26kawesome-feature-engineering. A curated list of resources dedicated to Feature Engineering Techniques for Machine Learning
★ 598awesome-open-mlops. The Fuzzy Labs guide to the universe of open source MLOps
★ 482grobid. A machine learning software for extracting information from scholarly documents
★ 5kAutoDeploy. AutoDeploy is a single configuration deployment library
★ 40ludwig. Low-code framework for building custom LLMs, neural networks, and other AI models
★ 12kRECON. This is the code for the paper 'RECON: Relation Extraction using Knowledge Graph Context in a Graph Neural Network'.
★ 40FlowQA. Implementation of conversational QA model: FlowQA (with slight improvement)
★ 196minecraft-dialogue-models. Python
★ 6ruletaker. Python
★ 56milvus. Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
★ 45knmslib. Non-Metric Space Library (NMSLIB): An efficient similarity search library and a toolkit for evaluation of k-NN methods for generic non-metric spaces.
★ 3.6kpet. This repository contains the code for "Exploiting Cloze Questions for Few-Shot Text Classification and Natural Language Inference"
★ 1.6kbort. Repository for the paper "Optimal Subarchitecture Extraction for BERT"
★ 470pytorch_geometric. Graph Neural Network Library for PyTorch
★ 24kariadne. Jupyter Notebook
★ 10serve. ☁️ Build multimodal AI applications with cloud-native stack
★ 22klong-range-arena. Long Range Arena for Benchmarking Efficient Transformers
★ 788quest_qa_labeling. Google QUEST Q&A Labeling. Improving automated understanding of complex question answer content
★ 249kaggle_tweet. Kaggle Tweet Sentiment Extraction Competition: 1st place solution (Dark of the Moon team)
★ 72onnxruntime. ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
★ 21kConvBert. Python
★ 257RL-MMR. Code for "Multi-document Summarization with Maximal Marginal Relevance-guided Reinforcement Learning", EMNLP 2020
★ 13TextAttack. TextAttack 🐙 is a Python framework for adversarial attacks, data augmentation, and model training in NLP https://textattack.readthedocs.io/en/master/
★ 3.5kMHGRN. Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question Answering (EMNLP 2020)
★ 256tag-based-multi-span-extraction. The official implementation of EMNLP 2020, "A Simple and Effective Model for Answering Multi-span Questions".
★ 158self_talk. Code and data for the paper: "Unsupervised Common Sense Question Answering with Self-Talk"
★ 79KERMIT. 🐸 KERMIT - A lightweight library to encode and interpret Universal Syntactic Embeddings
★ 57low-resource-text-classification-framework. Research framework for low resource text classification that allows the user to experiment with classification models and active learning strategies on a large number of sentence classification datasets, and to simulate real-world scenarios. The framework is easily expandable to new classification models, active learning strategies and datasets.
★ 101abs_pretraining. Python
★ 10techqa. The TechQA dataset -- http://ibm.biz/Tech_QA
★ 31diseaseBERT. Code and dataset of EMNLP 2020 paper "Infusing Disease Knowledge into BERT for Health Question Answering, Medical Inference and Disease Name Recognition"
★ 69longformer. Longformer: The Long-Document Transformer
★ 2.2ktf-explain. Interpretability Methods for tf.keras models with Tensorflow 2.x
★ 1kinterpret-text. A library that incorporates state-of-the-art explainers for text-based machine learning models and visualizes the result with a built-in dashboard.
★ 431interpretability-tutorial-emnlp2020. Materials for the EMNLP 2020 Tutorial on "Interpreting Predictions of NLP Models"
★ 199SelectiveMasking. Source code for "Train No Evil: Selective Masking for Task-Guided Pre-Training"
★ 70rasa. 💬 Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants
★ 21kru-gpts. Russian GPT3 models.
★ 2.1kFlexNeuART. Flexible classic and NeurAl Retrieval Toolkit
★ 224TensorPipe. High Performance Tensorflow Data Pipeline with State of Art Augmentations and low level optimizations.
★ 85TableBank. TableBank: A Benchmark Dataset for Table Detection and Recognition
★ 1.1krecommenders. Best Practices on Recommendation Systems
★ 22ksentence-transformers. State-of-the-Art Embeddings, Retrieval, and Reranking
★ 19kdocker-python. Kaggle Python docker image
★ 2.7kmodels. Officially maintained, supported by PaddlePaddle, including CV, NLP, Speech, Rec, TS, big models and so on.
★ 6.9kERNIE. Source code and dataset for ACL 2019 paper "ERNIE: Enhanced Language Representation with Informative Entities"
★ 1.4klanguage. Shared repository for open-sourced projects from the Google AI Language team.
★ 1.8kkedro. Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.
★ 11kMTM. MTM
★ 143mmner. Massively Multilingual Transfer for NER
★ 86pyod. A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic workflow for AI agents.
★ 9.9kmachine-learning-systems-design. A booklet on machine learning systems design with exercises. NOT the repo for the book "Designing Machine Learning Systems", which is `dmls-book`
★ 10kclip-as-service. 🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
★ 13kuda. Unsupervised Data Augmentation (UDA)
★ 2.2ktvm. Open Machine Learning Compiler Framework
★ 14kHalide. a language for fast, portable data-parallel computation
★ 6.6kTASO. The Tensor Algebra SuperOptimizer for Deep Learning
★ 743probability. Probabilistic reasoning and statistical analysis in TensorFlow
★ 4.4kgoogle-research. Google Research
★ 38kPLMpapers. Must-read Papers on pre-trained language models.
★ 3.4kulmfit_keras. Keras wikipedia-based Language Model
★ 21transformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kawd-lstm-lm. LSTM and QRNN Language Model Toolkit for PyTorch
★ 2kTabNine. AI Code Completions
★ 11knni. An open source AutoML toolkit for automate machine learning lifecycle, including feature engineering, neural architecture search, model compression and hyper-parameter tuning.
★ 14kscipy. SciPy library main repository
★ 15kkenlm. KenLM: Faster and Smaller Language Model Queries
★ 2.8kscibert. A BERT model for scientific text.
★ 1.7knlpaug. Data augmentation for NLP
★ 4.7kAutoPhrase. AutoPhrase: Automated Phrase Mining from Massive Text Corpora
★ 1.2kLM-LSTM-CRF. Empower Sequence Labeling with Task-Aware Language Model
★ 848LightNER. Inference with state-of-the-art models (pre-trained by LD-Net / AutoNER / VanillaNER / ...)
★ 118SegPhrase. C++
★ 263AutoNER. Learning Named Entity Tagger from Domain-Specific Dictionary
★ 484multihead_joint_entity_relation_extraction. Implementation of our papers Joint entity recognition and relation extraction as a multi-head selection problem (Expert Syst. Appl, 2018) and Adversarial training for multi-context joint entity and relation extraction (EMNLP, 2018).
★ 413LSTM-ER. Implementation of End-to-End Relation Extraction using LSTMs on Sequences and Tree Structures in ACL2016.
★ 227emnlp2017-relation-extraction. Context-Aware Representations for Knowledge Base Relation Extraction
★ 289tacred-relation. PyTorch implementation of the position-aware attention model for relation extraction
★ 359awesome-domain-adaptation. A collection of AWESOME things about domain adaptation
★ 5.5kawesome-transfer-learning. Best transfer learning and domain adaptation resources (papers, tutorials, datasets, etc.)
★ 1.8kawesome-quantum-machine-learning. Here you can get all the Quantum Machine learning Basics, Algorithms ,Study Materials ,Projects and the descriptions of the projects around the web
★ 3.6kpytorch-acnn-model. code of Relation Classification via Multi-Level Attention CNNs
★ 178Relation-Classification-using-Bidirectional-LSTM-Tree. TensorFlow Implementation of the paper "End-to-End Relation Extraction using LSTMs on Sequences and Tree Structures" and "Classifying Relations via Long Short Term Memory Networks along Shortest Dependency Paths" for classifying relations
★ 189RESIDE. EMNLP 2018: RESIDE: Improving Distantly-Supervised Neural Relation Extraction using Side Information
★ 249FewRel. A Large-Scale Few-Shot Relation Extraction Dataset
★ 747OpenNRE. An Open-Source Package for Neural Relation Extraction (NRE)
★ 4.5kNREPapers. Must-read papers on neural relation extraction (NRE)
★ 1kawesome-relation-extraction. 📖 A curated list of awesome resources dedicated to Relation Extraction, one of the most important tasks in Natural Language Processing (NLP).
★ 1.2kbran. Full abstract relation extraction from biological texts with bi-affine relation attention networks
★ 131awesome-bert. bert nlp papers, applications and github resources, including the newst xlnet , BERT、XLNet 相关论文和 github 项目
★ 1.8kgpt-2. Code for the paper "Language Models are Unsupervised Multitask Learners"
★ 25keda_nlp. Data augmentation for NLP, presented at EMNLP 2019
★ 1.7kawesome-annotation. List of online / computer-based annotation tools
★ 18kaggle_salt_bes_phalanx. Winning solution for the Kaggle TGS Salt Identification Challenge.
★ 356pytorch2keras. PyTorch to Keras model convertor
★ 862BERT-keras. Keras implementation of BERT with pre-trained weights
★ 813Russian-Language-Model. Jupyter Notebook
★ 56Kaggle-Carvana-3rd-Place-Solution. 3rd place solution ( Carvana Image Masking Challenge )
★ 218kaggle_carvana_segmentation. Code for the 1st place model in Carvana Image Masking Challenge
★ 448data_science_bowl_2018. My 5th place (out of 816 teams) solution to The 2018 Data Science Bowl organized by Booz Allen Hamilton
★ 157kaggle_2018_data_science_bowl_solution. 4th place solution for 2018 Data Science Bowl kaggle competition
★ 45DSB_2018. Python
★ 622018DSB. 2018 Data Science Bowl 2nd Place Solution
★ 103dsb2018_topcoders. DSB2018 [ods.ai] topcoders
★ 418TGS-Salt-Identification. TGS Salt Identification
★ 344toxic. Toxic Comment Classification Challenge, 12th place solution https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge
★ 78KagglePlanetPytorch. 9th place solution for https://www.kaggle.com/c/planet-understanding-the-amazon-from-space
★ 23Kaggle-Planet-Understanding-the-Amazon-from-Space. 3rd place solution
★ 63pseudo_label_keras. Implementation of the pseudo-label algorithm with keras
★ 18Keras-RetinaNet-for-Open-Images-Challenge-2018. Code for 15th place in Kaggle Google AI Open Images - Object Detection Track
★ 269pytorch-saltnet. Kaggle | 9th place single model solution for TGS Salt Identification Challenge
★ 281argus-tgs-salt. Kaggle | 14th place solution for TGS Salt Identification Challenge
★ 77100-nlp-papers. 100 Must-Read NLP Papers
★ 3.8ktqdm. :zap: A Fast, Extensible Progress Bar for Python and CLI
★ 31khyperopt. Distributed Asynchronous Hyperparameter Optimization in Python
★ 7.6ktest-tube. Python library to easily log experiments and parallelize hyperparameter search for neural networks
★ 736dtreeviz. A python library for decision tree visualization and model interpretation.
★ 3.2ksacred. Sacred is a tool to help you configure, organize, log and reproduce experiments developed at IDSIA.
★ 4.4kcatalyst. Accelerated deep learning R&D
★ 3.4kLovaszSoftmax. Code for the Lovász-Softmax loss (CVPR 2018)
★ 1.4kfeaturetools. An open source python library for automated feature engineering
★ 7.7kawesome-monorepo. A curated list of awesome Monorepo tools, software and architectures.
★ 5.8kGeospatial_Data_with_Python. Introduction to Geospatial Data with Python
★ 182selective_search_py. Python-based implementation of the Selective Search for Object Recognition.
★ 374autokeras. AutoML library for deep learning
★ 9.3kTernausNet. UNet model with VGG11 encoder pre-trained on Kaggle Carvana dataset
★ 1.1kcrfasrnn. This repository contains the source code for the semantic image segmentation method described in the ICCV 2015 paper: Conditional Random Fields as Recurrent Neural Networks. http://crfasrnn.torr.vision/
★ 1.3kunet. unet for image segmentation
★ 4.9ksalesforce-einstein-sentiment-analysis. Demo for Salesforce Einstein Sentiment Analysis (Natural Language Processing)
★ 18Image-OutPainting. 🏖 Keras Implementation of Painting outside the box
★ 1.1kcatboost. A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
★ 9kalbumentations. Fast and flexible image augmentation library. Paper about the library: https://www.mdpi.com/2078-2489/11/2/125
★ 15kbpemb. Pre-trained subword embeddings in 275 languages, based on Byte-Pair Encoding (BPE)
★ 1.2ksentencepiece. Unsupervised text tokenizer for Neural Network-based text generation.
★ 12kNLP-progress. Repository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the most common NLP tasks.
★ 23kkaggle-avazu. Code for the 3rd place finish for Avazu Click-Through Rate Prediction
★ 87ChatterBot. ChatterBot is a machine learning, conversational dialog engine for creating chat bots
★ 15kgated_cnn. Keras implementation of “Gated Linear Unit ”
★ 23DeepPavlov. An open source library for deep learning end-to-end dialog systems and chatbots.
★ 7kkeras-contrib. Keras community contributions
★ 1.6kNLP_Datasets. My NLP datasets for Russian language
★ 393anchor. Code for "High-Precision Model-Agnostic Explanations" paper
★ 816lime. Lime: Explaining the predictions of any machine learning classifier
★ 12kDetectron. FAIR's research platform for object detection research, implementing popular algorithms like Mask R-CNN and RetinaNet.
★ 26kjekyll. :globe_with_meridians: Jekyll is a blog-aware static site generator in Ruby
★ 52ktorch-rnn. Efficient, reusable RNNs and LSTMs for torch
★ 2.6kchar-rnn-tensorflow. Multi-layer Recurrent Neural Networks (LSTM, RNN) for character-level language models in Python using Tensorflow
★ 2.7kawesome-rnn. Recurrent Neural Network - A curated list of resources dedicated to RNN
★ 6.2kawesome-computer-vision. A curated list of awesome computer vision resources
★ 23kawesome-deep-vision. A curated list of deep learning resources for computer vision
★ 11kconvnetjs. Deep Learning in Javascript. Train Convolutional Neural Networks (or ordinary ones) in your browser.
★ 11kneuraltalk. NeuralTalk is a Python+numpy project for learning Multimodal Recurrent Neural Networks that describe images with sentences.
★ 5.5karxiv-sanity-preserver. Web interface for browsing, search and filtering recent arxiv submissions
★ 5.8ktestcontainers-scala. Docker containers for testing in scala
★ 667awesome-nlp. :book: A curated list of resources dedicated to Natural Language Processing (NLP)
★ 19kfastText. Library for fast text representation and classification.
★ 27kkafka. Apache Kafka - A distributed event streaming platform
★ 33kpandas. Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more
★ 49kbazel. a fast, scalable, multi-language and extensible build system
★ 26kconda. A system-level, binary package and environment manager running on all major operating systems and platforms.
★ 7.5kkubernetes. Production-Grade Container Scheduling and Management
★ 124k