This is your work, valued
Kaggle-Ensemble-Guide. Code for the Kaggle Ensembling Guide Article on MLWave
★ 1.6kextremely-simple-one-shot-learning. Extremely simple one-shot learning in Python
★ 179kaggle-criteo. Kaggle Criteo https://www.kaggle.com/c/criteo-display-ad-challenge
★ 97hodor-autoML. Hodor AutoML: Brute-Bandit fast good-enough solutions to a wide range of machine learning problems.
★ 82Online-Random-Bit-Regression-FTRL. Online Random Bit Regression with FTRL-Proximal in Python
★ 75Black-Boxxy. Some experiments into explaining complex black box ensemble predictions.
★ 74koolmogorov. Koolmogorov is a Python library based on CompLearn
★ 69kaggle_acquire-valued-shoppers-challenge. Code for the Kaggle acquire valued shoppers challenge
★ 66RGF-sklearn. Scikit-learn API toy wrapper for Regularized Greedy Forests
★ 44online-learning-perceptron. An online learning perceptron benchmark for Kaggle movie review competition
★ 26Kaggle_Rotten_Tomatoes. Code to munge data between Kaggle .tsv Rotten Tomatoes Sentiment Analysis data set and Vowpal Wabbit
★ 24normalized-compression-neighbors. Document or binary file vectorization with Normalized Compression Distance in Python.
★ 17Kaggle-decoding-the-human-brain. Python code to beat the benchmark for the Kaggle competition "Decoding the human brain" using Vowpal Wabbit
★ 13Kaggle_Connectomics. Python code for the Pearson Correlation Benchmark with Discretization to estimate brain connectivity from neuron activity
★ 13NCD_Chess. Normalized Compression Distance and Chess Games
★ 8Kaggle-Papirusy-z-Edhellond. Winning code for the "Papirusy z Edhellond" Kaggle in Class competition
★ 8umap. Uniform Manifold Approximation and Projection
★ 3auto-sklearn. Python
★ 3notebooks. general notebook repository
★ 3sofia-ml. Automatically exported from code.google.com/p/sofia-ml
★ 2stop-words. Automatically exported from code.google.com/p/stop-words
★ 1kepler-mapper-fork.
★ 1classic-libqsearch. The qsearch Quartet Tree Search library for use with classic-complearn
★ 1flair. A very simple framework for state-of-the-art Natural Language Processing (NLP)
★ 14kNLP-Models-Tensorflow. Gathers machine learning and Tensorflow deep learning models for NLP problems, 1.13 < Tensorflow < 2.0
★ 1.8ktf-quant-finance. High-performance TensorFlow library for quantitative finance.
★ 5.5kFakeNewsNet. This is a dataset for fake news detection research
★ 1.3kWordMoversEmbeddings. WordMoversEmbeddings(WME) is a simple code for generating the vector representation of sentence/document for text classification and clustering.
★ 83Math. Virtual Math Lab
★ 72AIX360. Interpretability and explainability of data and machine learning models
★ 1.8kDeepGBM. SIGKDD'2019: DeepGBM: A Deep Learning Framework Distilled by GBDT for Online Prediction Tasks
★ 661enstop. Ensemble topic modelling with pLSA
★ 113SOHU_competition. Sohu's 2018 content recognition competition 1st solution(搜狐内容识别大赛第一名解决方案)
★ 227transfer-learning-conv-ai. 🦄 State-of-the-Art Conversational AI with Transfer Learning
★ 1.8kinterpret. Fit interpretable models. Explain blackbox machine learning.
★ 6.9kmicrosoft-graph-toolkit. Authentication Providers and UI components for Microsoft Graph 🦒
★ 1.1ksingle-parameter-fit. Real numbers, data science and chaos: How to fit any dataset with a single parameter
★ 650fklearn. fklearn: Functional Machine Learning
★ 1.5kakutan. A distributed knowledge graph store
★ 1.7kfusenet. Network inference by fusing data from diverse distributions
★ 14fast_knn_gpu. Tensorflow and Theano implementations of fast K-nearest neighbour search on GPU. ~3000x faster than sklearn KNeighborsClassifier!
★ 18EconML. ALICE (Automated Learning and Intelligence for Causation and Economics) is a Microsoft Research project aimed at applying Artificial Intelligence concepts to economic decision making. One of its goals is to build a toolkit that combines state-of-the-art machine learning techniques with econometrics in order to bring automation to complex causal inference problems. To date, the ALICE Python SDK (econml) implements orthogonal machine learning algorithms such as the double machine learning work of Chernozhukov et al. This toolkit is designed to measure the causal effect of some treatment variable(s) t on an outcome variable y, controlling for a set of features x.
★ 4.7kPyTorch-BigGraph. Generate embeddings from large-scale graph-structured data.
★ 3.5ktetrad. Repository for the Tetrad Project, www.phil.cmu.edu/tetrad.
★ 444kepler-mapper-backend-example. Demonstration of how to use Kepler Mapper as backend.
★ 2MapperInteractive. Mapper Interactive is a customizable visualization framework for the analysis and visualization of high-dimensional point cloud data using the Mapper algorithm.
★ 27fuzzilli. A JavaScript Engine Fuzzer
★ 2.3kKaggle-Ensemble-Guide. Code for the Kaggle Ensembling Guide Article on MLWave
★ 1.6kcobaltstrike-extraneous-space. Historical list of {Cobalt Strike,NanoHTTPD} servers
★ 121tsne-cuda. GPU Accelerated t-SNE for CUDA with Python bindings
★ 1.9kspreadingvectors. Open source implementation of "Spreading Vectors for Similarity Search"
★ 322recommenders. Best Practices on Recommendation Systems
★ 22kinfo8004-advanced-machine-learning. Lectures for INFO8004 Advanced Machine Learning, ULiège
★ 121ludwig. Low-code framework for building custom LLMs, neural networks, and other AI models
★ 12kBiMPM. BiMPM: Bilateral Multi-Perspective Matching for Natural Language Sentences
★ 434merf. Mixed Effects Random Forest
★ 244thundergbm. ThunderGBM: Fast GBDTs and Random Forests on GPUs
★ 715Spark-RSVD. Randomized SVD of large sparse matrices on Spark
★ 78semantic_information. Semantic information -- simple model of a food-seeking agent
★ 7DANMF. A sparsity aware implementation of "Deep Autoencoder-like Nonnegative Matrix Factorization for Community Detection" (CIKM 2018).
★ 206info8006-introduction-to-ai. Lectures for INFO8006 Introduction to Artificial Intelligence, ULiège
★ 414clip-as-service. 🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
★ 13kknnFeat. Python Implementation of Feature Extraction with K-Nearest Neighbor
★ 64NimbusML. Python machine learning package providing simple interoperability between ML.NET and scikit-learn components.
★ 292walklets. A lightweight implementation of Walklets from "Don't Walk Skip! Online Learning of Multi-scale Network Embeddings" (ASONAM 2017).
★ 105leidenalg. Implementation of the Leiden algorithm for various quality functions to be used with igraph in Python.
★ 787spider. scripts and baselines for Spider: Yale complex and cross-domain semantic parsing and text-to-SQL challenge
★ 1.1kapricot. apricot implements submodular optimization for the purpose of selecting subsets of massive data sets to train machine learning models quickly. See the documentation page: https://apricot-select.readthedocs.io/en/latest/index.html
★ 534petastorm. Petastorm library enables single machine or distributed training and evaluation of deep learning models from datasets in Apache Parquet format. It supports ML frameworks such as Tensorflow, Pytorch, and PySpark and can be used from pure Python code.
★ 1.9kAIF360. A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models.
★ 2.8ktreelite. Universal model exchange and serialization format for decision tree forests
★ 826actionable-recourse. python tools to check recourse in linear classification
★ 77openTSNE. Extensible, parallel implementations of t-SNE
★ 1.6kconformal. Tools for conformal inference in regression
★ 253botnets. This is a collection of #botnet source codes, unorganized. For EDUCATIONAL PURPOSES ONLY
★ 3.3kFIt-SNE. Fast Fourier Transform-accelerated Interpolation-based t-SNE (FIt-SNE)
★ 607scikit-tda. Topological Data Analysis for Python🐍
★ 574ripser.py. A Lean Persistent Homology Library for Python
★ 338swifter. A package which efficiently applies any function to a pandas dataframe or series in the fastest available manner
★ 2.6kImage-OutPainting. 🏖 Keras Implementation of Painting outside the box
★ 1.1kwolpert. A stacked generalization framework. Built on top of scikit learn
★ 58abstract-reasoning-matrices. Progressive matrices dataset, as described in: Measuring abstract reasoning in neural networks (Barrett*, Hill*, Santoro*, Morcos, Lillicrap), ICML2018
★ 187PHATE. PHATE (Potential of Heat-diffusion for Affinity-based Transition Embedding) is a tool for visualizing high dimensional data.
★ 557fairlearn. A Python package to assess and improve fairness of machine learning models.
★ 2.3klore. Lore makes machine learning approachable for Software Engineers and maintainable for Machine Learning Researchers
★ 1.5kgrf. Generalized Random Forests
★ 1.1kalgo-ds. Course material for Algorithms and Data Structures (TU Delft TI3110TU)
★ 11device_detector. Python
★ 146Sentiment-and-Style-Transfer. Roff
★ 248pix-plot. A WebGL viewer for UMAP or TSNE-clustered images
★ 649spark-tda. SparkTDA is a package for Apache Spark providing Topological Data Analysis Functionalities.
★ 46xlearn. High performance, easy-to-use, and scalable machine learning (ML) package, including linear model (LR), factorization machines (FM), and field-aware factorization machines (FFM) for Python and CLI interface.
★ 3.1kWord2Bits. Quantized word vectors that take 8x-16x less space than regular word vectors
★ 753knusperli. A deblocking JPEG decoder
★ 473Tutoriais-de-AM. Algoritmos de aprendizado de máquina criados manualmente para maior compreensão das suas funcionalidades
★ 76siamese-triplet. Siamese and triplet networks with online pair/triplet mining in PyTorch
★ 3.2kcuckoo. a memory-bound graph-theoretic proof-of-work system
★ 860pynndescent. A Python nearest neighbor descent for approximate nearest neighbors
★ 970sql-differential-privacy. Dataflow analysis & differential privacy for SQL queries. This project is deprecated and not maintained.
★ 402OpenEnsembles. Code for ensemble clustering
★ 103inspect_word2vec. Python code for checking out Google's pre-trained, 3M word Word2Vec model
★ 319six. Python 2 and 3 compatibility library
★ 1knocode. The best way to write secure and reliable applications. Write nothing; deploy nowhere.
★ 66kskift. scikit-learn wrappers for Python fastText.
★ 234anchor. Code for "High-Precision Model-Agnostic Explanations" paper
★ 816fg-data-profiling. 1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
★ 14kshap. A game theoretic approach to explain the output of any machine learning model.
★ 26kmath-as-code. a cheat-sheet for mathematical notation in code form
★ 15krulefit. Python implementation of the rulefit algorithm
★ 445gif_animations. 📼 GIF animations created in Python
★ 19pixelsound. 🖼️🎵 Hide grayscale image in spectrogram by encoding pixels with sine waves
★ 4decision_boundary_viz. 📈 Interactive decision boundary visualizer
★ 11Learn-Pandas. Tutorials on how to use pandas effectively to do data analysis
★ 1.1kskorch. A scikit-learn compatible neural network library that wraps PyTorch
★ 6.2kTrueNgine. A N-Dimensional Renderer
★ 12cartographer. Flexible, extensible and fully scikit-compatible Mapper/TDA algorithm implementation
★ 10TSNE-UMAP-Embedding-Visualisation. A Simple and easy to use way to Visualise Embeddings!
★ 259tauthon. Fork of Python 2.7 with new syntax, builtins, and libraries backported from Python 3.
★ 687fastai. The fastai deep learning library
★ 28kword2vec-slim. word2vec Google News model slimmed down to 300k English words
★ 217StarSpace. Learning embeddings for classification, retrieval and ranking.
★ 4kpredict-taxi-trip-duration. Predict taxi trip duration based on historical trips using automated feature engineering
★ 62bounter. Efficient Counter that uses a limited (bounded) amount of memory regardless of data size.
★ 930dionysus. Library for computing persistent homology
★ 160universe. The fastest way to query and explore multivariate datasets
★ 156facets. Visualizations for machine learning datasets
★ 7.3knumpy. The fundamental package for scientific computing with Python.
★ 32kfast-style-transfer-deeplearnjs. Demo of in-browser Fast Neural Style Transfer with deeplearn.js library
★ 1.4klibfm_python. Python
★ 10catboost. A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
★ 9kumap. Uniform Manifold Approximation and Projection
★ 8.2kPredictionAPI. Tutorial on deploying machine learning models to production
★ 58scikit-plot. An intuitive library to add plotting functionality to scikit-learn objects.
★ 2.4kxcessiv. A web-based application for quick, scalable, and automated hyperparameter tuning and stacked ensembling in Python.
★ 1.3kGenericObjectDecoding. Demo code for Horikawa and Kamitani (2017) Generic decoding of seen and imagined objects using hierarchical visual features. Nat Commun https://www.nature.com/articles/ncomms15037.
★ 165project-rhubarb. predicting mortality in England using air quality data
★ 9scikit-wrapRs. Python
★ 13deepcredit. How to predict credit defaulting?
★ 94rafflers. A collection of fun, funky, esoteric rafflers
★ 1BrotliHaxe. BrotliHaxe - hand ported decoder & encoder in haxe => JavaScript, PHP, Python, Java, C# (Project dev: 2013-2015 release: 2017)
★ 83nim-rmad. Autograd (backpropagation, reverse-mode auto differentiation) in Nim
★ 12theZoo. A repository of LIVE malwares for your own joy and pleasure. theZoo is a project created to make the possibility of malware analysis open and available to the public.
★ 13kStackNet. StackNet is a computational, scalable and analytical Meta modelling framework
★ 1.3kbattle-top. Responsive single-page web application to assist in bookkeeping things like Characters, Initiative, Hit Points, Conditions, et cetera, during D20-based table top RPG sessions.
★ 2python-pmap. Clojure's pmap implementation for Python.
★ 15hyperband. Tuning hyperparams fast with Hyperband
★ 599fnc-1-baseline. A baseline implementation for FNC-1
★ 138pycon2016_talk. The notes and slides from my PyCon Ireland 2016 PyData talk an introduction to gradient boosting
★ 18Competitive-Feature-Learning. Online feature-extraction and classification algorithm that learns representations of input patterns.
★ 31ml-dashboard. HTML
★ 8skip-thoughts. Sent2Vec encoder and training code from the paper "Skip-Thought Vectors"
★ 2kre-interval. re-frame library to work with intervals, using core.async
★ 4fnc-1. Python
★ 300pyhunspell. (Official repo for pypi package) Python bindings for the Hunspell spellchecker engine
★ 190bayesian-optimization. Python code for bayesian optimization using Gaussian processes
★ 316movie-plots-by-genre. Movie plots by genre tutorial at PyData Berlin 2016
★ 263eli5. A library for debugging/inspecting machine learning classifiers and explaining their predictions
★ 2.8kvecstack. Python package for stacking (machine learning technique)
★ 700keras2cpp. This is a bunch of code to port Keras neural network model into pure C++.
★ 681fractal-compression. Python code of fractal compression
★ 5caretEnsemble. caret models all the way down :turtle:
★ 229mltools. Exploratory and diagnostic machine learning tools for R
★ 74stack-overflow-import. Import arbitrary code from Stack Overflow as Python modules.
★ 3.7kmeterstick. A concise syntax to describe and execute routine data analysis tasks
★ 117VowpalWabbitIntro. Vowpal Wabbit (VW) is a fast open source program and library for machine learning. This introductory talk will cover installation; basic use; and some clever tricks like feature hashing, and progressive validation. These neat tricks can be used to train models with billions of features from data sets that are both tall and wide. The goal of this talk is for attendees to leave having everything they need to start working with VW and applying it to their own problems and datasets.
★ 8Multicore-TSNE. Parallel t-SNE implementation with Python and Torch wrappers.
★ 1.9kLightGBM. A fast, distributed, high performance gradient boosting (GBT, GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other machine learning tasks.
★ 19kberkeley-doc-summarizer. The Berkeley Document Summarizer is a learning-based, single-document summarization system that extracts source document content, exploits syntactic information to compress it, and uses coreference constraints to ensure clarity.
★ 745reinforcement-learning. Implementation of Reinforcement Learning Algorithms. Python, OpenAI Gym, Tensorflow. Exercises and Solutions to accompany Sutton's Book and David Silver's course.
★ 22kmodels. Models and examples built with TensorFlow
★ 78kstackmc. Repository for doing Stacked Monte Carlo on datasets
★ 8deconvfaces. Generating faces with deconvolution networks
★ 890FastBDT. Stochastic Gradient Boosted Decision Trees as Standalone, TMVAPlugin and Python-Interface
★ 248cyavro. Cython based wrapper for libavro
★ 25fast-wavenet. Speedy Wavenet generation using dynamic programming :zap:
★ 1.8knnvm. C++
★ 1.6kCNTK. Microsoft Cognitive Toolkit (CNTK), an open source deep-learning toolkit
★ 18kheamy. A set of useful tools for competitive data science.
★ 549scikit-optimize. Sequential model-based optimization with a `scipy.optimize` interface
★ 2.8kstacked_generalization. Library for machine learning stacking generalization.
★ 119document-scanner. An OpenCV based document scanner
★ 827mwt-ds. Umbrella repository for projects related to the MWT Decision Service
★ 187Scumblr. Web framework that allows performing periodic syncs of data sources and performing analysis on the identified results
★ 2.6kLargeVis. C++
★ 712rgf. Home repository for the Regularized Greedy Forest (RGF) library. It includes original implementation from the paper and multithreaded one written in C++, along with various language-specific wrappers.
★ 382kaggle-ndsb. Winning solution for the National Data Science Bowl competition on Kaggle (plankton classification)
★ 566kaggle_ndsb2. 3rd place solution for the second kaggle national datascience bowl
★ 92seqlearn. Sequence learning toolkit for Python
★ 706papers. Collection of papers
★ 468django-landing. Landing, a simple A/B testing application for Django
★ 11atari-ai. C++
★ 1ktffm. TensorFlow implementation of an arbitrary order Factorization Machine
★ 777graffiti-monkey. Goes around tagging things
★ 2kaggle_Otto. R script for Otto Group Production Classification on Kaggle
★ 30gym. A toolkit for developing and comparing reinforcement learning algorithms.
★ 37kK_sa. Kagge shelter animal competition
★ 6mwt-ds-explore. [Deprecated]: Exploration library
★ 17python-graph. New official repository: https://github.com/Shoobx/python-graph
★ 274mapper. Continuing work on mapper.py. Leyda Almodovar and Isabel Darcy. Mapper.py was originally written by Daniel Muellner and Aravindakshan Babu
★ 6ensemble_amazon. Code to share different ensemble techniques with focus on meta-stacking , using data from Amazon.com - Employee Access Challenge kaggle competition
★ 227ftrl_proximal_lr. Multithreaded Asynchronous FTRL Proximal Implementation
★ 126blog. for code created as part of http://studywolf.wordpress.com
★ 266kepler-mapper. Kepler Mapper: A flexible Python implementation of the Mapper algorithm.
★ 653deer. DEEp Reinforcement learning framework
★ 488GitXiv. GitXiv - Collaborative Open Computer Science.
★ 269crfasrnn. This repository contains the source code for the semantic image segmentation method described in the ICCV 2015 paper: Conditional Random Fields as Recurrent Neural Networks. http://crfasrnn.torr.vision/
★ 1.3kBPL. Bayesian Program Learning model for one-shot learning
★ 986DSB2. Jupyter Notebook
★ 9Embarrassingly-simple-ZSL. This repository contains the code for the real data experiments presented in our paper “An embarrassingly simple approach to zero-shot learning”, presented at ICML 2015.
★ 68tfdeploy. Deploy tensorflow graphs for fast evaluation and export to tensorflow-less environments running numpy.
★ 353SimulatedAnnealing. Python
★ 45MTH594_MachineLearning. The materials for the course MTH 594 Advanced data mining: theory and applications (Dmitry Efimov, American University of Sharjah)
★ 385superset. Apache Superset is a Data Visualization and Data Exploration Platform
★ 74kneural-doodle. Turn your two-bit doodles into fine artworks with deep neural networks, generate seamless textures from photos, transfer style from one image to another, perform example-based upscaling, but wait... there's more! (An implementation of Semantic Style Transfer.)
★ 9.9kcriteo. A potential 22nd rank solution to Criteo Labs Display Advertising Challenge on Kaggle
★ 25image-analogies. Generate image analogies using neural matching and blending.
★ 3.5kdboost. Stochastic Dummy Boosting
★ 25semisup-learn. Semi-supervised learning frameworks for python, which allow fitting scikit-learn classifiers to partially labeled data
★ 504kaggletils. Python
★ 379Kaggle-Ensemble-Guide. Code for the Kaggle Ensembling Guide Article on MLWave
★ 5hpelm. High performance implementation of Extreme Learning Machines (fast randomized neural networks).
★ 203kaggle-stackoverflow. Predicting closed questions on Stack Overflow
★ 43SqueezeNet. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters
★ 2.2k