This is your work, valued
BaiduTraffic. This repo includes introduction, code and dataset of our paper Deep Sequence Learning with Auxiliary Information for Traffic Prediction (KDD 2018).
★ 246KG4ZeroShotText. Source code of the paper 'Integrating Semantic Knowledge to Tackle Zero-shot Text Classification. NAACL-HLT 2019. '
★ 50Semantic-HPO. Source code of "Unsupervised Annotation of Phenotypic Abnormalities via Semantic Latent Representations on Electronic Health Records". BIBM 2019.
★ 5AminerKnowledgeGraph. HTML
★ 4MicrobitController. Micro:bit as a Controller
★ 4knowledge-work-plugins. Open source repository of plugins primarily intended for knowledge workers to use in Claude Cowork
★ 23kPaddleOCR. Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
★ 87kpydicom. Read, modify and write DICOM files with python code
★ 2.2khealthcare-agent-orchestrator. Facilitates creating modular specialized agents that coordinate across diverse data types and tools like M365 and Teams to assist multi-disciplinary healthcare workflows—such as cancer care.
★ 205LLaVA-Med. Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.
★ 2.2kvllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kA2A. Agent2Agent (A2A) is an open protocol enabling communication and interoperability between opaque agentic applications.
★ 25kdify. Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
★ 151kdocling. Get your documents ready for gen AI
★ 64khelm. Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
★ 2.9kNeMo-Agent-Toolkit. The NVIDIA NeMo Agent toolkit is an open-source library for efficiently connecting and optimizing teams of AI agents.
★ 2.5kFHIR-Converter. Conversion utility to translate legacy data formats into FHIR
★ 523seismometer. AI model evaluation with a focus on healthcare
★ 269meditron. Meditron is a suite of open-source medical Large Language Models (LLMs).
★ 2.2kpromptflow. Build high-quality LLM apps - from prototyping, testing to production deployment and monitoring.
★ 11kgpt-crawler. Crawl a site to generate knowledge files to create your own custom GPT from a URL
★ 22kgenerative-ai-for-beginners. 21 Lessons, Get Started Building with Generative AI
★ 114kswagger-ui. Swagger UI is a collection of HTML, JavaScript, and CSS assets that dynamically generate beautiful documentation from a Swagger-compliant API.
★ 29kSpeech. A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
★ 18kiBKH. iBKH: The integrative Biomedical Knowledge Hub
★ 519connect. The swiss army knife of healthcare integration.
★ 1.2kazure-health-AI-services-samples. This repository contains different samples applications and sample code to help you get started with the different Health-AI services. Based on these you will learn how to use our products, and accelerate your implementations.
★ 50Baichuan-13B. A 13B large language model developed by Baichuan Intelligent Technology
★ 2.9kBaichuan-7B. A large-scale 7B pretraining language model developed by BaiChuan-Inc.
★ 5.7kneovim. Vim-fork focused on extensibility and usability
★ 101ktd-vimrc. Thomas' vimrc enhanced for pythoners
★ 3llm-foundry. LLM training code for Databricks foundation models
★ 4.4kLMFlow. An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
★ 8.5kColossalAI. Making large AI models cheaper, faster and more accessible
★ 41kllama. Inference code for Llama models
★ 60kPrompt-Engineering-Guide. 🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.
★ 77kazure-orphan-resources. Centralize orphan resources in Azure environments
★ 763flower. Flower: A Friendly Federated AI Framework
★ 7.1kchaosmonkey. Chaos Monkey is a resiliency tool that helps applications tolerate random instance failures.
★ 17kPClean. A domain-specific probabilistic programming language for scalable Bayesian data cleaning
★ 231scrapy. Scrapy, a fast high-level web crawling & scraping framework for Python.
★ 63kresponsible-ai-toolbox. Responsible AI Toolbox is a suite of tools providing model and data exploration and assessment user interfaces and libraries that enable a better understanding of AI systems. These interfaces and libraries empower developers and stakeholders of AI systems to develop and monitor AI more responsibly, and take better data-driven actions.
★ 1.8knlpaug. Data augmentation for NLP
★ 4.7kfibo. The Financial Industry Business Ontology (FIBO) defines the sets of things that are of interest in financial business applications and the ways that those things can relate to one another. In this way, FIBO can give meaning to any data (e.g., spreadsheets, relational databases, XML documents) that describe the business of finance.
★ 647awesome-mlops. :sunglasses: A curated list of awesome MLOps tools
★ 5.2kumap. Uniform Manifold Approximation and Projection
★ 8.2kmlflow. The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
★ 27ksemgrep. Lightweight static analysis for many languages. Find bug variants with patterns that look like source code.
★ 16kray. Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
★ 43kdvc. 🦉 Data Versioning and ML Experiments
★ 16kmodin. Modin: Scale your Pandas workflows by changing a single line of code
★ 10kms2. Python
★ 68deepmind-research. This repository contains implementations and illustrative code to accompany DeepMind publications
★ 15kmesh-transformer-jax. Model parallel transformers in JAX and Haiku
★ 6.4kauto-sklearn. Automated Machine Learning with scikit-learn
★ 8.1ktesseract. Tesseract Open Source OCR Engine (main repository)
★ 76kdask. Parallel computing with task scheduling
★ 14kgpt-neo. An implementation of model parallel GPT-2 and GPT-3-style models using the mesh-tensorflow library.
★ 8.3kshap. A game theoretic approach to explain the output of any machine learning model.
★ 26kopenai-python. The official Python library for the OpenAI API
★ 31kCommonDataModel. Definition and DDLs for the OMOP Common Data Model (CDM)
★ 1.1kpyarmor. A tool used to obfuscate python scripts, bind obfuscated scripts to fixed machine or expire obfuscated scripts.
★ 5.2konelinerizer. Shamelessly convert any Python 2 script into a terrible single line of code
★ 1.5khuman-phenotype-ontology. Ontology for the description of human clinical features
★ 365medal. Large medical text dataset curated for abbreviation disambiguation, designed for natural language understanding pre-training in the medical domain
★ 285covid-chestxray-dataset. We are building an open database of COVID-19 cases with chest X-ray or CT images.
★ 3.1kgpt-3. GPT-3: Language Models are Few-Shot Learners
★ 16kACTdb. Annotated Clinical Texts from MIMIC
★ 6SyferText. A privacy preserving NLP framework
★ 198thymedata. This is a repository for annotation data for the THYME Project, a clinical natural language processing project dedicated to extracting useful temporal relations from the clinical narrative.
★ 36phenotyping. This repository contains data and code for the paper"Comparing deep learning and concept extraction based methods for patient phenotyping".
★ 21AttentionExplanation. Jupyter Notebook
★ 323doccano. Open source annotation tool for machine learning practitioners.
★ 11ktqdm. :zap: A Fast, Extensible Progress Bar for Python and CLI
★ 31knegex. Automatically exported from code.google.com/p/negex
★ 36examples. A set of examples around pytorch in Vision, Text, Reinforcement Learning, etc.
★ 24kpegasus. Python
★ 1.7kgoogle-research. Google Research
★ 38ktips_for_interview. 我的一些面试心得;自学CS历程分享;找工作求职经验分享
★ 4ktrax. Trax — Deep Learning with Clear Code and Speed
★ 8.3kpoincare-embeddings. PyTorch implementation of the NIPS-17 paper "Poincaré Embeddings for Learning Hierarchical Representations"
★ 1.8kNLPzoo. A collection of the most popular Natural Language Processing algorithms, frameworks and applications (inspired by RLzoo).
★ 3AmsterdamUMCdb. AmsterdamUMCdb - Freely Accessible ICU database. Please access our Open Access manuscript at https://doi.org/10.1097/CCM.0000000000004916
★ 211phenopacket-schema. Repository for the GA4GH phenopacket schema
★ 98wenyan. 文言文編程語言 A programming language for the ancient Chinese.
★ 20kPLMpapers. Must-read Papers on pre-trained language models.
★ 3.4kBullshitGenerator. Needs to generate some texts to test if my GUI rendering codes good or not. so I made this.
★ 16kCogStack-Pipeline. Distributed, fault tolerant batch processing for Natural Language Applications and Search, using remote partitioning
★ 48jax. Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
★ 36kfhir. FHIR Protocol Buffers
★ 946d2l-zh. 《动手学深度学习》:面向中文读者、能运行、可讨论。中英文版被70多个国家的500多所大学用于教学。
★ 79kd2l-en. Interactive deep learning book with multi-framework code, math, and discussions. Adopted at 500 universities from 70 countries including Stanford, MIT, Harvard, and Cambridge.
★ 29kresearch_tao. NLP研究入门之道
★ 2.1kml_paper_club. A repository of papers that have been presented at nPlan's machine learning paper club
★ 269biograkn. BioGrakn Knowledge Graph
★ 184transformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kpytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
★ 102kNeuralCR. Python
★ 25terminal. The new Windows Terminal and the original Windows console host, all in the same place!
★ 104kpgmpy. Python Toolkit for Causal and Probabilistic Reasoning
★ 3.3ktext. Making text a first-class citizen in TensorFlow.
★ 1.3kseq2seq-chatbot. Chatbot in 200 lines of code using TensorLayer
★ 837ai-deadlines. :alarm_clock: AI conference deadline countdowns
★ 6kdragon. A Computation Graph Virtual Machine based ML Framework
★ 108996.ICU. Repo for counting stars and contributing. Press F to pay respect to glorious developers.
★ 277kmimic-code. MIMIC Code Repository: Code shared by the research community for the MIMIC family of databases
★ 3.3kstanza. Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languages
★ 7.9kNLP-progress. Repository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the most common NLP tasks.
★ 23kvaderSentiment. VADER Sentiment Analysis. VADER (Valence Aware Dictionary and sEntiment Reasoner) is a lexicon and rule-based sentiment analysis tool that is specifically attuned to sentiments expressed in social media, and works well on texts from other domains.
★ 5knltk. NLTK Source
★ 15kgenerating-reviews-discovering-sentiment. Code for "Learning to Generate Reviews and Discovering Sentiment"
★ 1.5kallennlp. An open-source NLP research library, built on PyTorch.
★ 12kmedical-question-answer-data. Medical question and answer dataset gathered from the web.
★ 125ReadingGroup. ImperialNLP Reading Group
★ 4tensor2tensor. Library of deep learning models and datasets designed to make deep learning more accessible and accelerate ML research.
★ 17ktrfl. TensorFlow Reinforcement Learning
★ 3.1kbert. TensorFlow code and pre-trained models for BERT
★ 40kEffectiveTensorflow. TensorFlow tutorials and best practices.
★ 8.6kDALI. A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
★ 5.7khorovod. Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.
★ 15ktensorlayer2-design. Python
★ 5HyperPose. Library for Fast and Flexible Human Pose Estimation
★ 1.3kRocket.Chat. The Secure CommsOS™ for mission-critical operations
★ 46klectures. Oxford Deep NLP 2017 course
★ 16kconference_call_for_paper. 2021-2022 International Conferences in Artificial Intelligence, Machine Learning, Computer Vision, Data Mining, Natural Language Processing and Robotics
★ 852Solutions-ECL-Training. ECL
★ 18awesome-text-summarization. A curated list of resources dedicated to text summarization
★ 1.5kwiki-reading. This repository contains the three WikiReading datasets as used and described in WikiReading: A Novel Large-scale Language Understanding Task over Wikipedia, Hewlett, et al, ACL 2016 (the English WikiReading dataset) and Byte-level Machine Reading across Morphologically Varied Languages, Kenter et al, AAAI-18 (the Turkish and Russian datasets).
★ 271GloVe. Software in C and data files for the popular GloVe model for distributed word representations, a.k.a. word vectors or embeddings
★ 7.2ktensorlayer-tricks. How to use TensorLayer
★ 348Crepe. Character-level Convolutional Networks for Text Classification
★ 848nlp_tasks. Natural Language Processing Tasks and References
★ 3ktypedb. TypeDB: Built for systems, not records
★ 4.4kmit-deep-learning-book-pdf. MIT Deep Learning Book in PDF format (complete and parts) by Ian Goodfellow, Yoshua Bengio and Aaron Courville
★ 14kAnnotated-Semantic-Relationships-Datasets. A collections of public and free annotated datasets of relationships between entities/nominals (Portuguese and English)
★ 710conv_arithmetic. A technical report on convolution arithmetic in the context of deep learning
★ 15kgym. A toolkit for developing and comparing reinforcement learning algorithms.
★ 37kawesome-public-datasets. A topic-centric list of HQ open datasets.
★ 78knlp-datasets. Alphabetical list of free/public domain datasets with text data for use in Natural Language Processing (NLP)
★ 6kEasyPR-python. EasyPR-python
★ 72openalpr. Automatic License Plate Recognition library
★ 11kSRGAN. Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network
★ 3.5ksilentarmy. Zcash miner optimized for AMD & Nvidia GPUs
★ 344AirSim. Open source simulator for autonomous vehicles built on Unreal Engine / Unity, from Microsoft AI & Research
★ 18kdarwin-xnu. Legacy mirror of Darwin Kernel. Replaced by https://github.com/apple-oss-distributions/xnu
★ 11krelational-networks. Pytorch implementation of "A simple neural network module for relational reasoning" (Relational Networks)
★ 816fold. Deep learning with dynamic computation graphs in TensorFlow
★ 1.8kacademicpages.github.io. Github Pages template based upon HTML and Markdown for personal, portfolio-based websites.
★ 17kfacets. Visualizations for machine learning datasets
★ 7.3kstock_market_reinforcement_learning. This project provides a stock market environment using OpenGym with Deep Q-learning and Policy Gradient.
★ 795YouCompleteMe. A code-completion engine for Vim
★ 26kganhacks. starter from "How to Train a GAN?" at NIPS2016
★ 12kstatsmodels. Statsmodels: statistical modeling and econometrics in Python
★ 12kAdversarialNetsPapers. Awesome paper list with code about generative adversarial nets
★ 6.6kiGAN. Interactive Image Generation via Generative Adversarial Networks
★ 4ktmux-resurrect. Persists tmux environment across system restarts.
★ 13kblaze. Blaze runtime system that support efficient accelerator integration for big data.
★ 24deep-visualization-toolbox. DeepVis Toolbox
★ 4.1kVQA. Python
★ 392nmn2. Neural module networks
★ 401TensorLayer. Deep Learning and Reinforcement Learning Library for Scientists and Engineers
★ 7.4kflask. The Python micro framework for building web applications.
★ 72kthuthesis. LaTeX Thesis Template for Tsinghua University
★ 5.4kwebsocketpp. C++ websocket client/server library
★ 7.7ktensorflow. An Open Source Machine Learning Framework for Everyone
★ 197kHPCC-Platform. HPCC Systems (High Performance Computing Cluster) is an open source, massive parallel-processing computing platform for big data processing and analytics.
★ 612jaspervdj. Source code of my personal home page.
★ 123my-personal-kanban. This is a one page HTML/JavaScript application for people who would like to use simple and basic Kanban board for their personal stuff
★ 820The-Personal-Page. This simple one-page website is a way for people to have a very quick and easy personable website that aggregates your activity and positions a simple logo, a portrait and some description text in a nicely-formatted manner. This is licensed under the MIT and GPL licenses.
★ 760vndn. V-NDN: an implementation of NDN for vehicular networks
★ 22xgboost. Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C++ and more. Runs on single machine, Hadoop, Spark, Dask, Flink and DataFlow
★ 29kcropper. ⚠️ [Deprecated] No longer maintained, please use https://github.com/fengyuanchen/jquery-cropper
★ 7.7kSemantic-UI. Semantic is a UI component framework based around useful principles from natural language.
★ 51ksass. Sass makes CSS fun!
★ 15kjQCloud. jQuery plugin for drawing neat word clouds that actually look like clouds
★ 648spark. A simple expressive web framework for java. Spark has a kotlin DSL https://github.com/perwendel/spark-kotlin
★ 9.7kd3. Bring data to life with SVG, Canvas and HTML. :bar_chart::chart_with_upwards_trend::tada:
★ 113kreveal.js. The HTML Presentation Framework
★ 72kohmyzsh. 🙃 A delightful community-driven (with 2,500+ contributors) framework for managing your zsh configuration. Includes 300+ optional plugins (rails, git, macOS, hub, docker, homebrew, node, php, python, etc), 140+ themes to spice up your morning, and an auto-update tool that makes it easy to keep up with the latest updates from the community.
★ 189kshadowsocks-iOS. Removed according to regulations.
★ 8.1kdeeplearning4j. Suite of tools for deploying and training deep learning models using the JVM. Highlights include model import for keras, tensorflow, and onnx/pytorch, a modular and tiny c++ library for running math code and a java based math library on top of the core c++ library. Also includes samediff: a pytorch/tensorflow like library for running deep learn...
★ 14ksnownlp. Python library for processing Chinese text
★ 6.6kthefuck. Magnificent app which corrects your previous console command.
★ 98k