This is your work, valued
I am an an Assistant Professor working on medical image analysis and machine learning at Shanghai Jiao Tong University.
PMC-LLaMA. The official codes for "PMC-LLaMA: Towards Building Open-source Language Models for Medicine"
★ 679RadFM. The official code for "Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data".
★ 562Finetune_LLAMA. 简单易懂的LLaMA微调指南。
★ 412GPT-4V_Medical_Evaluation.
★ 44Check_PMC_QA. Python
★ 3VQA-Med-2019. Visual Question Answering in the Medical Domain VQA-Med 2019
★ 2openai-proxy. OpenAI/ChatGPT 免翻墙代理
★ 2voxelmorph. Unsupervised Learning for Image Registration
★ 1AGXNet. Python
★ 1MedKLIP. Python
★ 1Build-PMC-OA. Python
★ 1DxEvolver. Python
★ 1MedSR-Copilot. Python
★ 2superpowers-zh. 🦸 AI 编程超能力 · 中文增强版 — superpowers(116k+ ⭐)完整汉化 + 6 个中国原创 skills,让 Claude Code / Copilot CLI / Hermes Agent / Cursor / Windsurf / Kiro / Gemini CLI 等 16 款 AI 编程工具真正会干活
★ 7.4kscience-superpowers. Composable computational-science methodology skills for AI research agents — pre-registration over TDD. A science-domain reimplementation of Superpowers.
★ 276MedSP1000. Python
★ 18OligoGym. OligoGym is a python package that streamlines processes involving featurization, training and evaluation of predictive models of oligonucleotide properties. Oligonucleotides include antisense oligonucleotides (ASOs) and small interfering RNAs (siRNAs).
★ 19evolver. The GEP-powered self-evolving engine for AI agents. Auditable evolution with Genes, Capsules, and Events. | evomap.ai
★ 8.9kSciAgent-Skills. 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon.
★ 291hermes-agent. The agent that grows with you
★ 223kMAP. Python
★ 16OpenSeeker. OpenSeeker: A search agent with open-source data and models
★ 764PyTrial. PyTrial: A Comprehensive Platform for Artificial Intelligence for Drug Development
★ 127SkillRL. SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
★ 916Exomiser. A Tool to Annotate and Prioritize Exome Variants
★ 260Agent-Memory-Paper-List. The paper list of "Memory in the Age of AI Agents: A Survey"
★ 2.3kAgent-Skills-for-Context-Engineering. A comprehensive collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, or debugging agent systems that require effective context management.
★ 18klangextract. A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
★ 38kdspy-doc-zh. DSPy中文文档
★ 52PhenoLIP. Python
★ 14MemBrain. Python
★ 281AgentEHR. Agentic System, Tool Use, Electronic Health Record, Large Language Models, Clinical Nature Language Processing
★ 24PageIndex. 📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
★ 35kmcp-toolbox. MCP Toolbox for Databases is an open source MCP server for databases.
★ 16kToolUniverse. Democratizing AI scientists with ToolUniverse
★ 1.6kawesome-claude-skills. A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows
★ 71khelm. Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
★ 2.9kscientific-agent-skills. Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 170,000+ scientists worldwide. 158 ready-to-use skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.
★ 32kTrialMind-SLR. Wang, Z., Cao, L., Danek, B. et al. Accelerating clinical evidence synthesis with large language models. npj Digit. Med. 8, 509 (2025). https://doi.org/10.1038/s41746-025-01840-7
★ 19rank_llm. RankLLM is a Python toolkit for reproducible information retrieval research using rerankers, with a focus on listwise reranking.
★ 611TREQS. Text-to-SQL Generation for Question Answering on Electronic Medical Records
★ 136EhrAgent. [EMNLP'24] EHRAgent: Code Empowers Large Language Models for Complex Tabular Reasoning on Electronic Health Records
★ 137Clinical-Tool-Learning. Python
★ 27deepscholar. build and benchmark deep research
★ 244Memori. Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori works with the data infrastructure you already run, no rip-and-replace, and deploys across managed cloud, single-tenant cloud, VPC, and on-premises.
★ 16kMiroThinker. MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.
★ 8.4kawesome-medical-mcp-servers. A collection of Medical MCP servers.
★ 70M3Builder. The official codes for "M^3Builder: A Multi-Agent System for Automated Machine Learning in Medical Imaging"
★ 45ZeroSearch. ZeroSearch: Incentivize the Search Capability of LLMs without Searching
★ 1.3kDeepResearch. Tongyi Deep Research, the Leading Open-source Deep Research Agent
★ 20kO1-Journey. O1 Replication Journey
★ 2kInnovatorBench. [ICLR 2026]InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
★ 16EHR-R1. Python
★ 37biochatter. Backend library for conversational AI in biomedicine
★ 216claude-agent-sdk-python. Python
★ 7.8kEHRFormer. Python
★ 40DeepAnalyze. DeepAnalyze is the first agentic LLM for autonomous data science. 🎈你的AI数据分析师,自动分析大量数据,一键生成专业分析报告!
★ 4.4kPrimeKG. Precision Medicine Knowledge Graph (PrimeKG)
★ 799PathPT. [Nature Communications, 2026] The official code for "Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping"
★ 28Context-Engineering. "Context engineering is the delicate art and science of filling the context window with just the right information for the next step." — Andrej Karpathy. A frontier, first-principles handbook inspired by Karpathy and 3Blue1Brown for moving beyond prompt engineering to the wider discipline of context design, orchestration, and optimization.
★ 9.2kDeep-DxSearch. An agentic RL framework to enhance retreival-augmented reasoning in Diagnostic Policy
★ 103agent-lightning. The absolute trainer to light up AI agents.
★ 17kDiagGym. A virtual clinical environment for self‑evolving LLM diagnostic agents.
★ 108DeepRare. Code implementation of DeepRare (Nature 2026)
★ 272MedAgentGym. [ICLR'26] MedAgentGYM: Training LLM Agents for Code-Based Medical Reasoning at Scale
★ 125ChestX-Reasoner. Python
★ 39ethos-paper. Jupyter Notebook
★ 96MedXpertQA. [ICML 2025] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
★ 171ReasoningEval. Official repo of Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains.
★ 43INSPECT_public. INSPECT dataset/benchmark paper, accepted by NeurIPS 2023
★ 53simple-evals. Python
★ 4.6ksynthea. Synthetic Patient Population Simulator
★ 3.3kADAS. [ICLR 2025] Automated Design of Agentic Systems
★ 1.6kautogen. A programming framework for agentic AI
★ 60kMetaGPT. 🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming
★ 70kMedReason. MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
★ 280textgrad. TextGrad: Automatic ''Differentiation'' via Text -- using large language models to backpropagate textual gradients. Published in Nature.
★ 3.7kRAG-Gym. Official repository for RAG-Gym
★ 126DAPO. An Open-source RL System from ByteDance Seed and Tsinghua AIR
★ 1.9kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kSearch-R1. Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL
★ 5.2kDDx_CaseReports. impact of lab test results on the accuracy of differential diagnosis (DDx) generated by large language models (LLMs) using clinical case reports
★ 4MedRBench. [Nature Communications] The official code for "Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases".
★ 70PokeLLMon. Python
★ 205pokechamp. Official repository of the spotlight ICML 2025 paper, PokeChamp: an Expert-level Minimax Language Agent.
★ 176GeneGPT. Code and data for GeneGPT.
★ 428smolagents. 🤗 smolagents: a barebones library for agents that think in code.
★ 29kRadABench. The official codes for "Can Modern LLMs Act as Agent Cores in Radiology Environments?"
★ 29Scrapegraph-ai. Python scraper based on AI
★ 29kMIMIC-Clinical-Decision-Making-Dataset. Code repository to create the MIMIC-CDM Dataset.
★ 48GenerateCT. ECCV 2024 & GenerateCT: Text-Conditional Generation of 3D Chest CT Volumes
★ 194ehrapy. Electronic Health Record Analysis with Python.
★ 356AgentTorch. large population models
★ 633BMTools. Tool Learning for Big Models, Open-Source Solutions of ChatGPT-Plugins
★ 2.8kToolLearningPapers.
★ 923ToolBench. [ICLR'24 spotlight] An open platform for training, serving, and evaluating large language model for tool learning.
★ 5.7kMedS-Ins. [npj digital medicine] The official codes for "Towards Evaluating and Building Versatile Large Language Models for Medicine"
★ 79ReXrank. HTML
★ 28RP3D-Diag. Code implementation of RP3D-Diag
★ 17RaTEScore. [EMNLP 2024] RaTEScore: A Metric for Radiology Report Generation
★ 67KEP. [ECCV 2024 Oral] Knowledge-enhanced pretraining for computational pathology
★ 50FlagData. Python
★ 366MedLane. Python
★ 1long-form-factuality. Benchmarking long-form factuality in large language models. Original code for our paper "Long-form factuality in large language models".
★ 692robustlearn. Robust machine learning for responsible AI
★ 508langchain. The agent engineering platform.
★ 143kMedAgents. [ACL 2024 Findings] MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning https://arxiv.org/abs/2311.10537
★ 363AgentGPT. 🤖 Assemble, configure, and deploy autonomous AI Agents in your browser.
★ 36kAscle. A Python Natural Language Processing Toolkit for Medical Text Generation
★ 84UltraMedical. [NeurIPS 2024 D&B Track, Spotlight] UltraMedical: Building Specialized Generalists in Biomedicine
★ 96Taiyi-LLM. Taiyi 2, Biomedical LLM, A Bilingual (Chinese and English) Fine-Tuned Large Language Model for Diverse Biomedical Tasks
★ 168Med-PaLM. Towards Generalist Biomedical AI
★ 432gpt_bionlp_benchmark.
★ 25BLUE_Benchmark. BLUE benchmark consists of five different biomedicine text-mining tasks with ten corpora.
★ 298caml-mimic. multilabel classification of EHR notes
★ 309esm. Evolutionary Scale Modeling (esm): Pretrained language models for proteins
★ 4.2kmimic3-benchmarks. Python suite to construct benchmark machine learning datasets from the MIMIC-III 💊 clinical database.
★ 891MEDIQA-M3G-2024. Python
★ 8CT-CLIP. Developing Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
★ 411clinical-trial-outcome-prediction. benchmark dataset and Deep learning method (Hierarchical Interaction Network, HINT) for clinical trial approval probability prediction, published in Cell Patterns 2022.
★ 166MemDPC. [ECCV'20 Spotlight] Memory-augmented Dense Predictive Coding for Video Representation Learning. Tengda Han, Weidi Xie, Andrew Zisserman.
★ 167Existing-Medical-QA-Datasets. Multimodal Question Answering in the Medical Domain: A summary of Existing Datasets and Systems
★ 317MIMIC-Diff-VQA. Python
★ 73FastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kInsTag. InsTag: A Tool for Data Analysis in LLM Supervised Fine-tuning
★ 288Awesome-LLMs-Datasets. Summarize existing representative LLMs text datasets.
★ 1.5kFLAN. Python
★ 1.6kmedalign. MedAlign is a clinician-generated dataset for instruction following with electronic medical records.
★ 102LLMDataHub. A quick guide (especially) for trending instruction finetuning datasets
★ 3.4kAIONER. AIONER
★ 70Prompt-Engineering-Guide. 🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.
★ 77kMMedLM. [Nature Communications] The official codes for "Towards Building Multilingual Language Model for Medicine"
★ 284discq. Python
★ 19GenMedicalEval.
★ 86OpenMoE. A family of open-sourced Mixture-of-Experts (MoE) Large Language Models
★ 1.7kpaper-qa. High accuracy RAG for answering questions from scientific documents with citations
★ 9kroco-dataset. Radiology Objects in COntext (ROCO): A Multimodal Image Dataset
★ 250.github.
★ 1pydicom-seg. Python package for DICOM-SEG medical segmentation file reading and writing
★ 85nekton. A python package for DICOM to NifTi and NifTi to DICOM-SEG and GSPS conversion
★ 12dcmrtstruct2nii. DICOM RT-Struct to mask
★ 104SAT. [npj Digital Medicine] The official repository for "Large-Vocabulary Segmentation for Medical Images with Text Prompts"
★ 307RP3D-Diag. Code implementation of RP3D-Diag
★ 79meditron. Meditron is a suite of open-source medical Large Language Models (LLMs).
★ 2.2kFlexLLMGen. Running large language models on a single GPU for throughput-oriented scenarios.
★ 9.4ktorchio. Medical imaging processing for AI applications.
★ 2.4kMedusa. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2.8kcleanlab. Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
★ 12kAwesome-LLM-Healthcare. The paper list of the review on LLMs in medicine - "Large Language Models Illuminate a Progressive Pathway to Artificial Healthcare Assistant: A Review".
★ 269knowledge-graphs. A collection of research on knowledge graphs
★ 1.8krgrg. Code for the CVPR paper "Interactive and Explainable Region-guided Radiology Report Generation"
★ 214LLaVA-Plus-Codebase. LLaVA-Plus: Large Language and Vision Assistants that Plug and Learn to Use Skills
★ 770SegWithDistMap. How Distance Transform Maps Boost Segmentation CNNs: An Empirical Study
★ 403MultiTalent. Implemention of the Paper "MultiTalent: A Multi-Dataset Approach to Medical Image Segmentation"
★ 69BioMedLM. Python
★ 643natural-instructions. Expanding natural instructions
★ 1kGPT-4V_Medical_Evaluation.
★ 44pixel-lesion-patient-network. Python
★ 32RLIPv2. [ICCV 2023] RLIPv2: Fast Scaling of Relational Language-Image Pre-training
★ 136RLIP. [NeurIPS 2022 Spotlight] RLIP: Relational Language-Image Pre-training and a series of other methods to solve HOI detection and Scene Graph Generation.
★ 78fiftyone. Refine high-quality datasets and visual AI models
★ 11kBDR-main. Python
★ 6DiRA. Official PyTorch Implementation for DiRA: Discriminative, Restorative, and Adversarial Learning for Self-supervised Medical Image Analysis - CVPR 2022
★ 105awesome-llm-powered-agent. Awesome things about LLM-powered agents. Papers / Repos / Blogs / ...
★ 2deepbgc. BGC Detection and Classification Using Deep Learning
★ 159MedCPT. Code for MedCPT, a model for zero-shot biomedical information retrieval.
★ 271niicat. This is a tool to quickly preview nifti images on the terminal
★ 72pyradiomics. Open-source python package for the extraction of Radiomics features from 2D and 3D images and binary masks. Support: https://discourse.slicer.org/c/community/radiomics
★ 1.4kViewers. OHIF zero-footprint DICOM viewer and oncology specific Lesion Tracker, plus shared extension packages
★ 4.3kRadImageNet. RadImageNet, a pre-trained convolutional neural networks trained solely from medical imaging to be used as the basis of transfer learning for medical imaging applications.
★ 491MedQuAD. Medical Question Answering Dataset of 47,457 QA pairs created from 12 NIH websites
★ 459VQA-Med-2019. Visual Question Answering in the Medical Domain VQA-Med 2019
★ 95CODER. CODER: Knowledge infused cross-lingual medical term embedding for term normalization. [JBI, ACL-BioNLP 2022]
★ 80LLMsNineStoryDemonTower. 【LLMs九层妖塔】分享 LLMs在自然语言处理(ChatGLM、Chinese-LLaMA-Alpaca、小羊驼 Vicuna、LLaMA、GPT4ALL等)、信息检索(langchain)、语言合成、语言识别、多模态等领域(Stable Diffusion、MiniGPT-4、VisualGLM-6B、Ziya-Visual等)等 实战与经验。
★ 2.2kanatomy_informed_DA. Python
★ 21UniBrain. An offcial implementation for UniBrain: Universal Brain MRI Diagnosis with Hierarchical Knowledge-enhanced Pre-training
★ 39gradient-checkpointing. Make huge neural nets fit in memory
★ 2.8kRedPajama-Data. The RedPajama-Data repository contains code for preparing large datasets for training large language models.
★ 5kSAM-Med2D. SAM-Med2D: Bridging the Gap between Natural Image Segmentation and Medical Image Segmentation
★ 76medical-data.
★ 6kOpenBioLink. OpenBioLink is a resource and evaluation framework for evaluating link prediction models on heterogeneous biomedical graph data.
★ 161SAM-Med2D. Official implementation of SAM-Med2D
★ 1.1kLibriSQA.
★ 39Awesome-Medical-Healthcare-Dataset-For-LLM. A curated list of popular Datasets, Models and Papers for LLMs in Medical/Healthcare
★ 330StoryGen. [CVPR 2024] Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models
★ 270Webpilot. Vue
★ 1.9kTotalSegmentator. Tool for robust segmentation of >100 important anatomical structures in CT and MR images
★ 2.9kRadFM. The official code for "Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data".
★ 562med-flamingo. Python
★ 452LVM-Med. [NeurIPS 2023] Release LMV-Med pre-trained models
★ 217fastmoe. A fast MoE impl for PyTorch
★ 1.9kMaskFormer. Per-Pixel Classification is Not All You Need for Semantic Segmentation (NeurIPS 2021, spotlight)
★ 1.5kLLaVA-Med. Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.
★ 2.2kSegLossOdyssey. A collection of loss functions for medical image segmentation
★ 4kAbdomenCT-1K. The official repository of "AbdomenCT-1K: Is Abdominal Organ Segmentation A Solved Problem?"
★ 267MAML. [MICCAI 2021] The official code for "Modality-aware Mutual Learning for Multi-modal Medical Image Segmentation"
★ 52chatgpt-retrieval-plugin. The ChatGPT Retrieval Plugin lets you easily find personal or work documents by asking questions in natural language.
★ 21kGPT4RoI. (ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
★ 555LLMSurvey. The official GitHub page for the survey paper "A Survey of Large Language Models".
★ 12kscispacy. A full spaCy pipeline and models for scientific/biomedical documents.
★ 2kMedSAM. Segment Anything in Medical Images
★ 4.4kKAD. Python
★ 158MedLSAM. MedLSAM: Localize and Segment Anything Model for 3D Medical Images
★ 522Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kXrayGPT. [BIONLP@ACL 2024] XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models.
★ 530Youku-mPLUG. Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
★ 307ChatRWKV. ChatRWKV is like ChatGPT but powered by RWKV (100% RNN) language model, and open source.
★ 9.5kmedicat. Dataset of medical images, captions, subfigure-subcaption annotations, and inline textual references
★ 175MING. 明医 (MING):中文医疗问诊大模型
★ 1.2kPMC-Patients. PMC-Patients
★ 113Awesome-Visual-Transformer. Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
★ 3.6k