This is your work, valued
Contributor of @apache Doris, @apache InLong, @apache DolphinScheduler @apache IoTDB, and etc.
springboot-datax. 使用springboot启动datax,方便以web方式使用
★ 225awsome-programming-note. 记录👨💻程序员常用笔记
★ 15hive-hook-plugin. hive插件获取表级别、字段级别血缘关系
★ 11bigData-starter. spark-starter , hive-starter , hbase-starter
★ 11Awsome-jdbc. 使用jdbc开发通用的服务模块
★ 4Java-Spider. 使用Java针对各大型网站的爬虫实战 🕷
★ 2Sparklens. Scala
★ 1fluss. Apache Fluss is a streaming storage built for real-time analytics.
★ 2kXcodesApp. The easiest way to install and switch between multiple versions of Xcode - with a mouse click.
★ 8.5kGPT_API_free. Free ChatGPT&DeepSeek API Key,免费ChatGPT&DeepSeek API。免费接入DeepSeek API和GPT4 API,支持 gpt | deepseek | claude | gemini | grok 等排名靠前的常用大模型。
★ 39kcloudstack. Apache CloudStack is an opensource Infrastructure as a Service (IaaS) cloud computing platform
★ 3klx-music-desktop. 一个基于 Electron 的音乐软件
★ 52knightingale. Nightingale is to monitoring and alerting what Grafana is to visualization.
★ 13kgo-mitmproxy. mitmproxy implemented with golang. 用 Golang 实现的中间人攻击(Man-in-the-middle),解析、监测、篡改 HTTP/HTTPS 流量。
★ 1.6kpuppeteer-extra. 💯 Teach puppeteer new tricks through plugins.
★ 7.4krod. A Chrome DevTools Protocol driver for web automation and scraping.
★ 7kChatGPT-Java-FunAi. ChatGPT Java 基于SpringBoot的后端开源web学习项目,FunAi。支持OpenAI官方所有接口。无限轮聊天 + 带上下文逻辑 + 流式输出 / 普通输出。PDF解析 + Embedding API+ 递归分词文段抽取 + 文本向量化 + 向量语义匹配 + 召回知识库相似文本匹配。接入文生图模型MidJourney / Stable Diffusion Model。智能客服/企业级知识库。APIKey额度精准查询 + 失效检测。AI游戏 + 专属于AI的社交平台
★ 946one-api. LLM API 管理 & 分发系统,支持 OpenAI、Azure、Anthropic Claude、Google Gemini、DeepSeek、字节豆包、ChatGLM、文心一言、讯飞星火、通义千问、360 智脑、腾讯混元等主流模型,统一 API 适配,可用于 key 管理与二次分发。单可执行文件,提供 Docker 镜像,一键部署,开箱即用。LLM API management & key redistribution system, unifying multiple providers under a single API. Single binary, Docker-ready, with an English UI.
★ 36kFastGPT. FastGPT is a knowledge-based platform built on the LLMs, offers a comprehensive suite of out-of-the-box capabilities such as data processing, RAG retrieval, and visual AI workflow orchestration, letting you easily develop and deploy complex question-answering systems without the need for extensive setup or configuration.
★ 29kdify. Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
★ 151kShortGPT. 🚀🎬 ShortGPT - Experimental AI framework for youtube shorts / tiktok channel automation
★ 7.7kdataCompare. big data comparison and data profiling platform: low code,data comparison and data profiling
★ 280Jpom. 【dromara】🚀简而轻的低侵入式在线构建、自动部署、日常运维、项目监控软件
★ 1.9kali-dbhub. 已迁移新仓库,此版本将不再维护
★ 8kstable-diffusion-webui-colab. stable diffusion webui colab
★ 16kdatavines. Know your data better!Datavines is Next-gen Data Observability Platform, support metadata manage and data quality.
★ 754QinSQL. AI 时代的智能数据库
★ 221laf. Laf is a vibrant cloud development platform that provides essential tools like cloud functions, databases, and storage solutions. It enables developers to quickly unleash their creativity and bring innovative ideas to life with ease.
★ 7.6kspeechgpt. 💬 SpeechGPT is a web application that enables you to converse with ChatGPT.
★ 2.8kNextChat. ✨ Light and Fast AI Assistant. Support: Web | iOS | MacOS | Android | Linux | Windows
★ 89kawesome-chatgpt-prompts-zh. ChatGPT 中文调教指南。各种场景使用指南。学习怎么让它听你的话。
★ 61kthe-algorithm-ml. Source code for Twitter's Recommendation Algorithm
★ 11kCloudEon. CloudEon uses Kubernetes to install and deploy open-source big data components, enabling the containerized operation of an open-source big data platform. This allows you to reduce your focus on underlying resource management and maintenance.
★ 493the-algorithm. Source code for the X Recommendation Algorithm
★ 74kMidJourney-Styles-and-Keywords-Reference. A reference containing Styles and Keywords that you can use with MidJourney AI. There are also pages showing resolution comparison, image weights, and much more!
★ 12ksqlchat. Chat-based SQL Client and Editor for the next decade
★ 5.8kChatGLM-6B. ChatGLM-6B: An Open Bilingual Dialogue Language Model | 开源双语对话语言模型
★ 41kMIGPT. 基于API流式对话的低延迟版MIGPT
★ 575carrot. AI 工具导航大全,帮你快速筛选免费、实用、高效的网站资源
★ 17kgo-ldap-admin. 🌉 基于Go+Vue实现的openLDAP后台管理项目
★ 2kliteflow. Lightweight, fast, stable, programmable component-based rule engine — where AI Agents orchestrate just like ordinary components. Uniquely designed DSL: component reuse, sync/async & dynamic orchestration, multi-language scripting, nested rules, hot deployment and smooth refresh. If you can orchestrate LiteFlow, you can orchestrate AI.
★ 3.8karthas. Alibaba Java Diagnostic Tool Arthas/Alibaba Java诊断利器Arthas
★ 37kHiBench. HiBench is a big data benchmark suite.
★ 1.5kxterm.js. A terminal for the web
★ 21kpinpoint. APM, (Application Performance Management) tool for large-scale distributed systems.
★ 14kchinese-independent-blogs. 中文独立博客列表
★ 24ktutorials. Getting Started with Spring Boot 3:
★ 37knomad. Nomad is an easy-to-use, flexible, and performant workload orchestrator that can deploy a mix of microservice, batch, containerized, and non-containerized applications. Nomad is easy to operate and scale and has native Consul and Vault integrations.
★ 17ksqlflow_public. Document, sample code and other materials for SQLFlow
★ 1kceleborn. Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.
★ 1.1kgo-awesome. Go 语言优秀资源整理,为项目落地加速🏃
★ 6.6kbitsail. BitSail is a distributed high-performance data integration engine which supports batch, streaming and incremental scenarios. BitSail is widely used to synchronize hundreds of trillions of data every day.
★ 1.7kawesome-data-catalogs. 📙 Awesome Data Catalogs and Observability Platforms.
★ 1.1kkestra. Event Driven Orchestration & Scheduling Platform for Mission Critical Applications
★ 28kakhq. Kafka GUI for Apache Kafka to manage topics, topics data, consumers group, schema registry, connect and more...
★ 3.8kspark-bench. Benchmark Suite for Apache Spark
★ 242tview. Terminal UI library with rich, interactive widgets — written in Golang
★ 14kgox. A dead simple, no frills Go cross compile tool
★ 4.6kzio. ZIO — A type-safe, composable library for async and concurrent programming in Scala
★ 4.4kpinot. Apache Pinot - A realtime distributed OLAP datastore
★ 6.1kgitpod. The developer platform for on-demand cloud development environments to create software faster and more securely.
★ 14kDockerfiles. 50+ DockerHub public images for Docker & Kubernetes - DevOps, CI/CD, GitHub Actions, CircleCI, Jenkins, TeamCity, Alpine, CentOS, Debian, Fedora, Ubuntu, Hadoop, Kafka, ZooKeeper, HBase, Cassandra, Solr, SolrCloud, Presto, Apache Drill, Nifi, Spark, Consul, Riak
★ 1.4kOpenMetadata. The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
★ 15kCloudShuffleService. Cloud Shuffle Service(CSS) is a general purpose remote shuffle solution for compute engines, including Spark/Flink/MapReduce.
★ 261MYDB. 一个简单的数据库实现
★ 1.2kkubebuilder. Kubebuilder - SDK for building Kubernetes APIs using CRDs
★ 9.3kamoro. Apache Amoro(incubating) is a Lakehouse management system built on open data lake formats.
★ 1.2kk9s. 🐶 Kubernetes CLI To Manage Your Clusters In Style!
★ 34kgo-zero. A cloud-native Go microservices framework with cli tool for productivity.
★ 33kkratos. Your ultimate Go microservices framework for the cloud-native era.
★ 26kalldata. 🔥🔥 AllData可定义数据中台,以数据平台为底座,以数据中台为桥梁,以机器学习平台为工厂,以大模型应用为上游产品,提供全链路数字化解决方案。产品正式演示体验、社群咨询、商务采购:https://docs.qq.com/doc/DVHlkSEtvVXVCdEFo
★ 3.1kgitflow-avh. AVH Edition of the git extensions to provide high-level repository operations for Vincent Driessen's branching model
★ 5.5khive. Apache Hive
★ 6kgo-gin-example. An example of gin
★ 7.2kngods-stocks. New Generation Opensource Data Stack Demo
★ 456tidb. TiDB is built for agentic workloads that grow unpredictably, with ACID guarantees and native support for transactions, analytics, and vector search. No data silos. No noisy neighbors. No infrastructure ceiling.
★ 40kawesome-database-learning. A list of learning materials to understand databases internals
★ 11kcube. 📊 Cube Core is open-source semantic layer for AI, BI and embedded analytics
★ 21kflinkful. flink endpoint for open world
★ 28scaleph. Open data platform based on Kubernetes. Scaleph supports SeaTunnel、Flink and Doris backended by SeaTunnel on Flink engine、Flink Kubernetes Operator and Doris operator.
★ 395uniffle. Uniffle is a high performance, general purpose Remote Shuffle Service.
★ 453flink-cdc. Flink CDC is a streaming data integration tool
★ 6.5kdatabend. Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.
★ 9.4kmongo-spark. The MongoDB Spark Connector
★ 730tispark. TiSpark is built for running Apache Spark on top of TiDB/TiKV
★ 888spark-oracle. On the fly, translation of Spark programs to run natively on your Oracle DB. Your Spark programs require no changes.
★ 36cassandra-spark-connector. Apache Spark to Apache Cassandra connector
★ 2kDatasourceX. Java
★ 69zdh_web. 大数据采集,抽取平台,zdh_web是zdh系列服务的可视化管理平台,包含数据采集,调度,权限,审批流,私域营销等模块
★ 536chengying-schema. Shell
★ 7chengying. 一款支持标准化schema定义、自动化部署产品包的软件,旨在对产品包下每个服务进行部署、升级、卸载、配置等操作,解放人工运维成本。
★ 199PowerJob. Enterprise job scheduling middleware with distributed computing ability.
★ 7.8kYCSB. Yahoo! Cloud Serving Benchmark
★ 5.2kjepsen. A framework for distributed systems verification, with fault injection
★ 7.5kbraft. An industrial-grade C++ implementation of RAFT consensus algorithm based on brpc, widely used inside Baidu to build highly-available distributed systems.
★ 4.2kCommonDataModel. Definition and DDLs for the OMOP Common Data Model (CDM)
★ 1.1kSREWorks. Cloud Native DataOps & AIOps Platform | 云原生数智运维平台
★ 2kDorisParser. DorisDB SQL解析器Java实现;Clickhouse SQL解析器Java实现
★ 103SparkCube. SparkCube is an open-source project for extremely fast OLAP data analysis. SparkCube is an extension of Apache Spark.
★ 136LarkMidTable. LarkMidTable 是一站式开源的数据中台,实现中台的 基础建设,数据治理,数据开发,监控告警,数据服务,数据的可视化,实现高效赋能数据前台并提供数据服务的产品。
★ 2.1kpf4j. Plugin Framework for Java (PF4J)
★ 2.7kframeless. Expressive types for Spark.
★ 898spark-clickhouse-connector. Spark ClickHouse Connector build on DataSourceV2 API
★ 217godlp. sensitive information protection toolkit
★ 1kdataease. 🔥 人人可用的开源 BI 工具,数据可视化神器。An open-source BI tool alternative to Tableau.
★ 24kesProc. esProc SPL is a JVM-based programming language designed for structured data computation, serving as both a data analysis tool and an embedded computing engine.
★ 4.7kiotdb. Apache IoTDB
★ 6.4kspark-terasort. Spark Terasort
★ 121XSQL. Unified SQL Analytics Engine Based on SparkSQL
★ 211datart. Datart is a next generation Data Visualization Open Platform
★ 2.3kfilling. 非常易用,高性能、支持实时流式和离线批处理的海量数据处理产品,架构于 Apache Flink之上。
★ 216dinky. Dinky is a real-time data development platform based on Apache Flink, enabling agile data development, deployment and operation.
★ 3.7kcoral. Coral is a translation, analysis, and query rewrite engine for SQL and other relational languages.
★ 906docusaurus. Easy to maintain open source documentation websites.
★ 66kHow-To-Ask-Questions-The-Smart-Way. 本文原文由知名 Hacker Eric S. Raymond 所撰寫,教你如何正確的提出技術問題並獲得你滿意的答案。
★ 35kReal-Time-Rendering-4th-Bibliography-Collection. Real-Time Rendering 4th (RTR4) 参考文献合集典藏 | Collection of <Real-Time Rendering 4th (RTR4)> Bibliography / Reference
★ 3.9ksql-runner. Scala
★ 17datafusion. Apache DataFusion SQL Query Engine
★ 9kstreampark. Make stream processing easier! Easy-to-use streaming application development framework and operation platform.
★ 4.3khop. Hop Orchestration Platform
★ 1.4kdebezium. Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
★ 13kflow. 🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
★ 956pulsar. Apache Pulsar - distributed pub-sub messaging system
★ 15kairbyte. Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
★ 22kdbt-core. dbt enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications.
★ 14kinlong. Apache InLong - a one-stop, full-scenario integration framework for massive data
★ 1.5kflink-spark-submiter. 从本地IDEA提交Flink/Spark任务到Yarn/k8s集群
★ 167FATE. An Industrial Grade Federated Learning Framework
★ 6.1ksparkMeasure. This repository contains the development code for sparkMeasure, an Apache Spark performance analysis and troubleshooting library. It simplifies collecting, aggregating, and exporting Spark task/stage metrics, and is designed for practical use by developers and data engineers in interactive analysis, testing, and production monitoring workflows.
★ 827docker-hadoop. Apache Hadoop docker image
★ 2.3kbistoury. Bistoury是去哪儿网的java应用生产问题诊断工具,提供了一站式的问题诊断方案
★ 4.1kspline-spark-agent. Spline agent for Apache Spark
★ 207tis. Support agile DataOps Based on Flink, DataX and Flink-CDC, Chunjun with Web-UI
★ 1.3ksealos. Deploy real projects from GitHub or your AI coding agent, then keep them running with AI-powered operations.
★ 18kspline. Data Lineage Tracking And Visualization Solution
★ 665kcat. Generic command line non-JVM Apache Kafka producer and consumer
★ 5.8kOpenLineage. An Open Standard for lineage metadata collection
★ 2.6kegeria. Egeria core
★ 918marquez. Collect, aggregate, and visualize a data ecosystem's metadata
★ 2.3kmy-life. 算是简历吧....
★ 597spark-fast-tests. Apache Spark testing helpers (dependency free & works with Scalatest, uTest, and MUnit)
★ 457doris. Apache Doris is a real-time analytics and hybrid search database for AI agents.
★ 16khetu-core. Java
★ 571OpenRefine. OpenRefine is a free, open source power tool for working with messy data and improving it
★ 12kDataCleaner. The premier open source Data Quality solution
★ 651flinkStreamSQL. 基于开源的flink,对其实时sql进行扩展;主要实现了流与维表的join,支持原生flink SQL所有的语法
★ 2.1kckman. This is a tool which used to manage and monitor ClickHouse database
★ 485kubeflow. Machine Learning Toolkit for Kubernetes
★ 16kssb-kylin. Star Schema Benchmark Tool for Apache Kylin
★ 96airpal. Web UI for PrestoDB.
★ 2.7ktabix. Tabix.io UI
★ 2.3klogkit. Very powerful server agent for collecting & sending logs & metrics with an easy-to-use web console.
★ 1.4kredash. Make Your Company Data Driven. Connect to any data source, easily visualize, dashboard and share your data.
★ 29kaliyun-emapreduce-datasources. Extended datasource support for Spark/Hadoop on Aliyun E-MapReduce.
★ 170javalin. A simple and modern Java and Kotlin web framework
★ 8.3kcdap. An open source framework for building data analytic applications.
★ 789hudi-resources. 汇总Apache Hudi相关资料
★ 556SynapseML. Simple and Distributed Machine Learning
★ 5.2klearning-spark. Example code from Learning Spark book
★ 3.9kamundsen. Amundsen is a metadata driven application for improving the productivity of data analysts, data scientists and engineers when interacting with data.
★ 4.8kTransmogrifAI. TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library for building modular, reusable, strongly typed machine learning workflows on Apache Spark with minimal hand-tuning
★ 2.3kangel. A Flexible and Powerful Parameter Server for large-scale machine learning
★ 6.8kdeeplearning4j. Suite of tools for deploying and training deep learning models using the JVM. Highlights include model import for keras, tensorflow, and onnx/pytorch, a modular and tiny c++ library for running math code and a java based math library on top of the core c++ library. Also includes samediff: a pytorch/tensorflow like library for running deep learn...
★ 14kTensorFlowOnSpark. TensorFlowOnSpark brings TensorFlow programs to Apache Spark clusters.
★ 3.8kipex-llm. Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discrete GPU such as Arc, Flex and Max); seamlessly integrate with llama.cpp, Ollama, HuggingFace, LangChain, LlamaIndex, vLLM, DeepSpeed, Axolotl, etc.
★ 8.9kdrill. Apache Drill is a distributed MPP query layer for self describing data
★ 2kyanagishima. Web UI for Trino, Hive and SparkSQL
★ 632moonbox. Moonbox is a DVtaaS (Data Virtualization as a Service) Platform
★ 505QStreaming. A simplified, lightweight ETL pipeline framework for build stream/batch processing applications on top of Apache Spark
★ 104deequ. Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.
★ 3.6kdcm4cheSystem. 对dicom文件进行归档整理的医疗影像平台
★ 27djl. An Engine-Agnostic Deep Learning Framework in Java
★ 4.8kmetorikku. A simplified, lightweight ETL Framework based on Apache Spark
★ 588VisualDL. Deep Learning Visualization Toolkit(『飞桨』深度学习可视化工具 )
★ 4.9kmetabase. The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with data :bar_chart:
★ 48kEFAK. A AI-Driven, Distributed and high-performance monitoring system, for comprehensive monitoring and management of kafka cluster.
★ 3.2kfraud-detection-demo. Repository for Advanced Flink Application Patterns series
★ 349KnowStreaming. 一站式云原生实时流数据平台,通过0侵入、插件化构建企业级Kafka服务,极大降低操作、存储和管理实时流数据门槛
★ 7.2kspark-authorizer. A Spark SQL extension which provides SQL Standard Authorization for Apache Spark | This repo is contributed to Apache Kyuubi | 项目已迁移至 Apache Kyuubi
★ 183spark-daria. Essential Spark extensions and helper methods ✨😲
★ 767flink-streaming-platform-web. 基于flink的实时流计算web平台
★ 1.9kAddax. A fast and versatile ETL tool that can transfer data between RDBMS and NoSQL seamlessly
★ 1.4knode-crawler. Web Crawler/Spider for NodeJS + server-side jQuery ;-)
★ 6.8kchunjun. A data integration framework
★ 4.1kmetacat. Java
★ 1.7kdatahub. The Context Platform for your Data and AI Stack
★ 12kkylin. Apache Kylin
★ 3.8kAwesome-AI4Med. A curated list of medical LLMs, multimodal systems, datasets, benchmarks, and more. 🏥
★ 2.9kgobblin. A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.
★ 2.3kides. 智能数据探索服务(Intelligent Data Exploration Service),一站式Data + AI数据解决方案!
★ 36hera. hera 分布式任务调度系统 大数据任务调度系统 任务调度 (数据部门专用)
★ 378h5-Dooring. H5 Page Maker, H5 Editor, LowCode. Make H5 as easy as building blocks. | 让H5制作像搭积木一样简单, 轻松搭建H5页面, H5网站, PC端网站,LowCode平台.
★ 10knifi. Apache NiFi
★ 6.2kkylo. Kylo is a data lake management software platform and framework for enabling scalable enterprise-class data lakes on big data technologies such as Teradata, Apache Spark and/or Hadoop. Kylo is licensed under Apache 2.0. Contributed by Teradata Inc.
★ 1.1ksylph. Stream computing platform for bigdata
★ 406AthenaX. SQL-based streaming analytics platform at scale
★ 1.2kdr-elephant. Dr. Elephant is a job and flow-level performance monitoring and tuning tool for Apache Hadoop and Apache Spark
★ 1.4kfacebook-hive-udfs. Facebook's Hive UDFs
★ 275vue-data-board. A Data Analysis Board in Vue.
★ 1.3kIQL. An ad hoc query service based on the spark sql engine.(基于spark sql引擎的即席查询服务)
★ 376Exchangis. Exchangis is a lightweight,highly extensible data exchange platform that supports data transmission between structured and unstructured heterogeneous data sources
★ 461LightGBM. A fast, distributed, high performance gradient boosting (GBT, GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other machine learning tasks.
★ 19kInfoSpider. INFO-SPIDER 是一个集众多数据源于一身的爬虫工具箱🧰,旨在安全快捷的帮助用户拿回自己的数据,工具代码开源,流程透明。支持数据源包括GitHub、QQ邮箱、网易邮箱、阿里邮箱、新浪邮箱、Hotmail邮箱、Outlook邮箱、京东、淘宝、支付宝、中国移动、中国联通、中国电信、知乎、哔哩哔哩、网易云音乐、QQ好友、QQ群、生成朋友圈相册、浏览器浏览历史、12306、博客园、CSDN博客、开源中国博客、简书。
★ 8.2kQualitis. Qualitis is a one-stop data quality management platform that supports quality verification, notification, and management for various datasource. It is used to solve various data quality problems caused by data processing. https://github.com/WeBankFinTech/Qualitis
★ 766DataSphereStudio. DataSphereStudio is a one stop data application development& management portal, covering scenarios including data exchange, desensitization/cleansing, analysis/mining, quality measurement, visualization, and task scheduling.
★ 3.3kdatax-web. DataX集成可视化页面,选择数据源即可一键生成数据同步任务,支持RDBMS、Hive、HBase、ClickHouse、MongoDB等数据源,批量创建RDBMS数据同步任务,集成开源调度系统,支持分布式、增量同步数据、实时查看运行日志、监控执行器资源、KILL运行进程、数据源信息加密等。
★ 6kfake2db. create custom test databases that are populated with fake data
★ 2.4k