This is your work, valued
Iceberg-Dbt-Trino-Hive-modern-open-source-data-stack. To provide a deeper understanding of how the modern, open-source data stack consisting of Iceberg, dbt, Trino, and Hive operates within a music streaming platform, let’s delve into the detailed workflow and benefits of each component.
47car-price-predictor. Predicting Car Prices with FastAPI, Streamlit, MLflow, Kafka, and Debezium: A Practical Demonstration
25modern-data-pipeline. reating a modern data pipeline using a combination of Terraform, AWS Lambda and S3, Snowflake, DBT, Mage AI, and Dash.
15ETL-Data-Pipeline-RDBMS-TO-HDFS-using-Airflow-Apache-Sqoop-Spark-Postgres-and-Hive. This project aims to move the data from a Relational database system (RDBMS) to a Hadoop file system (HDFS)
11stream-ingestion-redpanda-minio. In this article, you will learn how to set up a real-time data processing and analytics environment using Docker, MySQL, Redpanda, MinIO, and Apache Spark.
11Scalable-RSS-Feed-Pipeline. In this article, we'll walk through how to build a scalable ETL pipeline using Apache Airflow, Kafka, and Python, Mongo and Flask
9projet_data. Utilizing of open source technologies for the implementation of a data pipeline
9Nifi-ETL-Data-Pipeline. This post will demonstrate the creation of a containerized data engineer environment using Docker Stacks.
7ModernDataEngineerPipeline. Building a Robust Data Pipeline: Integrating Proxy Rotation, Kafka, MongoDB, Redis, Logstash, Elasticsearch, and MinIO for Efficient Web Scraping
6Kafka-pipeline. In the following post, we will learn how to build a data pipeline using a combination of open-source software (OSS), including Debezium, Apache Kafka, Kafka Connect.
6railway-station-streaming. Python
6Uber_projet. Unveiling the true cost of your ride-sharing and food delivery habits with an ELT data pipeline, PostgreSQL, dbt, and Power BI.
5Real-time-Data-Processing-and-Analysis-with-Kafka-Connect-KSQL-Elasticsearch-and-Flask. The project aims to demonstrate how to work with real-time data using Kafka, KSQL, Elasticsearch, and Flask. It shows how to perform joins on Kafka topics, ingest data into Elasticsearch using Kafka Connect, and build a REST API to provide real-time metrics to end-users.
5Free-Real-time-Flight-Status-Pipeline. real-time flight status data pipeline using a myriad of technologies such as Kafka, Schema Registry, Avro, GraphQL, Postgres, and React.
4etl_onaws_deploy_with_terraform. The objective of this guide is to demonstrate how to automate the deployment of a data pipeline on AWS using Terraform. The pipeline will utilize AWS services such as Lambda, Glue, Crawler, Redshift, and
3docSearch. Our project is a testament to this need, offering a comprehensive solution that combines modern technologies and architectures to create a powerful document search engine. This engine is not just a tool but a sophisticated ecosystem designed to handle complex data processing and retrieval tasks.
3datawarehouse. Python
3investissement. Jenkins Delta pipeline
2Real-Time-Data-Pipeline-Snake-Game. Dynamic Snake Game: Unleashing Real-Time Streaming Analytics with Redis, Kafka, Flink, ClickHouse & Chart.js in an Online Snake Game via Flask API
2Big-O-Algorithm. we’ll explain Big O notation an real-world Python examples to illustrate how it can be applied to various time complexities.
2Gmail-to-MongoDB-Script. This script facilitates the automation of fetching emails from a user's Gmail account and storing them into a MongoDB database. The emails fetched are filtered by specific labels such as Promotions, Social, Updates, and Forums. The script is intended to run continuously, checking for new emails every minute.
2build_api_devops_pipeline. Dockerfile
2open-source-data. This repository contains structured datasets in various categories
2eventmusic. EventMusic Producer is a Dockerized application designed to read data and output them to a Kafka topic, using Avro schemas for data serialization. It integrates seamlessly with Kafka and the Schema Registry to manage the flow of event data linked to music event information.
1MongoElasticMigrator. This tool migrates data from MongoDB collections to Elasticsearch indices. It's built using Rust and supports configurable migrations.
1Master-ElasticSearch. Setting Up and Querying Elasticsearch with Python and Streamlit
1legal-document-analyzer. Python
1terraform_snowflake_devops. Develop a scalable and secure data infrastructure, Integrate diverse data sources into Snowflake.
1build_api_auth2.0. Python
1IA_Data_Pipeline. The goal is to develop an intuitive platform where users can search for Airbnb apartments based on a target city, budget, and duration of stay, all powered by the intelligent language model, GPT-3.
1airflow_etl. The Pipeline for updating data between OLTP and OLAP environments
1Visualizing-Bitcoin-Pipeline. Visualize the exchange rate of Bitcoin to USD using FastAPI, Prometheus, Grafana Docker And Jenkins
1mcp-ml-platform. Python
1myUberEats_dataPipeline. Building a Modern Uber Eats Data Pipeline
1master-airflow. big data, the ability to extract, transform, and store data from the web into different storage systems like MongoDB, PostgreSQL, MinIO, and Elasticsearch is a crucial skill for developers and data scientists
1Retention_Analysis_Pipeline. Python
1Stefen-Taime. Config files for my GitHub profile.
1