Montréal

Stefen

Elite
@Stefen-Taime

Data Engineer & Ops

Iceberg-Dbt-Trino-Hive-modern-open-source-data-stack. To provide a deeper understanding of how the modern, open-source data stack consisting of Iceberg, dbt, Trino, and Hive operates within a music streaming platform, let’s delve into the detailed workflow and benefits of each component.

47

car-price-predictor. Predicting Car Prices with FastAPI, Streamlit, MLflow, Kafka, and Debezium: A Practical Demonstration

25

modern-data-pipeline. reating a modern data pipeline using a combination of Terraform, AWS Lambda and S3, Snowflake, DBT, Mage AI, and Dash.

15

ETL-Data-Pipeline-RDBMS-TO-HDFS-using-Airflow-Apache-Sqoop-Spark-Postgres-and-Hive. This project aims to move the data from a Relational database system (RDBMS) to a Hadoop file system (HDFS)

11

stream-ingestion-redpanda-minio. In this article, you will learn how to set up a real-time data processing and analytics environment using Docker, MySQL, Redpanda, MinIO, and Apache Spark.

11

Scalable-RSS-Feed-Pipeline. In this article, we'll walk through how to build a scalable ETL pipeline using Apache Airflow, Kafka, and Python, Mongo and Flask

9

projet_data. Utilizing of open source technologies for the implementation of a data pipeline

9

Nifi-ETL-Data-Pipeline. This post will demonstrate the creation of a containerized data engineer environment using Docker Stacks.

7

ModernDataEngineerPipeline. Building a Robust Data Pipeline: Integrating Proxy Rotation, Kafka, MongoDB, Redis, Logstash, Elasticsearch, and MinIO for Efficient Web Scraping

6

Kafka-pipeline. In the following post, we will learn how to build a data pipeline using a combination of open-source software (OSS), including Debezium, Apache Kafka, Kafka Connect.

6

railway-station-streaming. Python

6

Uber_projet. Unveiling the true cost of your ride-sharing and food delivery habits with an ELT data pipeline, PostgreSQL, dbt, and Power BI.

5

Real-time-Data-Processing-and-Analysis-with-Kafka-Connect-KSQL-Elasticsearch-and-Flask. The project aims to demonstrate how to work with real-time data using Kafka, KSQL, Elasticsearch, and Flask. It shows how to perform joins on Kafka topics, ingest data into Elasticsearch using Kafka Connect, and build a REST API to provide real-time metrics to end-users.

5

Free-Real-time-Flight-Status-Pipeline. real-time flight status data pipeline using a myriad of technologies such as Kafka, Schema Registry, Avro, GraphQL, Postgres, and React.

4

etl_onaws_deploy_with_terraform. The objective of this guide is to demonstrate how to automate the deployment of a data pipeline on AWS using Terraform. The pipeline will utilize AWS services such as Lambda, Glue, Crawler, Redshift, and

3

docSearch. Our project is a testament to this need, offering a comprehensive solution that combines modern technologies and architectures to create a powerful document search engine. This engine is not just a tool but a sophisticated ecosystem designed to handle complex data processing and retrieval tasks.

3

datawarehouse. Python

3

investissement. Jenkins Delta pipeline

2

Real-Time-Data-Pipeline-Snake-Game. Dynamic Snake Game: Unleashing Real-Time Streaming Analytics with Redis, Kafka, Flink, ClickHouse & Chart.js in an Online Snake Game via Flask API

2

Big-O-Algorithm. we’ll explain Big O notation an real-world Python examples to illustrate how it can be applied to various time complexities.

2

Gmail-to-MongoDB-Script. This script facilitates the automation of fetching emails from a user's Gmail account and storing them into a MongoDB database. The emails fetched are filtered by specific labels such as Promotions, Social, Updates, and Forums. The script is intended to run continuously, checking for new emails every minute.

2

build_api_devops_pipeline. Dockerfile

2

open-source-data. This repository contains structured datasets in various categories

2

eventmusic. EventMusic Producer is a Dockerized application designed to read data and output them to a Kafka topic, using Avro schemas for data serialization. It integrates seamlessly with Kafka and the Schema Registry to manage the flow of event data linked to music event information.

1

MongoElasticMigrator. This tool migrates data from MongoDB collections to Elasticsearch indices. It's built using Rust and supports configurable migrations.

1

Master-ElasticSearch. Setting Up and Querying Elasticsearch with Python and Streamlit

1

legal-document-analyzer. Python

1

terraform_snowflake_devops. Develop a scalable and secure data infrastructure, Integrate diverse data sources into Snowflake.

1

build_api_auth2.0. Python

1

IA_Data_Pipeline. The goal is to develop an intuitive platform where users can search for Airbnb apartments based on a target city, budget, and duration of stay, all powered by the intelligent language model, GPT-3.

1

airflow_etl. The Pipeline for updating data between OLTP and OLAP environments

1

Visualizing-Bitcoin-Pipeline. Visualize the exchange rate of Bitcoin to USD using FastAPI, Prometheus, Grafana Docker And Jenkins

1

mcp-ml-platform. Python

1

myUberEats_dataPipeline. Building a Modern Uber Eats Data Pipeline

1

master-airflow. big data, the ability to extract, transform, and store data from the web into different storage systems like MongoDB, PostgreSQL, MinIO, and Elasticsearch is a crucial skill for developers and data scientists

1

Retention_Analysis_Pipeline. Python

1

Stefen-Taime. Config files for my GitHub profile.

1
37
Apply