Data Engineer and Data Science Working on several Big Data Technologies like Apache Spark, Hadoop, Kafka, Imapala, Apache Beam, Hive and others
around-dataengineering. A Data Engineering & Machine Learning Knowledge Hub
1.1kdata-machinelearning-the-boring-way. Build & Learn Data Engineering,Machine Learning over Kubernetes. No Shortcut approach.
57VectorVerse. Explore Multiple Vector Databases and chat with documents on Multiple LLM models, private LLM models
48streamlit-healthcare-ML-App. Streamlit example showing Scikit Learn & Pyspark ML over Healthcare data ! Its simple !!
32snowflakeGPT. A Snowflake GPT Demo using SqlAlchemy
23SmileDetection. This is an Android Based Smile Detection Project , using OpenCV and JavaCV. It works well with android and to make it work you need to install the open cv libraries in your Android Phone.
16Awesome_Algorithm. Collection of Interesting Algorithms
16Kubectl-GPT. Kubernetes cli (kubectl) powered by GPT
15mlx-video-qa. Explore the capabilities of the MLX library and leverage the genAI stack on MacOS to interact with any video.
9dataengineering-agent. Data Engineering Agent Using Open AI Function Call
9mimic-ai. MIMICAI - Exploring MIMIC Data Using LLM and Open-WebUI
8cspaper-ai. Tool is to serve as an AI for Computer Science Papers, capable of referencing and extracting additional information
6SQL_GPT. A quick dirty code to generate sql code using chatGPT
6data-engineer-roadmap. Roadmap to becoming a data engineer in 2020
5fbmoviesuggestion. In Facebook, user can like movies, so based on that , here its extracting those movies from each of friend's user and then it fetches the details of each of the movie using API , then this recommends the best movie to each of the user within the movies in friends' circle.
3MachineLearning-using-R. CodesMachine Learning Source files
3evolveML. Intention is to use different algorithms of Machine Learning in R-Programming and Python to work with various dimension and range of data. The implementation will be based on BigData framework and main point of attraction will be Spark and Hive includinh hadoop
3hue. Let’s Big Data. Hue is an open source Web interface for analyzing data with Apache Hadoop.
2Algorithms. Data Structures and Algorithms in Python
2healthtracker. This is an iOS app built on top of Swift. App has an UI for monitor Health over iPhone as well as in Apple Watch
2awesome-k8s-resources. A curated list of awesome Kubernetes tools and resources.
2kaggle_facebook_recruiting_human_or_bot. Code for kaggle competition at https://www.kaggle.com/c/facebook-recruiting-iv-human-or-bot
2transfusion-webui. Transfusion UI based on Open-Web-UI | Experiment
1incubator-ignite. Mirror of Apache Ignite (Incubating)
1spark-cs190.1x. Working of CS190.1x, Scalable Machine Learning
1spark-docker. Spark Docker Environment for testing Purpose
1spark. Mirror of Apache Spark
1twitter-sentiment-analysis-tutorial-201107. Code to reproduce the simple sentiment analysis from my presentation
1monger. Monger is an idiomatic Clojure MongoDB driver for a more civilized age: with sane defaults, batteries included, well documented, very fast
1Flink-Examples. Apache Flink work and Examples
1GeoSpark. A Cluster Computing System for Processing Large-Scale Spatial Data
1spark_hungarian. Hungarian Method using Apache Spark
1kafka-example. Simple example for reading and writing into Kafka
1myvagrant. edX: Introduction to Big Data with Apache Spark
1data-science-from-scratch. code for Data Science From Scratch book
1