This is your work, valued
Founder, Safe Link Network. NYU Data Science PhD. Always curious :-)
malware-discoverer. Proactive malware campaign discovery system
★ 21informationtracer. Python client to interact with Information Tracer
★ 11cross-platform. cross platform analysis of the spread of information
★ 3fakenews-detection. Python
★ 2python-exercise-app. simple python app
★ 2twitterpythontoolkit. a nice python package that enables 12+ functions to do Twitter data analysis with Python and MongoDB
★ 1capstone-topic-discovery. NYU 2023 Capstone project to discover topics from large social media corpus using NLP, Clustering, LLM.
★ 1python-louvain. Louvain Community Detection
★ 1search-tweets-python. Python client for the Twitter search endpoints (v2/Labs/premium/enterprise). Now supports Twitter API v2 recent search.
★ 1Facebook-SSL-Pinning-Bypass. APKs and libraries to analyze Facebook network traffic and bypass SSL pinning on Android devices for educational and security testing purposes.
★ 61Facebook-SSL-Pinning-Bypass. Bypass Facebook SSL/TLS certificate pinning on Android to intercept & analyze HTTPS traffic — root & no-root, works with Burp Suite, mitmproxy & Reqable (2026).
★ 19tmvis. An archival thumbnail visualization server
★ 10Deep-Live-Cam. real time face swap and one-click video deepfake with only a single image
★ 95kclaw-code. An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
★ 195klast30days-skill. AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary
★ 55kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kyoutube_trans. Python
★ 5TikTok-Content-Scraper. TikTok Content Scraper \\ No API-Key needed with minimal dependencies and citable | Download videos (MP4), slides (JPEGs+MP3) and metadata of users, music, file, hashtags, content, interactions etc.
★ 105Telegram-OSINT. In-depth repository of Telegram OSINT resources covering, tools, techniques & tradecraft.
★ 1.9klinkedin_scraper. A library that scrapes Linkedin for user data
★ 4.4knoVNC. VNC client web application
★ 14kzeeschuimer. A browser extension to collect social media data with.
★ 391prefect. Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
★ 24kyt-dlp. A feature-rich command-line audio/video downloader
★ 181kpolymarket-subgraph. Polymarket's public subgraph manifest for indexing on-chain trade, volume, user, liquidity and market data.
★ 215glances. Glances an Eye on your system. A top/htop alternative for GNU/Linux, BSD, Mac OS and Windows operating systems.
★ 33kSpider_XHS. 小红书爬虫数据采集,小红书全域运营解决方案
★ 7.1kinstagram-users-scraper. Instagram Scraper. Scrape Instagram followers, following list, and post authors. Download CSV files with Instagram users from followers, following, tag and location pages.
★ 151facebook-group-members-scraper. Facebook Group Members Extractor. Download Facebook group members in CSV.
★ 322echarts-wordcloud. Word Cloud extension based on Apache ECharts and wordcloud2.js
★ 1.8kpython-genai. Google Gen AI Python SDK provides an interface for developers to integrate Google's generative models into their Python applications.
★ 3.9kecharts. Apache ECharts is a powerful, interactive charting and data visualization library for browser
★ 67kTGcollector. Web GUI for collecting messages from Telegram channels
★ 118label-studio. Label Studio is a multi-type data labeling and annotation tool with standardized output format
★ 28kapt-proxy. Lightweight package cache proxy for APT, YUM and APK with automatic mirror acceleration.
★ 266ddclient. ddclient updates dynamic DNS entries for accounts on a wide range of dynamic DNS services.
★ 3.5kcodex. Lightweight coding agent that runs in your terminal
★ 103kgraphistry-cli. Graphistry admin docs: launch, configure, use, & debug
★ 30nvidia-container-toolkit. Build and run containers leveraging NVIDIA GPUs
★ 4.5kERNIE. The official repository for ERNIE 4.5 and ERNIEKit – its industrial-grade development toolkit based on PaddlePaddle.
★ 7.7ktraktok. The goal of traktok is to provide easy access to TikTok data.
★ 114graphistry-js. API for controlling and reacting to embedded visualizations
★ 48DirectedLouvain. C++
★ 43instagram-private-api. NodeJS Instagram private API SDK. Written in TypeScript.
★ 6.5kchardet. Python character encoding detector
★ 2.6kcdnjs. 🤖 CDN assets - The #1 free and open source CDN built to make life easier for developers.
★ 11kcudf. cuDF - GPU DataFrame Library
★ 9.7kantibot_blog_archives. Dear obfio archives. About Twitter X-Client-Transaction-Id.
★ 15twitter-tid-deobf-fork. twitter-tid-deobf cleanup and autorun. About Twitter X-Client-Transaction-Id.
★ 76twitter_openapi_python. Implementation of Twitter internal API (Twitter graphql API) in Python with data validation by pydantic
★ 95XClientTransaction. Twitter X-Client-Transaction-Id generator written in python.
★ 263twitter-tid-generator. Full generator for `X-Client-Transaction-Id` header on Twitter. Made for my blog https://antibot.blog
★ 64Political-Science. 政治
★ 2.4kjupyterlite. Wasm powered Jupyter running in the browser 💡
★ 4.9kZeno. State-of-the-art web crawler 🔱
★ 421EasyOCR. Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
★ 30kpytesseract. A Python wrapper for Google Tesseract
★ 6.4kfree-llm-api-resources. A list of free LLM inference resources accessible via API.
★ 29kAwesome-LLM-in-Social-Science. Awesome papers involving LLMs in Social Science.
★ 639gridjs. Advanced table plugin
★ 4.7kopenmoji. Open source emojis for designers, developers and everyone else!
★ 4.5kDeepSeek-R1.
★ 92ktavily-python. The Tavily Python SDK allows for easy interaction with the Tavily API, offering the full range of our search, extract, crawl, map, and research functionalities directly from your Python programs. Easily integrate smart search, content extraction, and research capabilities into your applications, harnessing Tavily's powerful features.
★ 1.3kbrowser_cookie3. This is a fork of browser_cookie
★ 1.1kawesome-emails. ✉️ An awesome list of resources to build better emails.
★ 2.7kijson. Iterative JSON parser with Pythonic interfaces
★ 1.1kicann-rdap. ICANN implementation of the Registry Data Access Protocol (RDAP)
★ 4512025-youtube-podcast-men-for-trump. Data from the Bloomberg News analysis on streamers and podcasters on YouTube
★ 26zhplot. 一行代码搞定 Python 图表中文展示
★ 71DeepSeek-V3. Python
★ 104kFacebook-API-Params-Generator. Generate Facebook api parameters. 生成Facebook接口参数。
★ 25mprofile. A low-overhead sampling memory profiler for Python 3
★ 31facebook_page_scraper. Scrapes facebook's pages front end with no limitations & provides a feature to turn data into structured JSON or CSV
★ 1crawl4ai. 🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
★ 76kgoogle-drive-ocamlfuse. FUSE filesystem over Google Drive
★ 5.9kWeiboSpider. 持续维护的新浪微博采集工具🚀🚀🚀
★ 4.1katproto. The AT Protocol (🦋 Bluesky) SDK for Python 🐍
★ 654cracking4crawling. 一些爬虫相关的签名、验证码破解,目前已涉及:小红书。
★ 154Crawler_Illegal_Cases_In_China. Collection of China illegal cases about web crawler 本项目用来整理所有中国大陆爬虫开发者涉诉与违规相关的新闻、资料与法律法规。致力于帮助在中国大陆工作的爬虫行业从业者了解我国相关法律,避免触碰数据合规红线。
★ 4.7krequestly. Community hub for Requestly API Client — bugs, feature requests, and roadmap. The privacy-first Postman alternative.
★ 6.7kpython-for-data-and-media-communication-gitbook. An open source book on Python tailed for communication students with zero background
★ 124bpc_chrome_support.
★ 5.3kmacOCR. Get any text on your screen into your clipboard.
★ 2.4kpython-for-data-and-media-communication. Sample codes and datasets for COMM7780/ JOUR7280 @ HKBU
★ 13gathertool. gathertool是golang脚本化开发库,目的是提高对应场景程序开发的效率;轻量级爬虫库,接口测试&压力测试库,DB操作库等。
★ 55playwright_stealth. playwright stealth
★ 981TikTokDownloader. TikTok 发布/喜欢/合辑/直播/视频/图集/音乐;抖音发布/喜欢/收藏/收藏夹/视频/图集/实况/直播/音乐/合集/评论/账号/搜索/热榜数据采集工具/下载工具
★ 15kPyWebIO. Write interactive web app in script way.
★ 4.8kTikHub-API-Python-SDK. High-performance asynchronous Douyin(抖音) TikTok Xiaohongshu(小红书) Kuaishou(快手) Weibo(微博) Instagram YouTube(油管) Twitter(X) Captcha Solver(验证码解决器) Temp Mail(临时邮箱) API(接口).
★ 831MediaCrawler. 小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫
★ 59kWeiBoCrawler. 微博数据采集,后续会加上知乎,贴吧,小红书,抖音,快手等主流媒体内容
★ 149SinaSpider. 新浪微博爬虫(Scrapy、Redis)
★ 3.3kweiboSpider. 新浪微博爬虫,用python爬取新浪微博数据
★ 9.7kweibo-crawler. 新浪微博爬虫,用python爬取新浪微博数据,并下载微博图片和微博视频
★ 4.6kweibo-scraper. Simple Weibo Scraper
★ 107weibo_terminater. Final Weibo Crawler Scrap Anything From Weibo, comments, weibo contents, followers, anything. The Terminator
★ 2.3kflowerss-bot. A telegram bot for rss reader. 一个支持应用内阅读的 Telegram RSS Bot。
★ 1.8kcaddy-grafana. Monitoring Caddy Server with Grafana (Prometheus + Loki) on Debian
★ 140grafana-infinity-datasource. API datasource for grafana. Visualize data from JSON / CSV / TSV / XML / GraphQL endpoints
★ 1.1kcaddy-logger-graphite. Go
★ 2acad-homepage.github.io. AcadHomepage: A Modern and Responsive Academic Personal Homepage
★ 2.9kgarss. Github Actions采集RSS, 打造无广告内容优质的头版头条超赞宝藏页
★ 1.4kdistributed-rss. Distributed system for reading RSS feeds.
★ 5news-feed-list-of-countries. News feed list of countries
★ 127RSSHub. 🧡 Everything is RSSible
★ 45kawesome. 😎 Awesome lists about all kinds of interesting topics
★ 491kALL-about-RSS. A list of RSS related stuff: tools, services, communities and tutorials, etc.
★ 5.9kfeedparser. Parse feeds in Python
★ 2.4kawesome-rss-feeds. Awesome RSS feeds - A curated list of RSS feeds (and OPML files) used in Recommended Feeds and local news sections of Plenary - an RSS reader, article downloader and a podcast player app for android
★ 2.7ktrafilatura. Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
★ 6.4kcancel-culture. Tools for fighting abuse on Twitter
★ 433hassreden-tracker. Hassreden-Tracker
★ 46memory.lol. memory.lol
★ 716jax. Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
★ 36kec2instances.info. Amazon EC2 instance comparison site
★ 5.7kun_resolutions_voting. Voting Data of United Nation's Resolutions
★ 1LayerZero-v2. Move
★ 788Chart.js. Simple HTML5 Charts using the <canvas> tag
★ 68kairflow. Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
★ 46kacademicpages.github.io. Github Pages template based upon HTML and Markdown for personal, portfolio-based websites.
★ 17kgraphblas-algorithms. Graph algorithms written in GraphBLAS
★ 95ruff. An extremely fast Python linter and code formatter, written in Rust.
★ 49kuv. An extremely fast Python package and project manager, written in Rust.
★ 88kDataToVideoEncoderDecoder. Python
★ 155graph-app-kit. Go from graph data to a secure and interactive visual graph app in 15 minutes. Batteries-included self-hosting of graph data apps with Streamlit, Graphistry, RAPIDS, and more!
★ 246pygraphistry. PyGraphistry is a Python library to quickly load, shape, embed, and explore big graphs with the GPU-accelerated Graphistry visual graph analyzer
★ 2.5kplaywright. Playwright is a framework for Web Testing and Automation. It allows testing Chromium, Firefox and WebKit with a single API.
★ 94knetworkit. NetworKit is a growing open-source toolkit for large-scale network analysis.
★ 870connected-component. Map Reduce Implementation of Connected Component on Apache Spark
★ 85ClickHouse. ClickHouse® is a real-time analytics database management system
★ 49khtml-screen-capture-js. A tiny, highly-customizable, single-function javascript/typescript library that captures a webpage and returns a new lightweight, self-contained HTML document. The library removes all external file dependencies while preserving the original appearance of the page. At only 12KB, it offers unparalleled speed and peerless reliability.
★ 244Leaflet.markercluster. Marker Clustering plugin for Leaflet
★ 4.2kLeaflet.glify. fully functional, ridiculously fast web gl renderer plugin for leaflet
★ 538Leaflet.PixiOverlay. Bring Pixi.js power to Leaflet maps
★ 510filelock. A platform-independent file lock for Python.
★ 973PhantomCrawler. PhantomCrawler is a Python-based web testing and research tool that simulates website interactions from multiple proxy IP addresses to analyze traffic behavior, access controls, and response patterns under different network conditions.
★ 80twitter-api-client. Implementation of X/Twitter v1, v2, and GraphQL APIs
★ 1.9kobfuscator-io-deobfuscator. A deobfuscator for scripts obfuscated by Obfuscator.io
★ 789us-news-domains. A list of over 5000 US news domains and their social media accounts
★ 48gpt-researcher. An autonomous agent that conducts deep research on any data using any LLM providers
★ 29kwormhole. A reference implementation for the Wormhole blockchain interoperability protocol.
★ 1.9ksybil-detection.
★ 271docker-nginx-forward-proxy. For Centrally manage application export traffic, A/B testing, etc. Perhaps the smallest nginx forward proxy server docker images.
★ 32ngx_http_proxy_connect_module. A forward proxy module for CONNECT request handling
★ 2ktutorials. Getting Started with Spring Boot 3:
★ 37kvaderSentiment. VADER Sentiment Analysis. VADER (Valence Aware Dictionary and sEntiment Reasoner) is a lexicon and rule-based sentiment analysis tool that is specifically attuned to sentiments expressed in social media, and works well on texts from other domains.
★ 5kyoutube-transcript-api. This is a python API which allows you to get the transcript/subtitles for a given YouTube video. It also works for automatically generated subtitles and it does not require an API key nor a headless browser, like other selenium based solutions do!
★ 8kbuster. Captcha solver extension for humans, available for Chrome, Edge and Firefox
★ 9.2kFlareSolverr. Proxy server to bypass Cloudflare protection
★ 15kAwesome-Black-Friday-Cyber-Monday. Awesome apps, software, and SaaS deals on Black Friday.
★ 7.5kInformation_tracer. Here is a collections of blogs and codes I contributed to Information Tracer
★ 2Forbidden-Buster. A tool designed to automate various techniques in order to bypass HTTP 401 and 403 response codes and gain access to unauthorized areas in the system. This code is made for security enthusiasts and professionals only. Use it at your own risk.
★ 251internetarchive. A Python and Command-Line Interface to Archive.org
★ 1.9kweibo-search. 获取微博搜索结果信息,搜索即可以是微博关键词搜索,也可以是微博话题搜索
★ 2.3kchat-intents. Clustering sentence embeddings to extract message intent
★ 175aleph. Search and browse documents and data; find the people and companies you look for.
★ 2.4kdata-journalism-courses. List of data journalism courses and programmes from universities and higher education institutions around the world
★ 76capstone-topic-discovery. NYU 2023 Capstone project to discover topics from large social media corpus using NLP, Clustering, LLM.
★ 1invid-verification-plugin. Code of the InVID EU project plugin for video and image verification
★ 47opencti. Open Cyber Threat Intelligence Platform
★ 9.7kseamless_communication. Foundational Models for State-of-the-Art Speech and Text Translation
★ 12karchivetoday. Unofficial API and CLI for archive.today.
★ 353proxy. 3proxy - tiny free proxy server
★ 5.4k3proxy-docker. 🥷 Powerful and lightweight proxy server (3proxy) in a minimal docker image
★ 264proxy-list. 🚀 Gain access to an always up-to-date compilation of highly anonymous public HTTP, SOCKS4, and SOCKS5 proxies, guaranteeing utmost relevance and accuracy, refreshed EVERY 10 MINUTES!!!
★ 170Metallic. A powerful web proxy built for speed and customization.
★ 191demergi. A proxy server that helps to bypass the DPI systems implemented by various ISPs.
★ 147numpyro. Probabilistic programming with NumPy powered by JAX for autograd and JIT compilation to GPU/TPU/CPU.
★ 2.7knpc_gzip. Code for Paper: “Low-Resource” Text Classification: A Parameter-Free Classification Method with Compressors
★ 1.8kllama_index. LlamaIndex is the leading document agent and OCR platform
★ 51kmodin. Modin: Scale your Pandas workflows by changing a single line of code
★ 10kearthly. Super simple build framework with fast, repeatable builds and an instantly familiar syntax – like Dockerfile and Makefile had a baby.
★ 12kReal-ESRGAN. Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
★ 36kjuno. Python
★ 28xorbits. Scalable Python DS & ML, in an API compatible & lightning fast way.
★ 1.2kwayback-machine-downloader. Download an entire website from the Wayback Machine.
★ 5.9ktwscrape. Python library and CLI for X/Twitter scraping with multi-account rotation and built-in rate-limit handling.
★ 2.6kipyplot. IPyPlot is a small python package offering fast and efficient plotting of images inside Python Notebooks. It's using IPython with HTML for faster, richer and more interactive way of displaying big numbers of images.
★ 428itables. Python DataFrames as Interactive DataTables
★ 969click-to-deploy. Source for Google Click to Deploy solutions listed on Google Cloud Marketplace.
★ 774social-network-url-clustering. Jupyter Notebook
★ 3SocialMediaAnalysis. codebase for social media research seminar: twitter
★ 1camel_tools. A suite of Arabic natural language processing tools developed by the CAMeL Lab at New York University Abu Dhabi.
★ 566boolean-retrieval-engine. Python implementation of a Boolean search engine
★ 26curl. A command line tool and library for transferring data with URL syntax, supporting DICT, FILE, FTP, FTPS, GOPHER, GOPHERS, HTTP, HTTPS, IMAP, IMAPS, LDAP, LDAPS, MQTT, MQTTS, POP3, POP3S, RTSP, SCP, SFTP, SMB, SMBS, SMTP, SMTPS, TELNET, TFTP, WS and WSS. libcurl offers a myriad of powerful features
★ 42knitter. Alternative Twitter front-end
★ 13kapi. Pushshift API
★ 1.4kmilvus. Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
★ 45ktwitter-scraper-selenium. Python's package to scrap Twitter's front-end easily
★ 342caddy. Fast and extensible multi-platform HTTP/1-2-3 web server with automatic HTTPS
★ 74kselenium-stealth. Trying to make python selenium more stealthy.
★ 739undetected-testing. Bypassing CAPTCHAs and bot-detection in GitHub Actions with SeleniumBase / Python.
★ 110arrow-python-js-ipc-example. Example showing how to send Arrow RecordBatches from a Python backend to a web browser.
★ 12duckdb. DuckDB is an analytical in-process SQL database management system
★ 40kpinduoduo_backdoor_detailed_report. Maybe the most detailed analysis of pdd backdoors
★ 1.8ktesseract. Tesseract Open Source OCR Engine (main repository)
★ 76kkmeans_pytorch. kmeans using PyTorch
★ 536img2vec. :fire: Use pre-trained models in PyTorch to extract vector embeddings for any image
★ 625transformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kTelepathy-Community. Public release of Telepathy, an OSINT toolkit for investigating Telegram chats.
★ 1.2kcpython. The Python programming language
★ 74kcodon. A high-performance, zero-overhead, extensible Python compiler with built-in NumPy support
★ 17kdomain-quality-ratings. Comprehensive database of ratings for 11k news domains
★ 29orjson. Fast, correct Python JSON library supporting dataclasses, datetimes, and numpy
★ 8.2kpysentimiento. A Python multilingual toolkit for Sentiment Analysis and Social NLP tasks
★ 658dtaidistance. Time series distances: Dynamic Time Warping (fast DTW implementation in C)
★ 1.2kdtw-python. Python port of R's Comprehensive Dynamic Time Warp algorithms package
★ 343BLIP. PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
★ 5.7ktiktok-signature. Generate tiktok signature token using node
★ 1k