coreml-llm-cli. CLI to demonstrate running a large language model (LLM) on Apple Neural Engine.
131more-ane-transformers. Run transformers (incl. LLMs) on the Apple Neural Engine.
63sublime-spotify. Control Spotify from Sublime Text 2 or 3.
51ModernBERT-AppleNeuralEngine. ModernBERT model optimized for Apple Neural Engine.
38icloud-tabs. Update iCloud tabs from Chrome.
36CoreMLInspect. See the device (CPU/GPU/ANE) and estimated cost for every layer in your CoreML model.
25ue-speaker-app. App to enable Siri/Shortcuts support for UE speakers.
12apple-silicon-4bit-quant. Supporting code for "LLMs for your iPhone: Whole-Tensor 4 Bit Quantization"
11token-recycling. Unofficial implementation of Token Recycling self-speculative decoding method.
9norm-tweaking. Post post-training-quantization (PTQ) method for improving LLMs. Unofficial implementation of https://arxiv.org/abs/2309.02784
8ganq. GPU Adaptive Non-Uniform Quantization (GANQ) Unofficial Implementation
4mlx-squeezellm-gradients. SqueezeLLM-style gradients/Fisher Information collection in MLX
4alfred-darkmode. An Alfred workflow to toggle Yosemite's dark and light modes.
4zed-flatbuffers. zed.dev extension with language support for FlatBuffers
4netron. Visualizer for neural network, deep learning, and machine learning models
3esp-lights. ESP8266-controlled WS2812s
3flatbuffers-language-server. A language server implementation for Google FlatBuffers.
3QuaRot. Code for QuaRot, an end-to-end 4-bit inference of large language models.
2GPTQModel-1. Production ready LLM model compression/quantization toolkit with hw accelerated inference support for both cpu/gpu via HF, vLLM, and SGLang.
2macmon. 🦀⚙️ Sudoless performance monitoring for Apple Silicon processors. CPU / GPU / RAM usage, power consumption & temperature 🌡️
2WhisperKit. Swift native on-device speech recognition with Whisper for Apple Silicon
1Spec-Bench. Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)
1tree-sitter-flatbuffers. tree-sitter grammar for FlatBuffers
1litgpt. Hackable implementation of state-of-the-art open-source LLMs based on nanoGPT. Supports flash attention, 4-bit and 8-bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training. Apache 2.0-licensed.
1whisperkittools. Python tools for WhisperKit: Model conversion, optimization and evaluation
1swift-transformers. Swift Package to implement a transformers-like API in Swift
1