Stephen Panaro

Expert
@smpanaro

coreml-llm-cli. CLI to demonstrate running a large language model (LLM) on Apple Neural Engine.

131

more-ane-transformers. Run transformers (incl. LLMs) on the Apple Neural Engine.

63

sublime-spotify. Control Spotify from Sublime Text 2 or 3.

51

ModernBERT-AppleNeuralEngine. ModernBERT model optimized for Apple Neural Engine.

38

icloud-tabs. Update iCloud tabs from Chrome.

36

CoreMLInspect. See the device (CPU/GPU/ANE) and estimated cost for every layer in your CoreML model.

25

ue-speaker-app. App to enable Siri/Shortcuts support for UE speakers.

12

apple-silicon-4bit-quant. Supporting code for "LLMs for your iPhone: Whole-Tensor 4 Bit Quantization"

11

token-recycling. Unofficial implementation of Token Recycling self-speculative decoding method.

9

norm-tweaking. Post post-training-quantization (PTQ) method for improving LLMs. Unofficial implementation of https://arxiv.org/abs/2309.02784

8

ganq. GPU Adaptive Non-Uniform Quantization (GANQ) Unofficial Implementation

4

mlx-squeezellm-gradients. SqueezeLLM-style gradients/Fisher Information collection in MLX

4

alfred-darkmode. An Alfred workflow to toggle Yosemite's dark and light modes.

4

zed-flatbuffers. zed.dev extension with language support for FlatBuffers

4

netron. Visualizer for neural network, deep learning, and machine learning models

3

esp-lights. ESP8266-controlled WS2812s

3

flatbuffers-language-server. A language server implementation for Google FlatBuffers.

3

QuaRot. Code for QuaRot, an end-to-end 4-bit inference of large language models.

2

GPTQModel-1. Production ready LLM model compression/quantization toolkit with hw accelerated inference support for both cpu/gpu via HF, vLLM, and SGLang.

2

macmon. 🦀⚙️ Sudoless performance monitoring for Apple Silicon processors. CPU / GPU / RAM usage, power consumption & temperature 🌡️

2

WhisperKit. Swift native on-device speech recognition with Whisper for Apple Silicon

1

Spec-Bench. Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)

1

tree-sitter-flatbuffers. tree-sitter grammar for FlatBuffers

1

litgpt. Hackable implementation of state-of-the-art open-source LLMs based on nanoGPT. Supports flash attention, 4-bit and 8-bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training. Apache 2.0-licensed.

1

whisperkittools. Python tools for WhisperKit: Model conversion, optimization and evaluation

1

swift-transformers. Swift Package to implement a transformers-like API in Swift

1
26
Apply