This is your work, valued

AEON-7

@AEON-7

I build things. Some of them hum. Neural nets since '13, security always, and a soft spot for problems the docs warned us about.

Qwen3.6-27B-AEON-Ultimate-Uncensored-DFlash. Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.

426

Qwen3.6-35B-A3B-heretic-NVFP4-DFlash. Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehensive deployment guide.

131

vllm-ultimate-dgx-spark. AEON vLLM Ultimate — vLLM 0.25.0 built from source for DGX Spark / Blackwell (sm_121a/GB10). One image serves the whole AEON fleet (Gemma-4-26B-A4B, Qwen3.6-27B, Qwen3.6-35B-A3B) with DFlash spec-decode on a pinned V1 runner, Triton NVFP4-KV, FP8 KV, NVFP4 swizzled-scale decode fix, FlashInfer 0.6.13, TP=2-ready.

100

comfyui-aeon-spark. Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs + Flux 2 Dev / LTX 2.3 22B / ACE-Step v1.5 XL Turbo pre-bundled with abliterated text-encoder paths.

71

Ornith-1.0-35B-AEON-Ultimate-Uncensored. Uncensored/abliterated Ornith-1.0-35B (AEON Ultimate): 0% refusal, 0 coding-capability loss. BF16 + FP8 for vLLM.

68

vllm-dflash. DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding

53

Gemma-4-26B-A4B-it-Uncensored-NVFP4. NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding), up to 158 single (Extraction); beats prior v2 by 20-50%.

42

AEON-7. Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and the AEON Media Production toolchain · ☕ Tips welcome

37

Qwen3.6-27B-AEON-Ultimate-Uncensored-MLX. Native MLX quants of Qwen3.6-27B-AEON-Ultimate-Uncensored for Apple Silicon (Metal): 8-bit, FP4, and MTP self-speculation. Preserves the Gated-DeltaNet SSM, vision tower, and MTP head in BF16.

36

Gemma-4-31B-Uncensored-NVFP4-DFlash. DGX Spark / GB10 vLLM image for Gemma 4 31B Deckard Heretic Uncensored NVFP4 with z-lab DFlash speculative decoding.

33

Aeon-Bench-Pod. Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed25519-signed attested submit.

21

matrix-voip-agent. Headless Matrix WebRTC voice AND video agent — auto-answers calls, bridges audio to any AI agent via PipeWire, optional camera-frame vision for multimodal LLMs

17

unifi-ai-network-management. Python

15

aeon-magick-ai-computer-control. Universal AI computer-control KVM (REST API + MCP) that drives any machine as plain USB hardware — nothing installed on the target. Built-in Tor/I2P/DNSCrypt/VPN stack to help keep connection and correspondence private.

14

aeon-movie-maker. Fast cinematic video via LTX 2.3 22B. Single clips, full screenplays with character continuity, and sidechain-mixed final cuts.

13

supergemma4-26b-abliterated-multimodal-nvfp4. NVFP4 AWQ Full quantization of SuperGemma4-26B-Abliterated-Multimodal for Blackwell GPUs — pre-built vLLM container + patches included

9

Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored. Nemotron-3-Nano-Omni AEON Ultimate Uncensored: 12-D abliterated multimodal reasoning model (BF16 + NVFP4) for DGX Spark / Blackwell. Source-built vLLM v0.20.0 image + 4 patches + bench + DGX Spark deployment guide.

8

aeon-music-maker. ACE Step 1.5 XL music generation with dynamics-preserving mastering. Part of the AEON Media Production family.

8

Gemma-4-31B-DECKARD-HERETIC-Uncensored-NVFP4. NVFP4-quantized Gemma 4 31B DECKARD HERETIC (dense, uncensored, thinking) for DGX Spark. AWQ_FULL + SVDQuant variants.

7

qwen3-asr-server. OpenAI-compatible /v1/audio/transcriptions server for Qwen3-ASR-0.6B on DGX Spark — vLLM-native, flash-attn 2 (sm_120), RTF 16x real-time

7

aeon-radio-drama. Full-pipeline radio drama / audiobook production. Dialogue (Qwen3-TTS) + music (ACE) + SFX (MMAudio/SAO/ACE) + sidechain mix in one command.

5

Qwen3.6-27B-AEON-Ultimate-Uncensored-DDTree. Experimental DDTree-on-vLLM research track for Qwen3.6 AEON Ultimate on DGX Spark / GB10.

5

create-agentic-personas. Shell

5

qwen3-tts-server. OpenAI-compatible /v1/audio/speech server for Qwen3-TTS-12Hz-1.7B (VoiceDesign + VoiceClone) on the faster-qwen3-tts CUDA-graph streaming engine — DGX Spark, streams PCM while generating, first audio ~0.4 s, ~1.7x realtime

4

aeon-music-video. Audio-reactive music video builder. librosa-driven beat / onset / RMS detection drives ffmpeg filter chains for synced visual effects.

4

gemma4-aeon-abliterated-mlx-toolkit. Apple Silicon MLX toolkit + OpenAI server for the Gemma-4-12B AEON Abliterated MLX quant grid (MLXFP4 / MLX-8bit)

2

vllm-ultimate-deepseek-v4-gb10. Experimental DeepSeek-V4 / DSpark GB10 vLLM TP2 image for DGX Spark

2

regex-builder. A simple and elegant RegEg builer

1

cosmic-mind. A security and resiliency focused deployment of the Quartz web app. A place to build your second mind and share it with the world.

1

Gemma-4-E4B-DECKARD-HERETIC-Uncensored-NVFP4. EAGLE E4B speculative decoding drafter for Gemma 4 31B DECKARD HERETIC Uncensored NVFP4 — optimized for NVIDIA DGX Spark

1

Gemma-4-E4B-it-Uncensored-NVFP4. EAGLE E4B speculative decoding drafter for Gemma 4 26B MoE (TrevorJS Uncensored) — NVFP4 AWQ, optimized for NVIDIA DGX Spark

1

modelopt-fast-moe. Python

1

presto-cockpit. Python

1