This is your work, valued

Mia's AI Lab

Developing
@MiaAI-Lab

Local AI, LLMs, tech thinker & builder

DeepSeek-v4-Flash-DSpark-2x-DGX-Spark. Python

174

sparkDash. sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard

88

DeepSeek-V4-Flash-Dual-DGX-Spark-1M-Context. Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache

84

Laguna-S-2.1-DGX-Spark-RTX-6000-PRO. vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO

65

Qwen3.6-27B-NVFP4-vLLM. Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference

56

GLM-5.2-NVFP4-AQLM-Triple-DGX-Sparks. GLM-5.2 NVFP4+AQLM on 3× DGX Spark — 380k context MTP serve stack

42

Qwen3.6-35B-A3B-NVFP4-vLLM. Self-hosted vLLM inference for Qwen3.6-35B-A3B-NVFP4

31

Best-Local-Model_Agentic-Workflows_2026. Head-to-head comparison of local LLMs for agentic workflows (Hermes Agent, tool-eval-bench)

31

HermesGW-Desktop-setup. Setup wizard for Hermes Gateway Desktop

30

Unsloth-Qwen3.6-35b-NVFP4-DGX-Spark. vLLM deployment for Unsloth Qwen3.6-35B-A3B-NVFP4-Fast on NVIDIA DGX Spark

28

Qwen3.6-35B-A3B-UD-Q8_K_XL_DGX-Spark-Recipe. llama-server start/stop scripts for Qwen3.6-35B-A3B UD-Q8_K_XL GGUF on DGX Spark

26

DeepSeek-v4-Flash-DSpark-Abliterated-Uncensored. DeepSeek V4 Flash DSpark Abliterated (Uncensored) — 2× DGX Spark serving

21

DGX_Spark_Qwen_3.6_27b_35b_GGUF_start_script. Shell

16

MiMo-V2.5-vLLM-Dual-DGX-Sparks. Mia's MiMo-V2.5 Dual DGX Spark start script — Omni MTP1 + NVFP4-KV TP=2 on 2× GB10

12

Ternary-Bonsai-27B-tool-eval-bench-results. tool-eval-bench results for Prism ML Ternary-Bonsai-27B (Q2_0) — 8 trials, score 85, deployability 80

12

Gemma-4-26B-A4B-DGX-Spark-18-concurrencies. Gemma 4 26B vLLM launch scripts for DGX Spark concurrency testing

11

DeepSeek-v4-Flash-vs-Step-3.7-Flash-Tool-Call-Benchmark. Head-to-head comparison of DeepSeek-V4-Flash vs Step-3.7-Flash on tool-eval-bench v2.0.6 (69 scenarios). Full results, summary, and analysis.

7

Unsloth-Qwen3.6-27B-UD-Q8_K_XL_vs_nvidia-Qwen3.6-27B-NVFP4_tools_eval. tool-eval-bench: Qwen3.6-27B GGUF Q8_K_XL vs NVIDIA NVFP4 — head-to-head tool-calling quality comparison

7

Dual-DGX-Spark-Step-3.7-Flash-NVFP4. Shell

6

dgx-spark-dashboard. Real-time system monitoring dashboard for single and multi-unit NVIDIA DGX Spark (GB10) setups. Track GPU metrics, CPU, memory, storage, network, and more via a sleek dark-themed web UI.

6

Qwen3.6-35b-vs-DSF4-tool-eval-bench. Tool evaluation benchmarks comparing Qwen3.6-35B vs DeepSeek V4 Flash DSpark

6

Gemma-4-31B-IT-NVFP4-Recipe. Serve nvidia/Gemma-4-31B-IT-NVFP4 via vLLM with MTP speculative decoding, tool calling, and thinking/reasoning support

6

Nemotron-Labs-3-Puzzle-75B-DGX-Spark. Serve NVIDIA Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 on a single DGX Spark (GB10) node with vLLM 0.24 in Docker

6

Hy3-Dual-DGX-Spark. Hy3-295B NVFP4 serving on 2x DGX Spark (TP=2, Ray, vLLM)

6

Leanstral-1.5-119B-A6B_Review. Independent evaluation of Mistral Leanstral-1.5 for Lean 4 formal proof engineering

5

Slate. A fast, light-weight OLED-friendly MD/text editor.

2

Qwopus-3.6-27b_Tests. HTML

2

Orinth-1.0-35b-tool-eval-bench-hardmode. HTML

2

Laguna-XS-3.1-Q8_0_vs_Qwen3.6-35B-A3B-UD-Q8_K_XL. Head-to-head tool-eval-bench comparison: Laguna-XS-2.1-Q8_0 vs Qwen3.6-35B-A3B-UD-Q8_K_XL

2

Agents-A1_vs_Qwen3.6-35B_tools_eval. Tool-eval-bench comparison: Agents-A1-Q8_0.gguf vs Qwen3.6-35B-A3B-UD-Q8_K_XL

2

GLM5.2_vs_DS4F_vs_MM3_Beautiful_Page. HTML

1

Leanstral-1.5-119B-A6B_2x_DGX_Sparks_vLLM. HTML

1