Up to ×ばつ faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
-
Updated
Sep 1, 2026 - Python
Up to ×ばつ faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability
DeepSeek V4 Flash 284B on AMD Strix Halo (gfx1151) — up to 32 tok/s decode & ~250 tok/s prefill via ROCmFPX, DSpark & ROCm 7.2
Optimized two-node DGX Spark deployment recipe for DeepSeek V4 Flash with vLLM, DSpark, and NVFP4 KV cache
DeepSeek-V4-Flash-0731 on 8x RTX 3090 (SM86) with vLLM — verified FP8 serving, benchmarks, build guide, and reproducible release.
DeepSeek-V4-Flash + DSpark speculative decoding on a pair of NVIDIA DGX Sparks (vLLM TP=2 over RoCE) — tuned recipe, overlays that halve multi-turn TTFT, contamination-guarded benchmarks, ops runbook
MiniMax M3 NVFP4 on 4x DGX Spark with NVIDIA DSpark, native multi-node vLLM TP=4, reasoning, and tool calling
Code to train and evaluate DSpark draft models for speculative decoding using the speculators library and vllm
Lossless inference speedup benchmarking suite for local LLMs using DeepSeek DSpark speculative draft heads.
Source-only offline deployment, testing, and operations toolkit for DeepSeek-V4-Flash-0731 on 4x/8x NVIDIA A100 GPUs, pinned to a reviewed community vLLM R1 stack.
DeepSeek-V4-Flash-Vision-Exp on 2x Jetson AGX Thor (SM110): Thor-adapted vLLM port, verified + benchmarked. Native image input with DSpark speculative decoding, TP=2 dual-node.
DSpark-style speculative decoding for OCR VLMs on Apple Silicon (MLX). Faster local OCR for DeepSeek-OCR-2, GLM-OCR, Unlimited-OCR.
Code to train DSpark draft models for Llama-3.2-1B-Instruct using the speculators library and vllm
Verified IQ2_XXS plus native DSpark recipe for DeepSeek V4 Vision on one DGX Spark.
Universal Dual-Engine & Speculative Curation Configuration Kit for Antigravity, Cursor, Claude Code, Windsurf & Grok (FastMCP)
Agent-Level Speculative Orchestration & Formal Dual-Engine Code Generation (Rust, FastMCP, DeepSeek)
Verified two-node GB10 DeepSeek-V4-Flash-0731 DSpark Graph-8 deployment and benchmarks
To associate your repository with the dspark topic, visit your repo's landing page and select "manage topics."