GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
-
Updated
Sep 7, 2026 - Python
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights
A tiered-memory system design for workloads that don't fit in RAM: measure the working set, pin the hot tier, stream the cold tier from flash. Ships the residency calculator, measurement harnesses, and the build recipes behind it. Predictions validated against public benchmarks.
High-performance runtime extensions for vLLM.
LLM inference server for ExLlamaV3 / EXL3, with OpenAI- and Anthropic-compatible APIs optimized for Agent workloads.
DeepSeek-V4-Flash-Vision-Exp (EXL3 MixedK, 256 experts, uncensored) on one NVIDIA DGX Spark with vLLM + sparkinfer: 245,760 context, vision + DSpark speculative decoding, CUDA graphs. Recipe, overlay patches, benchmarks, receipts.
Containerized private AI lab: TabbyAPI EXL3 + SillyTavern + Open WebUI + Ollama + SearXNG
GLM-5.2 753B MoE at EXL3-TR3 3.0bpw on 4x RTX PRO 6000 (SM120): digest-pinned serving stack, validation suite, results. Upstream: vllm#139 / sparkinfer#49
GLM-5.3-Flash EXL3 4bpw on 2x NVIDIA DGX Spark (GB10): production recipe, boot ladder, quality gate, benchmarks and lessons learned — reproducible from CLAUDE.md/AGENTS.md
Run GLM-5.3-Flash, a 320B-parameter model, on two NVIDIA DGX Spark desktops as a private OpenAI-compatible API with 1.3M-token context.
Deploy GLM-5.3 Flash with DFlash2 speculative decoding on dual RTX PRO 6000 Blackwell 96GB GPUs, enabling one-million-token context and 16-image prompts via an OpenAI-compatible API.
To associate your repository with the exl3 topic, visit your repo's landing page and select "manage topics."