GLM-5.2-NVFP4-REAP-469B serving on SM120 (×ばつ RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
-
Updated
Jun 19, 2026 - Shell
GLM-5.2-NVFP4-REAP-469B serving on SM120 (×ばつ RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
An LLM server for a single RTX 5090, built for agent workloads: tool calls, long conversations, reasoning, and many requests at once. One of the fastest engines on this card, at batch 1 and at dozens of concurrent streams, with the numbers in the repo.
Reproducible SGLang recipe + public prebuilt image (ghcr.io) for DeepSeek-V4-Flash-0731 on 4x RTX PRO 6000 Blackwell (SM120): TP4/DP4/EP4, 1M ctx, benchmarks, and the DSPARK draft-depth corruption boundary
Rust + CUDA LLM inference engine for Blackwell (Tuned specifically on RTX PRO 6000, RTX 5090, B200): OpenAI-compatible (+converse and ant) serving, per-model X hardware exactness gates. NVFP4/mixed (fp8 hybrid, 4o6, etc - correctness, performance, hardware specific adapted) main quant support.
180-226 tok/s single-stream decode for Qwen3.8-Flash-Next NVFP4 on one RTX PRO 6000 Blackwell, at full 262K context with unchanged quantization. Config, benchmark harness, and the FlashInfer autotune correctness bug that silently corrupts output.
Systematic 24-hour benchmark study of Qwen3.6-27B inference on dual NVIDIA RTX PRO 6000 Blackwell SM120 (TP=2). 8 experiments comparing repne/vllm fork vs upstream vLLM across FP8/BF16/NVFP4/Q8_0 quants and MTP/DFlash speculative decoding. Peak: 2,083 tok/s at c=32. Quality: KLD vs BF16 = 0.0018 (noise floor).
Fish Audio OpenAudio S2-Pro on vLLM-Omni. low-latency ~100ms TTFA, OpenAI-compatible, runs on NVIDIA Blackwell (RTX 5090 / RTX PRO 6000). Self-hosted streaming TTS & voice cloning.
MiniMax-M3 (428B MoE) running on ×ばつ RTX PRO 6000 Blackwell at TP=3 with 240K context, FP8 KV cache, and working multimodal vision input. Includes dist_utils.py patch for non-divisible attention heads.
Image-to-3D-Video-Asset-Generator is an all-in-one generative 3D pipeline that transitions smoothly from textual concepts or reference images into fully realized 3D mesh assets (.glb), dynamic camera movements in 5-second MP4 videos, and clean bundle exports (.zip).
Production-grade FlashAttention FP8 e4m3 forward kernel for NVIDIA Blackwell consumer GPUs (sm_120a, e.g. RTX PRO 6000). 647–652 TFLOPS at hd=128, sl=8192. Multi-kernel dispatcher, C library with Go and Python bindings
Hub for ongoing Qwen inference benchmarks on NVIDIA Blackwell. Indexes all studies, hosts the rolling SOTA leaderboard, points to the toolchain.
Prolepsis is a speculative decoding implementation for Qwen3 draft-target models with Hugging Face and vLLM backends. On an RTX PRO 6000 at batch size 1, it measured 1.72x throughput with vLLM FP8 and 1.32x with Hugging Face BF16, with complete latency and response artifacts.
GLM-5.2-504B NVFP4 at 250K context on 4x RTX PRO 6000 Blackwell (sm_120) using STOCK vLLM — no fork, no Docker, no CUDA 13.2. One ~126-line patch. Documents the 3 upstream bugs that block you, with exact error strings and fixes.
Deploy the GLM-5.2-469B model on four RTX PRO 6000 Blackwell GPUs using a turnkey vLLM Docker configuration to enable high-speed sparse attention and inference.
Prebuilt spconv v2.3.8 wheels for CUDA 12.8 / 13.0 with native Blackwell (RTX 50-series, sm_120) kernels — for default PyPI torch (cu130) or torch +cu128
To associate your repository with the rtx-pro-6000 topic, visit your repo's landing page and select "manage topics."