GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
-
Updated
Sep 8, 2026 - Python
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.
NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.
Agent-assisted and full-agent reproducibility package for MLSys 2026 FlashInfer AI Kernel Generation Contest submissions: kernels, agent workflows, skills, configs, writeup, benchmark artifacts, and full optimization records.
Reproducible SGLang recipe + public prebuilt image (ghcr.io) for DeepSeek-V4-Flash-0731 on 4x RTX PRO 6000 Blackwell (SM120): TP4/DP4/EP4, 1M ctx, benchmarks, and the DSPARK draft-depth corruption boundary
180-226 tok/s single-stream decode for Qwen3.8-Flash-Next NVFP4 on one RTX PRO 6000 Blackwell, at full 262K context with unchanged quantization. Config, benchmark harness, and the FlashInfer autotune correctness bug that silently corrupts output.
Production runbook for Qwen3.5-122B hybrid INT4+FP8 on NVIDIA DGX Spark GB10 — optimization stack, PD firmware wedge diagnosis, bench results
vLLM + FlashInfer source integration for GLM-5.3-Flash on SM120 / RTX PRO 6000 Blackwell
🚀 Accelerate attention mechanisms with FlashMLA, featuring optimized kernels for DeepSeek models, enhancing performance through sparse and dense attention.
🚀 从零手写的 Qwen3 高性能推理引擎 —— 1,500 行纯 Python 实现 Paged KV Cache · 连续批处理 · FlashAttention-2 · CUDA Graph,零框架依赖(不基于 vLLM/SGLang),单卡 256 并发 1,200+ tok/s
Run GLM-5.3 Flash NVFP4 on 2x NVIDIA DGX Spark with vLLM TP2, FlashInfer sparse MLA, FP8 KV and MTP3
DeepSeek-V4-Flash-DSpark on ×ばつ DGX/ASUS Spark (GB10, SM120) using the stock spark-vllm container - recipes, a faster SM120 topk fix, and one-command run scripts.
Production-grade fine-tuning & LoRA toolkit for Chatterbox-Flash zero-shot TTS models. Combines parallel block diffusion and FlashInfer acceleration with smart placeholder vocabulary extension supporting languages. Features offline feature preprocessing, Silero VAD silence trimming, zero-padding leakage prevention, and fast voice cloning for custom
Single-GPU LLM decode research prototype: paged KV cache, Triton attention, CUDA append, scheduling, shared prefixes, and multi-layer transactions.
Community wishlist for reproducible kernel definitions, workloads, and validated optimized solutions.
⚡ Optimize attention mechanisms with FlashMLA, a library of advanced sparse and dense kernels for DeepSeek models, improving performance and efficiency.
Reproducible TP=2 deployment, container build, and experiment report for DeepSeek V4 Flash on two DGX Sparks
Correctness-first microbenchmarks for LLM attention and sampling kernels.
To associate your repository with the flashinfer topic, visit your repo's landing page and select "manage topics."