Skip to content

Navigation Menu

Sign in
Sign up
@tonyd2wild
tonyd2wild
Follow
Creator & Builder. 1M+ on YouTube. Building AI tools that give agents real memory. Founder of 2Wild Agency.

Sponsoring

Block or report tonyd2wild

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. DeepSeek-v4-Flash-Vision-Exp-DSpark-1M-NVFP4-KV-2x-DGX-Spark DeepSeek-v4-Flash-Vision-Exp-DSpark-1M-NVFP4-KV-2x-DGX-Spark Public

    DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark

    Python 475 59

  2. GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark Public

    GLM-5.3-Flash (NVFP4) on 2x NVIDIA DGX Spark - vLLM TP2, 262K context, MTP. World-first deploy recipe: 7 day-0 bugs found and fixed, patched sm121 image, probes and full report.

    Python 158 19

  3. Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX Public

    Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — Docker image, launch scripts, RDMA/NCCL setup, and the gotchas.

    Shell 81 12

  4. GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s Public

    Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster

    Python 77 9

  5. Qwen3.8-Flash-Next-NVFP4-DGX-Spark Qwen3.8-Flash-Next-NVFP4-DGX-Spark Public

    Qwen3.8-Flash-Next (NVFP4) on DGX Spark in vLLM: one Spark 43.9 tok/s with our disk-backed n-gram table patch, staged gather and reduced-vocab MTP draft; TP2 SPEED 53.7; TP4 CONTEXT 9M KV pool. Lau...

    Python 76 10

  6. MiniMax-M3-2x-DGX-Spark-36-tok-s MiniMax-M3-2x-DGX-Spark-36-tok-s Public

    MiniMax-M3 (428B, no pruning) at 36 tok/s on ×ばつ NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three serving lanes: speed / balanced / long-context.

    46 4

AltStyle によって変換されたページ (->オリジナル) /