fsdp
Here are 79 public repositories matching this topic...
Language: All
Sort: Most stars
Best practices & guides on how to write distributed pytorch training code
-
Updated
Oct 22, 2025 - Python
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
-
Updated
Jun 5, 2026 - Python
Repo for Qwen Image Finetune
-
Updated
Aug 22, 2026 - Jupyter Notebook
Open-source performance diagnostics for PyTorch training runs.
-
Updated
Sep 7, 2026 - Python
Write once, run anywhere; ezpz 🍋
-
Updated
Aug 31, 2026 - Python
A comprehensive hands-on guide to building production-grade distributed applications with Ray - from distributed training and multimodal data processing to inference and reinforcement learning.
-
Updated
Feb 12, 2026 - Jupyter Notebook
Research platform for model training, evaluation, and experimentation across architectures, benchmarks, and recipes.
-
Updated
Sep 3, 2026 - Python
META LLAMA3 GENAI Real World UseCases End To End Implementation Guide
-
Updated
Sep 24, 2024 - Jupyter Notebook
Fast and easy distributed model training examples.
-
Updated
Nov 26, 2024 - Python
Interactive exercises for understanding compute, communication, and dependencies in distributed LLM training traces.
-
Updated
Aug 26, 2026 - JavaScript
Forge kernels — fused Triton kernels for faster, leaner LLM fine-tuning. One-call patching into Hugging Face models via forge.patch(model), with FSDP2 multi-GPU support and a growing set of kernels and architectures. Every speedup backed by a committed benchmark. Apache-2.0.
-
Updated
Aug 15, 2026 - Python
A script for training the ConvNextV2 on CIFAR10 dataset using the FSDP technique for a distributed training scheme.
-
Updated
Dec 11, 2023 - Python
Simple and efficient implementation of 671B DeepSeek V3 that trainable with FSDP+EP and minimal requirement of 256x A100/H100, targeted for HuggingFace ecosystem
-
Updated
Jan 15, 2026 - Python
Minimal yet high performant code for pretraining llms. Attempts to implement some SOTA features. Implements training through: Deepspeed, Megatron-LM, and FSDP. WIP
-
Updated
Feb 6, 2024 - Python
Add this topic to your repo
To associate your repository with the fsdp topic, visit your repo's landing page and select "manage topics."