Skip to content

Navigation Menu

Sign in
Sign up
@Eddie-Wang1120
Eddie-Wang1120
Follow

Eddie-Wang Eddie-Wang1120

🎯
Focusing
  • Peking University
  • Beijing
  • 20:10 (UTC +08:00)

Block or report Eddie-Wang1120

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Eddie-Wang1120 /README.md

👋 Hi, I'm Jinheng Wang

A passionate AI Infra researcher focused on high-performance computing and model-and-system co-design.

🚀 What I Build

Core developer of:

BitNet — official inference framework for 1-bit LLMs.

  • 6.25x faster than full-precision and 2.32x faster than low-bit baselines
  • lossless inference for BitNet b1.58 via the TL / I2_S ternary mpGEMM kernels
  • measured on edge CPUs: Intel i7-13700H, Apple M2 Ultra

BitNet inference timeline

TeraMoE — cross-node expert-parallel MoE training library.

  • 1.30x speedup over DeepEP + TE in communication-bound sparse-expert regimes, 1.24x under 3.0x expert load imbalance
  • 28% less activation memory than Megatron
  • one cooperative persistent kernel overlapping dispatch, expert compute and combine
  • measured on SM100 GPUs, EP16–64 over RDMA

TeraMoE SM role timeline

📄 Research

🔧 Also Contributing To

llama.cpp · Paddle · TensorRT-LLM

🏆 GitHub Trophies

trophy

Pinned Loading

  1. microsoft/BitNet microsoft/BitNet Public

    Official inference framework for 1-bit LLMs

    C++ 40.2k 3.7k

  2. PFCCLab/TeraMoE PFCCLab/TeraMoE Public

    TeraMoE: A cross-node expert-parallel MoE training library that uses a cooperative persistent kernel to overlap dispatch, expert compute, and combine.

    Cuda 12 2

  3. HPC-Learning-Notes HPC-Learning-Notes Public

    高性能计算相关知识学习笔记,包含学习笔记和相关知识的代码demo,在持续完善中。 如果有帮助的话请Star一下,对作者帮助很大,谢谢!

    Jupyter Notebook 476 40

  4. Professional-CUDA-C-Programming-Code-and-Notes Professional-CUDA-C-Programming-Code-and-Notes Public

    CUDA C 编程权威指南代码实现 包含了书上第二章到第八章的大部分代码实现和作者笔记,全由作者本人手动实现,难免有错误的地方,请大家谨慎参考,非常欢迎对错误的指正。 如果有帮助的话请Star一下,对作者帮助很大,谢谢!

    Cuda 388 29

AltStyle によって変換されたページ (->オリジナル) /