Kai Shi Albresky
-
Beijing University of Posts and Telecommunications
- Beijing, China
-
18:30
(UTC +08:00) - https://www.albresky.cn
Highlights
- Pro
Stars
🚀 Efficient implementations for emerging model architectures
A library for efficient similarity search and clustering of dense vectors.
High-performance LLM operator library built on TileLang.
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
[ICML 2024] Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
desktop app to browse and analyze your Claude Code conversation history
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
[EMNLP 2025 Demo] PDF scientific paper translation with preserved formats - 基于 AI 完整保留排版的 PDF 文档全文双语翻译,支持 Google/DeepL/Ollama/OpenAI 等服务,提供 CLI/GUI/MCP/Docker/Zotero
LLM Inference analyzer for different hardware platforms
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
Run LLMs on AMD RyzenTM AI NPUs in minutes; purpose-built and deeply optimized for the AMD NPUs.
A framework for few-shot evaluation of language models.
VS Code extension for debugging Python and C++ together. Support CppVSgdb, GDB, CodeLLDB.
A Fusion Code Generator for NVIDIA GPUs (commonly known as "nvFuser")
TorchBench is a collection of open source benchmarks used to evaluate PyTorch performance.
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (V...
A debugging and profiling tool that can trace and visualize python code execution
A high-throughput and memory-efficient inference and serving engine for LLMs
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor...
NVIDIA® TensorRTTM is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.