Stars
Measuring and evolving with the frontier of agent work
Qwen-AgentWorld: Language World Models for General Agents
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo...
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
A minimal yet professional single agent demo project that showcases the core execution pipeline and production-grade features of agents.
Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
"RAG-Anything: All-in-One RAG Framework"
一个可复现的DPO微型项目:基于Qwen2-0.5B-Instruct模型与ultrafeedback_binarized,包含TensorBoard训练曲线及评估/调试脚本。A reproducible DPO (TRL) mini-project: Qwen2-0.5B-Instruct + ultrafeedback_binarized, TensorBoard curves, and...
GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Official Project Page for Deep Delta Learning (https://arxiv.org/abs/2601.00417)
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Ongoing research training transformer models at scale
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
An Open-source RL System from ByteDance Seed and Tsinghua AIR
Repository hosting code for "Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations" (https://arxiv.org/abs/2402.17152).
A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
An open-source AI agent that brings the power of Gemini directly into your terminal.
基于双流检索(Graph + Vector)的智能代码助手,专为复杂代码库(如 verl)设计。本Agentic GraphRAG系统,结合了Neo4j (知识图谱)和Qdrant (向量数据库),不仅能读懂代码的语义,还能理清代码的结构。快来试试吧!
A LoRA+DPO finetuned Role-Play LLM based on Qwen2.5
Fast and memory-efficient exact attention