A high-performance request router for vLLM — smart load balancing, response caching, prefill/decode disaggregation, semantic routing, Anthropic/OpenAI API translation, and operational tooling out of the box.
-
Updated
Aug 1, 2026 - Rust
A high-performance request router for vLLM — smart load balancing, response caching, prefill/decode disaggregation, semantic routing, Anthropic/OpenAI API translation, and operational tooling out of the box.
Testing Function Calling while running llama3.1 locally using ollama
Fuzzy Logic Toolbox with GUI and Files Analysis.
Local-first control platform for running, managing, chatting with, monitoring, and benchmarking GGUF language models on your own hardware.
Python codes generation from latex expressions. Using synthetic dataset and CodeT5-base model.
Classification of sperm heads based on its morphological quality via Mask-RCNN
An LLM inferencing benchmark tool focusing on device-specific latency and memory usage
To associate your repository with the inferencing topic, visit your repo's landing page and select "manage topics."