CPU inference for Kimi K3, a 2.78T-parameter MoE LLM, in pure Rust. No GPU, no BLAS, no PyTorch. Streams the checkpoint from disk. Byte-identical port of kimi-k3-in-c.
-
Updated
Aug 14, 2026 - Rust
CPU inference for Kimi K3, a 2.78T-parameter MoE LLM, in pure Rust. No GPU, no BLAS, no PyTorch. Streams the checkpoint from disk. Byte-identical port of kimi-k3-in-c.
🎙️ AI-powered Telegram bot for voice-to-text transcription using OpenAI Whisper. CPU-only, no GPU required, privacy-focused with local processing.
Collama - Run Ollama Models on Google Colab
Pixel art without GPU. Any text-only LLM can draw. Size + prompt → self-contained HTML. Claude Code skill / MCP compatible.
A biomorphic neuromorphic inference engine inspired by cricket auditory neuroscience — performing real-time temporal pattern recognition via delay-line coincidence detection, without matrix multiplication.
DeepSeek V4 Flash prompt engineering + Flux free image generation. Zero cost, no GPU, China-friendly AI image workflow.
No-GPU 3D scanning from phone rotation videos — pure geometry, no neural networks, no cloud
A 3D software rendering engine built from scratch in Rust. Implements a custom graphics pipeline, linear algebra library, and scene graph without hardware acceleration APIs.
🎬 PhantomRec: The Ultimate Low-End Screen Recorder — Smooth 60 FPS on Weak PCs, Dual-Core CPUs, No GPU Required
Genesis 2 — Cascade MoE Neural Network | CPU-only AI: 10,800 experts, 100% accuracy, 18ms inference | No GPU | Self-hosted alternative to GPT/LLaMA | Affiliate program: 40% commission
low-ends games web
CPU-trained reasoning model pipeline. LoRA SFT + DPO on SmolLM2-360M, GSM8K math reasoning, single-laptop deployment.
A renderer that doesn't sample. It solves. Noise-free CPU global illumination — no Monte Carlo, no denoiser, no GPU.
CPU-native inference runtime. Local-propagation paradigm: the active region pays the cost, not the field. Bit-exact across architectures. Validated for streaming anomaly detection and audio VAD.
CPU-only LLM inference engine in pure Rust — 4-bit quantized models, hand-written AVX2 kernels, speculative decoding, and a local HTTP/MCP server. No GPU, no Python runtime.
To associate your repository with the no-gpu topic, visit your repo's landing page and select "manage topics."