SHA-1, SHA-256 and SHA-512 compression functions using Intel, ARMv8 and Power8 SHA intrinsics
-
Updated
Jan 8, 2026 - C
SHA-1, SHA-256 and SHA-512 compression functions using Intel, ARMv8 and Power8 SHA intrinsics
LLM infrastructure cost reduction via NUMA-aware weight banking: 147 t/s (8.8x stock llama.cpp) on refurbished enterprise POWER8. Self-hosted inference, no cloud APIs. Part of the Proof of Physical AI stack.
llama.cpp optimizations for IBM POWER8: vec_perm non-bijunctive collapse, PSE hardware entropy, DCBT prefetch. Sovereign inference. Part of the Proof of Physical AI stack.
Native ppc64le modules for llama.cpp webui (lightningcss, tailwindcss-oxide) - Built on IBM POWER8
AES encryption function using Intel, ARMv8 and Power8 intrinsics
Non-bijunctive attention collapse for LLM inference — POWER8 hardware AES (vcipher) + AltiVec vec_perm. Hebbian path selection, cross-head diffusion, O(1) KV prefiltering.
RPI (Resonant Permutation Inference) — Zero-multiply text generation. 18K tok/s. 868 KB models. Standalone or as speculative draft engine for LLMs. Runs on N64, POWER8, x86, ARM.
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Radical LLM cost reduction: no cloud API, no GPU farm. Tiny engine, immense model. 🐦
Personal blog - debene.dev | Hugo + PaperMod | Self-hosted on K8s via Cloudflare Tunnel
Accelerate LLM inference by collapsing attention paths with hardware-optimized selective pruning using POWER8 vector instructions and crypto operators.
To associate your repository with the power8 topic, visit your repo's landing page and select "manage topics."