Linux host observability toolkit for AI/GPU infrastructure, exposing Prometheus metrics for memory pressure, RDMA/NIC health, PCIe/VFIO, NUMA, GPUs, and kernel events.
linux performance-engineering kernel gpu prometheus nvidia rdma infiniband sre numa observability pcie node-exporter vfio gpu-monitoring linux-monitoring ai-ops ai-infrastructure mlx5 rdma-monitoring
-
Updated
Jun 10, 2026 - Shell