powerinferTiiny-AI
ADDEDPowerInfer is a local LLM inference engine that uses activation locality to keep frequently used “hot” neurons on the GPU while computing input-dependent “cold” neurons on the CPU. It supports hybrid or CPU-only inference, serving, batching, perplexity evaluation, INT4 quantization, and PowerInfer GGUF models.