About this project
LLM Systems Manager is a self-hosted operations platform for LLM infrastructure, covering monitoring, remote control, tuning, routing and alerting in one place. It integrates llama.cpp, vLLM, LM Studio, stable-diffusion.cpp and OpenClaw session telemetry, while its agent reports general host metrics for any Linux or macOS machine; Ollama integration is listed as roadmap.
Installation paths include an interactive script installer (the preferred route, handling prerequisites, InfluxDB, config, TLS and agents), native .deb/.rpm packages, Docker Compose for a containerized control plane, Homebrew formulas for macOS Apple Silicon and Linux, and a standalone agent binary tarball for hosts without Python. The installer deploys the latest GitHub Release with SHA-256 verification and enables systemd units without starting services automatically.
Headline capabilities described in the README:
- Model Autopilot: declared-state model placement gated on host memory (VRAM or RAM), failover when a host drops out, and replica scaling with demand. Off by default; it proposes and the operator approves.
- OpenAI-compatible inference gateway: one endpoint fronting llama.cpp, LM Studio and vLLM, with a merged /v1/models catalog and routing by per-model pin, pool round-robin, or failover.
- GPU Report Card: a standardized benchmark producing comparable cards with time-to-first-token, prefill and generation throughput, tokens/joule, measured $/Mtok and the GPU used.
- Energy and cost intelligence: measured $/Mtok from real power draw, monthly savings against hosted-API pricing, idle-power accounting, and a per-host performance manager that adjusts CPU governor and cooling profiles to load.
- Benchmarking and autotuning: library-wide benchmarks plus autotuners for llama.cpp context/slot configuration and vLLM max-model-len.
- Model management: Hugging Face downloads and pruning, plus named profiles (chat/code/general) for one-click reload.
- Remote control without SSH: run servers, hot-swap models, edit configs, update llama.cpp, tail logs, and open an in-browser terminal; a Discord bot exposes similar commands.
- LLM-aware telemetry and alerting: slots, tokens/sec, prompt processing, KV cache and context alongside GPU, PSU, UPS and cooling stats; a standalone alarm engine persists samples to InfluxDB, evaluates threshold and anomaly rules, notifies via email/toast/webhook/Discord, buffers through outages, and correlates related alerts into incidents.
Additional features include a Tools launcher (Report Card, Benchmark, Autotune with a fleet-wide run ledger), layout and appearance options (Grid or Flow engines, role presets, compact density, seven themes), an installable PWA phone companion with push alerts, multi-user roles with an admin audit log, encrypted scheduled backups, OpenClaw cost/budget analytics, an image generation tab, and TLS/mTLS on all connections with support for bringing your own certificate.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.