About this project

Jano is a minimal, OpenAI-compatible HTTP router designed for users running multiple local LLMs on a single GPU. It reduces overhead by implementing greedy batching: instead of swapping models on every request change, it serves all pending requests for the currently loaded model first, only swapping when necessary. This can reduce model swaps from multiple to just one per burst, saving significant time (30 seconds to 3 minutes) on local hardware. Jano supports both local backends (like llama-server or vLLM) and remote APIs (e.g., DeepSeek), with automatic model name rewriting for remote providers. It includes built-in telemetry via /health, /status, and Prometheus metrics endpoints, tracking queue depth, swap duration, token throughput, and backend health. Configuration is handled through environment variables and a models.json file, where users define model names, URLs, aliases, and optional swap scripts. The swap script contract is simple: given a model name, it must activate that backend and exit cleanly; optional status reporting helps detect initial state. Jano serializes all requests per backend to keep logic simple, making it ideal for personal or small-scale setups rather than multi-tenant services. It does not manage GPU temperature or system resources — those remain outside its scope. Designed as a thin layer atop existing tools, Jano offers explicit control over model switching, compatibility with any OpenAI-style backend, and transparent passthrough of streaming and tool calls. Not needed if you’re using Ollama or a single model; best suited for users who already have custom backends and want fine-grained control over routing and swapping.