About this project

Lemonade is a local AI server aimed at running AI models on a user's own hardware instead of cloud APIs. It is distributed in two forms: Lemonade Server, which installs a service exposing standard OpenAI, Anthropic and Ollama compatible APIs so existing apps can connect, and Embeddable Lemonade, a portable binary intended to be bundled into other applications to provide multi-modal local AI that adapts to the user's PC. The project supports text generation, speech-to-text, text-to-speech, audio generation, image generation, 3D generation and text classification. It works with GGUF, FLM and ONNX model formats and includes a model catalog plus a Model Manager for browsing and downloading models. Custom GGUF or ONNX models can be pulled from Hugging Face or ModelScope. A CLI provides commands such as `lemonade run` to launch and chat with a model, `lemonade launch claude` for coding workflows, `lemonade list` and `lemonade pull` for model management, `lemonade backends` to inspect available inference backends on the machine, and `lemonade alias` for environment-independent model naming and active-standby failover. Hybrid setups can also route requests to OpenAI-compatible cloud providers alongside local models, described as experimental. Inference is handled by multiple engines with different hardware requirements. For text generation these include llamacpp with system, Metal, CUDA, Vulkan, ROCm and CPU backends, plus experimental llamacpp-hrx, flm and ryzenai-llm for XDNA2 NPUs, and experimental vllm and ds4 for AMD Strix Halo. Speech-to-text uses whispercpp and moonshine; text-to-speech uses kokoro and experimental openmoss; audio generation uses experimental thinksound and acestep; image generation uses sd-cpp and experimental thenoise; 3D generation uses experimental trellis; and text classification uses experimental onnxruntime. Supported operating systems include Windows 11, several Linux distributions (Arch, Debian, Fedora, Ubuntu, Snap, Docker) and macOS, with packages such as .msi, .deb, .rpm, .pkg and container images. Integration is straightforward for OpenAI-compatible clients: point the client at the local server base URL. The README lists client libraries for Python, C++, Java, C#, Node.js, Go, Ruby, Rust and PHP, and shows a Python example. A marketplace lists compatible applications including Claude Code, Open WebUI, AnythingLLM, Dify, n8n, OpenHands, GitHub Copilot and others. Mobile apps for iOS and Android are also referenced. The project is community-built with optimizations by AMD engineers for Ryzen AI, Radeon and Strix Halo PCs, and is sponsored by AMD. It is licensed under Apache 2.0 and builds on tools such as llama.cpp, whisper.cpp, stable-diffusion.cpp, Kokoros, OnnxRuntime GenAI, Hugging Face Hub and ModelScope. The README states that the program does not transfer information to networked systems unless requested, and that model downloads or registry lookups may contact Hugging Face Hub or ModelScope depending on the selected source.