About this project

Atomic Chat is a local AI application and inference engine aimed at running open-weight large language models on your own machine, with an emphasis on privacy and offline operation. It ships as desktop builds for macOS, Windows and Linux, plus iOS and Android apps. Core capabilities described in the README: - Local model execution: run open-weight models from Hugging Face, including Llama, Gemma, Qwen, Mistral and Phi families. - OpenAI-compatible local API: a server at http://localhost:1337/v1 acts as a drop-in replacement for the OpenAI SDK, so existing agents, CLIs and IDE plugins can point at it by changing only the base URL. It binds to 127.0.0.1 by default and can be exposed on a LAN by setting host to 0.0.0.0. - Multiple inference backends: a llama.cpp fork with TurboQuant KV-cache options, upstream llama.cpp, and MLX-VLM for Apple Silicon vision-language models. The README states clients do not need to know which backend is active. - Decoding and performance options: Multi-Token Prediction speculative decoding, DFlash block-diffusion decoding, Flash Attention toggles, reasoning-context tracking, and automatic context-window expansion. - Cloud model support: built-in providers such as OpenAI, Anthropic, Mistral, Groq, MiniMax, Qwen and Moonshot, with bring-your-own-key and per-chat model switching. - Tools and integrations: one-click launching of external agents (for example Claude Code, Codex CLI, Cline, OpenCode, Goose, OpenHands, Copilot CLI, Kilo Code, Zed), an Artifacts preview panel for HTML/CSS/JS, MCP server connections, custom assistants with system prompts, and projects with a conversation tree view. The README also documents building from source with Node.js 20+, Yarn, Make and Rust, system requirements per platform, Linux AppImage usage, and troubleshooting via GitHub issues or Discord. Performance figures quoted in the README (such as throughput boosts or KV-cache size reductions) are the project's own claims and are not independently verified here.