About this project

TensorFold is a local inference server that exposes language models through an OpenAI-compatible HTTP API. It targets two backends: Apple Silicon via MLX and NVIDIA GPUs via CUDA. Each supported model family ships its own kernels and draft-verification logic rather than relying on a single generic path. Installation is from the Git repository with pip, and the CLI provides `serve`, `pull`, `models`, `info` and `update` commands. A typical start is `tensorfold serve <checkpoint>`, after which clients point at `http://127.0.0.1:8080/v1` and use the model ID reported by `/v1/models`. Chat completions, completions and the Responses API are all served. Python 3.11+ is required, and MLX 0.32.2+ on macOS. The README lists a table of supported model families and checkpoints, including Nemotron 3.5 Lightning, Qwen3.8-27B, Qwen3.8 Flash Next, GLM-5.3-Flash, Gemma 4 26B-A4B, DeepSeek-V4-Flash and Ternary Bonsai 2 27B, with notes on which backend and drafting method each uses. Quantization support varies by family: Qwen3.8-27B reads MLX affine 2- through 8-bit checkpoints including mixed layer formats; Flash Next requires 4-bit/group-32 weights; Nemotron CUDA requires 4-bit/group-64 plus an MTP head; GLM reads 4-bit/group-64 and mixed-bit conversions. Some CUDA paths accept NVFP4 or EXL3 conversions, marked experimental. A central design claim is exact decoding: a speculative draft is accepted only when it equals the token the same engine would produce serially, with sampling tied to prompt or seed, absolute position and token ID. The README is explicit that exactness is relative to the same engine, weights, runtime and settings, and does not imply identical output across MLX and CUDA, different quantizations or different tensor-parallel rank counts. Users can compare a request with `"draft": false` to check drafted versus serial output. Serving options cover host/port, advertised model name, context size, reply limits, sampling defaults, thinking toggles and reasoning effort, backend selection, parallelism, draft control and, on MLX, prompt caching and memory tuning such as retained prefixes, disk spill, snapshot directories and a reusable buffer cache. CUDA-only options include KV cache dtype for Flash Next, an MTP confidence threshold and two-rank tensor-parallel execution. Memory handling is documented in detail: MLX defaults to a 70% process budget of RAM, overridable via an environment variable, with admission accounting for weights, cache growth, reply tokens and prefill workspace. A memory-class table lists qualification slots from 32 GB to 256 GB, but most cells are marked TBD; only a 64 GB M5 Pro row has published figures, attributed to a community run on an earlier version. The README states these are qualification slots, not minimum-memory promises. Prompt caching on MLX uses chunk plans derived from the rendered token sequence, with resume points at assistant-message starts and the second message start. Snapshots carry model, runtime, kernel and chunk-plan identity so prefixes can survive restarts. CUDA engines keep their own prompt and reply state and do not use the MLX disk-snapshot or retained-prefix options. For NVIDIA hardware the README recommends NVIDIA's PyTorch container, since the package has no CUDA installation extra. Qwen3.8-27B, Flash Next and Nemotron support one or two CUDA ranks, while GLM requires two. Rank 0 serves HTTP. Vision input is opt-in: installing the vision extra and starting a compatible Qwen3.5/3.8 dense checkpoint with `--vision` allows image and text content parts through the same engine. The project is MIT licensed, with model weights keeping their own licenses and some optional draft checkpoints carrying non-commercial terms noted in third-party notices.