About this project
# Cactus
Cactus is a hybrid edge-cloud AI engine designed for mobile devices and wearables. It provides a complete stack for on-device AI inference, including quantization, kernels, a computation graph, and a runtime engine.
## Architecture
The project is structured into four main layers:
- **Cactus Engine**: Provides OpenAI-compatible APIs for text, speech, and vision tasks, including chat completion, streaming, tool calling, transcription, embeddings, RAG, and cloud handoff.
- **Cactus Graph**: A zero-copy computation graph for tensor operations like matrix multiplication, attention, normalization, and activation functions.
- **Cactus Kernels**: CPU/GPU kernels optimized for Apple, Samsung, Pixel, and other mobile platforms, using ARM NEON SIMD.
- **Cactus Quants**: A custom rotation-based quantization technique supporting 4-bit to 1-bit precision.
## Quick Start (Mac)
Install via Homebrew and run:
```bash
brew install cactus-compute/cactus/cactus
cactus run
```
## Key Features
- **Hybrid Edge-Cloud**: Automatically routes hard queries to the cloud based on local model confidence.
- **Model Support**: Works with Liquid, Gemma, Whisper, Parakeet, and Qwen model families. Any HuggingFace model can be converted experimentally.
- **Multiple Bindings**: Swift, Kotlin, Flutter, React Native, Python, and Rust.
- **On-Device Tool Calling**: Includes a 26M parameter model called "Needle" for on-device tool calling.
- **CLI Tools**: Commands for running models, transcription, serving, conversion, benchmarking, and testing.
## Performance
Benchmarks show strong performance on Apple silicon and mobile devices. For example, on a Mac M5 Max, the engine achieves 2964 tokens/sec prefill and 154 tokens/sec decode for a 1k-context LLM benchmark. On iPhone 17 Pro, it achieves 729 tokens/sec prefill and 37 tokens/sec decode.
## Output Quality
Quantization quality is evaluated across multiple benchmarks (ARC, HellaSwag, WinoGrande, MMLU, GPQA, GSM8K, HumanEval, BFCL). The CQ4 quantization maintains quality close to the original FP16 model on most tasks.
## Usage
The CLI provides commands for:
- `cactus run`: Run a model with options for quantization bits, backend, image/audio input, tools, and more.
- `cactus transcribe`: Live microphone transcription or file transcription.
- `cactus download`: Download prebuilt model bundles.
- `cactus convert`: Convert HuggingFace models to Cactus CQ weights.
- `cactus serve`: Start an OpenAI-compatible local HTTP server.
- `cactus code`: Run an AI coding agent.
- `cactus benchmark`: Run benchmark suites.
- `cactus test`: Run test suites.
## Documentation
Detailed documentation is available for each component:
- Cactus Engine (C API)
- Cactus Graph (C++)
- Cactus Kernels (C++)
- Cactus Quants (C++)
- Cactus Hybrid (C/Python)
- Python Package
## Citation
If you use Cactus in research, please cite:
```bibtex
@software{cactus,
title = {Cactus: AI Inference Engine for Phones & Wearables},
author = {Ndubuaku, Henry and Cactus Team},
url = {https://github.com/cactus-compute/cactus},
year = {2025}
}
```
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.