About this project
AirLLM enables inference and fine-tuning of billion-parameter language models on consumer-grade GPUs with as little as 2 GB of VRAM. It works by keeping only a single transformer layer on the GPU at any moment — weights stream from disk layer by layer — so GPU memory requirements depend on layer size, not total model size.
Key capabilities:
- Inference for models ranging from ~8B to 2.8T parameters on a single low-end GPU
- Optional block-wise quantization (4-bit or 8-bit) for up to 3× faster inference
- LoRA fine-tuning of huge models on small VRAM; frozen base weights stream from disk while adapters stay resident
- Support for MoE, Flash, FP8, and native vision (VL) architectures
- Cross-platform: Linux/CUDA and macOS (Apple Silicon via MLX)
Supported model families include Llama (2/3/3.1/3.3/4), Qwen (1/2/2.5/3/3.5/3.8, dense and MoE), DeepSeek (V2/V3/R1), Mistral & Mixtral, Phi, Gemma, ChatGLM, Baichuan, InternLM, Yi, and Kimi K3.
Usage example:
```python
from airllm import AutoModel
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")
input_tokens = model.tokenizer(text, return_tensors="pt", truncation=True, max_length=128)
generation_output = model.generate(input_tokens['input_ids'].cuda(), max_new_tokens=20)
```
With optional compression:
```python
model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct", compression='4bit')
```
Training example (LoRA):
```bash
python air_llm/examples/train_qwen38_flash_next_lora.py --data my_data.jsonl --seq-len 512 --save-adapter adapter.pt
```
Configuration options: compression (none/4bit/8bit), profiling_mode, layer_shards_saving_path, hf_token, prefetching (default on), delete_original. The original model is split and cached layer-wise on first run, so sufficient disk space in the Hugging Face cache directory is required.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.