这个项目能做什么
AirLLM 可在显存低至 2 GB 的消费级 GPU 上对数十亿参数的语言模型进行推理和微调。其原理是任意时刻仅在 GPU 上保留单个 transformer 层——权重从磁盘逐层流式加载——因此 GPU 显存需求取决于单层大小,而非模型总大小。
主要能力:
- 在单块低端 GPU 上对约 8B 至 2.8T 参数的模型进行推理
- 可选的块级量化(4-bit 或 8-bit),推理速度最高提升 3 倍
- 在小显存上对超大模型进行 LoRA 微调;冻结的基础权重从磁盘流式加载,适配器常驻显存
- 支持 MoE、Flash、FP8 以及原生视觉(VL)架构
- 跨平台:Linux/CUDA 与 macOS(通过 MLX 支持 Apple Silicon)
支持的模型系列包括 Llama(2/3/3.1/3.3/4)、Qwen(1/2/2.5/3/3.5/3.8,稠密与 MoE)、DeepSeek(V2/V3/R1)、Mistral 与 Mixtral、Phi、Gemma、ChatGLM、Baichuan、InternLM、Yi 以及 Kimi K3。
使用示例:
```python
from airllm import AutoModel
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")
input_tokens = model.tokenizer(text, return_tensors="pt", truncation=True, max_length=128)
generation_output = model.generate(input_tokens['input_ids'].cuda(), max_new_tokens=20)
```
使用可选压缩:
```python
model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct", compression='4bit')
```
训练示例(LoRA):
```bash
python air_llm/examples/train_qwen38_flash_next_lora.py --data my_data.jsonl --seq-len 512 --save-adapter adapter.pt
```
配置选项:compression(none/4bit/8bit)、profiling_mode、layer_shards_saving_path、hf_token、prefetching(默认开启)、delete_original。首次运行时原始模型会被拆分并按层缓存,因此 Hugging Face 缓存目录中需要足够的磁盘空间。
评论
0 评分人数达到10人后显示
登录后参与讨论。