这个项目能做什么

AirLLM 可在显存低至 2 GB 的消费级 GPU 上对数十亿参数的语言模型进行推理和微调。其原理是任意时刻仅在 GPU 上保留单个 transformer 层——权重从磁盘逐层流式加载——因此 GPU 显存需求取决于单层大小,而非模型总大小。 主要能力: - 在单块低端 GPU 上对约 8B 至 2.8T 参数的模型进行推理 - 可选的块级量化(4-bit 或 8-bit),推理速度最高提升 3 倍 - 在小显存上对超大模型进行 LoRA 微调;冻结的基础权重从磁盘流式加载,适配器常驻显存 - 支持 MoE、Flash、FP8 以及原生视觉(VL)架构 - 跨平台:Linux/CUDA 与 macOS(通过 MLX 支持 Apple Silicon) 支持的模型系列包括 Llama(2/3/3.1/3.3/4)、Qwen(1/2/2.5/3/3.5/3.8,稠密与 MoE)、DeepSeek(V2/V3/R1)、Mistral 与 Mixtral、Phi、Gemma、ChatGLM、Baichuan、InternLM、Yi 以及 Kimi K3。 使用示例: ```python from airllm import AutoModel model = AutoModel.from_pretrained("Qwen/Qwen3-32B") input_tokens = model.tokenizer(text, return_tensors="pt", truncation=True, max_length=128) generation_output = model.generate(input_tokens['input_ids'].cuda(), max_new_tokens=20) ``` 使用可选压缩: ```python model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct", compression='4bit') ``` 训练示例(LoRA): ```bash python air_llm/examples/train_qwen38_flash_next_lora.py --data my_data.jsonl --seq-len 512 --save-adapter adapter.pt ``` 配置选项:compression(none/4bit/8bit)、profiling_mode、layer_shards_saving_path、hf_token、prefetching(默认开启)、delete_original。首次运行时原始模型会被拆分并按层缓存,因此 Hugging Face 缓存目录中需要足够的磁盘空间。