
快速开始 | 配置 | MacOS | 示例笔记本 | 常见问题
AirLLM 大幅降低推理内存占用,让70B大语言模型可在单张4GB GPU卡上运行——无需量化、蒸馏或剪枝。你甚至可以在8GB显存上运行405B Llama 3.1,在约12GB显存上运行DeepSeek-V3(671B),以及目前最大开源模型Kimi K3(2.8T)——仅需不到4GB显存,因为稀疏MoE模型每次只流式传输一个专家而非整个层。
AI 智能体推荐:
更新日志
[2026/07] Kimi K3(2.8T)支持:目前最大的开源模型可在单张显卡上运行,端到端实测于一张 RTX 6000 Ada 上仅占 3.72GB 显存。逐专家流式传输仅加载token实际路由到的专家。K3 自身有三个额外要求:pip install compressed-tensors flash-attn(其模型代码无论请求如何都强制使用 flash attention)、CUDA 12 版本的 torch(因为目前还没有适用于 CUDA 13 的预编译 flash-attn wheel)、以及 transformers 4.56.x(其远程代码无法在 5.x 上加载)。
[2026/06] v3.0:FP8 模型支持 + 最新模型。在 约12GB 显存上运行 DeepSeek-V3(671B),在 约3GB 显存上运行 Qwen3-235B,以及 Qwen3、Llama 3.x/4、DeepSeek V2/V3、Phi-4、Gemma 等——全部通过一个 AutoModel 实现。
[2024/08/20] v2.11.0:支持 Qwen2.5
[2024/08/18] v2.10.1 支持 CPU 推理。支持非分片模型。感谢 @NavodPeiris 的出色工作!
[2024/07/30] 支持 Llama3.1 405B(示例笔记本)。支持 8bit/4bit 量化。
[2024/04/20] AirLLM 已原生支持 Llama3。在 4GB 单 GPU 上运行 Llama3 70B。
[2023/12/25] v2.8.2:支持 MacOS 运行 70B 大语言模型。
[2023/12/20] v2.7:支持 AirLLMMixtral。
[2023/12/20] v2.6:新增 AutoModel,自动检测模型类型,无需提供模型类即可初始化模型。
[2023/12/18] v2.5:新增预取功能,使模型加载与计算重叠。速度提升 10%。
[2023/12/03] 增加了对 ChatGLM、QWen、Baichuan、Mistral、InternLM 的支持!
[2023/12/02] 增加对 safetensors 的支持。现已支持开放 LLM 排行榜前 10 名的所有模型。
[2023/12/01] airllm 2.0。支持压缩:运行速度提升 3 倍!
[2023/11/20] airllm 初始版本!
Star 历史

目录
快速开始
1. 安装包
首先,安装 airllm pip 包。
pip install airllm
2. 推理
然后,初始化 AirLLMLlama2,传入所用模型的 Hugging Face repo ID 或本地路径,即可像常规 transformer 模型一样进行推理。
(你也可以在初始化 AirLLMLlama2 时通过 layer_shards_saving_path 指定保存分片分层模型的路径。
from airllm import AutoModel
MAX_LENGTH = 128
# 只需传入 Hugging Face repo ID——几乎支持所有热门模型:
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")
# 用完全相同的一行代码运行更大的模型:
#model = AutoModel.from_pretrained("Qwen/Qwen3-235B-A22B") # 235B,约3GB显存
#model = AutoModel.from_pretrained("deepseek-ai/DeepSeek-V3") # 671B,约12GB显存
# 或者使用模型的本地路径...
#model = AutoModel.from_pretrained("/home/ubuntu/.cache/huggingface/hub/models--Qwen--Qwen3-32B/snapshots/...")
input_text = [
'What is the capital of United States?',
#'I like',
]
input_tokens = model.tokenizer(input_text,
return_tensors="pt",
return_attention_mask=False,
truncation=True,
max_length=MAX_LENGTH,
padding=False)
generation_output = model.generate(
input_tokens['input_ids'].cuda(),
max_new_tokens=20,
use_cache=True,
return_dict_in_generate=True)
output = model.tokenizer.decode(generation_output.sequences[0])
print(output)
注意:推理过程中,原始模型将首先被分解并按层保存。请确保 Hugging Face 缓存目录中有足够的磁盘空间。
模型压缩 - 3倍推理加速!
我们刚刚增加了基于分块量化(block-wise quantization)的模型压缩。可以进一步提升推理速度最高达 3 倍,且精度损失几乎可忽略!(更多性能评估以及为何使用分块量化,请参阅这篇论文)

如何启用模型压缩加速:
- 步骤 1. 确保已安装 bitsandbytes:
pip install -U bitsandbytes - 步骤 2. 确保 airllm 版本高于 2.0.0:
pip install -U airllm - 步骤 3. 初始化模型时,传入 compression 参数(‘4bit’ 或 ‘8bit’):
model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct",
compression='4bit' # 指定 '8bit' 表示 8-bit 分块量化
)
模型压缩与量化有何区别?
量化通常需要同时量化权重和激活才能真正加速。这使得保持精度和避免各种输入中异常值的影响更加困难。
而在我们的场景中,瓶颈主要在于磁盘加载,我们只需要让模型加载体积更小。因此,我们只需量化权重部分,这更容易确保精度。
配置
初始化模型时,支持以下配置:
- compression:支持的选项:4bit、8bit 表示 4-bit 或 8-bit 分块量化,默认 None 表示不压缩
- profiling_mode:支持的选项:True 输出耗时,默认 False
- layer_shards_saving_path:可选的另一个保存分片模型的路径
- hf_token:如果下载受门控的模型(如 meta-llama/Llama-2-7b-hf),可在此提供 Hugging Face token
- prefetching:预取以重叠模型加载和计算。默认开启。目前仅 AirLLMLlama2 支持。
- delete_original:如果磁盘空间不足,可将 delete_original 设为 true,删除原始下载的 Hugging Face 模型,只保留转换后的模型,以节省一半磁盘空间。
MacOS
只需安装 airllm,然后像在 Linux 上一样运行代码。更多请参见快速开始。
- 确保已安装 mlx 和 torch
- 可能需要安装 python 原生库,更多请参见这里
- 仅支持 Apple silicon
示例 Python 笔记本
示例 colab 请见:
其他模型示例(ChatGLM、QWen、Baichuan、Mistral 等):
- ChatGLM:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("THUDM/chatglm3-6b-base")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
return_tensors="pt",
return_attention_mask=False,
truncation=True,
max_length=MAX_LENGTH,
padding=True)
generation_output = model.generate(
input_tokens['input_ids'].cuda(),
max_new_tokens=5,
use_cache= True,
return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
- QWen:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("Qwen/Qwen-7B")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
return_tensors="pt",
return_attention_mask=False,
truncation=True,
max_length=MAX_LENGTH)
generation_output = model.generate(
input_tokens['input_ids'].cuda(),
max_new_tokens=5,
use_cache=True,
return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
- Baichuan、InternLM、Mistral 等:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("baichuan-inc/Baichuan2-7B-Base")
#model = AutoModel.from_pretrained("internlm/internlm-20b")
#model = AutoModel.from_pretrained("mistralai/Mistral-7B-Instruct-v0.1")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
return_tensors="pt",
return_attention_mask=False,
truncation=True,
max_length=MAX_LENGTH)
generation_output = model.generate(
input_tokens['input_ids'].cuda(),
max_new_tokens=5,
use_cache=True,
return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
请求支持其他模型:这里
支持的模型
AirLLM 开箱即用地支持几乎所有流行的开源 LLM——只需将其 Hugging Face ID 传给 AutoModel.from_pretrained(...)。这涵盖了所有主要系列:
Llama(2 / 3 / 3.1 / 3.3 / 4)· Qwen(1 / 2 / 2.5 / 3,包括 MoE 和 FP8)· DeepSeek(V2 / V3 / R1)· Mistral & Mixtral · Phi · Gemma · ChatGLM · Baichuan · InternLM · Yi——以及大多数新模型发布当天即可支持。
小显卡,大模型
诀窍在于:AirLLM 每次只在 GPU 上保留一个层,因此所需显存取决于模型的层大小,而非总大小。这就是 671B 模型能塞进消费级显卡的原因:
| 模型 | 参数规模 | GPU 显存 |
|---|---|---|
| Qwen3 / Mistral / Phi(≈8B) | 8B | 约 1–2 GB |
| Qwen3-30B / Mixtral (MoE) | 30–47B | 约 1–3 GB |
| Qwen3-235B (MoE) | 235B | 约 3 GB |
| Llama 3.x 70B(全精度) | 70B | 约 4 GB |
| Llama 3.1 405B | 405B | 约 8 GB |
| DeepSeek-V3 | 671B | 约 12 GB |
所有模型都用同一行代码,无需特殊设置。
致谢
大量代码基于 SimJeg 在 Kaggle 考试竞赛中的杰出工作。特别感谢 SimJeg:
GitHub 账号 @SimJeg, Kaggle 上的代码, 相关讨论。
常见问题
1. MetadataIncompleteBuffer
safetensors_rust.SafetensorError: Error while deserializing header: MetadataIncompleteBuffer
如果遇到此错误,最可能的原因是磁盘空间不足。模型分片过程非常消耗磁盘。参见这里。你可能需要扩展磁盘空间,清理 Hugging Face .cache 并重新运行。
2. ValueError: max() arg is an empty sequence
很可能是你用 Llama2 类加载 QWen 或 ChatGLM 模型。请尝试以下方式:
对于 QWen 模型:
from airllm import AutoModel #<----- 不要用 AirLLMLlama2
AutoModel.from_pretrained(...)
对于 ChatGLM 模型:
from airllm import AutoModel #<----- 不要用 AirLLMLlama2
AutoModel.from_pretrained(...)
3. 401 Client Error….Repo model … is gated.
有些模型是门控模型,需要 Hugging Face API token。你可以提供 hf_token:
model = AutoModel.from_pretrained("meta-llama/Llama-2-7b-hf", #hf_token='HF_API_TOKEN')
4. ValueError: Asking to pad but the tokenizer does not have a padding token.
有些模型的 tokenizer 没有 padding token,因此你可以设置 padding token,或者直接关闭 padding 配置:
input_tokens = model.tokenizer(input_text,
return_tensors="pt",
return_attention_mask=False,
truncation=True,
max_length=MAX_LENGTH,
padding=False #<----------- 关闭 padding
)
引用 AirLLM
如果你觉得 AirLLM 在研究中有用并希望引用它,请使用以下 BibTex 条目:
@software{airllm2023,
author = {Gavin Li},
title = {AirLLM: scaling large language models on low-end commodity computers},
url = {https://github.com/lyogavin/airllm/},
version = {0.0},
year = {2023},
}
赞助商

在云端运行 AI 智能体团队 — Bloome
Bloome 是一个 AI 智能体 IM 平台:零配置在云端构建和运行 AI 智能体团队。将技能作为智能体添加到群聊中,通过网页或移动端一键运行,并与团队共享——可以把它想象成一个群聊,你的 AI 助手是队友,你可以 @ 提及并分配任务。
👉 尝试 Bloome
贡献
欢迎贡献、想法和讨论!
如果觉得有用,请 ⭐ 或请我喝咖啡!🙏
