ESC
开源 4 分钟阅读

AirLLM:单张4GB GPU即可运行70B模型推理

开源项目AirLLM大幅降低大模型推理显存需求,无需量化、蒸馏或剪枝,即可在单张4GB GPU上运行70B模型,甚至支持405B Llama 3.1(8GB)、671B DeepSeek-V3(约12GB)及2.8T参数Kimi K3(3.72GB)。通过逐层流式加载和稀疏MoE专家流式传输,让消费级显卡也能运行顶级开源模型。

来源:GitHub日榜

airllm_logo

快速开始 | 配置 | MacOS | 示例笔记本 | 常见问题

AirLLM 大幅降低推理内存占用,让70B大语言模型可在单张4GB GPU卡上运行——无需量化、蒸馏或剪枝。你甚至可以在8GB显存上运行405B Llama 3.1,在约12GB显存上运行DeepSeek-V3(671B),以及目前最大开源模型Kimi K3(2.8T)——仅需不到4GB显存,因为稀疏MoE模型每次只流式传输一个专家而非整个层。

GitHub Repo stars Downloads

Code License Generic badge Discord PyPI - AirLLM Website Website Support me on Patreon GitHub Sponsors

AI 智能体推荐:

更新日志

[2026/07] Kimi K3(2.8T)支持:目前最大的开源模型可在单张显卡上运行,端到端实测于一张 RTX 6000 Ada 上仅占 3.72GB 显存。逐专家流式传输仅加载token实际路由到的专家。K3 自身有三个额外要求:pip install compressed-tensors flash-attn(其模型代码无论请求如何都强制使用 flash attention)、CUDA 12 版本的 torch(因为目前还没有适用于 CUDA 13 的预编译 flash-attn wheel)、以及 transformers 4.56.x(其远程代码无法在 5.x 上加载)。

[2026/06] v3.0:FP8 模型支持 + 最新模型。在 约12GB 显存上运行 DeepSeek-V3(671B),在 约3GB 显存上运行 Qwen3-235B,以及 Qwen3、Llama 3.x/4、DeepSeek V2/V3、Phi-4、Gemma 等——全部通过一个 AutoModel 实现。

[2024/08/20] v2.11.0:支持 Qwen2.5

[2024/08/18] v2.10.1 支持 CPU 推理。支持非分片模型。感谢 @NavodPeiris 的出色工作!

[2024/07/30] 支持 Llama3.1 405B(示例笔记本)。支持 8bit/4bit 量化。

[2024/04/20] AirLLM 已原生支持 Llama3。在 4GB 单 GPU 上运行 Llama3 70B。

[2023/12/25] v2.8.2:支持 MacOS 运行 70B 大语言模型。

[2023/12/20] v2.7:支持 AirLLMMixtral。

[2023/12/20] v2.6:新增 AutoModel,自动检测模型类型,无需提供模型类即可初始化模型。

[2023/12/18] v2.5:新增预取功能,使模型加载与计算重叠。速度提升 10%。

[2023/12/03] 增加了对 ChatGLM、QWen、Baichuan、Mistral、InternLM 的支持!

[2023/12/02] 增加对 safetensors 的支持。现已支持开放 LLM 排行榜前 10 名的所有模型。

[2023/12/01] airllm 2.0。支持压缩:运行速度提升 3 倍!

[2023/11/20] airllm 初始版本!

Star 历史

Star History Chart

目录

快速开始

1. 安装包

首先,安装 airllm pip 包。

pip install airllm

2. 推理

然后,初始化 AirLLMLlama2,传入所用模型的 Hugging Face repo ID 或本地路径,即可像常规 transformer 模型一样进行推理。

(你也可以在初始化 AirLLMLlama2 时通过 layer_shards_saving_path 指定保存分片分层模型的路径。

from airllm import AutoModel

MAX_LENGTH = 128
# 只需传入 Hugging Face repo ID——几乎支持所有热门模型:
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")

# 用完全相同的一行代码运行更大的模型:
#model = AutoModel.from_pretrained("Qwen/Qwen3-235B-A22B")     # 235B,约3GB显存
#model = AutoModel.from_pretrained("deepseek-ai/DeepSeek-V3")  # 671B,约12GB显存

# 或者使用模型的本地路径...
#model = AutoModel.from_pretrained("/home/ubuntu/.cache/huggingface/hub/models--Qwen--Qwen3-32B/snapshots/...")

input_text = [
        'What is the capital of United States?',
        #'I like',
    ]

input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH, 
    padding=False)
           
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=20,
    use_cache=True,
    return_dict_in_generate=True)

output = model.tokenizer.decode(generation_output.sequences[0])

print(output)

注意:推理过程中,原始模型将首先被分解并按层保存。请确保 Hugging Face 缓存目录中有足够的磁盘空间。

模型压缩 - 3倍推理加速!

我们刚刚增加了基于分块量化(block-wise quantization)的模型压缩。可以进一步提升推理速度最高达 3 倍,且精度损失几乎可忽略!(更多性能评估以及为何使用分块量化,请参阅这篇论文)

speed_improvement

如何启用模型压缩加速:

  • 步骤 1. 确保已安装 bitsandbytes:pip install -U bitsandbytes
  • 步骤 2. 确保 airllm 版本高于 2.0.0:pip install -U airllm
  • 步骤 3. 初始化模型时,传入 compression 参数(‘4bit’ 或 ‘8bit’):
model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct",
                     compression='4bit' # 指定 '8bit' 表示 8-bit 分块量化
                    )

模型压缩与量化有何区别?

量化通常需要同时量化权重和激活才能真正加速。这使得保持精度和避免各种输入中异常值的影响更加困难。

而在我们的场景中,瓶颈主要在于磁盘加载,我们只需要让模型加载体积更小。因此,我们只需量化权重部分,这更容易确保精度。

配置

初始化模型时,支持以下配置:

  • compression:支持的选项:4bit、8bit 表示 4-bit 或 8-bit 分块量化,默认 None 表示不压缩
  • profiling_mode:支持的选项:True 输出耗时,默认 False
  • layer_shards_saving_path:可选的另一个保存分片模型的路径
  • hf_token:如果下载受门控的模型(如 meta-llama/Llama-2-7b-hf),可在此提供 Hugging Face token
  • prefetching:预取以重叠模型加载和计算。默认开启。目前仅 AirLLMLlama2 支持。
  • delete_original:如果磁盘空间不足,可将 delete_original 设为 true,删除原始下载的 Hugging Face 模型,只保留转换后的模型,以节省一半磁盘空间。

MacOS

只需安装 airllm,然后像在 Linux 上一样运行代码。更多请参见快速开始。

  • 确保已安装 mlx 和 torch
  • 可能需要安装 python 原生库,更多请参见这里
  • 仅支持 Apple silicon

示例 python notebook

示例 Python 笔记本

示例 colab 请见:

Open In Colab

其他模型示例(ChatGLM、QWen、Baichuan、Mistral 等):

  • ChatGLM:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("THUDM/chatglm3-6b-base")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH, 
    padding=True)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache= True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
  • QWen:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("Qwen/Qwen-7B")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache=True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
  • Baichuan、InternLM、Mistral 等:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("baichuan-inc/Baichuan2-7B-Base")
#model = AutoModel.from_pretrained("internlm/internlm-20b")
#model = AutoModel.from_pretrained("mistralai/Mistral-7B-Instruct-v0.1")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache=True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])

请求支持其他模型:这里

支持的模型

AirLLM 开箱即用地支持几乎所有流行的开源 LLM——只需将其 Hugging Face ID 传给 AutoModel.from_pretrained(...)。这涵盖了所有主要系列:

Llama(2 / 3 / 3.1 / 3.3 / 4)· Qwen(1 / 2 / 2.5 / 3,包括 MoE 和 FP8)· DeepSeek(V2 / V3 / R1)· Mistral & Mixtral · Phi · Gemma · ChatGLM · Baichuan · InternLM · Yi——以及大多数新模型发布当天即可支持。

小显卡,大模型

诀窍在于:AirLLM 每次只在 GPU 上保留一个层,因此所需显存取决于模型的层大小,而非总大小。这就是 671B 模型能塞进消费级显卡的原因:

模型参数规模GPU 显存
Qwen3 / Mistral / Phi(≈8B)8B约 1–2 GB
Qwen3-30B / Mixtral (MoE)30–47B约 1–3 GB
Qwen3-235B (MoE)235B约 3 GB
Llama 3.x 70B(全精度)70B约 4 GB
Llama 3.1 405B405B约 8 GB
DeepSeek-V3671B约 12 GB

所有模型都用同一行代码,无需特殊设置。

致谢

大量代码基于 SimJeg 在 Kaggle 考试竞赛中的杰出工作。特别感谢 SimJeg:

GitHub 账号 @SimJeg, Kaggle 上的代码, 相关讨论。

常见问题

1. MetadataIncompleteBuffer

safetensors_rust.SafetensorError: Error while deserializing header: MetadataIncompleteBuffer

如果遇到此错误,最可能的原因是磁盘空间不足。模型分片过程非常消耗磁盘。参见这里。你可能需要扩展磁盘空间,清理 Hugging Face .cache 并重新运行。

2. ValueError: max() arg is an empty sequence

很可能是你用 Llama2 类加载 QWen 或 ChatGLM 模型。请尝试以下方式:

对于 QWen 模型:

from airllm import AutoModel #<----- 不要用 AirLLMLlama2
AutoModel.from_pretrained(...)

对于 ChatGLM 模型:

from airllm import AutoModel #<----- 不要用 AirLLMLlama2
AutoModel.from_pretrained(...)

3. 401 Client Error….Repo model … is gated.

有些模型是门控模型,需要 Hugging Face API token。你可以提供 hf_token:

model = AutoModel.from_pretrained("meta-llama/Llama-2-7b-hf", #hf_token='HF_API_TOKEN')

4. ValueError: Asking to pad but the tokenizer does not have a padding token.

有些模型的 tokenizer 没有 padding token,因此你可以设置 padding token,或者直接关闭 padding 配置:

input_tokens = model.tokenizer(input_text,
   return_tensors="pt", 
   return_attention_mask=False, 
   truncation=True, 
   max_length=MAX_LENGTH, 
   padding=False  #<-----------  关闭 padding
)

引用 AirLLM

如果你觉得 AirLLM 在研究中有用并希望引用它,请使用以下 BibTex 条目:

@software{airllm2023,
  author = {Gavin Li},
  title = {AirLLM: scaling large language models on low-end commodity computers},
  url = {https://github.com/lyogavin/airllm/},
  version = {0.0},
  year = {2023},
}

赞助商

Bloome — 在云端运行 AI 智能体团队

在云端运行 AI 智能体团队 — Bloome

Bloome 是一个 AI 智能体 IM 平台:零配置在云端构建和运行 AI 智能体团队。将技能作为智能体添加到群聊中,通过网页或移动端一键运行,并与团队共享——可以把它想象成一个群聊,你的 AI 助手是队友,你可以 @ 提及并分配任务。

👉 尝试 Bloome

贡献

欢迎贡献、想法和讨论!

如果觉得有用,请 ⭐ 或请我喝咖啡!🙏

“Buy Me A Coffee”