下拉刷新
Repository Details
Shared bynavbar_avatar
repo_avatar
HelloGitHub Rating
10.0
2 ratings
Running 70B Large Models with Just 4GB VRAM
FreeApache-2.0
Claim
Collect
Share
34.4k
Stars
No
Chinese
Jupyter Notebook
Language
Yes
Active
10
Contributors
153
Issues
No
Organization
None
Latest
4k
Forks
Apache-2.0
License
More
airllm image
This is a Python library that drastically reduces inference VRAM usage via hierarchical loading. It can run 70B large models with only 4GB VRAM without quantization, distillation, or pruning. It supports mainstream models such as Llama 3.x, Qwen3, DeepSeek, and offers prefetch acceleration, 8bit/4bit quantization, and macOS support.

Comments

Rating:
No comments yet