Repository Details
Shared by
HelloGitHub Rating
0 ratings
Free•Apache-2.0
Claim
Discuss
Collect
Share
32.7k
Stars
No
Chinese
Jupyter Notebook
Language
Yes
Active
10
Contributors
144
Issues
No
Organization
None
Latest
3k
Forks
Apache-2.0
License
More

This is a Python library that drastically reduces inference VRAM usage via hierarchical loading. It can run 70B large models with only 4GB VRAM without quantization, distillation, or pruning. It supports mainstream models such as Llama 3.x, Qwen3, DeepSeek, and offers prefetch acceleration, 8bit/4bit quantization, and macOS support.
Comments
Rating:
No comments yet