airllm—Running 70B Large Models with Just 4GB VRAM
2
airllm—Running 70B Large Models with Just 4GB VRAM
2
This is a Python library that drastically reduces inference VRAM usage via hierarchical loading. It
lyogavin
2.7k
- That's all for now, only these are available at the moment -