Running on <4gb Vram

#31
by Jit2024 - opened

Hi, I found out a way to run this model under 4 gb vram. Without any quantization.
repo: https://github.com/Jit-Roy/WeeLLM

Sign up or log in to comment