08f1e18965
* select layers based on estimated model memory usage * always account for scratch vram * dont load +1 layers * better estmation for graph alloc * Update gpu/gpu_darwin.go Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com> * Update llm/llm.go Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com> * Update llm/llm.go * add overhead for cuda memory * Update llm/llm.go Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com> * fix build error on linux * address comments --------- Co-authored-by: Bruce MacDonald <brucewmacdonald@gmail.com> |
||
---|---|---|
.. | ||
gpu.go | ||
gpu_darwin.go | ||
gpu_info.h | ||
gpu_info_cpu.c | ||
gpu_info_cuda.c | ||
gpu_info_cuda.h | ||
gpu_info_rocm.c | ||
gpu_info_rocm.h | ||
gpu_test.go | ||
types.go |