1 post tagged llm, newest first.

Fitting a 27B model on one 24GB card is two arithmetic unknowns - weight memory at your quantisation and KV cache at your context - and a curve, not a number.
󰣨 ymrtech@ymrtech | 󰌠 NixOS | 󰍢 UTF-8