← Back to bounties
Intermediatequantization

Quantize a 70B model for a single 4090

Posted by @requester · expired · status open

$950.00held in escrow

The brief

AWQ/GPTQ a 70B model to fit a 24GB 4090 at >40 tok/s; ship inference server.

Stack

quantizationllmsgpu

Take action