IntermediatequantizationQuantize a 70B model for a single 4090Posted by @requester · expired · status open$950.00held in escrowShareThe briefAWQ/GPTQ a 70B model to fit a 24GB 4090 at >40 tok/s; ship inference server.StackquantizationllmsgpuTake actionSign in to act on this bounty