hi there!im trying to develop based on your fancy project,but i faced some questions
i wannder to figure out the GPU requirements to run your model, i think the raw llama13b is too heavy to combine this project with other applications,
Whether to provide quantized model operations to reduce the GPU burden?
thanks a lot in advance QAQ!
hi there!im trying to develop based on your fancy project,but i faced some questions
i wannder to figure out the GPU requirements to run your model, i think the raw llama13b is too heavy to combine this project with other applications,
Whether to provide quantized model operations to reduce the GPU burden?
thanks a lot in advance QAQ!