Large Language Models like Llama 2 7B and Mistral 7B have been successfully optimized for on-device deployment, with techniques like Activation-aware Weight Quantization (AWQ) enabling 70B Llama-2 on mobile GPUs.
Open the full topic