To run Llama 3.2 3B locally at FP16 quantization, you need at minimum 9.2 GB of GPU VRAM.
Used market gem. Tight on VRAM but viable for this workload.
Strong inference GPU. Handles 7-13B models comfortably.
Best consumer GPU. Breeze through 13B models at any quantization.
Use the interactive calculator to compare Llama 3.2 3B across all available formats.
Open Live Calculator →