Library
huggingface.cosaved Jun 13, 2026published Aug 30, 2026archived

unsloth/gemma-4-E4B-it-qat-GGUF · Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co/unsloth/gemma-4-E4B-it-qat-GGUF

Archived text733 words

Read our How to Run Gemma 4 QAT Guide! See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. Jun 9 Update: Added MTP support. See our MTP Guide. Gemma 4 can now be run and fine-tuned in Unsloth Studio. Read our guide. See all versions of Gemma 4 QAT (GGUF, 16-bit etc.) in our collection. Example of Gemma 4 E4B (4-bit GGUF) running in Unsloth Studio with tool-calling: Run with MTP (speculative decoding) This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E4B-it.gguf, a near-lossless smart Q4_0). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: ./build/bin/llama-server \ -hf unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL \ --spec-type draft-mtp --spec-draft-n-max 4 \ -ngl 999 -fa off The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Hugging Face | GitHub | Launch Blog | Documentation License: Apache 2.0 | Authors: Google DeepMind This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q4_0): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q4_0): Ready-to-deploy formats for broad ecosystem compatibility. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B. Mobile-optimized (wNa8o8): A custom schema engineered explicitly for mobile hardware efficiency. It features targeted 2-bit decoding layers, optimized KV caches, and static activations to maximize VRAM savings. Available for Gemma 4 E2B and E4B. Compressed Tensors (w4a16): QAT checkpoints serialized in the compressed-tensors format for native, optimized inference with vLLM. Available for Gemma 4 E2B, E4B, 12B, and 31B. This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized…

crawled Sep 8, 2026

Related sites

Nearest neighbours by embedding distance, computed at index time.

okara.ai

AI CMO | Okara

Your AI Chief Marketing Officer that automates SEO, Reddit engagement, and content creation. Get actionable marketing insights…