daniloreddy commited on
Commit
bd681e4
·
verified ·
1 Parent(s): c5460d6

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +61 -3
README.md CHANGED
@@ -1,3 +1,61 @@
1
- ---
2
- license: apache-2.0
3
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: google/gemma-4-E2B-it
4
+ tags:
5
+ - llama.cpp
6
+ - gguf
7
+ - quantized
8
+ - text-generation
9
+ - lightweight
10
+ - lmstudio
11
+ - jan
12
+ - cobalt
13
+ - text-generation-webui
14
+ ---
15
+
16
+ # Gemma-4-E2B-IT - GGUF High-Quality Quantizations
17
+
18
+ This repository provides **GGUF** quantized versions of the [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it) model, optimized for local execution using `llama.cpp` and compatible ecosystems.
19
+
20
+ ## 📌 Version Notes
21
+ All quantizations were generated from the official **FP16** weights.
22
+ - **Target:** Efficient execution on consumer hardware, mobile/edge devices, and systems with limited memory.
23
+ - **Performance:** The output quality (reasoning, coherence, and accuracy) is strictly dependent on the base model's parameter scale (1B).
24
+
25
+ ## 📊 Quantization Table
26
+
27
+ | File | Method | Bit | Description |
28
+ | :--- | :--- | :--- | :--- |
29
+ | **fp16.gguf** | FP16 | 16-bit | **Original Weights.** No quantization applied. Maximum fidelity. |
30
+ | **Q8_0.gguf** | Q8_0 | 8-bit | **Near-lossless.** Practically identical to the original model with lower memory footprint. |
31
+ | **Q5_K_M.gguf** | Q5_K_M | 5-bit | **High Precision.** Minimizes quantization error for critical tasks. |
32
+ | **Q4_K_M.gguf** | Q4_K_M | 4-bit | **Recommended.** Best balance between speed and performance. |
33
+ | **Q4_K_S.gguf** | Q4_K_S | 4-bit | **Fast/Small.** Optimized for maximum throughput and low RAM usage. |
34
+
35
+ ## 🛠️ Technical Details
36
+ - **Quantization Date:** 2026-04-05
37
+ - **Tool used:** `llama-quantize` (llama.cpp)
38
+ - **Method:** K-Quantization (optimized for AVX2/AVX-512 and modern GPU architectures).
39
+
40
+ ## 🚀 How to Use
41
+ # Start a local OpenAI-compatible server with a web UI:
42
+
43
+ ### llama.cpp (CLI) using model from HuggingFace
44
+ ```bash
45
+ ./llama-cli -hf daniloreddy/gemma-4-E2B-it_GGUF:Q4_K_M -p "User: Hello! Assistant:" -n 512 --temp 0.7
46
+ ```
47
+
48
+ ### llama.cpp (CLI) using downloaded model
49
+ ```bash
50
+ ./llama-cli -m path/to/gemma-4-E2B-it_Q4_K_M.gguf -p "User: Hello! Assistant:" -n 512 --temp 0.7
51
+ ```
52
+
53
+ ### llama.cpp (SERVER) using model from HuggingFace
54
+ ```bash
55
+ ./llama-server -hf daniloreddy/gemma-4-E2B-it_GGUF:Q4_K_M --port 8080 -c 4096
56
+ ```
57
+
58
+ ### llama.cpp (SERVER) using downloaded model
59
+ ```bash
60
+ ./llama-server -m /path/to/gemma-4-E2B-it_Q4_K_M.gguf --port 8080 -c 4096
61
+ ```