GEV-26B-Decide-GGUF

autotrust/GEV-26B-Decide is an independent open-weights decision model from AutoTrust AI, built on Google's Gemma-4-26B-A4B-it (26B parameters, about 4B active per token; previously published as autotrust/JEV-Gemma4-26B-A4B), that pairs System 1 (a LoRA plus a 24-slot decision head returning calibrated probabilities for noul, choice with 2-256 options, or score questions in one forward pass of about 45 ms, over text and images) with System 2 (the unmodified Gemma-4 base in thinking mode). Its distinctive feature is opt-in adaptive thinking (thinking: "auto"): when System 1's leading option falls below 0.8 confidence, the base model reasons, and the answer distribution it produces is averaged equally with System 1's. On 1,754 questions from six public sets outside the Decision Index, this lifts accuracy from 73.3% to 83.4% (96% of the always-think gain, thinking on 48% of questions). The model reports a Decision Index 0.2.1 of 62.48 against TypeSafe Jev 1.13's 57.91, using adaptive thinking on Knowledge & Reasoning and System 1 on the other four areas, but the card says this is its own scoring, not a board entry, and that adaptive thinking exceeds the board's latency limit on that area. The gains are real on GPQA Diamond (42.9 to 78.6), MMLU-Pro and BBH, but HLE only returns to chance, chess does not benefit, and thinking can hurt classification (BANKING77 macro-F1 88.0 to 85.0). Its System 1 computer-use result is 95% on 60 browser tasks at about 85 ms per click, the same as JEV-27B-VL but 3x faster, while the robot-arm task completes only 40% of scenes versus 75% for JEV-27B-VL. It scores 78.4% on VL-RewardBench, though the decision head was trained on text, so image decisions are zero-shot. It is served through a patched vLLM server (serve_decide.py, with a bundled LoRA-on-tied-lm_head patch) exposing POST /v1/decide alongside the OpenAI endpoints. The adapter and head are Apache-2.0, with the base model under the Gemma 4 terms. The card also discloses that BANKING77 and CLINC150 training splits were in the training data.

Model Files

File Name Quant Type File Size File Link Description
GEV-26B-Decide.BF16.gguf BF16 50.5 GB Link Full BF16 weights. Highest quality, largest file size.
GEV-26B-Decide.Q3_K_L.gguf Q3_K_L 13.8 GB Link Lower quality but usable, good for low RAM availability.
GEV-26B-Decide.Q4_K_M.gguf Q4_K_M 16.8 GB Link Good quality, default size for most use cases, recommended.
GEV-26B-Decide.Q5_K_M.gguf Q5_K_M 19.1 GB Link High quality, recommended.
GEV-26B-Decide.mmproj-bf16.gguf mmproj-bf16 1.19 GB Link Multimodal projection file in BF16 format. Used for vision/language models.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Releases / v0.6.0 — https://github.com/ggml-org/llama.cpp/releases/tag/v0.6.0

Downloads last month
507
GGUF
Model size
25B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/GEV-26B-Decide-GGUF

Quantized
(7)
this model

Collections including prithivMLmods/GEV-26B-Decide-GGUF