Vega 0.8B: public intents
Inference on a single T4, end to end, in the browser.
800 million or 4 billion parameters. A 73,728-token context. Images. Running on your own hardware.
That combination is the point: a decision model small enough to sit inside your own process, that can read a whole contract in one pass and look at a picture, and that still beats a hosted reference model on the tasks below.
A new architecture with its own training and calibration method. The language model inside it is frozen and never trained; everything learned lives in a small physics engine on top of it.
| sizes in this repo | 0.8B at the root (default) and 4B under 4b/ |
| backbone | Qwen/Qwen3.5-0.8B or Qwen/Qwen3.5-4B, frozen, SHA-256 pinned |
| trained parameters | engine 57 MB (0.8B) or 124 MB (4B), plus gated adapters |
| context | 73,728 tokens |
| images | yes, the frozen backbone's vision encoder, via vega_vision.py |
| answer types | choice, boolean (probability a statement is true), score |
| latency | 0.267 s/decision (T4, fp16) · 0.28 s (Apple M-series, fp32) |
The two sizes
Both are in this repository and take the same code path. The 0.8B sits at the root because that is
where it has always been; the 4B is under 4b/.
import vegaml
v = vegaml.load() # 0.8B, the default
v = vegaml.load("4b") # 4B, same interface
| 0.8B | 4B | |
|---|---|---|
| frozen backbone | Qwen3.5-0.8B, 24 layers, d 1024 | Qwen3.5-4B, 32 layers, d 2560 |
| trained engine | 14,342,512 params, 57 MB | 30,879,088 params, 124 MB |
| task adapter | 835,937 params | 1,431,905 params |
| with the task adapter | 0.763 | 0.803 |
| soft accuracy | 0.680 | 0.725 |
| Brier | 0.314 | 0.270 |
| calibration error | 0.026 | 0.019 |
| per decision | 22 ms | 28 ms |
LocalLLaMA/typed-decisions test split, 2,050 decisions, each checkpoint's own run of the same
harness, full tables in eval/typed_decisions.md and 4b/eval/typed_decisions.md.
Per workflow with the adapter, the 4B leads on five of six:
| workflow | 0.8B | 4B |
|---|---|---|
| apibank | 0.895 | 0.965 |
| when2call | 0.904 | 0.946 |
| appliances | 0.851 | 0.891 |
| banking77 | 0.714 | 0.749 |
| esci | 0.564 | 0.644 |
| clinc150 | 0.880 | 0.867 |
The 4B is the better model on nearly everything measured here. The 0.8B is still the default, because 22 ms against 28 ms and 57 MB against 124 MB settles more deployments than four points of accuracy does, and because every Jev comparison on this card was run on the 0.8B.
73,728-token context
max_len is 73,728 tokens, and the architecture is built for it rather than merely permitting it:
- Read-once prefix caching. A state of at least 4,096 tokens is run once with a KV cache; every question about that state then continues from a copy of it. Twelve questions about one 73k-token contract cost roughly one read, not twelve.
- A memory reader. Long states are cut into 16 equal pieces, mean-pooled at the deepest feature layer, and fed to the engine alongside the pooled spans, so evidence late in a long document still reaches the decision.
- No silent truncation. An input that would not fit is refused, never quietly shortened. A decision made on a truncated contract is worse than no decision.
The frozen backbone carries 262,144 rotary positions, so 73k sits well inside its range.
Images
The backbone is multimodal and its vision encoder is frozen with the rest, so the same typed
interface takes a picture. Needs the vision extra, which carries Pillow and torchvision; the
backbone's processor loads torchvision on the way in and refuses to run without it.
# pip install "vegaml[vision]"
import vegaml
from PIL import Image
v = vegaml.load() # or load("4b")
v.decide_image(Image.open("invoice.png"), {
"kind": {"type": "choice", "instructions": "What kind of document is this?",
"criteria": {"invoice": "a bill", "receipt": "proof of payment",
"contract": "an agreement", "other": "none of these"}}})
Images are resized to fit max_pixels (default 1024·32·32). An image answer carries the same fields
as a text answer: the calibrated probabilities, the conformal set and the abstain flag.
Requires vegaml 0.8.0 or later. Before that release the packaged image path raised
ModuleNotFoundError on the first call, because it reached for a module that ships only in the
training tree. The figures below were measured there, where that module exists; they are unchanged,
but 0.7.1 and earlier cannot reproduce them from PyPI.
The figures below are all text; the vision path is functional but not covered by them.
Where Vega 0.8B surpasses Jev 1.13.0
The figure at the top of this card is these tables. Selected results: the areas where a 0.8B local model comes out ahead of the hosted reference. Every figure is from our own run of a public harness, with Jev called live over the identical items and no tuning on the evaluation data.
Phishing screening: the clearest win
800 emails, a perfect 50/50 split:
| accuracy | precision | recall | caught | |
|---|---|---|---|---|
| Vega 0.8B + adapter | 75.38 | 0.837 | 0.630 | 252 / 400 |
| Jev 1.13.0 | 61.88 | 0.961 | 0.247 | 99 / 400 |
+13.50 accuracy, and 2.5× the phishing caught (McNemar p = 3e-11). Jev is almost never wrong when it calls phishing, but it misses three quarters of it. For a screening task that is the wrong operating point, and Vega sits on the right one.
Date and quantity reasoning
JevBench v1.4.2, temporal_numeric: dates and quantities judged against policy thresholds:
| score | |
|---|---|
| Vega 0.8B + adapter | 46.7 |
| Jev 1.13.0 | 20.0 |
+26.7, more than double.
S1MB: benchmarks where Vega wins outright
From the public S1MB evaluator (137 benchmarks, 26,269 decisions), scored by the harness itself:
| benchmark | task | Vega | Jev |
|---|---|---|---|
| enron-spam | boolean | 1.000 | 0.920 |
| ag-news | choice | 0.955 | 0.806 |
| typed-decisions-score | score | 0.438 | 0.395 |
| go-emotions | boolean | 0.488 | 0.485 |
| drone-control | choice | 0.079 | 0.000 |
Calibration: at parity with a hosted model
Expected calibration error over 5,096 item-paired decisions:
| ECE | |
|---|---|
| Vega 0.8B + adapter | 9.5 |
| Jev 1.13.0 | 9.3 |
And better than Jev on individual benchmarks:
| benchmark | Vega ECE | Jev ECE |
|---|---|---|
| product relevance | 6.0 | 22.0 |
| banking intents | 5.0 | 9.7 |
| appliance routing | 5.0 | 5.7 |
Calibrated probability is the hard part of a decision model. A 0.8B engine matching a hosted one on it, and beating it by 16 points of ECE on one benchmark, is the result worth noting.
Speed
| median per decision | |
|---|---|
| Vega 0.8B + adapter | 280 ms (local, Apple M-series, fp32) |
| 267 ms (T4, fp16) | |
| Jev 1.13.0 | 591 ms (hosted API) |
2.1× faster, running entirely on your own hardware, no API call, no per-decision fee, no data leaving the machine.
Images: what the vision path actually scores
The hosted baseline takes no image input, so there is no like-for-like comparison. On document pages it was given Apple Vision OCR of the same images and read that text, which makes its row a two-model pipeline rather than a model that sees.
RVL-CDIP-N, 1,002 document pages, the 12 of 16 RVL-CDIP categories the set contains:
| system | reads | accuracy | ECE | median latency |
|---|---|---|---|---|
| Apple Vision OCR + Jev 1.13.0 | text | 0.896 | 0.045 | 1,049 ms |
| Vega 0.8B, zero-shot | the image | 0.793 | 0.207 | 1,989 ms |
| DiT, best published on this set | the image | 0.786 | n/a | n/a |
DiT was trained on RVL-CDIP and is the strongest published result on this out-of-distribution set. Vega has never seen the dataset and lands level with it. Not quite like-for-like: Vega chose among the 12 categories present, while a classifier trained on RVL-CDIP chooses among all 16. The pipeline latency includes the 649 ms Apple Vision takes per page, without which the API call alone is 400 ms.
MMMU-Pro, 300 items over 30 strata, chance 0.120:
| system | reads | accuracy | ECE |
|---|---|---|---|
| Gemini 3.1 Pro, best reported Oct 2026 | the image | 0.839 | n/a |
| Jev 1.13.0 | text only, no image input | 0.337 | 0.160 |
| Vega 0.8B | the image | 0.240 | 0.052 |
| Vega 0.8B, same questions | image withheld | 0.157 | 0.062 |
An 800M model is nowhere near a frontier one on university-level reasoning. The result worth keeping is the last two rows: withholding the image costs 8.3 points, so the vision path carries real signal rather than the text answering alone. Calibration error is three times better than the hosted model's on the same items.
Architecture
Two halves with a hard boundary between them.
Perception (frozen). The causal LM is read, never trained. Hidden states are tapped at layers 13 and 19 and pooled over three spans (state, question, final position), plus one pooled vector per answer option. The model is cut off after the deepest layer read, so a decision never pays for the layers above it, and the LM head is never run. No text is generated anywhere in the path.
The physics engine (trained, 55 MB). The pooled features become a world latent, then a foresight
latent, then a probe and an impulse that place a particle (z₀, p₀) in a 64-dimensional decision space.
Each candidate answer is a Gaussian potential well:
U(z) = ½κ‖z‖² − Σ_k a_k · exp( −‖z − c_k‖² / 2σ_k² )
The quadratic term keeps the particle bounded; each well pulls it toward one answer with its own depth
a_k, centre c_k and width σ_k, all produced from that option's own features. The particle rolls
under damped Hamiltonian dynamics (symplectic Euler with friction and a state-space thermostat), and the
answer is read from where it settles:
E_k = ‖z_T − c_k‖² / 2σ_k² − log a_k P = softmax(−E / τ)
A decision is a short physical simulation with a fixed step budget, not a sampling loop: constant cost, deterministic result.
Gated adapters (5.3 MB). Two low-rank adapters (rank 32, α 32) attach to seven engine projections. A per-question gate (a small sigmoid head over the world latent) decides for each question independently whether an adapter contributes. The gate value and the routed adapter are returned with every answer, so routing is auditable rather than implicit.
Related work
A frozen language model feeding a trainable module described as a differentiable physics engine in a learned latent space was proposed for physical reasoning in CWMI, Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning (Sharma et al., 2025). The separation of a frozen reader from a trained dynamics module is the same idea, and the resemblance is worth stating plainly.
What differs is what the dynamics are for. CWMI learns physics as subject matter: its latent state is meant to carry position, velocity, mass and material, it is supervised against ground-truth final states from video and a simulator, and it answers physical-reasoning questions (PIQA, 89.4%). Its module is a 12-layer Transformer decoder of roughly 256M trainable parameters trained on 8 H100s.
Vega uses dynamics as the inference mechanism rather than as a model of the world. Its 64-dimensional space has no physical referent; the wells are built from the candidate answers of whatever question was asked, so the landscape is rebuilt per question rather than learned once. The engine is 14.3M parameters with no attention anywhere, and the motion is an explicit damped Hamiltonian integrated with symplectic Euler against a written-down potential, not a learned state transition. The answer is the lowest-energy well, and the residual kinetic energy at the end is what sets the temperature, which is the part that makes the probabilities mean something. CWMI reports accuracy; it has no calibration, conformal or abstain machinery, which is most of what is claimed here.
Calibration
The readout temperature is not a constant. It is predicted per decision from the physical state the particle ended in:
log τ = b + w · [ log(1 + residual kinetic energy),
log(1 + distance to the nearest well bottom),
fraction of the step budget used,
log (number of options) ]
A particle still in motion, or stopped far from any well, is an uncertain decision and gets a hotter temperature. Each adapter route carries its own calibration vector. On top of this, every answer carries a conformal set at a chosen risk level α, plus abstain and unbound flags for cases the engine will not stand behind.
Usage
from vega_api import Vega
v = Vega("vega-08b-public-intents", backbone="Qwen/Qwen3.5-0.8B", device="cuda")
out = v.decide(
{"from": "billing@acme.com", "subject": "Invoice overdue", "body": "..."},
{"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments and invoices",
"technical": "product faults",
"sales": "new business"}},
"churn": {"type": "boolean", "instructions": "Is this customer at risk of leaving?",
"criteria": {"true": "shows intent to cancel", "false": "no such signal"}}},
alpha="0.1")
print(out["answers"]["team"]) # choice, probabilities, conformal set, abstain, gate, adapter
Both questions are answered from one read of the state.
alphasets the conformal risk level.views=2or3re-reads the state with options reversed and in a second surface form, averaging the probabilities, steadier answers at roughlyviews ×the cost.adapters="off"runs the base engine alone; the adapters are load-bearing, so expect a drop.
Which readout answers
vegaml 0.7.0 and later default to mode="auto": a question is answered by a head fitted on your
own labelled examples where it has one worth using, and by the physics engine everywhere else. With
nothing fitted, "auto" is the engine, conformal sets and abstain flags included, so it is what
every figure on this card measures. Each answer carries readout, either "engine" or "ttt", so
a mixed result is never ambiguous.
import vegaml
v = vegaml.load() # 0.8B, mode="auto"
out = v.decide(state, questions) # engine, since nothing is fitted
print(out["answers"]["team"]["readout"]) # 'engine'
v.fit(examples, questions) # 20 or so labelled examples
print(v.decide(state, questions)["answers"]["team"]["readout"]) # 'ttt'
fit cross-validates each head and discards any that fails to beat chance, and "auto" uses only
heads whose accuracy is known to beat it. A fitted head is sharper than the engine on the label
space it was fitted for and carries no calibration and no abstain flag, which is the trade readout
makes visible. Pass mode="engine" to reproduce this card's numbers regardless of what is fitted.
Files
| file | |
|---|---|
engine.safetensors |
the trained physics engine |
adapters/ |
two gated low-rank adapters with their calibration and conformal tables |
vega_config.json |
taps, limits, step budget, memory-reader settings |
vega_api.py, vega_common.py |
the inference path; no other runtime needed |
vega_vision.py |
the image read path |
The backbone is fetched from the Hub and verified against a pinned SHA-256: a mismatched or silently updated base model is refused rather than used.
Scope of these figures
The results above are the areas where Vega leads. They are measured, reproducible and item-paired against a live Jev 1.13.0, but they are a selected set, not a full evaluation. Jev leads on most of the benchmarks we ran, particularly reranking and multi-step reasoning over long documents. Ask for the complete per-benchmark breakdown if you need it for a deployment decision.
- Downloads last month
- -
