Zeno — the Surprise Intelligence Model
Does this forecast deserve its confidence? Give Zeno several forecasts for one question. It returns a corrected forecast with a 90% range and, where it has learned to, the chance that each confident forecaster is about to be confidently wrong.
Zeno never replaces the forecasters. For each source it serves a correction only where that correction beat the source's own reference on two later periods it never trained on. Everywhere else it returns the reference, clearly labelled.
Current training run: v0.57 (v057-main-l40sx1-20261007T1138Z).
Try it in your browser · Website · live API POST https://zenodivergent.dev/api/public/v1/check
One model, any domain. Zeno is a single network with shared weights, trained across every domain at once. Any question that comes as forecasts (ranges, probabilities or numbers) fits. What changes per domain is only permission: in proven mode (the default), Zeno corrects only where it beat the reference on held-out weeks. In experimental mode (mode="experimental"), it applies its full correction anywhere, including domains it has never seen, and labels the answer as untested there.
Quickstart (CPU, about 1 minute)
pip install huggingface_hub onnxruntime numpy
import json, sys
from huggingface_hub import snapshot_download
root = snapshot_download("ZenoDivergent/zeno-divergent-v057",
allow_patterns=["serving.json", "onnx/*.onnx", "zeno/*", "examples/*"])
sys.path.insert(0, root)
from zeno import Zeno
z = Zeno.load(root)
question = json.load(open(f"{root}/examples/weather.json")) # a real past question
print(json.dumps(z.assess(question), indent=1))
This prints the same output as examples/weather.expected.json.
Use this model
| Way | Code |
|---|---|
| Python, CPU | Zeno.load().assess(question) (Quickstart above) · add mode="experimental" for any domain |
| safetensors + PyTorch | from zeno.torch_model import load; net = load("model.safetensors") (full) or "model-no-track-record.safetensors" |
| ONNX | onnx/{lite,lite_nohist}-s{11,23,37}-best.onnx (proven) · onnx/experimental/lite-s*-best.onnx (full correction) |
| HTTP, no key | curl -X POST https://zenodivergent.dev/api/public/v1/check -H 'content-type: application/json' -d '{"question": {...}, "mode": "proven"}' |
| MCP (Claude, ChatGPT, Cursor) | server URL https://zenodivergent.dev/api/public/v1/mcp · tool zeno_check |
| AI agents | instructions + inventory: https://zenodivergent.dev/agents.md |
| Browser | https://huggingface.co/spaces/mbarbosa1/zeno-try-it |
Your own question
question = {
"source": "ensemble-forecasts", # which kind of forecasts these are (table below)
"decision_at": "2026-10-08T00:00:00Z", # when you need the answer
"question": {"site_id": "lisbon", "channel": "temperature_2m", "unit": "°C",
"target_at": "2026-10-09T00:00:00Z"},
"inputs": {"panel": [ # one entry per forecaster
{"forecaster_id": "ens:ecmwf", "available_at": "2026-10-07T23:00:00Z",
"forecast_type": "quantiles_05_50_95", "q05": 15.1, "q50": 16.4, "q95": 17.0},
{"forecaster_id": "ens:gefs", "available_at": "2026-10-07T23:00:00Z",
"forecast_type": "quantiles_05_50_95", "q05": 14.0, "q50": 16.9, "q95": 19.5},
{"forecaster_id": "ens:icon", "available_at": "2026-10-07T23:00:00Z",
"forecast_type": "quantiles_05_50_95", "q05": 15.5, "q50": 17.8, "q95": 20.1}
]}
}
z.assess(question)
{"served_as": "full",
"reference": {"value": 16.9, "rule": "panel median"},
"forecast": {"mean": 16.74, "sd": 1.49, "range90": [14.30, 19.19]},
"confident_forecasters": ["ens:ecmwf"],
"confidently_wrong": {"ens:ecmwf": 0.566, "ens:icon": 1e-06, "ens:gefs": 1e-06}}
In plain words: the three forecasters' middle value is 16.9 °C. Zeno moves it to 16.7 °C, with 90% between 14.3 and 19.2. ECMWF is the only confident forecaster (its range is narrower than the others'), and Zeno gives it a 57% chance that the outcome lands outside that range.
Input
| Field | Required | Meaning |
|---|---|---|
source |
yes | Which kind of forecasts these are. See the table below. Unknown sources return the reference. |
decision_at |
yes | ISO time. Anything published or resolved after this time is ignored. |
question |
no | Description: site, variable, unit, target time. Used as task context. |
inputs.panel[] |
yes | One per forecaster: forecaster_id, available_at, forecast_type, and values. |
forecast_type |
yes | quantiles_05_50_95 (give q05, q50, q95), binary_probability (give probability), or point (give q50). |
panel[].history |
no | The forecaster's track record: n, coverage_90, mean_signed_error, hours_since_last_resolution, and experiences[] (each with label_available_at, signed_error, covered_90, width_90, missed). Only records resolved before decision_at are used. See examples/ for full examples. |
inputs.pair_dependence.pairs[] |
no | How often two forecasters missed together: a, b, matched_outcomes, both_missed, a_missed, b_missed. |
Output
| Field | Meaning |
|---|---|
served_as |
full, no_track_record (the version that ignores track records), or reference (Zeno did not prove itself here; use the reference). |
reference |
What Zeno is compared with: the panel median (numbers) or the panel's mean probability (yes/no). |
forecast |
Numbers: mean, sd, range90 in the question's own units. Yes/no: probability_yes. |
confident_forecasters |
Forecasters whose range is narrower than the panel's median range (or yes/no forecasts at ≥85% / ≤15%). |
confidently_wrong |
For each forecaster, the chance its range misses the outcome under Zeno's forecast. Non-confident forecasters get 1e-06, because they can't be confidently wrong. Only given where this warning passed its test (ensemble weather). |
forecasters, kind, training_run |
The order used, number or yes/no, and the exact run. |
What is served, per source
source |
Served as | Result on later held-out weeks (lower error) |
|---|---|---|
crypto (or coinbase-*) |
full | distribution 11.7% better than the reference |
ensemble-forecasts |
full | distribution 5.6% better; warning 9.8% better than a tuned gradient-boosted model |
weather-deterministic |
no_track_record | distribution 6.2% better |
hubverse_cdc_covid19 |
no_track_record | distribution 11.8% better |
sea |
no_track_record | distribution 3.2% better |
hubverse_cdc_rsv, hubverse_flusight, markets, forecastbench, numinous |
reference | no proven gain, so the reference is returned |
Worked examples (real past questions, in examples/)
| File | What happened |
|---|---|
weather.json |
Sydney temperature, 24 h ahead. Panel middle 26.5 °C; Zeno 25.1 °C; actual 22.2 °C. Zeno flagged two confident forecasters. UKMO (66%) did miss. ECMWF (88%) did not: a false alarm. Shown as-is. |
crypto.json |
ETH hourly close. Panel middle 4,431; Zeno 4,417 (90%: 4,056–4,779); actual 4,614. |
market.json |
A prediction-market question. Zeno returns the market's own probability (served_as: reference). |
Model at a glance
| Parameters | 9,160,161 (9.16M) |
| Architecture | Transformer over the panel of forecasters (6 layers, width 320, 8 heads), with attention that knows which pairs miss together, plus a GRU that reads each forecaster's track record |
| Training | 1x NVIDIA L40S · 30 epochs behaviour pretraining (including 339,667 dated forecasts from 323 Numinous AI forecasters, learn-only) + 6 epochs task training · 3 seeds, combined as an equal mixture |
| Formats | model.safetensors, model-no-track-record.safetensors (seed 11; all seeds in research/weights/) · onnx/ (opset 17, used by zeno) · research/weights/*.pt |
| Checks | ONNX = PyTorch to < 5e-7 (onnx/export-check.json). zeno.assess reproduces the run's own backtest predictions on 40 real questions per source (research/parity-check.json). |
How it was evaluated
Every source is split in time: train → dev → confirmation → a sealed final period that has not been scored yet. A source is served only if Zeno beat the reference on both dev and confirmation (upper 95% bound of the paired block-bootstrap difference below 0). Numbers are scored by CRPS and yes/no questions by log loss. Warnings are scored by Brier score against the strongest comparator (here, a tuned gradient-boosted model). Pass/fail files: research/gates/; run certificate: research/provenance/run-certificate.json.
Intended use and limits
- For: checking whether a set of forecasts deserves its confidence; research on combining forecasts; early warning of confident failures in the sources listed above.
- Not for: sources it wasn't trained on (you get the reference), or as the only basis for financial, medical or safety decisions.
- No gain on markets, RSV or FluSight. ForecastBench and Numinous could not be scored.
- 90% ranges are too narrow. Across development and confirmation they covered about 81% of outcomes on crypto, 85% on ensemble weather and 87–89% on the other served sources. Better scores do not yet mean reliable uncertainty; the next run's gate requires coverage.
- The API and this package drop forecasts made after
decision_at, refuse probabilities outside 0–1, merge duplicate forecasters, and give yes/no questions the reference (no served source was evaluated on them). - The warning is proven only on ensemble weather, and it gives false alarms (see the example).
- This run cannot credit the warning gain to Numinous alone. A matched with/without-Numinous comparison will.
- Nothing here measures live performance. The sealed period is still unscored.
Advanced: raw ONNX
Inputs: X B×P×20, mask B×P, H B×P×K×7, HM B×P×K, J B×9, PR B×P×P×7, ref_logit B×P, sidx B (index in serving.json sources; last index = unseen source), ctx B (task-context hash bucket). Build them with Zeno.features(question); it calls the exact training feature code in zeno/features.py.
The served output is dist_delta (B×2). Mean = reference + dist_delta[0]·sigma0·scale; sd = sigma0·exp(clip(dist_delta[1], −4, 4))·scale, with sigma0 per source from serving.json. The second output, warn_logit, is a training-time head and is not the served warning. Served warnings come from the forecast distribution (see implied_miss in zeno/__init__.py).
Live hosted model: POST https://zenodivergent.dev/api/public/v1/check and MCP tool zeno_check. These run this exact package. If the model host is unreachable, they return the labelled panel reference.
Files
model.safetensors, model-no-track-record.safetensors
zeno/ one-call inference (numpy + onnxruntime); zeno/torch_model.py loads safetensors
assets/ card figures
onnx/ 6 proven ONNX files + onnx/experimental/ (3, full correction)
examples/ 3 real questions + expected outputs
serving.json served version per source, references, fingerprints
config.json architecture and sizes
research/ PyTorch/safetensors weights, backtest predictions, gates, provenance, exact training code
Citation
@misc{zeno2026surprise,
title = {Zeno: the Surprise Intelligence Model (training run v0.57)},
author = {Zeno Divergent},
year = {2026},
url = {https://huggingface.co/ZenoDivergent/zeno-divergent-v057}
}
Licence: see LICENSE (free for research and non-commercial use; commercial use by agreement).
- Downloads last month
- 43

