Zeno — the Surprise Intelligence Model

Does this forecast deserve its confidence? Give Zeno several forecasts for one question. It returns a corrected forecast with a 90% range and, where it has learned to, the chance that each confident forecaster is about to be confidently wrong.

Zeno never replaces the forecasters. For each source it serves a correction only where that correction beat the source's own reference on two later periods it never trained on. Everywhere else it returns the reference, clearly labelled.

Current training run: v0.57 (v057-main-l40sx1-20261007T1138Z).

Try it in your browser · Website · live API POST https://zenodivergent.dev/api/public/v1/check

How Zeno works

One model, any domain. Zeno is a single network with shared weights, trained across every domain at once. Any question that comes as forecasts (ranges, probabilities or numbers) fits. What changes per domain is only permission: in proven mode (the default), Zeno corrects only where it beat the reference on held-out weeks. In experimental mode (mode="experimental"), it applies its full correction anywhere, including domains it has never seen, and labels the answer as untested there.

Results

Quickstart (CPU, about 1 minute)

pip install huggingface_hub onnxruntime numpy
import json, sys
from huggingface_hub import snapshot_download

root = snapshot_download("ZenoDivergent/zeno-divergent-v057",
                         allow_patterns=["serving.json", "onnx/*.onnx", "zeno/*", "examples/*"])
sys.path.insert(0, root)
from zeno import Zeno

z = Zeno.load(root)
question = json.load(open(f"{root}/examples/weather.json"))   # a real past question
print(json.dumps(z.assess(question), indent=1))

This prints the same output as examples/weather.expected.json.

Use this model

Way Code
Python, CPU Zeno.load().assess(question) (Quickstart above) · add mode="experimental" for any domain
safetensors + PyTorch from zeno.torch_model import load; net = load("model.safetensors") (full) or "model-no-track-record.safetensors"
ONNX onnx/{lite,lite_nohist}-s{11,23,37}-best.onnx (proven) · onnx/experimental/lite-s*-best.onnx (full correction)
HTTP, no key curl -X POST https://zenodivergent.dev/api/public/v1/check -H 'content-type: application/json' -d '{"question": {...}, "mode": "proven"}'
MCP (Claude, ChatGPT, Cursor) server URL https://zenodivergent.dev/api/public/v1/mcp · tool zeno_check
AI agents instructions + inventory: https://zenodivergent.dev/agents.md
Browser https://huggingface.co/spaces/mbarbosa1/zeno-try-it

Your own question

question = {
  "source": "ensemble-forecasts",                 # which kind of forecasts these are (table below)
  "decision_at": "2026-10-08T00:00:00Z",          # when you need the answer
  "question": {"site_id": "lisbon", "channel": "temperature_2m", "unit": "°C",
               "target_at": "2026-10-09T00:00:00Z"},
  "inputs": {"panel": [                            # one entry per forecaster
    {"forecaster_id": "ens:ecmwf", "available_at": "2026-10-07T23:00:00Z",
     "forecast_type": "quantiles_05_50_95", "q05": 15.1, "q50": 16.4, "q95": 17.0},
    {"forecaster_id": "ens:gefs",  "available_at": "2026-10-07T23:00:00Z",
     "forecast_type": "quantiles_05_50_95", "q05": 14.0, "q50": 16.9, "q95": 19.5},
    {"forecaster_id": "ens:icon",  "available_at": "2026-10-07T23:00:00Z",
     "forecast_type": "quantiles_05_50_95", "q05": 15.5, "q50": 17.8, "q95": 20.1}
  ]}
}
z.assess(question)
{"served_as": "full",
 "reference": {"value": 16.9, "rule": "panel median"},
 "forecast": {"mean": 16.74, "sd": 1.49, "range90": [14.30, 19.19]},
 "confident_forecasters": ["ens:ecmwf"],
 "confidently_wrong": {"ens:ecmwf": 0.566, "ens:icon": 1e-06, "ens:gefs": 1e-06}}

In plain words: the three forecasters' middle value is 16.9 °C. Zeno moves it to 16.7 °C, with 90% between 14.3 and 19.2. ECMWF is the only confident forecaster (its range is narrower than the others'), and Zeno gives it a 57% chance that the outcome lands outside that range.

Input

Field Required Meaning
source yes Which kind of forecasts these are. See the table below. Unknown sources return the reference.
decision_at yes ISO time. Anything published or resolved after this time is ignored.
question no Description: site, variable, unit, target time. Used as task context.
inputs.panel[] yes One per forecaster: forecaster_id, available_at, forecast_type, and values.
forecast_type yes quantiles_05_50_95 (give q05, q50, q95), binary_probability (give probability), or point (give q50).
panel[].history no The forecaster's track record: n, coverage_90, mean_signed_error, hours_since_last_resolution, and experiences[] (each with label_available_at, signed_error, covered_90, width_90, missed). Only records resolved before decision_at are used. See examples/ for full examples.
inputs.pair_dependence.pairs[] no How often two forecasters missed together: a, b, matched_outcomes, both_missed, a_missed, b_missed.

Output

Field Meaning
served_as full, no_track_record (the version that ignores track records), or reference (Zeno did not prove itself here; use the reference).
reference What Zeno is compared with: the panel median (numbers) or the panel's mean probability (yes/no).
forecast Numbers: mean, sd, range90 in the question's own units. Yes/no: probability_yes.
confident_forecasters Forecasters whose range is narrower than the panel's median range (or yes/no forecasts at ≥85% / ≤15%).
confidently_wrong For each forecaster, the chance its range misses the outcome under Zeno's forecast. Non-confident forecasters get 1e-06, because they can't be confidently wrong. Only given where this warning passed its test (ensemble weather).
forecasters, kind, training_run The order used, number or yes/no, and the exact run.

What is served, per source

source Served as Result on later held-out weeks (lower error)
crypto (or coinbase-*) full distribution 11.7% better than the reference
ensemble-forecasts full distribution 5.6% better; warning 9.8% better than a tuned gradient-boosted model
weather-deterministic no_track_record distribution 6.2% better
hubverse_cdc_covid19 no_track_record distribution 11.8% better
sea no_track_record distribution 3.2% better
hubverse_cdc_rsv, hubverse_flusight, markets, forecastbench, numinous reference no proven gain, so the reference is returned

Worked examples (real past questions, in examples/)

File What happened
weather.json Sydney temperature, 24 h ahead. Panel middle 26.5 °C; Zeno 25.1 °C; actual 22.2 °C. Zeno flagged two confident forecasters. UKMO (66%) did miss. ECMWF (88%) did not: a false alarm. Shown as-is.
crypto.json ETH hourly close. Panel middle 4,431; Zeno 4,417 (90%: 4,056–4,779); actual 4,614.
market.json A prediction-market question. Zeno returns the market's own probability (served_as: reference).

Model at a glance

Parameters 9,160,161 (9.16M)
Architecture Transformer over the panel of forecasters (6 layers, width 320, 8 heads), with attention that knows which pairs miss together, plus a GRU that reads each forecaster's track record
Training 1x NVIDIA L40S · 30 epochs behaviour pretraining (including 339,667 dated forecasts from 323 Numinous AI forecasters, learn-only) + 6 epochs task training · 3 seeds, combined as an equal mixture
Formats model.safetensors, model-no-track-record.safetensors (seed 11; all seeds in research/weights/) · onnx/ (opset 17, used by zeno) · research/weights/*.pt
Checks ONNX = PyTorch to < 5e-7 (onnx/export-check.json). zeno.assess reproduces the run's own backtest predictions on 40 real questions per source (research/parity-check.json).

How it was evaluated

Every source is split in time: train → dev → confirmation → a sealed final period that has not been scored yet. A source is served only if Zeno beat the reference on both dev and confirmation (upper 95% bound of the paired block-bootstrap difference below 0). Numbers are scored by CRPS and yes/no questions by log loss. Warnings are scored by Brier score against the strongest comparator (here, a tuned gradient-boosted model). Pass/fail files: research/gates/; run certificate: research/provenance/run-certificate.json.

Intended use and limits

  • For: checking whether a set of forecasts deserves its confidence; research on combining forecasts; early warning of confident failures in the sources listed above.
  • Not for: sources it wasn't trained on (you get the reference), or as the only basis for financial, medical or safety decisions.
  • No gain on markets, RSV or FluSight. ForecastBench and Numinous could not be scored.
  • 90% ranges are too narrow. Across development and confirmation they covered about 81% of outcomes on crypto, 85% on ensemble weather and 87–89% on the other served sources. Better scores do not yet mean reliable uncertainty; the next run's gate requires coverage.
  • The API and this package drop forecasts made after decision_at, refuse probabilities outside 0–1, merge duplicate forecasters, and give yes/no questions the reference (no served source was evaluated on them).
  • The warning is proven only on ensemble weather, and it gives false alarms (see the example).
  • This run cannot credit the warning gain to Numinous alone. A matched with/without-Numinous comparison will.
  • Nothing here measures live performance. The sealed period is still unscored.

Advanced: raw ONNX

Inputs: X B×P×20, mask B×P, H B×P×K×7, HM B×P×K, J B×9, PR B×P×P×7, ref_logit B×P, sidx B (index in serving.json sources; last index = unseen source), ctx B (task-context hash bucket). Build them with Zeno.features(question); it calls the exact training feature code in zeno/features.py.

The served output is dist_delta (B×2). Mean = reference + dist_delta[0]·sigma0·scale; sd = sigma0·exp(clip(dist_delta[1], −4, 4))·scale, with sigma0 per source from serving.json. The second output, warn_logit, is a training-time head and is not the served warning. Served warnings come from the forecast distribution (see implied_miss in zeno/__init__.py).

Live hosted model: POST https://zenodivergent.dev/api/public/v1/check and MCP tool zeno_check. These run this exact package. If the model host is unreachable, they return the labelled panel reference.

Files

model.safetensors, model-no-track-record.safetensors
zeno/            one-call inference (numpy + onnxruntime); zeno/torch_model.py loads safetensors
assets/          card figures
onnx/            6 proven ONNX files + onnx/experimental/ (3, full correction)
examples/        3 real questions + expected outputs
serving.json     served version per source, references, fingerprints
config.json      architecture and sizes
research/        PyTorch/safetensors weights, backtest predictions, gates, provenance, exact training code

Citation

@misc{zeno2026surprise,
  title  = {Zeno: the Surprise Intelligence Model (training run v0.57)},
  author = {Zeno Divergent},
  year   = {2026},
  url    = {https://huggingface.co/ZenoDivergent/zeno-divergent-v057}
}

Licence: see LICENSE (free for research and non-commercial use; commercial use by agreement).

Downloads last month
43
Safetensors
Model size
9.16M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using ZenoDivergent/zeno-divergent-v057 1