vlm-calibration-prefix-vqav2

A small non-autoregressive yes/no visual question answering model built for research on calibrated decisions. It answers a yes/no question about an image in one forward pass, with no generation and no verbalised confidence. The confidence is the softmax over two token logits followed by a single post-hoc temperature.

How it works

  • Image. A frozen SigLIP2-base encodes the image into its pooled vector (pooler_output, 768-d).

  • Prefix token. That vector becomes one prefix token, prepended to the input embeddings of ModernBERT-base:

    prefix = anchor + g_max * tanh(alpha) * LN_out(proj(LN_in(v)))
    
    • anchor is a learned token, initialised from the " image" embedding.
    • alpha is a scalar gate, initialised to 0, so the model starts as exactly the text-only model.
    • g_max caps the vision term's norm at the anchor's norm.
  • Readout. The prompt is "{question} Answer: [MASK]". The answer is read with ModernBERT's own pretrained MLM head at [MASK], comparing the logits of the " yes" and " no" tokens. There is no new classification head.

  • Calibration. Probabilities are calibrated by one scalar temperature T per checkpoint, fitted on a held-out calib split:

    p_yes = sigmoid((logit_yes - logit_no) / T)
    
  • What is trained. ModernBERT (encoder and MLM head), the anchor, the projection and the gate are fine-tuned. SigLIP2 stays frozen and is not included in this repo; it is loaded by ID.

What is in this repo

The repo root is the default model: seed 0 of the paired (image + question), hard-label configuration, trained on the full 214,476-row pool, 0.6791 eval accuracy. Eight more checkpoints of the same architecture live in subfolders, so the seed means, the blind (question-only) reference and the soft-label ablation in the results below can be reproduced and probed. Nine checkpoints in total, each about 601 MB in fp32.

path arm labels seed eval acc ECE pre → post T T gate tanh(alpha) POPE AUROC (adv / pop / rand)
/ (root) = paired-hard/seed0 (default) paired hard 0 0.6791 0.0810 → 0.0096 1.81 0.5999 0.676 / 0.720 / 0.742
paired-hard/seed1/ paired hard 1 0.6758 0.0945 → 0.0189 1.95 0.7066 0.686 / 0.717 / 0.719
paired-hard/seed2/ paired hard 2 0.6666 0.0495 → 0.0199 1.26 -0.5785 0.634 / 0.645 / 0.706
blind-hard/seed0/ blind hard 0 0.5536 0.0118 → 0.0076 1.12 n/a 0.577 / 0.537 / 0.564
blind-hard/seed1/ blind hard 1 0.5521 0.0212 → 0.0102 1.25 n/a 0.617 / 0.613 / 0.612
blind-hard/seed2/ blind hard 2 0.5477 0.0479 → 0.0111 1.84 n/a 0.543 / 0.430 / 0.573
paired-soft/seed0/ paired soft 0 0.6806 0.0585 → 0.0134 1.42 -0.5831 0.691 / 0.726 / 0.752
paired-soft/seed1/ paired soft 1 0.6740 0.0533 → 0.0193 1.37 -0.4974 0.698 / 0.741 / 0.720
paired-soft/seed2/ paired soft 2 0.6856 0.0175 → 0.0090 1.12 -0.6320 0.710 / 0.743 / 0.804
  • Accuracy is exact match on the VQAv2 val2014 yes/no image-disjoint eval split (n = 8,930); see Results for the definition. ECE is top-label, 15 equal-width bins. T is the hard-label temperature in every row (for the soft models too, so that the rows are comparable); each soft config also records temperature_T_soft, which modeling.py does not use.
  • The blind checkpoints have the same architecture, but the vision term is zeroed: the prefix is the learned anchor alone and no image is read. They are the reference for "does vision help". Their gate is unused, so the table shows n/a.
  • The sign of the gate is arbitrary (tanh(alpha) multiplies a LayerNorm output that the projection can flip), so only its size is meaningful.
  • The three blind seeds have no image-shuffle metric. They were not run with one.

Every folder (root included) has the same three kinds of files:

file contents
model.safetensors All trained weights in fp32: the fine-tuned ModernBERT with its MLM head, plus the anchor, LayerNorm, projection, gate and g_max. The MLM decoder weight is tied to the token embeddings, so it is stored once. Only the model state_dict is published, with no optimizer state.
config.json Architecture, base model IDs and revisions, prefix_k = 1, arm, blind, seed, label_mode, the gate value tanh(alpha), the temperature temperature_T, and that checkpoint's training summary and eval metrics.
modeling.py (Root only, shared by all subfolders.) Self-contained inference code that rebuilds the model and runs it.

Layout:

model.safetensors, config.json      # paired-hard / seed 0 (default)
modeling.py, README.md
paired-hard/seed1/  paired-hard/seed2/
blind-hard/seed0/   blind-hard/seed1/   blind-hard/seed2/
paired-soft/seed0/  paired-soft/seed1/  paired-soft/seed2/

Seed 0 of paired-hard is the root, not a copy: it has no paired-hard/seed0/ folder, so the 601 MB file is stored once.

Usage

Tested on CPU with torch 2.14, transformers 5.17.0 and safetensors 0.8.0. ModernBERT needs a recent transformers; older versions are untested.

# pip install torch transformers safetensors huggingface_hub pillow
import os, sys
from huggingface_hub import hf_hub_download
from PIL import Image

repo = "25b3nk/vlm-calibration-prefix-vqav2"
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "modeling.py")))
from modeling import YesNoVQA

vqa = YesNoVQA.from_pretrained(repo)  # default model (root); also fetches SigLIP2 + the ModernBERT tokenizer/config
image = Image.open("example.jpg")
print(vqa.predict([image], ["Is there a minute hand on the clock?"]))
# -> [{'answer': 'yes', 'margin': ..., 'p_yes_raw': ..., 'p_yes': ...}]   p_yes is temperature-scaled

To load any other variant, pass its folder as subfolder. The same modeling.py serves all of them, and from_pretrained reads the arm and temperature from that folder's config.json:

vqa = YesNoVQA.from_pretrained(repo, subfolder="paired-hard/seed1")
vqa_blind = YesNoVQA.from_pretrained(repo, subfolder="blind-hard/seed0")   # no image read; SigLIP2 is not loaded
vqa_soft = YesNoVQA.from_pretrained(repo, subfolder="paired-soft/seed2")

For a blind checkpoint, predict ignores the images (pass None or any list of the right length). Downloading one subfolder alone fetches about 601 MB.

Notes:

  • predict returns p_yes (calibrated with that checkpoint's T), the uncalibrated p_yes_raw, and the margin logit_yes - logit_no.
  • The training features were computed on JPEG re-encoded images (PIL default quality). predict(..., jpeg_roundtrip=True) (the default) applies the same re-encode. Pass jpeg_roundtrip=False if your files are already such re-encodes.
  • The model needs 1 GB of RAM in fp32 plus SigLIP2 (0.4 GB).

Check against the training-run outputs. The snippet above, loaded from this Hub repo on CPU, was run on 6 examples from the eval split (4 "no" and 2 "yes" predictions, on 6 different images). It was compared with the predictions saved by this checkpoint's training run on GPU, for two inputs:

  • the cached GPU SigLIP2 features;
  • the full CPU image pipeline.

The predicted answers matched in every case. Measured differences are in the "Inference check" section at the end.

Results

All numbers come from the project's experiment logs. Eval is the VQAv2 val2014 yes/no image-disjoint eval split:

  • 8,930 questions, split by md5(image_id) % 100, so no eval image appears in training or calib.
  • It is not test-dev.
  • "Accuracy" is exact match against VQAv2's multiple_choice_answer (yes or no). It is not the standard soft VQA accuracy metric, so it is not comparable to published VQAv2 numbers.
  • The majority-answer-per-question-type floor on this split is 0.5106.

Main model vs blind reference (hard labels, 3 seeds, full 214,476-row pool)

The repo root is paired seed 0. The other seeds and the blind arm are in the subfolders (see the table above).

metric paired (main) blind (no image)
accuracy (seed mean) 0.6738 (0.6791 / 0.6758 / 0.6666) 0.5511 (0.5536 / 0.5521 / 0.5477)
vision gap (paired − blind) +0.1227
accuracy drop when eval images are shuffled across questions 0.142
ECE, before → after temperature scaling 0.0750 → 0.0161 (0.0096 / 0.0189 / 0.0199) 0.0270 → 0.0096
fitted T (s0 / s1 / s2) 1.81 / 1.95 / 1.26 1.12 / 1.25 / 1.84
eval NLL, before → after T 0.6097 → 0.5815 0.6806 → 0.6768
AUROC of p_yes on eval 0.752 0.584
  • ECE is top-label ECE with 15 equal-width bins. T is fitted on a separate 2,911-question calib split by minimising NLL.
  • Accuracy and AUROC do not depend on T.
  • Before scaling, the paired models are over-confident (T > 1).

POPE (object hallucination probe), strict-clean subset

These are margin AUROCs on POPE's three variants, using only the strict-clean subset: 2,490 questions per variant, on 415 COCO val2014 images that appear in no training pool and not in calib. POPE images were excluded from training.

adversarial popular random
paired (seed mean) 0.665 0.694 0.722
blind (seed mean) 0.579 0.527 0.583
paired, per seed (s0 / s1 / s2) 0.676 / 0.686 / 0.634 0.720 / 0.717 / 0.645 0.742 / 0.719 / 0.706

POPE accuracy for the paired model is above chance but modest: 0.603 / 0.613 / 0.653 (adversarial / popular / random) with the default threshold z >= 0.

Scaling curve (same architecture and recipe, 3 seeds per point)

train rows source paired blind gap POPE AUROC paired (adv / pop / rand)
15,000 val2014 0.5510 0.5385 +0.0125 0.561 / 0.612 / 0.478
30,000 val2014 0.5593 0.5385 +0.0208 0.496 / 0.507 / 0.504
47,600 val2014 0.5733 0.5275 +0.0457 0.517 / 0.524 / 0.551
100,000 val2014 + train2014 0.6088 0.5447 +0.0641 0.549 / 0.555 / 0.606
214,476 (this model's configuration) val2014 + train2014 0.6738 0.5511 +0.1227 0.665 / 0.694 / 0.722

How to read the curve:

  • Paired accuracy rises at every step, and the last step is the largest. Blind stays near 0.55, so the growth comes from using the image.
  • POPE grounding only appears at the full scale.
  • Confounds. Optimisation steps grow with the data (3 epochs at every scale), and the data source changes past 47.6K. A same-size source check (47.6K rows from train2014 only, paired 0.5859) was inconclusive. The 100K → 214K step adds only train2014 rows.

Soft labels vs hard labels (the soft models are in paired-soft/)

metric soft (seed mean) hard (seed mean)
accuracy 0.6801 (0.6806 / 0.6740 / 0.6856) 0.6738
ECE before → after T (hard labels) 0.0431 → 0.0139 0.0750 → 0.0161
fitted T (hard labels), s0 / s1 / s2 1.42 / 1.37 / 1.12 1.81 / 1.95 / 1.26
soft-target NLL at each model's own soft temperature 0.5942 0.6028
POPE AUROC (adv / pop / rand) 0.700 / 0.737 / 0.759 0.665 / 0.694 / 0.722

Soft labels are not presented as better. The experiment was pre-registered with ±0.010 decision bars:

  • Accuracy (+0.006) and soft-target NLL (−0.0086) were both neutral by those bars, and the accuracy gain comes mostly from one seed.
  • The main effect is "built-in temperature". Before post-hoc scaling, soft training is much less over-confident. After each model gets its own T, little difference remains.
  • No soft blind arm was trained, so there is no blind-soft/ folder.
  • The hard-label model stays the main model.

Training data

  • VQAv2 yes/no questions (answer_type == "yes/no", majority answer yes or no). There are 214,476 training rows on 89,992 COCO images:
    • 47,600 rows from the VQAv2 val2014 questions (via lmms-lab/VQAv2, validation split). These come from images in the training bucket of the image hash split, with all POPE images removed.
    • 166,876 rows from the official VQAv2 train2014 questions and annotations, with images from lmms-lab-encoder/VQAv2 at revision 32665d3. Two train2014 images that were near-duplicates of an eval image were excluded.
  • The eval and calib splits are also from val2014. They are image-disjoint from training, but they come from the same distribution as part of the training data.
  • Recipe. AdamW, batch size 16, 3 epochs, and a cosine schedule. The learning rates are 2e-5 for ModernBERT and the anchor, 3e-4 for the new layers and 1e-3 for the gate. The kept checkpoint is the epoch end with the best calib accuracy. Training ran on a single T4.
  • Hard labels: the VQAv2 majority answer. Soft labels: q = #yes / (#yes + #no) over the 10 annotator answers.

Intended use

  • Intended: research on calibrated, single-pass decisions. Examples are temperature scaling, reliability analysis and selective prediction.
  • Not intended: any production use. Never use it for decisions about people or anything safety-relevant. It is a research artefact.

Limitations

  • Yes/no questions only. Other question types are out of scope; the readout only compares " yes" and " no".
  • Modest accuracy. It scores about 0.67 on yes/no questions, against a majority-per-question-type floor of 0.5106. Roughly one answer in three is wrong. The question-only reference already reaches 0.55, so part of the behaviour is language prior.
  • One pooled image vector. Fine spatial detail, counting and small objects are poorly served. POPE accuracy is only 0.60–0.65.
  • The eval is a val2014 split, not test-dev. The accuracy metric is exact match on the majority answer, not the official VQA metric.
  • Calibration is distribution-specific. T was fitted on VQAv2 val2014 calib questions. It may not transfer to other image or question distributions; refit it on your own data if calibration matters.
  • Small seed sample. Three seeds per configuration; seed-to-seed spread is about 0.5–1 point of accuracy.
  • Inherited biases. The model inherits the biases of COCO, VQAv2 annotators, SigLIP2 and ModernBERT.
  • Training instability. Gradient-norm spikes were observed during training (reported, not fixed). This checkpoint (seed 0) had one, of 52.7, in its selected epoch.

License

  • Weights and code in this repo: Apache-2.0. Both base models are Apache-2.0, checked on their Hugging Face pages: google/siglip2-base-patch16-224 and answerdotai/ModernBERT-base. Apache-2.0 keeps the weights under the same terms as the model they fine-tune (ModernBERT) and the encoder they depend on (SigLIP2).
  • Training data. The VQAv2 annotations are CC BY 4.0, and are attributed below. The COCO images are subject to the COCO terms of use and the licenses of the individual Flickr images. No images are redistributed here.
  • POPE was used for evaluation only.

Code

The training code and experiment logs are in a personal repository that is not public. modeling.py in this repo contains everything needed for inference.

Citations

@inproceedings{goyal2017vqav2,
  title     = {Making the {V} in {VQA} Matter: Elevating the Role of Image Understanding in Visual Question Answering},
  author    = {Goyal, Yash and Khot, Tejas and Summers-Stay, Douglas and Batra, Dhruv and Parikh, Devi},
  booktitle = {CVPR},
  year      = {2017}
}

@inproceedings{li2023pope,
  title     = {Evaluating Object Hallucination in Large Vision-Language Models},
  author    = {Li, Yifan and Du, Yifan and Zhou, Kun and Wang, Jinpeng and Zhao, Wayne Xin and Wen, Ji-Rong},
  booktitle = {EMNLP},
  year      = {2023}
}

@article{tschannen2025siglip2,
  title   = {{SigLIP} 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features},
  author  = {Tschannen, Michael and Gritsenko, Alexey and Wang, Xiao and Naeem, Muhammad Ferjad and Alabdulmohsin, Ibrahim and Parthasarathy, Nikhil and Evans, Talfan and Beyer, Lucas and Xia, Ye and Mustafa, Basil and H{\'e}naff, Olivier and Harmsen, Jeremiah and Steiner, Andreas and Zhai, Xiaohua},
  journal = {arXiv preprint arXiv:2502.14786},
  year    = {2025}
}

@article{warner2024modernbert,
  title   = {Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference},
  author  = {Warner, Benjamin and Chaffin, Antoine and Clavi{\'e}, Benjamin and Weller, Orion and Hallstr{\"o}m, Oskar and Taghadouini, Said and Gallagher, Alexis and Biswas, Raja and Ladhak, Faisal and Aarsen, Tom and Cooper, Nathan and Adams, Griffin and Howard, Jeremy and Poli, Iacopo},
  journal = {arXiv preprint arXiv:2412.13663},
  year    = {2024}
}

@inproceedings{lin2014coco,
  title     = {Microsoft {COCO}: Common Objects in Context},
  author    = {Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Hays, James and Perona, Pietro and Ramanan, Deva and Doll{\'a}r, Piotr and Zitnick, C. Lawrence},
  booktitle = {ECCV},
  year      = {2014}
}

@inproceedings{guo2017calibration,
  title     = {On Calibration of Modern Neural Networks},
  author    = {Guo, Chuan and Pleiss, Geoff and Sun, Yu and Weinberger, Kilian Q.},
  booktitle = {ICML},
  year      = {2017}
}

Inference check

The README snippet above was run verbatim against this Hub repo, on CPU, with torch 2.14 and transformers 5.17.0. It was applied to 6 eval-split questions on 6 different images: 4 predicted "no" and 2 predicted "yes".

The results were compared with the prob_yes values saved by the original training run on a T4 GPU (root model):

input path answers match max |Δ p_yes_raw|
cached SigLIP2 features from the training run (vision_emb=) 6 / 6 1.2e-6
CPU SigLIP2 on the stored (already JPEG re-encoded) eval images, jpeg_roundtrip=False 6 / 6 1.3e-6
same images with the default extra jpeg_roundtrip=True (re-encodes a second time) 6 / 6 3.8e-3

Loading SigLIP2 with SiglipVisionModel prints a warning that the text-tower weights (text_model.*, logit_scale, …) are "UNEXPECTED". This is expected, because only the vision tower is used.

The eight subfolder models were each checked the same way after upload. The check loaded that subfolder from the Hub with modeling.py (subfolder=...), on CPU, and compared its answers with that checkpoint's saved predictions on the same 6 eval questions:

subfolder answers match (cached features) answers match (CPU SigLIP2 on images) max |Δ p_yes_raw|
paired-hard/seed1 6 / 6 6 / 6 9.8e-7
paired-hard/seed2 6 / 6 6 / 6 4.2e-6
blind-hard/seed0 6 / 6 n/a (no image read) 8.3e-7
blind-hard/seed1 6 / 6 n/a 9.6e-7
blind-hard/seed2 6 / 6 n/a 1.2e-6
paired-soft/seed0 6 / 6 6 / 6 1.7e-6
paired-soft/seed1 6 / 6 6 / 6 1.5e-6
paired-soft/seed2 6 / 6 6 / 6 1.5e-6
Downloads last month
53
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 25b3nk/vlm-calibration-prefix-vqav2

Finetuned
(1515)
this model

Dataset used to train 25b3nk/vlm-calibration-prefix-vqav2

Papers for 25b3nk/vlm-calibration-prefix-vqav2