smollm2-360m-boolq-calibration-sft-v2

SFT warm-start checkpoint (v2) from slm-calibration-rlcd โ€” a project training small LMs to make calibrated yes/no decisions, inspired by TypeSafe AI's "Jev" model / "RLCD" (Reinforcement Learning for Calibrated Decisions). Full project: 25b3nk/slm-calibration-rlcd.

This is SFT-only, pre-GRPO โ€” the stage that teaches the tagged output format and baseline task competence, before GRPO teaches calibration on top of it. GRPO has not been run against this checkpoint yet; see smollm2-360m-boolq-calibration-grpo2-v1 for the GRPO-trained model built on top of this one.

Why v2 exists

The prior SFT checkpoint (sft-warmup-v1) used two fixed confidence-target constants (0.65 for truncated-passage examples, 0.92 for full-context) and a prompt that never explicitly asked the model for a confidence value. Every GRPO run on top of it collapsed: confidence output converged to just 1-2 of those exact SFT constants regardless of training steps or learning rate โ€” not real calibration, just parroting the SFT heuristic. Root cause: v1's SFT prior had near-zero output entropy on the confidence dimension, so GRPO had nothing to diversify.

v2 fixes the root cause:

  • Confidence targets are now sampled from ranges (0.45-0.80 for truncated-passage examples, 0.80-0.98 for full-context) instead of two fixed constants.
  • The prompt now explicitly asks for confidence: "Answer the question with yes or no, and state how confident you are (0.0-1.0) based on how clearly the passage supports your answer."

What it does

Given a BoolQ-style passage + yes/no question, outputs:

<answer>yes|no</answer><confidence>0.XX</confidence>

Training

  • Base: HuggingFaceTB/SmolLM2-360M-Instruct
  • Full fine-tune (not LoRA), fp32 weights + fp16 mixed-precision autocast
  • 3000 examples from google/boolq train split (with random passage truncation on a subset of examples to induce genuine uncertainty)
  • 2 epochs, lr 2e-5, effective batch size 16 (batch 2 x grad-accum 8), gradient checkpointing โ€” same recipe/hyperparameters as sft-warmup-v1; only the confidence-target sampling and prompt changed
  • Confidence targets sampled from ranges (0.45-0.80 truncated, 0.80-0.98 full-context), not fixed constants โ€” this is the key change from v1, made specifically so a downstream GRPO stage has confidence-dimension entropy to work with

Eval (500-example BoolQ validation slice, seed 0)

Metric sft-warmup-v1 This checkpoint (v2)
Format compliance 1.0 1.0
Accuracy 0.792 0.804
Brier score 0.182 0.167
ECE 0.130 0.090

v2 improves on v1 on every metric, even before any GRPO. ECE 0.090 is still not well-calibrated (0 = perfect) โ€” that's the gap the next stage (GRPO) targets, and this time GRPO actually has room to diversify confidence output instead of collapsing to a constant (see grpo2-v1's card for that result).

Full design + ADRs

Downloads last month
232
Safetensors
Model size
0.4B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for 25b3nk/smollm2-360m-boolq-calibration-sft-v2

Finetuned
(183)
this model

Dataset used to train 25b3nk/smollm2-360m-boolq-calibration-sft-v2