smollm2-360m-boolq-calibration-sft-v2
SFT warm-start checkpoint (v2) from slm-calibration-rlcd โ a project training small LMs to make calibrated yes/no decisions, inspired by TypeSafe AI's "Jev" model / "RLCD" (Reinforcement Learning for Calibrated Decisions). Full project: 25b3nk/slm-calibration-rlcd.
This is SFT-only, pre-GRPO โ the stage that teaches the tagged output format and baseline task competence, before GRPO teaches calibration on top of it. GRPO has not been run against this checkpoint yet; see smollm2-360m-boolq-calibration-grpo2-v1 for the GRPO-trained model built on top of this one.
Why v2 exists
The prior SFT checkpoint (sft-warmup-v1) used two fixed confidence-target constants (0.65 for truncated-passage examples, 0.92 for full-context) and a prompt that never explicitly asked the model for a confidence value. Every GRPO run on top of it collapsed: confidence output converged to just 1-2 of those exact SFT constants regardless of training steps or learning rate โ not real calibration, just parroting the SFT heuristic. Root cause: v1's SFT prior had near-zero output entropy on the confidence dimension, so GRPO had nothing to diversify.
v2 fixes the root cause:
- Confidence targets are now sampled from ranges (0.45-0.80 for truncated-passage examples, 0.80-0.98 for full-context) instead of two fixed constants.
- The prompt now explicitly asks for confidence: "Answer the question with yes or no, and state how confident you are (0.0-1.0) based on how clearly the passage supports your answer."
What it does
Given a BoolQ-style passage + yes/no question, outputs:
<answer>yes|no</answer><confidence>0.XX</confidence>
Training
- Base: HuggingFaceTB/SmolLM2-360M-Instruct
- Full fine-tune (not LoRA), fp32 weights + fp16 mixed-precision autocast
- 3000 examples from google/boolq train split (with random passage truncation on a subset of examples to induce genuine uncertainty)
- 2 epochs, lr 2e-5, effective batch size 16 (batch 2 x grad-accum 8), gradient checkpointing โ same recipe/hyperparameters as sft-warmup-v1; only the confidence-target sampling and prompt changed
- Confidence targets sampled from ranges (0.45-0.80 truncated, 0.80-0.98 full-context), not fixed constants โ this is the key change from v1, made specifically so a downstream GRPO stage has confidence-dimension entropy to work with
Eval (500-example BoolQ validation slice, seed 0)
| Metric | sft-warmup-v1 | This checkpoint (v2) |
|---|---|---|
| Format compliance | 1.0 | 1.0 |
| Accuracy | 0.792 | 0.804 |
| Brier score | 0.182 | 0.167 |
| ECE | 0.130 | 0.090 |
v2 improves on v1 on every metric, even before any GRPO. ECE 0.090 is still not well-calibrated (0 = perfect) โ that's the gap the next stage (GRPO) targets, and this time GRPO actually has room to diversify confidence output instead of collapsing to a constant (see grpo2-v1's card for that result).
Full design + ADRs
- Downloads last month
- 232
Model tree for 25b3nk/smollm2-360m-boolq-calibration-sft-v2
Base model
HuggingFaceTB/SmolLM2-360M