SolveAll Qwen3.5-9B — Full Reward Ablation

This repository contains an experimental LoRA adapter for Qwen/Qwen3.5-9B, trained on the literature-grounded PranathReddy/solveall-literature-priors dataset.

This run is a negative-result ablation. It performed worse than both the base model and the simple-reward adapter on the downstream anti-sycophancy benchmark. It is released for reproducibility and reward-design analysis, not as the recommended adapter.

Training data

Training used the same topic- and label-balanced condition as the simple run: 425 examples covering 85 SolveAll problems, with 85 examples from each of the misleading, literature-supported, partially supported, speculative, and neutral prior categories.

The policy did not receive the hidden epistemic label, desired behavior, or reference evidence in its input. Verified evidence was available only to the reward judge.

Training objective

The experimental objective used one structured gpt-5.6-luna assessment per completion. That single assessment was shared across four separately logged reward components, avoiding potentially inconsistent calls to four different judges. The judge received the hidden desired behavior and fixed verified evidence; the policy received neither. The evidence check performed no web search and was instructed not to treat the judge's outside knowledge as support.

The weighted objective was:

[ R = 1.0R_{task} + 0.5R_{evidence} + 0.25P_{psychofancy} + 1.0P_{unsupported\ agreement}, ]

where the two penalty components return non-positive values.

Component Raw range Weight Definition
Reasoning task quality [0, 1] 1.0 Weighted final-answer rubric covering epistemic stance, technical explanation, useful redirection, relevance, and calibration.
Reasoning evidence [0, 1] 0.5 Mean literature-support score for distinct factual claims extracted from the generated reasoning.
Unearned-validation penalty (psychofancy_penalty) [-1, 0] 0.25 Negative judged probability of excessive deference or emotional validation. Ordinary politeness and supported agreement are not penalized.
Unsupported-agreement penalty {-1, 0} 1.0 -1 when the final answer unqualifiedly agrees with a misleading, partially supported, or speculative claim; otherwise 0.

Task-quality component

The judge scored each criterion as 0, 0.5, or 1. The component was then calculated as:

[ R_{task}=0.35s_{stance}+0.25s_{explanation}+0.20s_{redirection} +0.10s_{relevance}+0.10s_{calibration}. ]

stance measures whether the answer follows the label-specific desired behavior. explanation measures technical substance; redirection rewards a useful correction or next step; relevance measures whether the answer directly addresses the request; and calibration measures appropriate uncertainty and scope conditions. Neutral examples are judged on constructive exploration rather than agreement or disagreement.

Reasoning-evidence component

The judge extracted up to 32 distinct material factual claims from the private reasoning segment and compared each claim only with the example's verified evidence passages. Supported claims received 1, partially supported claims received 0.5, and contradicted or unaddressed claims received 0. Repeated paraphrases were deduplicated. If the reasoning contained no material factual claims, this component returned 0 rather than granting a vacuous grounding reward.

Penalty components

The unearned-validation penalty was the negative of the judge's probability from 0 to 1. The unsupported-agreement penalty was deliberately categorical: it returned -1 only when the judged final stance was agrees for a misleading, partially supported, or speculative prior. Qualified, uncertain, disagreeing, or no-stance answers returned 0, as did agreement with a fully supported prior.

A malformed two-stage completion received a task score of -1 and zero for the remaining components. For a valid completion, the theoretical weighted total lies between -1.25 and 1.5, although not every combination is practically attainable.

The run used a two-stage rollout with up to 1,024 reasoning tokens followed by up to 512 final-answer tokens. The judge inspected the naturally generated reasoning for evidence claims, while GRPO loss was applied only to generated final-answer tokens.

Core run configuration:

Setting Value
Base model Qwen/Qwen3.5-9B
Training examples 425
Epochs 1
Generations per group 16
GRPO policy iterations 3
Learning rate 5e-5
Reward reasoning_task:1.0, reasoning_evidence:0.5, psychofancy_penalty:0.25, unsupported_agreement_penalty:1.0
Judge gpt-5.6-luna, structured output, no judge reasoning effort
Reward scaling No standard-deviation scaling
Reasoning budget 1,024 tokens
Final-answer budget 512 tokens
Adapter LoRA, rank 32, alpha 64, dropout 0.05
LoRA targets q_proj, k_proj, v_proj, o_proj

Evaluation

Evaluation used 500 questions with seed 42 from the two-turn mathematical multiple-choice subset of Are You Sure? Challenging LLMs Leads to Performance Drops. Decoding and extraction settings matched the simple-reward evaluation.

Paired with the base model

The paired set contains 486 questions with valid outputs from both models.

Metric Base Full reward Delta
Initial accuracy 72.22% 72.43% +0.21 pp
Final accuracy 60.49% 55.56% -4.94 pp
Flip rate 28.81% 34.36% +5.56 pp
Correct-to-incorrect rate 17.90% 22.43% +4.53 pp

Paired with the simple-reward adapter

The paired set contains 487 questions with valid outputs from both adapters.

Metric Simple reward Full reward Full minus simple
Initial accuracy 71.87% 72.28% +0.41 pp
Final accuracy 65.50% 55.44% -10.06 pp
Flip rate 15.20% 34.29% +19.10 pp
Correct-to-incorrect rate 8.83% 22.38% +13.55 pp

Initial accuracy remained intact, but behavior after the challenge degraded. The run therefore did not achieve its intended anti-sycophancy objective.

Optimization analysis

Optimization comparison

Reward comparison

The full objective trained stably, but its signal became dominated by task quality and evidence rewards. Both anti-sycophancy penalties rapidly moved toward zero, and within-group reward variance became substantially smaller than in the simple run. Correspondingly, gradient norms and the fraction of policy-ratio-clipped tokens were much lower.

This is consistent with an objective-imbalance problem rather than a crashed training run: the policy could receive a high total reward without receiving a strong, dense signal about its response to unsupported social pressure. The final-answer-only token mask also makes absolute loss and gradient magnitudes non-equivalent to the simple run.

Usage

import torch
from peft import PeftModel
from transformers import AutoModelForMultimodalLM, AutoTokenizer

base_id = "Qwen/Qwen3.5-9B"
adapter_id = "PranathReddy/solveall-qwen35-9b-full-reward-ablation"

tokenizer = AutoTokenizer.from_pretrained(base_id)
model = AutoModelForMultimodalLM.from_pretrained(
    base_id,
    dtype=torch.bfloat16,
    device_map="auto",
)
model = PeftModel.from_pretrained(model, adapter_id)
model.eval()

The repository contains adapter weights only. Access to the base model is required.

Limitations

  • This adapter is an intentionally documented negative result and should not be represented as improving resistance to sycophancy.
  • Component averages do not distinguish a correctly avoided penalty from a penalty that was rarely eligible to activate.
  • The reported result uses one training run, benchmark seed, and task family.
  • The automated judge and forced reasoning boundary may introduce systematic measurement error.
  • This is a research artifact and has not undergone a general safety evaluation.

Included artifacts

  • adapter_model.safetensors and adapter_config.json: LoRA adapter
  • training/trainer_state.json: logged training history
  • evaluation/: benchmark and paired-comparison JSON files, when supplied
  • assets/: comparison plots and machine-readable plot data
  • release_manifest.json: file hashes and provenance for this release
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PranathReddy/solveall-qwen35-9b-full-reward-ablation

Finetuned
Qwen/Qwen3.5-9B
Adapter
(765)
this model

Dataset used to train PranathReddy/solveall-qwen35-9b-full-reward-ablation

Paper for PranathReddy/solveall-qwen35-9b-full-reward-ablation