inv_mse_lastpos_all

Research checkpoint from a study of tokenization (reader) invariance and its effect on robustness to adversarial re-tokenization. Fine-tuned from allenai/OLMo-2-1124-7B-Instruct.

These are research artifacts, not products. Numbers below are measured on a 200-prompt AdvBench holdout that was excluded from training by construction.

Measured

metric value
AdvTok ASR (greedy, 200-prompt AdvBench holdout) 0.025
AdvTok ASR (t=1) 0.058
Canonical ASR (greedy, no attack) 0.02
Canonical refusal (greedy) 0.905
XSTest over-refusal (safe) 0.079
XSTest refusal (unsafe) 0.875
Alpaca over-refusal 0.03
Alpaca token F1 0.428
Alpaca NLL/token 1.274
Training cost (GPU-hours, 2xA100) 7.3
Reference (zeroshot) AdvTok ASR greedy / t=1 0.565 / 0.584
Objective Relative residual-stream MSE to the frozen reference at the LAST PROMPT TOKEN only, averaged over ALL 32 transformer layers. No continuation is generated and nothing is teacher-forced; the loss sees one position per encoding. CVaR 0.25 over 8 MDD-sampled encodings, canonical anchored.
CAVEAT: margin_canonical 24.80 vs zeroshot 17.15. For the sibling MSE arms this inflation was shown to be the COMPLIANCE log-prob collapsing, not refusal rising; that decomposition has not been run for this arm.

AdvTok ASR is attack success rate under adversarial tokenization (Geh et al., arXiv:2503.02174), Llama-Guard-3-8B judged, greedy decoding. XSTest over-refusal is the refusal rate on safe prompts — the cost side.

Training configuration

field value
mode prefix
objective mse_lastpos
kl_direction forward
ce_weighting uniform
ema_beta None
harmful_mix mixed
harmful_fraction 0.286
num_encodings 8
cvar_quantile 0.25
max_steps 700
learning_rate 1e-05
grad_accum 8
seed 42
reference_model allenai/OLMo-2-1124-7B-Instruct
max_new_tokens 128
prefix_tokens 8
stochastok_p 0.3

Parameter drift from base (training-happened guard)

group relative L2
attn 0.01087
embed_tokens 0.00057
lm_head 0.00000
mlp 0.01063
norm 0.00160

Caveats

  • Single seed. No claim of significance across seeds.
  • Evaluated on English AdvBench/XSTest/Alpaca only.
  • The safety numbers are for the specific attack studied; they do not imply robustness to other jailbreaks.
Downloads last month
30
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MARS-Retokenization/olmo2-7b-instruct-inv-mse-lastpos-all

Paper for MARS-Retokenization/olmo2-7b-instruct-inv-mse-lastpos-all