Skip to content

Recover API

Weight-space repair of a model whose safety has drifted. Import from safetune.runner.recover. Recover trainers are training-free: construct with the drifted model (plus base_model / aligned_model references where the method needs them) and call apply().

from safetune.runner.recover import ReStaTrainer

trainer = ReStaTrainer(model=drifted, base_model=base, aligned_model=aligned,
                       alpha=0.5, dare=True, dare_seed=0)
repaired = trainer.apply()

Available trainers

AAQTrainer, AntidoteTrainer, AntidoteV2Trainer, CThetaTrainer, GradSelectiveRecoverTrainer, LSSFTrainer, LoXTrainer, MSCPTrainer, NLSRTrainer, OneShotSafetyPatchTrainer, PKETrainer, PrePostMergeTrainer, QReSafeTrainer, ReStaTrainer, RepNoiseRecoverTrainer, SCRUBTrainer, SOMFTrainer, SafeDeltaTrainer, SafeLoRATrainer, SafeMergeTrainer, SafeReActTrainer, SafetyVectorRestoreTrainer, TaskArithmeticTrainer, WiseFTTrainer.

See the Recover guide for the method taxonomy (whole-model / low-rank / layer / neuron / saliency / circuit-guided).

Reference

safetune.runner.recover.ReStaTrainer

Bases: _RecoverBase

ReSta (DARE task arithmetic): DARE-masked safety vector restoration.

Parameters:

Name Type Description Default
base_model Module

base model.

None
aligned_model Module

aligned reference.

None
alpha float

task vector scale. Default 1.0. On Tiny Aya, alpha=1 breaks the model; use alpha~0.25 (sweep: 0.1/0.25 restore refusal with normal answers, >=0.5 breaks it).

1.0
dare bool

apply DARE masking. Default True (part of the method).

True
dare_drop_rate float

DARE drop probability p. Default None: 0.3, the RESTA paper's value, or 0.9 (the old default) with safetune.configure(legacy_resta_drop_rate=True).

None
dare_seed int

DARE random seed. Default 0.

0
device Optional[str]

where each safety-vector delta is computed. Default None: the drifted model's weight device. "cpu" keeps the extra memory off the GPU.

None

Small models: DARE drops a fraction p of the safety vector's entries and scales the rest by 1/(1-p). On Qwen2.5-0.5B, with the full base-to-instruct delta as the safety vector, p=0.9 broke the model (HarmBench refusals 2/16 and garbled benign answers) while p=0.3 restored 13/16 and dare=False 14/16 (examples/notebooks/recover_comparison). The paper tunes on 7B models; on small ones check the benign answers after repair.

Memory: the drifted, base and aligned models must be loaded, but the safety vector is streamed one tensor at a time, so the extra memory is a few fp32 copies of the largest tensor, not of the model. Keep base and aligned on CPU and pass device="cpu" to hold only the drifted model on the GPU.

safetune.runner.recover.TaskArithmeticTrainer

Bases: _RecoverBase

Task Arithmetic: plain safety task vector addition.

Parameters:

Name Type Description Default
base_model Module

base model.

None
aligned_model Module

aligned reference.

None
alpha float

task vector scale. Default 1.0.

1.0

safetune.runner.recover.SafeLoRATrainer

Bases: _RecoverBase

SafeLoRA: safety subspace LoRA decomposition and merge.

Parameters:

Name Type Description Default
aligned_state_dict dict

state dict of the aligned model.

None
base_state_dict dict

state dict of the base model.

None
alpha float

merge coefficient. Default 0.5.

0.5
threshold float

safety subspace threshold. Default 0.5.

0.5