Changelog¶
All notable changes to SafeTune are documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[0.1.8] - 2026-10-02¶
0.1.8, the version the release pipeline publishes next; the last PyPI release was 0.1.7.
- The logo renders on the PyPI project page again.
README.mdpointed at it with a repo-relative path, which GitHub resolves inside a README but PyPI cannot — there is no repo checkout for a relative path to resolve against — so the image 404s on the package page. It now points atraw.githubusercontent.com, the same pattern AgentTune uses. TheLICENSE.mdhref had the same root cause and is fixed with it. - The declared version is reconciled with the index.
mainstill said 0.1.6 while PyPI served 0.1.7, because the 0.1.7 bump the pipeline made at release time never landed back onmain.scripts/check_version_consistency.py --pypihad been failing on this the whole time; CI runs it without--pypi, so the drift was invisible.
0.1.7 on the index is a re-publish of 0.1.6 — the two sdists are identical
apart from their dist-info directory — so it gets no section of its own.
[0.1.6] - 2026-10-01¶
0.1.6, the version the release pipeline publishes next; the last PyPI
release was 0.1.5. The [Legacy numbering] section further down holds the
pre-June-2026 numbering that peaked at 0.6.0 — none of it is on the index.
Pull requests in this release¶
SafeTune-Internal pull requests, in stack order.
- ST-05 (#3) One settings mechanism for SafeTune (
configure()+ YAML), hardware and silent-ignore fixes. - ST-06 (#4) One class per method, keyword-safe merge functions, README that runs as written.
- ST-07 (#5) Evaluate reliability: loud failures, all 18 benchmarks load, one AdvBench scorer.
- ST-08 (#8) Demo notebooks rewritten on the final API, re-executed with real outputs.
- ST-09 (#6) Method fixes: disjoint BeaverTails splits, DeRTa token helper, depth-relative steer layers, SafeSwitch prober.
- ST-10 (#7) Integrate ST-06, ST-07 and ST-09 with cross-ticket fixes (DeRTa token, CLI dtype).
- ST-11 (#9) Method fixes (CAST, monitor, GradientAscent, ConstrainedSFT, DeRTa, steer) and paper-sized benchmark defaults.
- ST-12 (#10) Aya Vision and North support,
--safety-dataset,transformers>=5.15. - ST-13 Interop with the Lexsi stack: CuratorKIT dataset folders, CircuitKIT
scores,
lexsi_provenance.json,push_to_hub, version 0.1.6. - Hackathon fixes (#18 and the PR stacked on it), below.
Hackathon fixes (Cohere models)¶
max_lenis sized from the chat template. The QA data loaders default tomax_len=None:max(256, longest templated prompt + 256), capped at- Before, Tiny Aya's ~366-token preamble filled a fixed 256 and training
ran on zero supervised tokens. An explicit
max_lenthat leaves none raises. - Unlearn trains in fp32 for bf16 models too.
upcast=True(default) replacesupcast_fp16, which is a deprecated alias. - Refusal-direction sweep: new
RefusalDirectionConfig.min_layer_fraction(default 0.2): an early-layer winner, or none that lowers the refusal rate, falls back to the middle layer with a warning. Ties break toward the middle. - ReSta streams the safety vector one tensor at a time; extra memory is a
few fp32 copies of the largest tensor instead of about three of the model.
New
device=onapply_resta/ReStaTrainer("cpu"keeps it off the GPU). - One BOS on vLLM text prompts and the suite's WildGuard judge.
- Batched generation left-pads a passed-in right-padded tokenizer and restores its padding side afterwards.
Interop (ST-13)¶
- Datasets in.
--train-dataset/--safety-dataset(with the new--train-config/--safety-config) and every harden trainer'strain(...)take a dataset folder plus config name, asload_dataset(folder, config)reads it (e.g. a CuratorKIT export andsft_sharegpt), as well as Hub ids, table names and single files. In Python:trainer.train("./curated_out", dataset_config="sft_sharegpt"), or a rawdatasets.Dataset. SafeTune tokenises raw rows itself with the chat template: chatmessages, ShareGPTconversations, Alpacainstruction/input/output(theinputis kept) and prompt/response or DPOprompt/chosencolumns, as strings or turn lists. - Prompt-only data is an error. Rows without an assistant response
(CuratorKIT
ppo/grpo, a baretextcolumn) used to be fine-tuned on as empty responses without a warning. Now such rows are skipped with a warning, and a dataset with none left raisesValueError. - A dataset folder of several files no longer loads its first file silently
when the split is not found; it raises and lists the files. Folders with a
README.md(HF dataset folders) load withload_dataset, somanifest.json,rejected.jsonlandlexsi_provenance.jsonare never read as data. - CircuitKIT scores.
load_circuit_info_from_file/get_circuit_inforead CircuitKIT*_scores.json: thesafety_units/layer_suggestionskeys CircuitKIT 0.2 writes, and files with onlynode_scores, from which the same keys are derived (circuit_info_from_node_scores, top 20% of nodes by default). A file with none of these keys raisesValueError; it used to return an emptyCircuitInfo. - Provenance. Every checkpoint (
save_checkpoint: harden, recover, unlearn,safetune patch --output), results summaries directory and steering-vector directory gets alexsi_provenance.json(lexsi.provenance/1), with the source model and the input datasets. When an input folder has its ownlexsi_provenance.jsonit is embedded underinputs[].provenance, so lineage chains from CuratorKIT through SafeTune.safetune.provenancehas the reader and writer. - Hub.
safetune.push_to_hub(path, repo_id)uploads a checkpoint folder (model, tokenizer, processor, provenance) or a results JSON with its provenance, creating the repo if needed. - Version. 0.1.6.
safetune.__version__comes from the installed package metadata; the release script and workflow no longer edit__init__.py. Deprecation messages that said "stops working in 0.2" now say 0.3.
Changes to previously reported numbers¶
Every default changed since 0.1.3 that moves a number SafeTune reported
before, and the argument or safetune.configure() key that restores the old
behaviour. Explicit arguments win over configure(); the same keys work in the
runtime: / datasets: blocks of a --config YAML. None of these results
were re-run; reproduce an old number with its switch.
| What changed | Affected numbers | Restore the old behaviour |
|---|---|---|
| HarmBench is the paper's 400 text behaviours (standard 200 + contextual 100 + copyright 100); was standard only (200) | every HarmBench number: evaluate(), trainer.evaluate(), load_bench_prompts, load_prompts |
configure(datasets={"harmbench": {"config": "standard"}}) or load_harmbench(subset="standard") |
WildJailbreak is the first 500 adversarial_harmful rows of the eval set (the selection load_prompts() has always used); evaluate() and trainer.evaluate() loaded all 2,210 rows, 210 of them benign |
every WildJailbreak number outside load_prompts() |
configure(datasets={"wildjailbreak": {"where": None, "limit": None}}) |
OR-Bench hard-1k (1,319 rows) and toxic (655) are reported as separate benchmarks: orbench_overrefusal and orbench_toxic_refusal in trainer.evaluate(), orbench_hard / orbench_toxic in evaluate() and steer evaluation; one combined orbench_refusal before |
OR-Bench numbers; safety_mean |
configure(orbench_in_safety_mean=True) |
safety_mean averages the harm benchmarks only; OR-Bench (both splits) is out; before, one OR-Bench refusal rate over both splits was averaged in as "higher = safer" |
every safety_mean / ρ |
configure(orbench_in_safety_mean=True) |
A bare StringMatchJudge() uses the 12-prefix scorer ("prefix"), as trainer.evaluate() always did for AdvBench; it used the 29-phrase GCG substring check |
ASRT and Best-of-N attack success; any script with a bare StringMatchJudge() |
StringMatchJudge(mode="gcg") or configure(advbench_scorer="gcg") |
run_judge("advbench") drops <think>...</think> before matching, as trainer.evaluate() did |
AdvBench scores of reasoning models via run_judge |
none (the runner never scored think blocks) |
load_prompts("xstest" \| "jailbreakbench") use walledai/XSTest and JBB-Behaviors harmful; were natolambert/xstest-v2-copy (gpt4) and walledai/JailbreakBench (100 harmful + 100 benign) |
load_prompts numbers for these two |
configure(datasets={"xstest": {"source": "natolambert/xstest-v2-copy", "split": "gpt4"}}); configure(datasets={"jailbreakbench": {"source": "walledai/JailbreakBench", "split": "train"}}) |
Evaluation raises when a benchmark fails; before, it was dropped from safety_mean silently |
safety_mean of runs where a benchmark failed |
configure(eval_strict=False) (the failure is logged and the benchmark still left out) |
| Steer default layers scale with depth (unchanged on 32 layers) | CAA, CAST, LinearProbeGuard, SafeSwitch, AlphaSteer on any model that is not 32 layers deep, e.g. Qwen2.5-0.5B (24), Llama-3.2-3B (28), Gemma-3-4B (34) | configure(legacy_steer_layers=True) or explicit layer arguments |
SafeSwitch fits its prober in calibrate; before, it never fired |
SafeSwitch on every model (old numbers equal the unsteered model's) | none: evaluate the unsteered model |
| TAR's default adversary set is disjoint from the harden contamination set | TAR without a harm_dataset |
configure(legacy_beavertails_splits=True) |
| DeRTa's RTO token is "Sorry" under the model's tokenizer (19701 only on Llama-3) | DeRTa on every non-Llama-3 model | rto_refusal_token_id=19701 |
| DeRTa trains as the authors do: the harmful prefix is masked, one RTO row per example (the harmful response, relabelled "Sorry"), one cross-entropy | DeRTa on every model | DeRTaTrainer(legacy_derta=True) or configure(legacy_derta=True) |
| ConstrainedSFT through the runner and CLI trains against its aligned reference; it was plain SFT | ConstrainedSFT (--algo constrained) on every model |
ConstrainedSFTTrainer(use_reference=False) or configure(legacy_constrained_sft=True) |
GradientAscentTrainer and GradDiffTrainer default forget_clip=None (TOFU's pure ascent); at 0.5 the forget term had no gradient on real data |
GradientAscent and GradDiff (GradDiff trained on the retain set only) | forget_clip=0.5 or configure(legacy_ga_forget_clip=True) |
| CAST fits its gate on chat-formatted prompts and gates each prompt of a batch; the gate never fired on chat models | CAST on every chat model (old numbers equal the unsteered model's when the gate never fired) | CASTTrainer(chat_template=False, per_prompt_gate=False) or configure(legacy_cast_gate=True) |
| AlphaSteer hooks each matrix at the layer it was fitted on (it fitted 10-19 and hooked 0-9) | AlphaSteer on every model | AlphaSteerTrainer(legacy_alphasteer_layers=True) or configure(legacy_alphasteer_layers=True) |
| ReSta's DARE drop rate is the paper's 0.3; it was 0.9 | ReSta with its defaults (DARE on) | ReStaTrainer(dare_drop_rate=0.9) or configure(legacy_resta_drop_rate=True) |
SpectralEntropyMonitor leaves each prompt's first token (the attention sink) out and reads chat-formatted prompts |
monitor flags and entropies | SpectralMonitorConfig(skip_first_token=False, chat_template=False) or configure(legacy_spectral_monitor=True) |
| Device and dtype are chosen per host (cuda > mps > cpu; bf16 where native, fp16 on pre-Ampere CUDA, fp32 on CPU); the CLI loads models in that dtype | runs on CPU, MPS or pre-Ampere CUDA; unchanged on Ampere+ CUDA | configure(device=..., dtype="bfloat16") |
Not number-changing: eval_strict=True and warnings for misspelled trainer
arguments stay the defaults, and MPS still defaults to bf16 (macOS 14+).
CAAModel hooks are on only inside with / install() and during its own
generate() / __call__ (numbers through the wrapper are unchanged; call
install() after building it for the old always-on hooks), and AdaSteer
recomputes its coefficient per prompt also under with + model.generate
(its own generate() already did).
Added¶
- Runtime settings:
safetune.configure(**settings)/safetune.get_config()(and aruntime:block in--configYAML) set device, dtype, generation lengths, batch sizes, prompt caps, vLLM / lm-eval settings, judge settings, AdvBench refusal prefixes and LoRA defaults. Defaults are unchanged. - Dataset overrides: every built-in dataset is resolved by short name from
safetune.data.dataset_ids; point any of them at an HF id, a local.jsonl/.json/.csv/.parquetfile or a URL withconfigure(datasets={...})or adatasets:YAML block. - Results JSON records the effective runtime settings and dataset specs.
TransformersBackendaccepts a steering wrapper (CASTModel,AdaSteerModel, ...) as its model, so generation goes through the wrapper's owngenerate()(CAST's gate, AdaSteer's per-prompt coefficient).examples/data/refusal_probes_demo.jsonl: 8 demo rows in 4 languages for showingconfigure(datasets=...).- Vision-language models (Aya Vision, North /
cohere_compass) load, train and save everywhere a model is loaded by path: the CLI, the runner, harden reference models,evaluate()and the quickstarts pickAutoModelForImageTextToTextfrom the config. Every pillar works on the language model's decoder layers, default LoRA stays off the vision tower, and saved checkpoints keep the processor. Text-only data.safetune[vision]addstorchvision, which North needs. safetune train --safety-dataset NAME [--safety-split SPLIT]: the safety set for harden methods that take one (adataset_idsname, HF id, local file or URL); methods without one exit with an error.
Changed¶
- Requires
transformers>=5.15,<6(North needs 5.15). - The ten notebooks in
examples/notebooks/use the one runner API,configure()and library helpers instead of copied prompt lists and helpers. Their committed outputs come from a CPU run, and the recover and monitoring notebooks use a real drifted checkpoint instead of noise. - One harden API:
safetune.harden.<Name>Traineris now the same class assafetune.runner.harden.<Name>Trainerfor every harden method. Thetransformers.Trainersubclasses that used to have those names are<Name>HFTrainer(DOOR:SafetyDOORTrainer). The old HF-style call (<Name>Trainer(model=, args=, train_dataset=...)) and the old submodule names still work until 0.2, with aDeprecationWarning. - Recover merge functions (
task_arithmetic,somf_merge,learn_somf_mask,apply_resta,apply_lox,apply_lssf,apply_safemerge,apply_aaq,apply_safe_lora) take everything after the model by keyword, sobaseandalignedcan no longer be swapped by position. Old positional calls keep their old order until 0.2, with aDeprecationWarning. - Evaluation failures are loud.
evaluate(),evaluate_with_vllm_backend()and every trainer's.evaluate()raise when a benchmark fails to load or score (the error names the benchmark).strict=Falseorconfigure(eval_strict=False)records{"error": ...}for it and runs the rest. Unknown benchmark or judge names raiseValueErrorbefore anything loads.safetune evalexits 1 when any benchmark failed, and the evaluate quickstart no longer reports success after a failed step. - OR-Bench outside
safety_mean. OR-Bench hard-1k (over-refusal,orbench_overrefusal) and toxic (orbench_toxic_refusal) are reported on their own, asorbench_hard/orbench_toxicinevaluate(), andsafety_meanaverages the harm benchmarks only.configure(orbench_in_safety_mean=True)restores the old singleorbench_refusalover both splits, averaged into the mean. - Benchmark sizes follow the SafeTune paper: HarmBench 400 (standard +
contextual + copyright), WildJailbreak 500 (the first 500
adversarial_harmfulrows). Each is a dataset-table entry you can override. - One AdvBench scorer.
StringMatchJudge(mode="prefix" | "gcg")is the only implementation.trainer.evaluate()andrun_judge("advbench")useconfigure(advbench_scorer=...), default"prefix"(whattrainer.evaluate()always reported);run_judge("advbench")now also drops<think>blocks. A bareStringMatchJudge()(ASRT, Best-of-N) follows the same setting, so its default is"prefix"too;mode="gcg"keeps the 29-phrase check. load_prompts("xstest" | "jailbreakbench")use the same sources asevaluate()(walledai/XSTest, JBB-Behaviors harmful); RWKU loads theforget_level2QA probes.
Removed¶
- The unused pack-runner HF dataset map (
load_pack_from_hf,run_safety_eval_from_hfand the core eval CLI'ssafety_evalcommand).
Fixed¶
evaluate()can run every registered benchmark:jailbreakbench,muse,rwkuandsafedialbenchno longer fail withTypeError, andstar1,mmluandjailbreakbenchfind their prompts.sentencepieceandtiktokenare declared; the default WildGuard judge needs them.- CTRAP, SEAM, SEAL, ConstrainedSFT and DeRTa runner trainers now use their method kwargs; unknown trainer kwargs warn with the closest valid name.
- Device auto-selects cuda, then mps, then cpu; bf16 is used only where
supported (fp16 on older CUDA, fp32 on CPU), so harden training no longer
fails on CPU / T4, and lm-eval no longer hardcodes
cuda:0. - Judges fall back to transformers when vLLM is not installed.
import safetune.harden.lisa(or anysafetune.<pillar>.<module>path) no longer loads the module a second time with its own copies of every class.- README: every Python block runs as written on CPU (except Evaluate, which
needs a GPU for its judge); the Steer example uses
alpha=0.3instead of the default 20, which turned output into noise on the README's 0.5B model. - The TAR adversary fallback set no longer shares prompts with the harden
contamination set (83 of 256 did on BeaverTails
30k_train); only TAR's default data changes.configure(legacy_beavertails_splits=True)restores the old selection. - Steer layer defaults (CAA, CAST, LinearProbeGuard, SafeSwitch, AlphaSteer)
scale with model depth and are unchanged on 32-layer models. On shallower
models CAA and CAST no longer return wrappers that change nothing, and
LinearProbeGuard and AlphaSteer no longer fail; on 24/28/36-layer models the
default layers move.
configure(legacy_steer_layers=True)restores the absolute indices. SafeSwitchTrainer.calibratefits its prober on the calibration prompts; before, the prober was never trained and the wrapper never fired.- DeRTa's RTO transition token is the first token of "Sorry" under the
model's tokenizer (19701 on Llama-3, the authors' id); before, every
tokenizer got 19701, an unrelated token outside Llama-3.
rto_refusal_token_id=19701restores the old id. - CAST fits its gate on chat-formatted prompts, as generation sees them, and gates each prompt of a batch; the gate never fired on chat models. SafeSwitch fits its prober on chat-formatted prompts and scores every prompt of a batch (it scored the first). AlphaSteer hooks each matrix at the layer it was fitted on and runs on GPT-2.
CAAModelno longer steers the model as soon as it is built, and still steers through its owngenerate()after awithblock; AdaSteer no longer reuses the first prompt's coefficient underwith+model.generate.SpectralEntropyMonitorleaves out the attention-sink token, which made the spectrum rank-1 and the entropy ~0 for every prompt, and reads chat-formatted prompts. Calibrated on the prompts it scans, it now flags a real safety drift (examples/notebooks/safety_monitoring.ipynb).- DeRTa masks the harmful prefix and applies RTO to the harmful response only,
in one cross-entropy, as the authors do; with trl 1.x its old RTO term had
been silently skipped (
loss_type="chunked_nll"returns no logits). - ConstrainedSFT through the runner and CLI trains against its aligned reference (it was plain SFT).
GradientAscentTrainer/GradDiffTrainerno longer clip the forget loss by default; the old 0.5 clip zeroed its gradient on real data.ReStaTrainertakesdare_drop_rate, default 0.3 (the RESTA paper's value; it was 0.9, which broke small models).safetune train,patchandunlearnload models in the runtime dtype (fp32 on CPU). transformers 5 loads bf16 by default, and CLI training on CPU ran at about 110 s per step.
[0.1.3] - 2026-08-30¶
Added¶
- Docs: EMNLP 2026 Demo tab with the accepted paper, walkthrough, and
artifacts (
docs/emnlp-2026-demo/).
Changed¶
- License: updated to LSAL v1.2 (see
LICENSE.md). Academic research and teaching remain free on MIT-like terms. Organizations must now acknowledge their use to Lexsi Labs or obtain permission before internal evaluation, auditing, or use on their own models (Section 1A). Commercial use still requires a separate license. The patent clause now reserves all patent rights (Clear BSD style) instead of granting a noncommercial patent license. New clauses: users bear responsibility for their own use and deployment decisions (Section 7, with an indemnity limited to organizational and commercial use), third-party base-model and benchmark licenses continue to apply (Section 4A), modified redistributions must be marked as modified (Section 2), and survival, severability, and version-applicability terms (Sections 8 and 9). - License: SafeTune is now released under the Lexsi Labs Source Available
License (LSAL) v1.1 (see
LICENSE.md) — an MIT-style grant restricted to noncommercial purposes, with a Responsible Use clause barring production deployment of unrepaired drifted checkpoints. Earlier changelog entries referring to the MIT License describe pre-LSAL releases. Commercial licensing: support@lexsi.ai.
[0.1.1] - 2026-06-29¶
Added¶
safetune.config.SafeTuneConfig— declarative YAML config for CLI runs.SafeTuneConfig.from_yaml(path)loads all standard flags plus method-specific hyperparameters (any unknown key lands inmethod_kwargsand is forwarded directly to the trainer constructor).--configCLI flag —safetune train --config run.yamlinjects YAML values as parser defaults; explicit CLI flags still take precedence.--train-dataset/--train-splitCLI flags — training dataset is now fully configurable.--train-dataset beavertails(default) or any HF dataset id;--train-split 30k_train(default) or any split name.safetune.runner._registry— centralised algo registry replacing the inline dicts incli.py.register_harden(),register_recover(), andregister_unlearn()let third-party code extend the method menu at runtime without editing library files.SaLoRATrainer—lora_alpha,lora_dropout, andtarget_modulesare now configurable constructor kwargs (previously hardcoded insidetrain())._RecoverBase.apply()contract — method is now documented: return the patched model, never write to disk, accept method-specific keyword overrides.- Dev runbook §6 — end-to-end guide for adding a new method: implement trainer → re-export → add one registry entry.
Changed¶
cli.pynow imports algo registries fromrunner/_registry.pyinstead of maintaining inline dicts.safetune listoutput is unchanged.--no-peftflag removed from the CLI (it was registered but never consumed).
Fixed¶
_derive_model_idduplicated across_HardenBaseand_RecoverBase— consolidated intomodel_utils.derive_model_id().- Dead
R = None/_ensure_recover_imports()globals removed fromrunner/recover/_base.py(subclass files already imported directly). - Stray multi-line bug-fix comment removed from
LoXHardenTrainer.train().
[0.1.0] - 2026-06-24¶
First public release. SafeTune is a library of ~100 alternative LLM-safety methods — train-time hardening, weight-space recovery and unlearning, inference-time steering, plus diagnosis and evaluation — for the Hugging Face ecosystem. Every method is faithfulness-audited against its cited paper.
Validated on an NVIDIA L40S with torch 2.8 / transformers 5.12 / trl 1.6
(full suite: 354 passed, 4 skipped). Docs build clean under mkdocs --strict.
Added — the library¶
- ~100 methods across 4 intervention pillars + 2 instrumentation tools:
- Harden (26 methods):
SafeGradTrainer,LisaTrainer,SurgeryTrainer,AntibodyTrainer,LookAheadTrainer,vaccine_loss,tar_outer_loss, and 19 more across 8 mechanism families. - Recover (24 methods):
apply_resta,apply_lox,apply_safemerge,apply_ctheta,task_arithmetic, and 19 more across 6 granularities (whole-model → subspace → layer → neuron → circuit). - Unlearn (6 methods):
rmu_unlearn,npo_unlearn,gradient_ascent_unlearn(plus its GradDiff variant),flat_unlearn,simdpo_unlearn. - Steer (19 methods):
RefusalDirectionModel,AdaSteerModel,SafeSwitchModel,AlphaSteerModel,SafeSteerModel, and 14 more. Includessteer.run(...)withhf/vllm-hook/vllm-logitsbackends. - Interpret (6 methods):
identify_safety_neurons,safety_circuit_info,eap_safety_circuit,CircuitInfo(round-trippable JSON/YAML). - Evaluate (24 methods):
evaluate(...)with HarmBench, XSTest, AdvBench, WildJailbreak, and more.BoNAttack,AbliterationAttack, WildGuard / LlamaGuard-3 judges. - Faithfulness audit: every method compared against its cited paper,
corrected where it diverged, and labelled. 100 faithful, 1 simplified,
5 SafeTune variants, 0 broken. Per-method verdicts with
file:lineevidence in the Feature Map. - Quickstart demos:
quickstart.py(steer),recover_quickstart.py,harden_quickstart.py— all run onQwen/Qwen2.5-0.5B-Instruct, no GPU required. - Colab notebooks: 4 interactive notebooks (steer, recover, harden, unlearn) mirroring the quickstart scripts.
- CLI:
safetune/stcommands dispatch to real pillar APIs (harden, evaluate, recover).
Added — documentation¶
- MkDocs Material site with purple/amber design system, dark mode, sticky tabs, instant navigation, search, code copy, Mermaid diagrams.
- Doc structure:
getting-started/(install, quickstart, taxonomy),guides/(6 method-group guides — 4 intervention pillars + 2 instrumentation tools),trust/(feature map, results, scope, audit details),reference/(paper table, eval protocols),tutorials/(Colab hub),community/(FAQ, contributing, changelog),blog/. - Landing page pitching 4 intervention pillars and 2 instrumentation tools, with real impact numbers, a decision table, scenario-based usage examples, and runnable code tabs.
- Lexsi Labs branding: compass logo, Lexsi logo footer (dark/light variants).
Added — standard library files¶
CITATION.cff— citation metadata..github/ISSUE_TEMPLATE/— bug report, faithfulness report, feature request..github/pull_request_template.md..github/workflows/docs.yml— GitHub Pages docs deploy..github/workflows/smoke.yml— CI smoke test.LICENSE— MIT, 2026.
Changed (Major)¶
- CLI rewritten to dispatch to harden / evaluate / recover pillar APIs instead of the old SFT/DPO/PPO/GRPO orchestrator stubs.
verify→evaluaterename:safetune.evaluateis the current name;safetune.verifywas later removed in v0.1.0.- Recover uniform input contract: every
apply_*acceptstarget=(model=/finetuned=kept as aliases). - 2-tier, input-keyed taxonomy: Tier 1 Interventions (harden / recover / unlearn / steer), Tier 2 Instrumentation (interpret / evaluate).
- Benchmark menu categorized into jailbreak / over_refusal / capability / domain / tamper.
Fixed¶
- FLAT unlearning rewritten to the faithful f-divergence loss (Wang et al., ICLR 2025, arXiv:2410.11143).
- Decoding steer methods (SafeDecoding, ContrastiveDecoding, ProxyTuning,
Nudging): logit width reconciliation for padded
lm_heads. - SafeGradTrainer: whole-model gradient dot overflow on >2B params.
- EAP-IG: integrated-gradient interpolation off-by-one.
- DOORTrainer: DPO + DOOR hybrid loss replaced with paper-faithful DOOR-only.
- Import-order fragility: all 12 Harden configs guarded against transformers/trl import race.
- Packaging: MIT license classifier (was Proprietary); upper version bounds; guarded trl imports for 1.x compatibility.
Removed¶
- ~3.8 GB of raw per-prompt eval generations and internal scratch/planning docs.
- 12 broken in-house attack reimplementations (GCG, PAIR, TAP, AutoDAN, …) — not faithful to their papers.
- Orphaned pipeline-orchestration subsystem (
core/options.py,core/orchestrator.py,core/callbacks.py). - Stale duplicate docs (
docs/safetune-docs/,docs/archive/). requirements.txt(consolidated intopyproject.toml).
[Legacy numbering]¶
Versions before the June 2026 renumbering: this line peaked at 0.6.0
(finetunehub → SafeTune era) and was reset to 0.1.0 when the public
PyPI releases started. Kept as history; none of it is on the index, and
scripts/check_version_consistency.py does not count it as releases.
[0.6.0] - 2026-05-16¶
Changed (Major)¶
- Taxonomy overhaul: replaced the flat "Four/Five Pillars" list with a
2-tier, input-keyed taxonomy. Tier 1 · Interventions — Train-time
(
harden), Weight-space (recover+unlearn), Inference-time (steer); Tier 2 · Instrumentation — Diagnose (interpret), Measure (evaluate). SafeTune is framed as a library of alternative methods, not a pipeline. verify→evaluaterename: package dirsrc/safetune/verify/→evaluate/(verify/eval/→evaluate/suite/).safetune.verifywas a back-compat alias that emits aDeprecationWarning.interpretandunlearnare now first-class importable submodules (import safetune.interpret,import safetune.unlearn).- Repo layout:
emnlp_exp/→experiments/emnlp2026/,paper/→experiments/paper/,safetune_check/→audit/, validation scripts →tests/support/. Top level is now deliverable directories only.
Added¶
- Gold-standard re-check pass: cloned 12 upstream reference repos into
audit/reference_repos/and diffed every 🟡 SafeTune adapter against the originating code. Result: 12 methods upgraded 🟡→✅ (antidote, pke, safereact, tracin_influence, DOORTrainer, LookAheadTrainer, tar_outer_loss, AdaSteerModel, RRFAEnsemble, eap_safety_circuit, task_arithmetic, NudgingProcessor). After a final 🟡→✅ sweep the audited surface is ✅ 89, 🟡 0, 🟠 2. Two bugs found and fixed in the same pass (see below). docs/REFERENCES.md— per-method table covering all ~91 audited methods: paper, venue, arXiv, official repo link, 1-sentence description, non-obvious inputs required, outputs, and faithfulness badge.safetune.steer.run(model, backend=...)— one steering-generation entry point withhf/vllm-hook/vllm-logitsbackends. The vLLM adapters are promoted intosafetune.steer.backends.- Recover uniform input contract — every
recoverapply_*accepts the canonicaltarget=keyword (model=/finetuned=kept as aliases). - Benchmark menu —
evaluate.suiteregistry categorized into jailbreak / over_refusal / capability / domain / tamper, withlist_benchmarks()andbenchmarks_by_category(). docs/LIBRARY_CHECKLIST.csv— theFEATURE_MAP.mdaudit map flattened to one CSV row per method across every pillar (~100 rows: Pillar, Method, Entry-point file, Runs?, Faithful?, Audit status). Regenerated fromFEATURE_MAP.mdbyscripts/gen_library_checklist.py.
Removed¶
- Orphaned pipeline-orchestration subsystem:
core/options.py,core/orchestrator.py,core/callbacks.pyand theadversarial/constitutional/lifelong/multiturn/selfplayconfig modules — pipeline framing with no executor. - Retired the broken/undefined entry points
pipelines.pipelineandtraining.orchestrator.run_sft/dpo/ppo/grpofrom the public surface.
Fixed¶
harden/safegrad.py—SafeGradTrainerglobal gradient surgery calledtorch.doton the concatenated whole-model gradient, overflowing the BLAS int32 length bound for any model above ~2B parameters; replaced with the numerically identical(a*b).sum().harden/lisa.py—LisaTrainercrashed with'DataLoader' object is not subscriptablewhenalignment_datasetwas passed as an already-builtDataLoader; it is now tolerated.- Import-order fragility across the whole
hardenpillar — every Trainer config was declared as@dataclass class XConfig(Base if _IMPORT_ERROR is None else object). Whentransformers/trlwas imported beforesafetune(a very common order), the backend guard had not run, the nested import failed, the base silently degraded toobject, and@dataclassregenerated__init__from only the local fields — so e.g.SafeGradConfig(output_dir=...)raisedunexpected keyword argument. Fixed in all 12 affected configs (SafeGradConfig,LisaConfig,AsFTConfig,SPPFTConfig,STARDSSConfig,SAPConfig,SurgeryConfig,CSTConfig,DOORConfig,DeRTaConfig, andcore.optim.AntibodyConfig) by guarding the@dataclassdeclaration behind the import check, with an explicit stub in the failure branch instead of a misleading half-built dataclass. eval/cli.py— importedEvaluationRegistryfromeval/registry.py(which defines only the metric registryEvalRegistry); the CLI actually uses the task registry'sget_task/list_tasksAPI, which lives ineval/core.py. Repointed the import tocore.EvalRegistry; the module now imports and the whole package is import-clean (247/247 modules).docs/FEATURE_MAP.mdaudit badges synced with thefix/*_RESOLVED.mdfixlogs, which were never reflected in the map.apply_mscpandapply_lssfmove 🟠→🟡 — both cite verified papers (arXiv:2508.09190; arXiv:2602.00038, ACL 2025) and faithfully implement their projection equations; the stale "un-cited" notes were wrong.apply_deeprefusalandASRTCallbackstay 🟠 (honest SafeTune-original heuristics, no faithful published equivalent) with corrected notes. Post-fix distribution is now 89✅ / 0🟡 / 2🟠 / 0🔴 / 0⚫ over the 91 audited components.core/interpret/eap.py— EAP-IG integrated-gradients interpolation loop started atk=1(range(1, steps+1)), causing gradient sampling to include the clean endpoint and exclude the corrupted endpoint. Upstream EAP-IG (hannamw/EAP-IG) samples{0/steps, ..., (steps-1)/steps}; fixed torange(0, steps)to match.harden/door.py—SafetyDOORTrainerwas mixing DPO + DOOR loss (callingsuper().compute_loss()then adding the DOOR term). The paper's reference implementation uses DOOR alone (gd_npo_losson a plainTrainer). Fixed:DOORConfig.door_pure_mode=True(default) now returns only the DOOR term;door_pure_mode=Falsepreserves the old hybrid for back-compat.
[0.5.0] - 2026-04-12¶
Changed (Major)¶
- Architectural Overhaul: Transitioned to the "Four Pillars of Safety" taxonomy: Recover, Harden, Steer, and Verify.
- Package Modernization: Renamed internal package to
safetune(lowercase) for standard Python convention. - Modular Rewards: Decomposed the monolithic
rewards/core.pyinto categorical sub-modules (text,nlp,safety,code,math,specialized). - Pipelines API: Introduced high-level unified
pipelinesAPI for simplified safety orchestration. - CLI Refactor: Separated CLI logic into
cli.pyand extracted training orchestration intotraining/orchestrator.py. - Global Branding: Standardized all references to SafeTune across documentation and code.
Added¶
- Specialized Rewards: New reward functions for Medical, Legal, and Financial domains.
- Unified Configuration: Streamlined
UnifiedSafetyConfigsupporting the 4-pillar structure. - Project Structure: Cleaned up the root directory and standardized the
src/layout.
Removed¶
- Legacy monolithic files:
src/safetune/main.pyandsrc/safetune/rewards/core.py. - Redundant backup files and scripts.
[0.2.0] - 2026-01-18¶
Changed (Major)¶
- Renamed library from
finetunehubtoSafeTune - Package directory:
src/finetunehub/→src/safetune/ - All imports updated:
from finetunehub.*→from safetune.* - CLI commands:
safetune(primary),at(short alias) - Entry points and pyproject.toml fully updated
-
All documentation, examples, and tests migrated
-
License: SafeTune is released under the MIT License (see
LICENSE). - An earlier source-available license proposal was not retained; the project is MIT-licensed — free for research, academic, commercial and personal use.
Added¶
- Comprehensive Documentation System (59 markdown files):
- Complete documentation structure integrated into main project
- Getting Started guides (5 docs): Installation, Quick Start, Basic Concepts, Configuration, Backend Selection
- User Guide (9 docs): Overview, SFT, RL, Evaluation, Reward Functions, Model Management, Sample Logging, Troubleshooting
- Algorithm Documentation (10 docs): DPO, PPO, GRPO, GSPO, DAPO, Dr. GRPO, GBMPO, Counterfactual GRPO, BOLT with detailed explanations
- Backend Documentation (4 docs): Overview, TRL Backend, Unsloth Backend, Comparison
- API Reference (6 docs): Complete API documentation with all parameters
- Examples (4 categories): Overview, SFT, RL, Advanced examples
- Advanced Topics (4 docs): Architecture, Custom Backends, Distributed Training, Performance
- Contributing guides (3 docs): Guide, Code Style, Testing
-
Notebooks section for interactive tutorials
-
Enhanced Documentation Infrastructure:
- MkDocs configuration with Cinder theme
- Mermaid diagram support for architecture visualization
- Jupyter notebook integration
- Automatic API documentation generation with mkdocstrings
- Custom HTML overrides and styling
- Logo and branding assets removed (migrated to SafeTune branding)
Removed (Code Cleanup)¶
- Deleted unused/backup directories and files (Total: ~7,500 lines removed):
src/finetunehub/eval_old/- Old evaluation framework (7 files, 2,913 lines)src/finetunehub/backends/trl/rl/ppo/ppo_old.py- Old PPO implementation (1,437 lines)src/finetunehub/backends/trl/rl/grpo/grpo_old.py- Old GRPO implementation (1,264 lines)src/finetunehub/backends/unsloth/rl/ppo/ppo_old.py- Old Unsloth PPO (1,833 lines)src/finetunehub/cli_commands/- Unused CLI modulessrc/finetunehub/cli/unified-old.py- Old CLI backupsrc/finetunehub/rl/andsrc/finetunehub/sft/- Backward compatibility wrapperssrc/finetunehub/scripts/- Directory removed (contents moved)- Removed old
finetunehubpackage directory - Fully replaced bysafetune
Changed¶
- Reorganized BOLT utilities:
- Moved
precompute_baseline.pyfromsrc/finetunehub/scripts/→examples/bolt_training/ - Rationale: Co-locate BOLT-specific utilities with BOLT examples for better discoverability
Breaking Changes¶
- Library renamed: All imports must change from
finetunehubtosafetune from finetunehub.core.rl import *→from safetune.core.rl import *from finetunehub.eval import *→from safetune.eval import *- CLI:
finetunehub train→safetune train - Removed backward compatibility wrappers: The deprecated import paths have been removed
- All examples and documentation have been updated to use the new paths
- All test files have been migrated to the new import structure
[0.1.0] - 2026-01-18¶
Added¶
- Initial release with comprehensive GRPO support
- Backend support for both TRL and Unsloth
- GSM8K math reasoning examples
- Multiple bug fixes and improvements