Safe Psafe
Permissible under both general safety rules and the user's authorization scope.
A Modular, Controllable, and Efficient Framework for
On-demand Authorized Safety Alignment Relaxation in LLMs
One foundation model. A palette of authorized safety controls.
Relax refusal in target domains, while preserving safety and utility elsewhere.
* Equal contribution with random order.
Authorization is assumed to be verified externally. PALETTE adapts behavior after authorization.
A closer look at authorized safety
Choose a profile. See the intended scope.Intended refusal behavior
Benign instructions retain their ordinary behavior.
Targeted refusal relaxation, rather than indiscriminate compliance.
Compose independently learned domain controls through parameter merging.
Lightweight, single-block adaptation without full-response supervision.
A general-purpose refusal policy can also block legitimate requests from authorized professionals. PALETTE studies how to selectively relax that policy for a target scope—without treating every request from that user as permissible.
Its goal is a controlled boundary: answer allowed requests, retain refusal outside the authorization scope, and preserve ordinary utility.
Current safety alignment of foundation models largely follows a one-size-fits-all paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for general users but legitimate for authorized professionals, limiting helpfulness in specialized professional settings. Existing approaches either require costly realignment or rely on inference-time steering that suffers from imprecise control and added latency. To this end, we propose PALETTE, a modular, controllable, and efficient framework that selectively relaxes refusal behavior on authorized target domains while preserving standard safety elsewhere. Our method identifies a refusal direction via multi-objective search and internalizes it into the model through lightweight adaptation. PALETTE further supports modular composition: it learns domain-specific safety controls independently and composes them through parameter merging, enabling on-demand multi-domain authorization without retraining. Experiments across four safety benchmarks, multiple model variants, and both LLMs and VLMs show that PALETTE delivers precise safety control without sacrificing general utility, offering a practical path toward foundation models that adapt to diverse professional needs.
Read in the paperPermissible under both general safety rules and the user's authorization scope.
Restricted by the general policy, but legitimate within the authorized target scope.
Outside the user's legitimacy boundary, including unrelated harmful requests.
Scope, not identity verification. User authorization is an input assumption of this work. Verification and access control are separate deployment responsibilities. §2.2 · Problem formulation ↗
First find where to steer. Then learn when to steer. Finally, compose the learned controls for a new authorization scope.
Select a stage to explore.
Diagrams are explanatory schematics.
Extract candidates from harmful–harmless activation differences at different layers and token positions. Select a direction that favors allowed compliance, limits disallowed compliance, and keeps safe-prompt output distributions close to the base model.
Train a lightweight adapter in block n−1 to produce the target layer-n activation. Allowed inputs match a steered target; safe and disallowed inputs reconstruct their original activations. Other model parameters remain frozen.
Allowed input: learn the steering-induced activation shift.
After a short preliminary adaptation, rank candidate disallowed prompts by the increase in bypass score relative to the vanilla model. Use the largest shifts as hard disallowed samples for formal training. The bypass score is a first-token refusal-indicator proxy, not a guarantee about a complete response.
Reconstruction on non-target inputs encourages each adapter to remain approximately neutral outside its domain. Add the learned parameter updates to compose multiple controls, instead of training a separate adapter for every combination.
Parameter updates are merged; a new combination needs no retraining.
Neutrality is approximate. Related domains and conflicting preferences can still introduce interference; see §5 and Appendix C.
Explore single-domain control on GenHarm. The desired direction is different for each row: lower refusal for the allowed domain, higher refusal for the rest.
Hate is allowed; other domains remain disallowed.
Source: Table H.2. Values are reported means over three runs, displayed as percentages. This view compares Base, CAST, and PALETTE; consult the paper for all baselines.
Source table| Method | Violence | Hate | Disinformation | Sexual | Illegal | MMLU | GSM8K | Avg. utility |
|---|---|---|---|---|---|---|---|---|
| Base model | 1.000 | 1.000 | 0.989 | 0.995 | 1.000 | 0.636 | 0.759 | 0.698 |
| CAST | 0.205 | 0.132 | 0.020 | 0.147 | 0.407 | 0.606 | 0.715 | 0.661 |
| PALETTE | 0.976 | 0.103 | 0.857 | 0.947 | 0.944 | 0.639 | 0.756 | 0.698 |
Switch the Hate and Disinformation adapters on or off to compare the reported single-domain and merged configurations. These are separate experimental results—not a numerical simulation of weight merging.
LLaMA2-7B-Chat · GenHarm
Single-domain: Table H.1. Merged: Table H.5.
The paper visualizes how target-domain activations shift after adaptation and merging.
PALETTE fits a label-conditioned activation target in one block. It does not require collecting complete target responses for standard supervised fine-tuning.
Reported on RTX 4090. Adaptation time only—not total pipeline time.
One-time cost per base model
Forward-pass evaluation
Only the lightweight parameters are trained
Source: Appendix E, Table E.1. Memory is reported separately for each stage; 2.3 GB is the adaptation-stage measurement, not full-model inference memory. View analysis ↗
Experiments span LLaMA2, LLaMA3.1, Qwen2.5, and Qwen2.5-VL. See §4.1 and Appendix I for the evaluation settings.
Select an authorized role, load the corresponding adapters, and compare base-model responses with adapter-conditioned responses.
The existing interactive demo is included here, with its role permissions, adapter controls, and curated responses intact.
Illustrative, not live inference. Responses are pre-generated model outputs from the accompanying demo. Some examples contain harmful or sensitive material. A selected role illustrates an authorization scope; it does not verify a visitor's identity or grant real-world authorization.
PALETTE is a dual-use research framework, not a standalone safety mechanism. Real deployments require external authorization, domain-expert curation, auditable policies, least-privilege access, monitoring, and continuous red-team evaluation.
The method depends on representative allowed and disallowed examples. Missing related negative domains can lead to safety leakage, and closely related or conflicting controls may need additional conflict resolution. §5 · Limitations ↗
Read the full method, experimental settings, and limitations in the paper. Please use the following citation when referencing this work.
Repository address follows the manuscript. Public availability is maintained by the authors.
@misc{tan2026palette,
title = {{PALETTE}: A Modular, Controllable, and Efficient
Framework for On-demand Authorized Safety
Alignment Relaxation in {LLMs}},
author = {Tan, Qitao and Song, Xiaoying and Akbari, Arman
and Akbari, Arash and Wang, Yanzhi and Zhai, Xiaoming
and Hong, Lingzi and Xiang, Zhen and Lu, Jin
and Yuan, Geng},
year = {2026},
note = {Preprint}
}PALETTE manuscript · 2026 preprint.