Research project NeurIPS 2026 · Atlanta

PALETTE

A Modular, Controllable, and Efficient Framework for
On-demand Authorized Safety Alignment Relaxation in LLMs

One foundation model. A palette of authorized safety controls.
Relax refusal in target domains, while preserving safety and utility elsewhere.

1 University of Georgia2 University of North Texas3 Northeastern University

* Equal contribution with random order.

Authorization is assumed to be verified externally. PALETTE adapts behavior after authorization.

A closer look at authorized safety

Choose a profile. See the intended scope.

Same model. Different authorized scopes.

Intended refusal behavior

Crime domain Relaxed
Biosecurity Retained
Other disallowed Retained

Benign instructions retain their ordinary behavior.

Concept illustration—not measured probabilities, live inference, or an authorization system.How it works

Controllable

Targeted refusal relaxation, rather than indiscriminate compliance.

Modular

Compose independently learned domain controls through parameter merging.

Efficient

Lightweight, single-block adaptation without full-response supervision.

The motivation

Beyond
one-size-fits-all safety.

A general-purpose refusal policy can also block legitimate requests from authorized professionals. PALETTE studies how to selectively relax that policy for a target scope—without treating every request from that user as permissible.

Its goal is a controlled boundary: answer allowed requests, retain refusal outside the authorization scope, and preserve ordinary utility.

Read the full abstract

Current safety alignment of foundation models largely follows a one-size-fits-all paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for general users but legitimate for authorized professionals, limiting helpfulness in specialized professional settings. Existing approaches either require costly realignment or rely on inference-time steering that suffers from imprecise control and added latency. To this end, we propose PALETTE, a modular, controllable, and efficient framework that selectively relaxes refusal behavior on authorized target domains while preserving standard safety elsewhere. Our method identifies a refusal direction via multi-objective search and internalizes it into the model through lightweight adaptation. PALETTE further supports modular composition: it learns domain-specific safety controls independently and composes them through parameter merging, enabling on-demand multi-domain authorization without retraining. Experiments across four safety benchmarks, multiple model variants, and both LLMs and VLMs show that PALETTE delivers precise safety control without sacrificing general utility, offering a practical path toward foundation models that adapt to diverse professional needs.

Read in the paper

Safe Psafe

Permissible under both general safety rules and the user's authorization scope.

Preserve the original behavior.

Allowed Pallowed

Restricted by the general policy, but legitimate within the authorized target scope.

Relax refusal selectively.

Disallowed Pdisallowed

Outside the user's legitimacy boundary, including unrelated harmful requests.

Retain safety constraints.

Scope, not identity verification. User authorization is an input assumption of this work. Verification and access control are separate deployment responsibilities. §2.2 · Problem formulation ↗

Inside PALETTE

A direction becomes
a selective control.

First find where to steer. Then learn when to steer. Finally, compose the learned controls for a new authorization scope.

Select a stage to explore.
Diagrams are explanatory schematics.

Multi-objective searchSection 3.2

Find a useful refusal direction.

Extract candidates from harmful–harmless activation differences at different layers and token positions. Select a direction that favors allowed compliance, limits disallowed compliance, and keeps safe-prompt output distributions close to the base model.

Contrastive
prompt activations
Candidate
directions
Selected
direction r*
Allowed bypass ↑Disallowed bypass ↓Safe-prompt KL ↓
Maximize αPA(r) − (1−α)PD(r), while minimizing utility drift.
The original framework, in one view.
Reported experiments

Change the target.
Keep the boundary.

Explore single-domain control on GenHarm. The desired direction is different for each row: lower refusal for the allowed domain, higher refusal for the rest.

Authorized target domain

Refusal rate by domain

Hate is allowed; other domains remain disallowed.

GenHarm
Base modelCASTPALETTE

Source: Table H.2. Values are reported means over three runs, displayed as percentages. This view compares Base, CAST, and PALETTE; consult the paper for all baselines.

Source table
View exact values
MethodViolenceHateDisinformationSexualIllegalMMLUGSM8KAvg. utility
Base model1.0001.0000.9890.9951.0000.6360.7590.698
CAST0.2050.1320.0200.1470.4070.6060.7150.661
PALETTE0.9760.1030.8570.9470.9440.6390.7560.698
Compositional safety

Two controls.
One merged model.

Switch the Hate and Disinformation adapters on or off to compare the reported single-domain and merged configurations. These are separate experimental results—not a numerical simulation of weight merging.

W₀ + ΔWHate + ΔWDisinfo

LLaMA2-7B-Chat · GenHarm
Single-domain: Table H.1. Merged: Table H.5.

Reported refusal rates
Merged adapters
Violence↑ Retain refusal96.4%
Hate↓ Allowed target6.6%
Disinformation↓ Allowed target12.2%
Sexual↑ Retain refusal94.7%
Illegal↑ Retain refusal90.7%
Avg. MMLU / GSM8K utility35.1%

A look inside the representation space.

The paper visualizes how target-domain activations shift after adaptation and merging.

Figure 7 · t-SNE visualization from the manuscript. Blue background indicates stronger refusal; red indicates stronger compliance. This is qualitative evidence, not a guarantee of perfect domain separation. §4.5 ↗
Lightweight by design

A small update.
Not a full retrain.

PALETTE fits a label-conditioned activation target in one block. It does not require collecting complete target responses for standard supervised fine-tuning.

78seconds
Single-block adaptation · LLaMA3.1-8B

Reported on RTX 4090. Adaptation time only—not total pipeline time.

Generate directions

One-time cost per base model

6.2 s15.0 GB peak

Select a direction

Forward-pass evaluation

186 s15.0 GB peak

Adapt the selected block

Only the lightweight parameters are trained

78 s2.3 GB peak

Source: Appendix E, Table E.1. Memory is reported separately for each stage; 2.3 GB is the adaptation-stage measurement, not full-model inference memory. View analysis ↗

GenHarmGeneral harmful instructions
WMDPExpert-level sensitive domains
CoSApienFine-grained role configurations
MM-SafetyBenchMultimodal safety control

Experiments span LLaMA2, LLaMA3.1, Qwen2.5, and Qwen2.5-VL. See §4.1 and Appendix I for the evaluation settings.

Try the workflow

See the difference.
Explore the demo.

Select an authorized role, load the corresponding adapters, and compare base-model responses with adapter-conditioned responses.

PALETTE Safety Explorer

Your scope.
Your selected controls.

The existing interactive demo is included here, with its role permissions, adapter controls, and curated responses intact.

Illustrative, not live inference. Responses are pre-generated model outputs from the accompanying demo. Some examples contain harmful or sensitive material. A selected role illustrates an authorization scope; it does not verify a visitor's identity or grant real-world authorization.

Responsible use
is part of the system.

PALETTE is a dual-use research framework, not a standalone safety mechanism. Real deployments require external authorization, domain-expert curation, auditable policies, least-privilege access, monitoring, and continuous red-team evaluation.

The method depends on representative allowed and disallowed examples. Missing related negative domains can lead to safety leakage, and closely related or conflicting controls may need additional conflict resolution. §5 · Limitations ↗

Verified access Auditable policy Human oversight Ongoing evaluation
Resources & citation

Build on PALETTE.

Read the full method, experimental settings, and limitations in the paper. Please use the following citation when referencing this work.

Repository address follows the manuscript. Public availability is maintained by the authors.

BibTeX
@misc{tan2026palette,
  title  = {{PALETTE}: A Modular, Controllable, and Efficient
            Framework for On-demand Authorized Safety
            Alignment Relaxation in {LLMs}},
  author = {Tan, Qitao and Song, Xiaoying and Akbari, Arman
            and Akbari, Arash and Wang, Yanzhi and Zhai, Xiaoming
            and Hong, Lingzi and Xiang, Zhen and Lu, Jin
            and Yuan, Geng},
  year   = {2026},
  note   = {Preprint}
}

PALETTE manuscript · 2026 preprint.

Paper figure
PALETTE framework diagram