Large Language Models (LLMs) have reached significant milestones in capability, yet they remain susceptible to jailbreak attacks designed to elicit harmful or unsafe outputs. Current safety alignment techniques, such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often result in "shallow" safety. These models frequently learn to recognize and refuse specific patterns rather than understanding the underlying safety principles, leading to vulnerabilities and the unintended refusal of benign queries.
Introducing SSRFT
To tackle these limitations, researchers have introduced SSRFT (Supervised Safe-Role Fine-Tuning). This framework reformulates safety alignment as the internalization of a predefined "safe role." Instead of training a model on what not to say, SSRFT focuses on training the model to embody a persona that adheres to safety-oriented values and principles.
The process involves constructing a Safe-Role Question-Answer (SRQA) dataset derived from:
- Psychometric questions.
- A limited set of jailbreak prompts.
- Detailed safe-role descriptions.
Performance and Generalization
Experimental results across multiple Base and Instruct models demonstrate that SSRFT achieves more robust and generalizable safety alignment compared to standard SFT. Notably, the framework shows superior resistance to prefilling attacks and generalizes better to unseen jailbreak domains. Furthermore, SSRFT effectively reduces over-refusal on benign queries while preserving the model's core performance, establishing role-internalization as a viable alternative to traditional refusal-centric methods.