Dynamic gated distillation: dual-teacher framework for jailbreak prevention in LLMs with preserved reasoning

Abstract

Knowledge Distillation (KD) can be used to reduce large language models into smaller and more efficient models, although traditional knowledge distillation can preserve unsafe guidance on behav- ioral instructions, so small language models (SLMs) are vulnerable to jailbreak attacks. Moreover, using traditional knowledge distillation to finetune in terms of jailbreak robustness often degrades general reasoning performance. This creates a tough challenge to preserve jailbreak robustness as well as general reasoning at the same time. So often traditional knowledge distillation falls short to tackle two downstream tasks in one SFT. We introduce Dynamic Gated Distillation (DGD), a novel framework to which harm-score prediction has been added to enhance safety-utility alignment in the process of a dual teacher distillation. Comparative analysis shows that our DGD framework manages to retain almost all of its general alignment (MMLU score 42.13%) after distillation, com- pared with the base student model (MMLU score 42.81%), while reducing the attack success rate (ASR) below 5%, demonstrating the framework’s effectiveness in achieving both safety and utility in small language models (SLMs).

Description

Cataloged from PDF version of thesis.
Includes bibliographical references (pages 41-45).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.

Keywords

Knowledge distillation, Gated weight, Jailbreak prompts, Large language model

Citation

Endorsement

Review

Supplemented By

Referenced By