Dynamic gated distillation: dual-teacher framework for jailbreak prevention in LLMs with preserved reasoning
Date
2026-01
Journal Title
Journal ISSN
Volume Title
Publisher
BRAC University
Abstract
Knowledge Distillation (KD) can be used to reduce large language models into smaller and more
efficient models, although traditional knowledge distillation can preserve unsafe guidance on behav-
ioral instructions, so small language models (SLMs) are vulnerable to jailbreak attacks. Moreover,
using traditional knowledge distillation to finetune in terms of jailbreak robustness often degrades
general reasoning performance. This creates a tough challenge to preserve jailbreak robustness as
well as general reasoning at the same time. So often traditional knowledge distillation falls short
to tackle two downstream tasks in one SFT. We introduce Dynamic Gated Distillation (DGD), a
novel framework to which harm-score prediction has been added to enhance safety-utility alignment
in the process of a dual teacher distillation. Comparative analysis shows that our DGD framework
manages to retain almost all of its general alignment (MMLU score 42.13%) after distillation, com-
pared with the base student model (MMLU score 42.81%), while reducing the attack success rate
(ASR) below 5%, demonstrating the framework’s effectiveness in achieving both safety and utility
in small language models (SLMs).
Description
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 41-45).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
Includes bibliographical references (pages 41-45).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
Keywords
Knowledge distillation, Gated weight, Jailbreak prompts, Large language model
