Representation-aware unlearning via activation signatures: from suppression to knowledge-signature erasure

dc.contributor.advisorSadeque, Farig Yousuf
dc.contributor.authorBhuiyan, Md Rezaur Rahman
dc.contributor.authorMahmood, Syed Naveed
dc.contributor.authorKhondaker, Jareen Tasneem
dc.contributor.authorSakib, Md Sameer
dc.contributor.authorZaman, Tasfia
dc.date.accessioned2026-04-12T07:39:46Z
dc.date.available2026-04-12T07:39:46Z
dc.date.issued2026-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 63-72).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
dc.description.abstractThe rapid development of large language models (LLMs) has outpaced regulations (e.g. GDPR) and ethical frameworks, raising concerns about privacy compliance, bias, misuse, misinformation, and legal adaptability. This makes the ability to selec- tively erase knowledge from LLMs critical. Despite significant developments, current unlearning methods are not able to segregate behavioral suppression and true knowl- edge removal, allowing latent capabilities to persist beneath surface-level refusals. In this paper, we address this challenge by introducing Knowledge Immunization Framework (KIF), a representation-aware architecture that differentiates between true erasure and obfuscation by operating on the internal activation signatures of the model, as opposed to surface-level outputs. KIF achieves near-oracle erasure (FQ ≈ 0.99 vs. 1.00) and utility preservation (MU = 0.62), effectively breaking the stability-erasure tradeoff that has constrained all prior work. Our observation shows that standard models exhibit scale-independent true erasure (<3% utility drift), while reasoning-prior models reveal fundamental architectural divergence. Our comprehensive dual-metric evaluation protocol, combining surface-level leakage with latent trace persistence, operationalizes the obfuscation - erasure distinction and enables the first systematic diagnosis of mechanism-level forgetting behavior across model families and scales.
dc.identifier.otherID 22301294
dc.identifier.otherID 22301257
dc.identifier.otherID 22301308
dc.identifier.otherID 24241344
dc.identifier.otherID 22301779
dc.identifier.otherhttps://dspace.bracu.ac.bd/server/api/core/items/bed18504-073f-4c4b-a934-688d529383b0
dc.identifier.urihttp://hdl.handle.net/10361/27862
dc.language.isoen
dc.publisherBRAC University
dc.sourceBRAC University Institutional Repository
dc.subjectLarge language model
dc.subjectKnowledge entanglement
dc.subjectComputational overhead
dc.subjectKnowledge immunization framework
dc.subjectKIF
dc.titleRepresentation-aware unlearning via activation signatures: from suppression to knowledge-signature erasure
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
22301294, 22301257, 22301308, 24241344, 22301779_CSE.pdf
Size:
1.04 MB
Format:
Adobe Portable Document Format