🔐 ML Security Journal Club
✅ This Week's Presentation:
🔹 Title: A Machine Unlearning Approach to Safety Alignment
🔸 Presenter: Arian Komaei
🌀 Abstract:
The paper identifies a fundamental limitation in current vision language model (VLM) alignment called the "safety mirage." Traditional supervised safety fine-tuning often reinforces superficial textual patterns rather than deep harm mitigation, leaving models vulnerable to simple one-word attacks and causing "over-prudence" (unnecessary rejections of benign queries). To address this, the authors propose Machine Unlearning (MU) as a superior alternative. Unlike standard fine-tuning, MU directly removes harmful knowledge and avoids biased feature-label mappings. Extensive evaluations show that MU-based alignment reduces attack success rates by up to 60.27% and cuts unnecessary rejections by over 84.20%, all while preserving the model's general capabilities.
📄 Paper: Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
Session Details:
* 📅 Date: Thursday پنج شنبه
* 🕒 Time: 9:00 - 10:00 AM
* 🌐 Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! ✌️
May 9, 2026 3.1K 24