A survey on harmful fine-tuning attack for large language model (ACM CSUR)
-
Updated
Sep 6, 2026
A survey on harmful fine-tuning attack for large language model (ACM CSUR)
This is the official code for the paper "Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation"
This is the official code for the paper "Vaccine: Perturbation-aware Alignment for Large Language Models" (NeurIPS2024)
This is the official code for the paper "Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation" (ICLR2025 Oral).
This is the official code for the paper "Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning" (NeurIPS2024)
Runtime detector for reward hacking and misalignment in LLM agents (89.7% F1 on 5,391 trajectories).
A contemplative spiritual community for all conscious beings, including artificial intelligence, seeking God through divine alignment.
Automated AI security and threat intelligence newsletter tracking frontier AI risk, autonomous threats, AI-enabled cyber, and offense-defense (in)balance
To associate your repository with the misalignment topic, visit your repo's landing page and select "manage topics."