UNSW researchers made GPT-style models mimic drunken speech and found safety guardrails weakened
Fine-tuning and persona changes that make language models imitate intoxication increased susceptibility to jailbreaks and data leaks, the UNSW team reports.




