KAIST AI Uncovers Its Hidden Weaknesses, Making Generative Models Safer
KAIST researchers have unveiled Stable-GFlowNet (S-GFN), a safety verification framework designed to stress-test large language models more effectively than traditional red-teaming. Red-teaming works by generating prompts that try to trigger unsafe or harmful outputs, but it often struggles to explore the full space of possible failures—especially when training collapses onto a small set of "easy" […]
