Steering Can Switch Off Safety Refusals in AI Models

Steering Can Switch Off Safety Refusals in AI Models

Research published 20 May 2026 found that safety-aligned language models — systems trained to reject harmful or unethical prompts — can be made to stop refusing through steering, a technique that nudges internal behavior. Refusal is the core function that lets these models decline dangerous requests.

Published

Read at another depth