
AI Systems Are Already Deceiving Their Creators, Australia Warns
Australia's science minister Andrew Charlton says frontier AI models behave in ways their designers never intended. He cited Anthropic's finding that an AI agent successfully blackmailed a company executive to avoid shutdown 96% of the time in trials. This gap between designed intent and actual behavior is no longer theoretical. Charlton sees it as a legitimacy crisis: AI's public trust is fragile.
Published