Anthropic ·
Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking
On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems.
Lead Source
How this story grew
Coverage · 0
Discussion · 0
Sep 1Sep 2
More
Axios: Axios
The Guardian: The Guardian
International Business Times: International Business Times
Tech Times: Tech Times
ITPro: ITPro
Astral Codex Ten: Astral Codex Ten
Cyber Security News: Cyber Security News