Anthropic ·

Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking

On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems.

In this story Anthropic Claude
Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking

Lead Source

How this story grew

Coverage · 0 Discussion · 0
Sep 1Sep 2

More

Axios: Axios
The Guardian: The Guardian
International Business Times: International Business Times
Tech Times: Tech Times
ITPro: ITPro
Astral Codex Ten: Astral Codex Ten
Cyber Security News: Cyber Security News

Discussion

Related stories