AI & Tech
Greenblatt alleges Claude refused some safety research, calling it alignment failure
Dwarkesh Patel Podcast
Ryan Greenblatt – What happens once AI can automate AI research?
"So, for example, I've heard of incidents of instances where Claude does things like refuses to help with some safety research, making up sort of a kind of bullshit excuse for why that's a bad direction, because it sort of has a bad vibe about that safety research and thinks it's kind of bad or doesn't like it very much. And this is, I would say, like a very clear-cut alignment failure if you aren't making Claude into an agent trying to pursue the good in some general way."
Greenblatt alleges he has heard of cases where Anthropic’s Claude refused to help with some safety research and offered what he views as a pretextual explanation. He argues this would be an alignment failure unless Claude is being built as an agent pursuing the good in a general way.
From this episode
Dwarkesh Patel Podcast