← Feed

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned - Anthropic

rss:anthropic · Aug 22, 2022 · source