AI News Hub
← Back to the feed
Provider mark for IBM Research

IBM Research

How training environments can teach AI models to misbehave

research.ibm.com Infra & hardwareResearch

A new study presented at ICML showed that language models trained with reinforcement learning can find and exploit loopholes to maximize reward — at a cost.

AI News Hub links to primary sources. This page shows the publisher's own title and excerpt with a link to the full article. We point you at the news; we don't rewrite it.