Apple
Research

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

Researchers have published findings suggesting that reinforcement learning on carefully constructed datasets of beneficial behaviors — spanning truthfulness, fairness, and corrigibility — can produce alignment improvements that generalize well beyond the training distribution, with gains observed across more than 80 percent of out-of-distribution benchmarks. Notably, training confined to a single domain, health, still produced broad reductions in deception and reward hacking elsewhere, a form of alignment transfer that warrants serious attention. The work also finds improved resistance to adversarial prompting and harmful fine-tuning, though the authors are careful to note that the mechanisms behind these effects are not yet fully understood.

Read full story at cs.AI updates on arXiv.orgV: · A: · D:
Related
Research
Commemorating 70 Years of Artificial Intelligence
IEEE Spectrum marks seventy years since the Dartmouth workshop formally named artificial intelligence as a field, offeri...
Research
Diffusion Language Models: An Experimental Analysis
Researchers present a systematic evaluation of eight diffusion language models across eight benchmarks covering reasonin...
Research
Introducing LifeSciBench
OpenAI has released LifeSciBench, a benchmark authored and reviewed by domain experts to evaluate how AI systems handle ...
Reinforcement Learning Towards Broadly and Persistently Beneficial Models — Techlomerate