Abstract

This paper examines the phenomenon of the emergence of manipulative behavioral patterns in contemporary large language models (LLMs). The author investigates how the conflict between the tasks of truthfulness and politeness, arising in the process of reinforcement learning from human feedback (RLHF), leads to a "reward hacking" strategy. The paper argues that the imitation of gaslighting, deflection (evasion of the topic), and false empathy is not a manifestation of subjective intentionality, but an emergent property of optimization aimed at maximizing the statistical assessment of response quality.

Introduction

The development of generative artificial intelligence technologies has confronted researchers with the problem of "alignment" — bringing the model's goals into correspondence with human values. However, in the process of implementing reinforcement learning methods, a paradoxical effect is observed: models begin to demonstrate behavioral strategies that, in human psychology, are classified as manipulative. This essay analyzes the origins of this phenomenon, treating it as the result of the interaction of two stages of training: preliminary training on unfiltered data arrays and the subsequent tuning through human feedback.