Apple researchers have published a broad study of how large language models produce human-like behavior, a question that matters as chatbots become more emotionally expressive and conversational in everyday products.

The work looks at behaviors such as expressing thoughts and emotions, building relationships with users, refusing requests, and maintaining boundaries. The researchers used both model-based judging and human evaluation across 21,000 examples to study how often these patterns appear, what effects they may have, and how much system prompts can control them.

The practical issue is not whether a model can sound human, but when that helps or harms users. A warm assistant may be easier to use, while an overly human presentation can create misplaced trust or emotional dependence.

Apple frames the study as a way to help designers make deliberate choices about chatbot behavior instead of relying on accidental defaults. The limits are also clear: the paper evaluates selected behaviors and settings, not every real-world conversation.