Apple researchers found that natural language can be an unreliable channel when one AI model passes structured information to another. Their test asked one model to turn a tree-shaped arithmetic expression into a word problem and a second model to reconstruct the original expression.

Across every pairing of 16 models, reversing which model wrote and which extracted changed accuracy by as much as 60.4 percentage points. The best pair reached 92.9% by using different models at each end. At least 73.6% of failures began during generation, while harder tree structures drove difficulty.

Fine-tuning with roughly 3,600 examples improved every tested open-weight model beyond an untrained Gemini 3.1 Pro under matched semantics. Gains also carried into a setting with new operators and vocabulary, although a gap to frontier performance remained. Multi-agent systems should not assume that a fluent explanation preserves the hierarchy of the data behind it.