OpenRouter published an analysis of image input detail levels for multimodal LLMs, using 1,730 visual reasoning questions across five models.

The test found that setting images to low detail can meaningfully hurt accuracy and does not always reduce cost; on GPT-5.5, OpenRouter says the bill increased. The more reliable cost lever was reasoning effort.

The result is useful for developers tuning multimodal applications because image compression, detail settings, and reasoning budgets can interact in non-obvious ways.