Microsoft Research introduced Lens, a text-to-image model with 3.8 billion parameters that performs competitively with much larger systems. The reported differentiator is training data: 800 million detailed captions generated by GPT-4.1 rather than sparse web alt-text, suggesting data quality can offset some scale demands.
Jun 8, 2026
Microsoft Lens shows caption quality can beat raw scale
Source
The Decoder