Stronger language models have not removed the engineering work needed to turn messy business files into information an AI agent can trust. Microsoft says production systems still need a dedicated extraction layer that preserves layout, normalizes fields and links answers back to their source.
A simple prompt may work for a few documents, but scale introduces PDFs, scans, images, tables and relationships that cross pages. Teams then have to manage chunking, confidence scores, date and currency normalization, error handling, token costs, security and audits. That gap separates a convincing demo from a system that can process millions of pages predictably.
Microsoft positions two Foundry services for different parts of the problem. Azure Document Intelligence uses purpose-built extraction for known formats such as invoices, receipts, identity documents and tax forms. Azure Content Understanding adds generative analysis across documents, images, audio and video, producing application-defined schemas with grounding and confidence signals.
The company is working on richer document understanding, lower model costs, synchronous read and layout APIs, framework integrations and human review for uncertain output. Building a custom pipeline can still make sense for a narrow workload with stable formats, but the team running it must own its evaluation and compliance burden whenever models or source material change.