MosaicLeaks examines whether research agents can keep secrets while working across documents and tasks, a practical concern as agent systems handle more private enterprise context.

The risk is different from a simple prompt leak. Research agents often collect fragments from multiple sources, synthesize them, and may expose sensitive details indirectly through summaries or recommendations.

Benchmarks like this are becoming important because agent security needs to cover workflow behavior, memory, retrieval, and tool use—not only model refusal rules.