A new arXiv paper tests how to ground a large language model in an industrial wastewater simulator so operators can ask causal questions such as why emissions are rising or what happens if aeration is reduced. The work compares live simulator access, structured parameter injection, and a plant-portable retriever.

On the authors’ 198-question benchmark, the live simulator oracle reached 99.5 percent accuracy, while structured injection and the retriever reached 79 percent and 75.8 percent. The strongest retrieval-augmented baseline reached 48 percent.

The result is research, not a ready plant-control product. Its importance is the deployment ladder: teams could trade accuracy, portability, and infrastructure needs depending on whether they can connect a model to a simulator, inject structured plant parameters, or use a smaller retriever trained per plant.