Cloudflare has released Clef-omni, an open-weight model that can evaluate audio, video, images and text together and return answers in a predefined structure. It is designed for classification and operational decisions rather than free-form writing.

A developer can provide several media files plus questions whose valid responses are specified in advance. Cloudflare’s example checks an appliance photo, recording and video to decide whether a label is visible, the machine sounds normal and its fan is turning. Handling all inputs in one model removes the need to transcribe audio or split a video into separate pipelines first.

Clef-omni uses the comprehension components of Qwen3-Omni-30B-A3B-Instruct but discards its text-to-speech output. Instead of generating a sentence token by token, it scores the permitted values for each question. The weights are available on Hugging Face, and the hosted model is accessible through Workers AI.

Cloudflare also says the original Clef now runs up to twice as fast and has reduced the hosted price of Clef-flash. These are vendor performance claims, so teams should test accuracy and latency on their own media and decision sets before replacing specialized pipelines.