Benchy is a proposed language and execution engine for describing AI benchmarks in a portable, reproducible form. It defines each evaluation as three explicit parts—a program, a scoring function and a dataset—kept separate from the AI system being tested.
Authors write benchmarks in canonical YAML, where each concept has one valid syntax and tasks are classified through a shared ontology. A deterministic compiler produces a canonical JSON representation without repairing invalid definitions or adding hidden defaults. External systems adapt to one runtime contract: named input fields go in and named output fields come back.
That separation could make it easier to compare systems without allowing integration details to silently change what an evaluation measures. It also makes failures part of the benchmark’s stated semantics rather than ad hoc runner behavior. The paper specifies the object model, validation rules, scoring and first implementation contract; it does not yet demonstrate that the language covers every evaluation style. Adoption will depend on tooling and whether benchmark authors find the stricter format practical.