AWS has released an open-source benchmark harness that compares AI models by the cost of a successful result rather than the list price of a million tokens. It tests GPT-5.6 Luna, Terra and Sol on Amazon Bedrock alongside GPT-5.4 Mini and Nano through one Responses API client.

In a 60-question mathematics sample, Sol answered 75 percent correctly versus Mini’s 37 percent. After July price reductions, Luna’s observed cost per correct answer was $0.0021, compared with $0.0139 for Mini. A 50-question web-research test exposed another factor: Mini averaged 7.6 turns and 114,000 input tokens because each turn resent growing context. Luna cost $0.05 per passing answer versus Mini’s $0.40 in that run.

On 48 professional-document tasks, Luna passed 27 and Mini passed 20, at observed costs of $0.010 and $0.030 per pass respectively. These are practical configurations, not controlled measures of intrinsic capability: reasoning was disabled for Bedrock models, sample sizes were small, and several longer outputs hit an 8,192-token cap. AWS records prompts and results for reproduction and recommends rerunning the harness on 50 to 100 representative tasks before choosing a model.