A new arXiv paper audits whether answer correctness is enough to judge legal language models when those models also mention legal authority. The authors tested four LLMs on 238 Taiwan bar-examination items with verified governing provisions.

The results show that correct final answers and correct authority grounding can diverge. In criminal law, 24.0% to 42.4% of valid responses were answer-correct but missed the gold authority, while 15.2% to 21.7% were answer-incorrect but cited it.

That mismatch matters because legal reasoning often depends on the statute or precedent behind an answer. A benchmark that scores only the final choice may treat an ungrounded answer as a full success.

The authors argue that authority grounding can be audited automatically when the governing provision is externally verifiable. The broader implication is that legal AI benchmarks need to measure both conclusion and support.