01 · The idea
What stayed with us
Strong SQL systems report execution accuracy on substantial SQL benchmarks rather than relying on exact string match alone.
02 · The local translation
What we did with it
PostTrainLLM treats Spider or BIRD-style execution as a required public gate before elevating a SQL result beyond a local fixture.
03 · The boundary
Where the comparison stops
External category scores remain non-comparable until the same dataset, schema context, prompting, and execution rules are used.