01 · The idea
What stayed with us
Ternary representation, packing overhead, activation transforms, kernel support, artifact size, runtime memory, and capability retention can rank a model differently depending on the deployment constraint.
02 · The local translation
What we did with it
The retained exercise is a same-Mac comparison sheet covering revisions, task gates, peak RAM, time to first token, prefill, decode throughput, and total task latency.
03 · The boundary
Where the comparison stops
File size is not peak runtime memory, and retention relative to a source model does not establish frontier parity.