01 · The idea
What stayed with us
Shape-specific Metal kernels, a model-specific speculative draft, packed weights, cache reuse, and a startup memory plan trade generic model support for Mac-local serving performance.
02 · The local translation
What we did with it
The retained exercise compares a specialized engine with a general engine on cold and cached latency, prefill, decode, concurrency, memory, output validity, and task completion.
03 · The boundary
Where the comparison stops
A model-specific runtime can be fast without being a general serving replacement, and tuning cost must be included beside runtime gains.