A note of thanks · Scale boundary

teale and Petals

These decentralized inference projects make the latency cost of sharding a model across machines tangible. The comparison helps explain why a single-Mac specialist is a different system.

01 · The idea

What stayed with us

These decentralized inference projects make the latency cost of sharding a model across machines tangible. The comparison helps explain why a single-Mac specialist is a different system.

02 · The local translation

What we did with it

Together these projects illustrate two shapes of decentralized inference: sharding work across machines and making complete models available across a network. The distinction sharpens our understanding of per-token latency and what locality buys a small specialist.

03 · The boundary

Where the comparison stops

We have not installed or benchmarked either system for PostTrainLLM. Their networked design is a comparison at the distributed boundary, not a substitute for a verified local decode and task-completion gate.