A note of thanks · Training loops

Castform

Castform's trace and reward framing helped us draw a more inspectable route from an agent failure to training data. The local translation remains scaffolding until a frozen evaluator proves a gain.

01 · The idea

What stayed with us

Composite rewards, trace-driven data loops, reasoning-depth classification, and explicit environment contracts make reinforcement-style fine-tuning more inspectable.

02 · The local translation

What we did with it

PostTrainLLM mapped those ideas into composite reward scaffolding, trace-to-data workflows, and reasoning-depth classification.

03 · The boundary

Where the comparison stops

The hosted dashboard and pay-per-compute product model were not copied, and the scaffolding is not proof that an RLVR run improves a specialist.