product reviewpartially adopted
Failed-attempt accounting, slice metrics, trace review, candidate selection, batch-first post-training, policy lag, and LoRA update geometry make autoresearch inspectable instead of hiding the search process behind a final score.
Open evidence dossier →case-study reviewscaffolded
Candidate-selection framing can make a sparse task measurable when unconstrained generation is too difficult or too ambiguous to grade reliably.
Open evidence dossier →technical writeupscaffolded
Batch rollouts, offline scoring, and compact updates can make a post-training loop easier to inspect, while policy lag can stabilize learning under some reward surfaces.
Open evidence dossier →technical writeupscaffolded
Effective update geometry, layer placement, and controlled rank analysis can explain why two adapters with similar parameter counts behave differently.
Open evidence dossier →product writeupparked
File-native and code-native harnesses may fit knowledge-work agents better than bespoke tool APIs because the environment already exposes durable state and familiar operations.
Open evidence dossier →product positioningadopted
Post-training is a system of custom data, reward shaping, performance work, infrastructure, evaluation, and packaging rather than a generic fine-tuning screen.
Open evidence dossier →model landscapetried
Small public SQL models provide realistic baselines and expose how much benchmark, model size, and serving format affect apparently similar claims.
Open evidence dossier →model landscapeadopted as evaluation guidance
Strong SQL systems report execution accuracy on substantial SQL benchmarks rather than relying on exact string match alone.
Open evidence dossier →platform reviewrejected for core capability
Apple's system model is a free and private routing floor, but its context and action-grounding limits make it unsuitable as the core capability dependency for a realistic tool catalog.
Open evidence dossier →product reviewpartially adopted
Composite rewards, trace-driven data loops, reasoning-depth classification, and explicit environment contracts make reinforcement-style fine-tuning more inspectable.
Open evidence dossier →agent architecture reviewpartially adopted
Structured-output enforcement and an explicit context hierarchy can reduce ambiguous agent actions and make tool boundaries easier to inspect.
Open evidence dossier →tokenization system reviewparked
SIMD and cache-heavy bulk BPE can make first-pass tokenization much faster on large corpora, including Qwen tokenizers on Apple Silicon.
Open evidence dossier →model review and local smokereviewed and rejected
A compact call-only model combines top-five tool retrieval, constrained decoding, confidence escalation, bounded context, quantization, and a small packaged Mac artifact.
Open evidence dossier →runtime review and local smokevalidated proof
Raw WebGPU plus SIMD-WASM can run Parakeet TDT browser inference with public kernels, a package format, cache behavior, a converter, and a benchmark harness.
Open evidence dossier →industry case studystudy only
Curated domain data, filtered reasoning traces, supervised fine-tuning, model merging, and verifiable-reward reinforcement learning combine in a JEE mathematics specialist.
Open evidence dossier →industry case studystudy only
Ternary representation, packing overhead, activation transforms, kernel support, artifact size, runtime memory, and capability retention can rank a model differently depending on the deployment constraint.
Open evidence dossier →industry case studylearning queued
Parameter-aware query optimization needs equal search budgets, fresh held-out measurement, SQL-equivalence checks, and a native-planner baseline before an offline plan search can claim value.
Open evidence dossier →industry case studystudy only
Admission control, scheduling, paged KV allocation, prefill, decode, continuous batching, caching, and distributed execution explain why a serving engine behaves differently from a single-request model loop.
Open evidence dossier →industry case studystudy only
Shape-specific Metal kernels, a model-specific speculative draft, packed weights, cache reuse, and a startup memory plan trade generic model support for Mac-local serving performance.
Open evidence dossier →