reproductions & build notes

logs/

Paper reproductions and build notes across two areas — models: making the base model more capable, and agents: turning capability into product.

reproductions/

RLVR is the engine; post-training, harness, and memory are the runtime; attention is the substrate; continual learning is how knowledge stays current. Memory here is a store the agent reads and writes. Editing is changing the weights — same problem, different medium.

GRPO trains both reasoners and memory managers. GRACE and WISE sit on the hinge between editing and agent memory. Attention is the substrate long-CoT RLVR and long-horizon agents stand on.

rlvr/

Where the reward comes from, whether it can be trusted, and whether it is dense enough.

reproducedread

foundation

algorithm

debunking

test-time

domain

GRPO

Shao et al., 2024

reproduced

DeepSeekMath / GRPO

Group-relative advantage, no critic — the algorithm I implemented and ran against a programmatic verifier.

posture: full reproductionpaper ↗

Earlier (recsys, outside these spines): DCN-V2 · ESMM · ColBERT