reproductions & build notes
logs/
Paper reproductions and build notes across two areas — models: making the base model more capable, and agents: turning capability into product.
reproductions/
RLVR is the engine; post-training, harness, and memory are the runtime; attention is the substrate; continual learning is how knowledge stays current. Memory here is a store the agent reads and writes. Editing is changing the weights — same problem, different medium.
GRPO trains both reasoners and memory managers. GRACE and WISE sit on the hinge between editing and agent memory. Attention is the substrate long-CoT RLVR and long-horizon agents stand on.
rlvr/
Where the reward comes from, whether it can be trusted, and whether it is dense enough.
reproducedread
foundation
algorithm
debunking
test-time
domain
GRPO
Shao et al., 2024
reproducedDeepSeekMath / GRPO
Group-relative advantage, no critic — the algorithm I implemented and ran against a programmatic verifier.
posture: full reproductionpaper ↗
Earlier (recsys, outside these spines): DCN-V2 · ESMM · ColBERT
models/
foundation-model capability: post-training/RFT, model editing, inference & serving
To Forge or Not to Forge: a pre-registered study of ten-example pilots
build note · 2026-08-31
Build note on my technical report: under one frozen LoRA protocol, pair a ten-example pilot with a full fine-tune on 61 tasks and ask whether the pilot can decide which tasks are worth fine-tuning at all. Full PDF embedded.
takeaway: the pilot is a ranker with budget value, not a go/no-go gate — and the negative result was publishable because the honesty clause was pre-registered.
Engram: writing user beliefs into model weights
build note · 2026-06-27
Build note on routing personal memory between RAG and model-weight edits, then proving attribution with retrieval disabled.
takeaway: memory systems need provenance, not just recall.
agents/
harness engineering, evaluation, multi-agent supervision