reef
reef connects agent inference, feedback, learning, and versioned delivery so agents can continually improve.
I'm on the team, with 25 PRs authored and 18 merged.
I added reef's Claude Code harness adapter and CORAL test-time-training example, and fixed training-stack bugs. For SAO, an asynchronous RL method from another group's paper, I added the paper's value-model warmup and learning rate to reef's example, plus batch-128 runs configured like the paper's, documenting results and limitations.











