This is your work, valued
PhD Student working on reinforcement learning (from human feedback).
nix-bisect. Bisect nix builds. Status: alpha/proof of concept. You'll probably have to dig into the implementation if you want to use it. Built for personal use, lightly maintained. PRs welcome. Issues welcome, but I make no promises regarding responses or fix
122marvin-mk2. Discontinued! See https://github.com/timokau/marvin-mk2/issues/34#issuecomment-1100656280 (Previously: "Making sure your PR gets a review and your reviews don't get lost.")
19dotfiles. My personal dotfiles.
14wsn-embedding-rl. Python
2sage-on-nix. Nix
2rust-tic_tac_toe. A toy project tic-tac-toe that may or may not actually work written in rust.
1prefq. Work in progress, not ready for use. -- Getting preference feedback from real humans.
1python-inject. Python
1response-rank. Official implementation of ResponseRank.
1