I am a PhD student in Computer Science at Stanford University. My research interests are in scaling up decision-making methods such as reinforcement learning.
understanding-rlhf. Learning from preferences is a common paradigm for fine-tuning language models. Yet, many algorithmic design decisions come into play. Our new work finds that approaches employing on-policy sampling or negative gradients outperform offline, maximum likelihood objectives.
32PTR. This repository contains the implementation of the PTR algorithm described in the paper: Pre-Training for Robots: Leveraging Diverse Multitask Data via Offline Reinforcement Learning.
32fewshot-preference-optimization. Few-Shot Preference Optimization (FSPO) personalizes LLMs by reframing reward modeling as a meta-learning problem, enabling rapid adaptation to user preferences with minimal labeled data, leveraging synthetic datasets for scalability, and achieving high success rates in personalized content generation across multiple domains.
16Personalized-Text-To-Image-Diffusion. Public Implementation of PPD
13OfflineRlWorkflow. This repository accompanies the following paper: A Workflow for Offline Model-Free Robotic RL
13DeepCriminalize. Project that uses GAN's to develop a sketch artist like representation of a criminal. Winners of the Cal Hack Fellowship 2019
2D4RL. A collection of reference environments for offline reinforcement learning
1byol_rl. Python
1epickitchensproc. Jupyter Notebook
1Release. Release of Supervision Search. Recently published findings in the IEEE Journal of Translational Engineering in Health and Medicine: A mobile application for keyword search in real-world scenes.
1antmaze_gen. Python
1Scheme. JavaScript
1roboverse. Python
1widowx_control. Python
1CVODE. Optimization of Stellarator Optimizer Package by Parallelizing the Ordinary Differential Equation Solver using CUDA on the GPU
1railrl_evalsawyer. Python
1gridworld_notebook. Jupyter Notebook
1coq_softwarefoundations. Work on Software Foundations Course
1scripts_tpus. Shell
1