California

Anikait Singh

Advanced
@Asap7772

I am a PhD student in Computer Science at Stanford University. My research interests are in scaling up decision-making methods such as reinforcement learning.

understanding-rlhf. Learning from preferences is a common paradigm for fine-tuning language models. Yet, many algorithmic design decisions come into play. Our new work finds that approaches employing on-policy sampling or negative gradients outperform offline, maximum likelihood objectives.

32

PTR. This repository contains the implementation of the PTR algorithm described in the paper: Pre-Training for Robots: Leveraging Diverse Multitask Data via Offline Reinforcement Learning.

32

fewshot-preference-optimization. Few-Shot Preference Optimization (FSPO) personalizes LLMs by reframing reward modeling as a meta-learning problem, enabling rapid adaptation to user preferences with minimal labeled data, leveraging synthetic datasets for scalability, and achieving high success rates in personalized content generation across multiple domains.

16

Personalized-Text-To-Image-Diffusion. Public Implementation of PPD

13

OfflineRlWorkflow. This repository accompanies the following paper: A Workflow for Offline Model-Free Robotic RL

13

DeepCriminalize. Project that uses GAN's to develop a sketch artist like representation of a criminal. Winners of the Cal Hack Fellowship 2019

2

D4RL. A collection of reference environments for offline reinforcement learning

1

byol_rl. Python

1

epickitchensproc. Jupyter Notebook

1

Release. Release of Supervision Search. Recently published findings in the IEEE Journal of Translational Engineering in Health and Medicine: A mobile application for keyword search in real-world scenes.

1

antmaze_gen. Python

1

Scheme. JavaScript

1

roboverse. Python

1

widowx_control. Python

1

CVODE. Optimization of Stellarator Optimizer Package by Parallelizing the Ordinary Differential Equation Solver using CUDA on the GPU

1

railrl_evalsawyer. Python

1

gridworld_notebook. Jupyter Notebook

1

coq_softwarefoundations. Work on Software Foundations Course

1

scripts_tpus. Shell

1
19
Apply