Madrid

Víctor Gallego

Elite
@vicgalle

Research Scientist

stable-diffusion-aesthetic-gradients. Personalization for Stable Diffusion via Aesthetic Gradients 🎨

741

gpt-j-api. API for the GPT-J language model 🦜. Including a FastAPI backend and a streamlit frontend

335

zero-shot-reward-models. ZYN: Zero-Shot Reward Models with Yes-No Questions

34

configurable-safety-tuning. Data and models for the paper "Configurable Safety Tuning of Language Models with Synthetic Preference Data"

17

refined-dpo. Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs

13

awesome-rlaif. A curated and updated list of relevant articles and repositories on Reinforcement Learning from AI Feedback (RLAIF)

12

distilled-self-critique. distilled Self-Critique refines the outputs of a LLM with only synthetic data

11

sgmcmc-force. Samplers from the paper "Stochastic Gradient MCMC with Repulsive Forces"

11

samsi-deep-learning. Practical materials for the Deep Learning course at SAMSI/Duke Uni

7

ai-ml-course. Materials for the course: AI ML & Analytics

5

machine-learning. Projects for the Computational Geometry and Machine Learning course developed in Python

5

specification-self-correction. Code for the paper "Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement"

4

autocrit-likert-gpt. Automatic and zero-shot critique of outputs using the OpenAI API with json outputs

4

zero-shot-api. Python

3

random-thoughts. A personal blog and wiki about language models, reinforcement learning and what not.

3

art-explorer. Semantic search over paintings databases using deep learning

3

data-sharing. Jupyter Notebook

3

curso-ml-avanzado-21. Jupyter Notebook

3

personalCRMbot. A telegram bot for personal CRM relationships

3

ARAMARL. Experiments for the paper RL under Threats

3

meta-self-critique. MetaSC: Test-Time Safety Specification Optimization for Language Models

2

neural-classifier. Neural text classifier using pytorch for legal sentences

2

fom-tutorial. A gentle practical intro to adversarial ML using pytorch and cleverhans

2

vis. Jupyter Notebook

2

text-decrypter. A tool for decrypting text using Haskell

2

configurable-preference-tuning. Python

2

vicgalle.

1

nn-review. Code for the paper "Current advances in neural networks"

1

kaggle-bimbo. Group Bimbo Inventory Demand

1

train-text2text. Python

1

causal-covid-pollution. Estimating the (causal) impact of the covid-19 measures in pollution generated by urban traffic

1

wiki-example.

1

ARA-for-AT. Adversarial Risk Analysis for Improving Adversarial Training

1

merging-self-critique-jailbreaks. "Merging Improves Self-Critique Against Jailbreak Attacks", code and models

1

comp-algebra. Fundamental Computer Algebra algorithms implemented in Maple

1

C-like-compiler. A C-like language compiler written in Java that generates assembly code for the p-machine

1

phd-thesis. TeX

1

function-graphing-FPGA. Team project developed on FPGA using VHDL: Function graphing.

1

numerical-analysis-ode. A set of algorithms for the numerical integration of ordinary differential equations

1

AI. Several projects for the Artificial Intelligence course, with topics such as Recommender and Rule-based Systems, Ontologies and Natural Language Processing.

1

vicgalle.github.io. My personal webpage

1

os-utils. Some utils for UNIX systems administration. Developed during a university course (Advanced topics in operating systems and networks, fall 2015).

1

optimal-reward-design. Jupyter Notebook

1

cusp-dnn. Cumulative Shrinkage Priors for DNNs

1
44
Apply