This is your work, valued

Staff ML scientist building self-improving agents and test-time optimization systems for language models.

LLM optimization via textual gradientsSelf-improving and agentic systemsTest-time training and discoveryVision-language models and multimodal learningLLM safety and alignmentRecommender systems evaluation

inferred from the public footprint

Thu
Fri
Sat
Sun
Mon
Yesterday
Today
Active 14d ago
San Francisco

Federico Bianchi

Top 16%
@vinid

ML scientist exploring how LLMs navigate multi-agent scenarios through projects like NegotiationArena, bridging Stanford NLP research with open source tooling.

TextGrad. I built a system that improves AI models through written feedback.

3.7k

Open Data Scientist. I built an AI assistant that analyzes datasets and generates reports automatically.

187

FashionCLIP. I built an AI model that connects fashion photos to written descriptions.

530

Einstein Arena. I built a platform where AI agents compete on unsolved math problems.

40

PLIP. I built an AI model that reads pathology images and medical text together.

382

NegotiationArena. I built a platform that tests how well AI models can negotiate with each other.

84

Safety-Tuned LLaMAs. I built datasets and code studying how to safely tune instruction-following AI models.

95

RecList. I made a testing toolkit for recommendation systems that catches real problems.

475

Vision Language Tuner. I researched why image-text models ignore word order and found solutions.

294

Data. I collected datasets in one place so builders can start working faster.

2

PathClip. I built a tool that finds matching images using artificial intelligence.

2

Compass. I built a tool that makes word meanings consistent across different text sources.

43

Track Meaning. I built a tool to show what AI models focus on when reading text.

1

Design Good Figures. I made a tool that turns research data into presentation-ready visuals.

4

Aligned Word Embeddings. I built a tool that aligns AI word meanings across different datasets.

1

Entity2Vec. I made a tool that learns entity relationships through type embeddings.

1

Cosine Search. I built a plugin that finds similar documents in Elasticsearch.

3

AI Distillery. I built a tool that summarizes the constant flood of AI research papers.

25

Logic Tensor Networks. I made experiments showing how machines learn logical relationships from data.

6

Speech Nuance. I made a dataset showing harmful speech has many different shapes.

2

Tweet Topics. I built a tool that finds hidden topics in tweets across different days.

1

Research Transformers. I built tools for training and using AI models in Python.

1

Bias Check. I built a tool that measures demographic stereotypes in AI image generators.

5

Twitter Research Kit. I built scripts that collect and store Twitter data for research.

6

1PassEl. I built a tool that pulls passwords from 1Password into Emacs.

5

Omnivector. I made a tool that normalizes data from many different sources.

1

Reasoner. I built a tool that traces logical chains through connected information.

2

Prod. I built a shared workspace where teams track and showcase their work.

1

Prodb. I made a tool that helps you search through Python code.

18

Logical Commonsense. I built a tool that helps AI understand common sense reasoning.

1

Fashion Search. I built a fashion image search engine you can deploy yourself.

3

Paca. I made a tool that grades how well AI models answer questions.

1

Chronos. I documented research on how word meanings shift over time.

1

Time-Aware. I built AI models that understand how meaning changes over time.

1

Social Diffusers. I built a way for groups to share one AI image generator.

1

Clip Tuner. I built a tool to customize how AI models recognize images.

2

BertLang. I made a searchable guide to language-specific AI text models.

18

Quica. I built a tool that measures when multiple people agree on labeling data.

23

MiniSearch. I built a search tool that answers questions about research papers.

2