calibration-on-disagreement-data. Code accompanying the EMNLP 2022 paper "Stop Measuring Calibration When Humans Disagree" in which we show problems with popular calibration metrics like ECE in settings where more than one answer is acceptable, and argue for several metrics that take into account the full human judgement distribution.

github.com/jsbaan/calibration-on-disagreement-data

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.