CLIcK. CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
48BenchHub. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
10CoTrace. "I didn’t Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration (COLM 2026)
8SCRIPTS. Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues (ACL 2026)
1