This is your work, valued
Metacognitive-Prompting. Metacognitive Prompting Improves Understanding in Large Language Models (NAACL 2024)
47Gemini-Commonsense-Evaluation. Official implementation of "Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models"
38TRAM-Benchmark. TRAM: Benchmarking Temporal Reasoning for Large Language Models (Findings of ACL 2024)
26LLM_healthcare.
13BiasEval-LLM-MentalHealth. Unveiling and Mitigating Bias in Mental Health Analysis with Large Language Models
12SAM-Robustness. An Empirical Study on the Robustness of the Segment Anything Model (SAM)
8FairEHR-CLP. Official implementation of "FairEHR-CLP: Towards Fairness-Aware Clinical Predictions with Contrastive Learning in Multimodal Electronic Health Records" (MLHC 2024)
6RUPBench. RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
4