Rare find

FlagEvalMM. Test AI models that work with images, video, and text all at once.

github.com/flageval-baai/FlagEvalMM

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

March 2026
  • fix text vqa (#126)
  • fix mmvet_v2 and upgrade judge model to gpt-5-mini (#125)
  • add skill for integrating new datasets into FlagEvalMM (#124)
  • add MeasureBench task for visual measurement reading evaluation (#123)
February 2026
  • add RelScene dataset (#121)
January 2026
  • refactor: pass annotation dict to prompt template functions (#120)
  • update tasks config (#119)
  • fix pyav video reader for FLV files (#117)
  • add Animalbench (#116)
  • fix: task name
  • Fix ARPGrounding ...
  • fix: remove incorrect folder name arpggrounding
  • Add ARPGrounding visual grounding task
  • fix: remove unused variable and format code with black
  • fix erqaplus name (#114)
December 2025
  • point norm1k (#111)
  • ERQAPlus (#108)
  • add revel evaluator (#106)
October 2025
  • remove decord in dependency (#103)
  • fix t2i evaluator (#102)