This is your work, valued
VSR-guided-CIC. Human-like Controllable Image Captioning with Verb-specific Semantic Roles.