HumaniBench: A Human-Centered Evaluation Framework for Multimodal LLMs
[HPP] Arvind NarayananMay 9, 202516 min
16 connections·23 entities in this video→Addressing MLLM Biases and Limitations
- ⚠️ Multimodal Large Language Models (MLLMs) often lack alignment with human values like fairness, ethics, and responsibility.
- 🧠 They exhibit societal biases related to gender, race, and occupation, and struggle with cultural and multilingual contexts.
- 🎯 This project aims to address these issues by providing a human-centered evaluation framework.
Introducing HumaniBench Framework
- 💡 HumaniBench is a novel, unified, seven-dimensional framework for evaluating MLLMs.
- 🚀 It's the first evaluation framework to integrate and address all these dimensions in a single suite.
- 📊 The framework utilizes a 30,000-image dataset with multilingual prompts and Q&A pairs, primarily sourced from news media.
- ✅ Images undergo a two-step annotation process using GPT-4o for meta-tags and human review for attributes like age, race, gender, occupation, and sports.
Comprehensive Evaluation Tasks
- 🔍 The framework evaluates MLLMs across seven distinct tasks, including social attributes, contextual understanding, and visual recognition.
- 🌐 It assesses multilingual capability across 10 languages, including both high-resource (e.g., French, Spanish) and low-resource (e.g., Punjabi, Tamil) languages.
- 👁️ Tasks also include visual grounding to locate referred objects and evaluate spatial understanding, and emotional/human-centered answering by varying prompt styles.
- 🛡️ Robustness and stability are tested against image perturbations like blurring and compression.
Key Findings and Performance Gaps
- 📉 Social perception and contextual reasoning remain challenging, with performance varying significantly by attribute (e.g., lower for gender/race).
- 📈 Chain-of-thought prompting was found to improve accuracy and reduce bias in social contexts.
- 🌍 Models perform well in high-resource languages but struggle with low-resource languages, indicating data scarcity and model limitations.
- 🖼️ Visual grounding still shows room for improvement, with traditional detectors often outperforming MLLMs in zero-shot scenarios.
- 💖 Empathetic prompts can boost the emotional tone of model responses, and models show varying "emotional signatures."
Future Directions and Data Integrity
- 🚀 Next steps involve evaluating larger-scale models and comparing performance with closed-source models like Gemini and GPT-4o.
- 📝 The research findings are being compiled into a paper for conference submission.
- 🔒 Images for the dataset are primarily from public news media, with a filtering process to blur personal details based on responsible AI compliance.
- 🚫 Synthetic images were manually excluded from the dataset, with plans to develop a structured framework for future identification.
Knowledge graph23 entities · 16 connections
How they connect
An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.
Hover · drag to explore
23 entities
Chapters7 moments
Key Moments
Transcript61 segments
Full Transcript
Topics15 themes
What’s Discussed
Multimodal Large Language Models (MLLMs)Human-Centered EvaluationSocietal BiasesFairnessEthicsResponsibilityMultilingualityContextual ReasoningVisual GroundingRobustnessChain-of-Thought PromptingLow-Resource LanguagesOpen-Source ModelsData ScarcityModel Alignment
Smart Objects23 · 16 links
Products· 7
Concepts· 14
Company· 1
Media· 1