中文聚焦VLM:视觉语言模型中的视线追踪与社会性注视预测基准测试
ENEyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models
视觉语言模型(VLM)虽能推理物理场景和社交情境,但理解人类注视和注意力的可靠性尚未明确。EyeVLM系统性地评估了VLM在注视分析任务中的表现,揭示了其优势与局限,为改进多模态行为理解提供了基准。
arXiv:2605.19859v2 Announce Type: replace Abstract: Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs could greatly benefit the analysis of human gaze and attention, a central task in human behavior understanding that requires reasoning about the physical scene as well as the activity, interactions, and social context. However, the extent to which VLMs can reliably understand human gaze and related attentional behaviors remains largely unexplored. In this work, we present EyeVLM, a systematic