LLM-wrapper：视觉语言基础模型的黑盒语义感知适配

arXiv - CS - Computer Vision and Pattern Recognition Pub Date : 2024-09-18 DOI:arxiv-2409.11919

Amaia Cardiel, Eloi Zablocki, Oriane Siméoni, Elias Ramzi, Matthieu Cord

{"title":"LLM-wrapper：视觉语言基础模型的黑盒语义感知适配","authors":"Amaia Cardiel, Eloi Zablocki, Oriane Siméoni, Elias Ramzi, Matthieu Cord","doi":"arxiv-2409.11919","DOIUrl":null,"url":null,"abstract":"Vision Language Models (VLMs) have shown impressive performances on numerous\ntasks but their zero-shot capabilities can be limited compared to dedicated or\nfine-tuned models. Yet, fine-tuning VLMs comes with limitations as it requires\n`white-box' access to the model's architecture and weights as well as expertise\nto design the fine-tuning objectives and optimize the hyper-parameters, which\nare specific to each VLM and downstream task. In this work, we propose\nLLM-wrapper, a novel approach to adapt VLMs in a `black-box' manner by\nleveraging large language models (LLMs) so as to reason on their outputs. We\ndemonstrate the effectiveness of LLM-wrapper on Referring Expression\nComprehension (REC), a challenging open-vocabulary task that requires spatial\nand semantic reasoning. Our approach significantly boosts the performance of\noff-the-shelf models, resulting in competitive results when compared with\nclassic fine-tuning.","PeriodicalId":501130,"journal":{"name":"arXiv - CS - Computer Vision and Pattern Recognition","volume":"65 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2024-09-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Foundation Models\",\"authors\":\"Amaia Cardiel, Eloi Zablocki, Oriane Siméoni, Elias Ramzi, Matthieu Cord\",\"doi\":\"arxiv-2409.11919\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Vision Language Models (VLMs) have shown impressive performances on numerous\\ntasks but their zero-shot capabilities can be limited compared to dedicated or\\nfine-tuned models. Yet, fine-tuning VLMs comes with limitations as it requires\\n`white-box' access to the model's architecture and weights as well as expertise\\nto design the fine-tuning objectives and optimize the hyper-parameters, which\\nare specific to each VLM and downstream task. In this work, we propose\\nLLM-wrapper, a novel approach to adapt VLMs in a `black-box' manner by\\nleveraging large language models (LLMs) so as to reason on their outputs. We\\ndemonstrate the effectiveness of LLM-wrapper on Referring Expression\\nComprehension (REC), a challenging open-vocabulary task that requires spatial\\nand semantic reasoning. Our approach significantly boosts the performance of\\noff-the-shelf models, resulting in competitive results when compared with\\nclassic fine-tuning.\",\"PeriodicalId\":501130,\"journal\":{\"name\":\"arXiv - CS - Computer Vision and Pattern Recognition\",\"volume\":\"65 1\",\"pages\":\"\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2024-09-18\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"arXiv - CS - Computer Vision and Pattern Recognition\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/arxiv-2409.11919\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"arXiv - CS - Computer Vision and Pattern Recognition","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/arxiv-2409.11919","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

视觉语言模型（VLM）在大量任务中表现出令人印象深刻的性能，但与专用模型或微调模型相比，它们的零拍摄能力可能有限。然而，对 VLM 进行微调也有其局限性，因为它需要 "白盒 "访问模型的架构和权重，还需要专家来设计微调目标和优化超参数，这些都是每个 VLM 和下游任务所特有的。在这项工作中，我们提出了LLM-wrapper，这是一种以 "黑箱 "方式调整VLM的新方法，通过利用大型语言模型（LLM）来对其输出进行推理。我们演示了 LLM-wrapper 在参考表达式理解（REC）上的有效性，这是一项具有挑战性的开放词汇任务，需要空间和语义推理。我们的方法大大提高了现成模型的性能，与传统的微调方法相比，结果极具竞争力。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Foundation Models

Vision Language Models (VLMs) have shown impressive performances on numerous tasks but their zero-shot capabilities can be limited compared to dedicated or fine-tuned models. Yet, fine-tuning VLMs comes with limitations as it requires `white-box' access to the model's architecture and weights as well as expertise to design the fine-tuning objectives and optimize the hyper-parameters, which are specific to each VLM and downstream task. In this work, we propose LLM-wrapper, a novel approach to adapt VLMs in a `black-box' manner by leveraging large language models (LLMs) so as to reason on their outputs. We demonstrate the effectiveness of LLM-wrapper on Referring Expression Comprehension (REC), a challenging open-vocabulary task that requires spatial and semantic reasoning. Our approach significantly boosts the performance of off-the-shelf models, resulting in competitive results when compared with classic fine-tuning.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

arXiv - CS - Computer Vision and Pattern Recognition

自引率

0.00%

发文量

期刊最新文献

Massively Multi-Person 3D Human Motion Forecasting with Scene Context Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution Precise Forecasting of Sky Images Using Spatial Warping JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation Applications of Knowledge Distillation in Remote Sensing: A Survey