메뉴 건너뛰기
소속 기관 / 학교 인증
인증하면 논문, 학술자료 등을  무료로 열람할 수 있어요.
한국대학교, 누리자동차, 시립도서관 등 나의 기관을 확인해보세요
(국내 대학 90% 이상 구독 중)
고객센터 ENG
주제분류

논문 기본 정보

저자정보
(이화여자대학교) (연세대학교) (홍익대학교) (동국대학교) (성균관대학교)
저널정보
대한전자공학회 대한전자공학회 학술대회 2024년도 대한전자공학회 추계학술대회 논문집
오류 신고하기

피인용 0

검색

    초록·키워드

    In Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowledge for accurate question answering, significant advancements have been made through the use of Large Language Models (LLMs). BLIP-2, a popular multimodal LLM, employs a single-layer Q-Former for visual feature extraction and cross-modal interactions but faces challenges with complex reasoning tasks.
    To overcome these limitations, we propose integrating the Multimodal Co-Attention Network (MCAN), which uses a multi-layered approach to enhance the interaction between visual and textual inputs. Additionally, we introduce Question-Aware Prompts during fine-tuning, combining Answer Candidates with confidence scores and Answer-Aware Examples from past cases. This improves the model's ability to interpret questions accurately and generate more contextually appropriate answers.
    Experimental results on KB-VQA datasets show a 6.9% improvement in accuracy compared to baseline models, demonstrating the effectiveness of our approach in handling complex multimodal reasoning tasks.

    최근 본 자료 전체보기