메뉴 건너뛰기
소속 기관 / 학교 인증
인증하면 논문, 학술자료 등을  무료로 열람할 수 있어요.
한국대학교, 누리자동차, 시립도서관 등 나의 기관을 확인해보세요
(국내 대학 90% 이상 구독 중)
고객센터 ENG
주제분류

논문 기본 정보

저자정보
(한성대학교) (한성대학교)
저널정보
대한전자공학회 대한전자공학회 학술대회 2025년도 대한전자공학회 추계학술대회 논문집
오류 신고하기

피인용 0

검색

    초록·키워드

    This study addresses the issue of data leakage in Large Language Model (LLM) evaluation and proposes a perturbation-based robustness test to assess true generalization ability. We construct semantically preserved but surface-level altered datasets by applying synonym substitution and label/schema renaming. Three recent small LLMs—Gemma-3-4b-it, Llama-3.2-3B- Instruct, and Qwen2.5-3B-Instruct—are evaluated on original and perturbed versions of AG News, IMDb, SST-2, DBpedia-14, and Titanic benchmarks. Results show that some models suffer up to 37% accuracy drops on perturbed data, revealing potential benchmark memorization, while Gemma-3-4b-it maintains stable performance, indicating stronger robustness. These findings demonstrate that conventional benchmark scores may overestimate model capability, and that perturbation-based evaluation provides a more reliable measure of LLM generalization and data leakage resilience.

    최근 본 자료 전체보기

      UCI(KEPA) : I410-151-26-02-095554639