생성-복원-재생성 폐루프 프로토콜을 통한 VLM 프롬프트 복원 성능 정량 평가

Quantitative Evaluation of VLM Prompt Recovery via a Generation-Restoration-Regeneration Closed-Loop Protocol

초록

This study proposes a Generation–Restoration–Regeneration closed-loop protocol to quantitatively evaluate Vision-Language Models(VLM) for prompt recovery from generated images. We designed 50 prompts across five categories—Spatial Arrangement(SA), Color/Texture/Count(CT), Pose/Orientation (PO), Relational Expression(RE), and Action/Context (AC)—with 10 per category. Four source images per prompt were generated using FLUX Schnell. Three VLMs(GPT-4o, Claude Sonnet 4.5, BLIP-2 opt-2.7b) restored single-sentence prompts from each image via a unified template. Restored prompts regenerated four images each with the same generator. Without cherry-picking, all images were assessed using CLIP-based Faithfulness(F) and Reproducibility(R), plus A+B constraint evaluation: checklist-based multimodal judging(A) and fixed two-question VQA(B) with dual-judge consensus (94% confirmation rate). Commercial VLMs outperformed BLIP-2(Claude: F=0.9256, R=0.9435; GPT-4o: F=0.9195, R=0.9342; BLIP-2: F=0.8188, R=0.8356). SA proved easiest to recover, while RE and AC remained challenging, with frequent failures in counts, attributes, and relations. Failure patterns and practical guidelines are provided.

키워드

Prompt recoveryReverse promptingVision-language modelCLIP similarityVQA verification
제목
생성-복원-재생성 폐루프 프로토콜을 통한 VLM 프롬프트 복원 성능 정량 평가
제목 (타언어)
Quantitative Evaluation of VLM Prompt Recovery via a Generation-Restoration-Regeneration Closed-Loop Protocol
저자
이금희김동호
DOI
10.9717/kmms.2026.29.6.915
발행일
2026-06
유형
Y
저널명
멀티미디어학회논문지
29
6
페이지
915 ~ 925