상세 보기
Field-Adaptive Dense Retrieval of Structured Documents
- SyedMudasir;
- AbdulWaheedAgha;
- SYED MUZAMIL HUSSAIN;
- 정선태
초록
This paper presents a novel field-adaptive methodology for dense retrieval of structured documents, tackling the persistent semantic gap between natural language queries and field-based content organization. As structured document repositories proliferate in enterprise environments, traditional dense retrieval methods face challenges due to the heterogeneous composition of fields and uneven semantic density. Our approach introduces three key innovations. First, we employ fine-tuned language models with similarity filtering to generate high-fidelity training data, addressing the scarcity of reliable query-document pairs. Second, we implement query-length-based adaptive field weighting, dynamically adjusting the contribution of titles, descriptions, and metadata during bi-encoder contrastive training. Third, we design a two-stage hybrid ranking strategy that combines the efficiency of bi-encoders with the precision of cross-encoders through optimized score integration. Extensive experiments on the Crello dataset, comprising over 25,000 structured documents, demonstrate a 33.8% improvement in Mean Reciprocal Rank (MRR) compared to the baseline, while maintaining inference efficiency. These results establish a scalable and domain-independent solution for structured document retrieval, offering both theoretical contributions and practical feasibility for real-world deployment.
키워드
- 제목
- Field-Adaptive Dense Retrieval of Structured Documents
- 저자
- SyedMudasir; AbdulWaheedAgha; SYED MUZAMIL HUSSAIN; 정선태
- 발행일
- 2025-08
- 유형
- Y
- 저널명
- 멀티미디어학회논문지
- 권
- 28
- 호
- 8
- 페이지
- 1001 ~ 1014