3D Vision-Language Gaussian Splatting - 김동욱 발표 > Seminar

3D Vision-Language Gaussian Splatting - 김동욱 발표

페이지 정보

작성자 최고관리자 댓글 조회 작성일 25-08-01 13:39

본문

Recent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches have naively embedded semantic representations into 3D reconstruction methods without striking a balance between visual and language modalities, which leads to unsatisfying semantic rasterization of translucent or reflective objects, as well as over-fitting on color modality. To alleviate these limitations, we propose a solution that adequately handles the distinct visual and semantic modalities, i.e., a 3D vision-language Gaussian splatting model for scene understanding, to put emphasis on the representation learning of language modality. We propose a novel cross-modal rasterizer, using modality fusion along with a smoothed semantic indicator for enhancing semantic rasterization. We also employ a camera-view blending technique to improve semantic consistency between existing and synthesized views, thereby effectively mitigating over-fitting. Extensive experiments demonstrate that our method achieves state-of-the-art performance in open-vocabulary semantic segmentation, surpassing existing methods by a significant margin.

첨부파일

2025-5-14 랩세미나_3D Vision-Language Gaussian Splatting.pptx (19.4M) 1회 다운로드 | DATE : 2025-08-01 13:39:35

이전글Pow3r, VGGT - 권흥찬 발표 25.08.01
다음글Diffusion for 3D Large Scene Generation - 이연규 발표 25.08.01

댓글목록

등록된 댓글이 없습니다.

Boards

Seminar

페이지 정보

본문

첨부파일

댓글목록