Related Experiment Videos
Hybrid Compression Method for Trained 3D Gaussian Splatting Models Based on VQ and HEVC
Dong-Ha Kim1, Byung-Yoon Choi1, Kwan-Jung Oh2
1Department of Smart Air Mobility, Korea Aerospace University, Goyang 10540, Republic of Korea.
Abstract:
3D Gaussian Splatting (3DGS) has recently emerged as an effective representation for immersive 3D scene rendering, providing high visual fidelity and real-time rendering efficiency. To support interoperable compression of trained 3DGS content, the Moving Picture Experts Group (MPEG) is exploring Gaussian Splat Coding (GSC), which mainly targets already trained 3DGS models following the INRIA reference format. The current video-based GSC anchor reorders 3DGS attributes into 2D attribute maps using Parallel Assignment Linear Sorting (PLAS) and compresses the resulting maps using High Efficiency Video Coding (HEVC). However, higher-order spherical harmonic coefficients (SH-AC) often remain irregular and exhibit low local spatial correlation even after PLAS reordering, limiting the coding efficiency of conventional video codecs. This paper proposes a VQ-HEVC hybrid compression framework that is structurally compatible with the video-based GSC anchor framework, in which SH-AC coefficients are represented by vector quantization (VQ) indices, while the remaining attributes are encoded using the same HEVC-based procedure as the GSC anchor. The proposed method adopts a two-stage VQ scheme that combines coarse VQ and product-quantization-based residual quantization, together with zero-masked residual VQ and flexible PQ grouping, to improve index-map coding efficiency across rate points. The generated VQ indices are packed into YUV400 index-map sequences and encoded using HEVC lossless coding, while the corresponding codebooks are transmitted as metadata. Experimental results on the Bartender and Cinema sequences of the MPEG GSC CTC demonstrate consistent rate-distortion improvements over the video-based GSC anchor across multiple objective quality metrics within the evaluated setting. In terms of RGB-PSNR, the proposed method achieves BD-rate reductions of 22.3% and 18.5% for the Bartender and Cinema datasets, respectively. These results suggest that, for the evaluated GSC CTC sequences, VQ-based SH-AC representation can effectively complement PLAS-based video coding while maintaining consistency with the existing GSC coding structure.
Related Concept Videos
Upsampling
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Divergence Theorem in 3D Space
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...