Special issue on realistic and immersive media technologies

Seung‐Yeol Lee, Donghyeon Cho, Yong Hyub Won, Munchurl Kim · ETRI Journal · 2022

Beyond the conventional media contents of planar displays, realistic and immersive media technologies have been researched actively for decades as one of the crucial technologies in the fourth industrial revolution. Immersive media, which is commonly referred to as augmented/virtual reality (AR/VR) and has since been extended to mixed/extended reality (MR/XR) and even the Metaverse, may be a key technology that will eventually change the way we live soon. Immersive media, like other key technologies of the fourth industrial revolution, is a convergence technology that requires not only communication/broadcasting techniques but also optics, electronics, and artificial intelligence techniques, as recently demonstrated. Immersive media includes a broad scope of techniques such as extracting the three-dimensional (3D) depth information, 3D contents formation, and optimization, quality assessments of virtual objects, adjusting the signals to the process of digital hologram patterns. In this special issue, we have selected eight publications that represent the current state-of-the-art in the field of realistic and immersive media technologies. Those publications, in particular, are primarily concerned with overcoming the current limitations of immersive media technologies for practical applications. For example, novel types of immersive rendering models enhanced pruning algorithms for high visual quality in moving pictures expert group (MPEG) immersive video, and digital hologram resolution adjustment techniques are introduced, as are techniques for synthesizing virtual view from light fields and providing innovative algorithms for estimating the depth of 3D object or camera pose to dramatically reduce computational resources. Such methods are highly demanded to improve the image quality of representative immersive media, so that those are essential technologies for commercial applications of AR/VR instruments and holographic display panels. The first paper “Camera pose estimation framework for array structured images” by Min-jung Shin and others demonstrates a novel algorithm for estimating the camera pose from array structured images that captured the 3D objects with different positions and angle conditions. The proposed method estimates the set of array structured images by using incremental reconstruction steps, which includes two major processes: 3D point outlier elimination and bundle adjustment with a constraint term for rotation vectors. Camera pose estimation from structured array images is critical for systematically reconstructing depth information of omnidirectional object scenes and fully understanding the given situation of a measuring camera system from captured images. In the second paper entitled “View synthesis with the sparse light field for 6DoF immersive video,” Sangwoon Kwak and others successfully report an effective method to synthesize the virtual view from the sparse light field which can apply to immersive video. Unlike the conventional media contents, immersive video needs realistic binocular disparity and smooth motion parallax to make the observer recognize the media as a 3D scene. The presented work proposed a distinctive blending architecture based on light field concepts that trace the ray's directions and the distributions of depth values. It has been demonstrated that using the proposed method yields improved virtual view synthesizing when compared to the well-designed synthesizers used in MPEG-I. This is also implemented on a full GPU-based system for synthesizing and rendering with fast computation speed. While the first and second papers have focused on synthesizing and analyzing the novel view of immersive media, the next paper “Recursive block splitting in feature-driven decoder-side depth estimation” by Błażej Szydełko and others proposed a decoder-side depth estimation algorithm using a recursive block splitting method. In contrast to conventional encoding methods, the proposed multiview video encoding method does not require the transmission of depth map information to the decoder side, instead of sending only a set of input views and their parameters. This paper also proposes a novel recursive block splitting method for feature extraction, which allows for a better fit to the edges of the objects, thus leading to improving the quality and decreasing the computational time for depth estimation. When the proposed method is used, the quality of synthesized views outperforms the quality achieved by the state-of-the-art approach. For synthesizing the immersive media contents, not only depth map information but also voxel-type information are often used. The fourth paper entitled “Voxel-wise UV parameterization and view-dependent texture synthesis for the immersive rendering of TSDF Scene model” by Soowoong Kim and others introduces a voxel-wise UV parameterization method that delegates a precomputed UV map to each voxel based on the UV map look-up table. Since the precomputed UV map is provided, the proposed method allows for fast, efficient, and high-quality texture mapping without requiring complex computation to obtain UV coordinates. To use eigenspace analysis of the separated specular maps, the proposed method employs a simple diffuse-specular map separation and view-dependent specular map estimation based on texture representation. Because increasing computation speed is one of the most pressing issues in the field of immersive media, the proposed work may have significant potential in the computation of voxel-wise 3D objects. Related to the computation time issue, real-time computation is required for the practical use of immersive media. The fifth paper “Real-time multi-GPU-based 8KVR stitching and streaming on 5G MEC/Cloud environments” by HeeKyung Lee and others also shows a great possibility of real-time computation of immersive media. Their presented work describes a multi-GPU-based 8KVR stitching system that runs in real-time on both local and cloud machine environments. The feasibility of the proposed 8KVR stitching system with stitching speed of up to 83.7 fps for six channel inputs and 62.7 fps for eight channel inputs was demonstrated in their test experiments on both local machine and cloud machine environments. Based on the high-speed performance of the proposed system, stable performances of 8K@30 fps in both indoor and outdoor environments can be achieved. The proposed work might be quite important for the broadcasting service of immersive media, which would accelerate its commercialization in near future. The sixth paper entitled “Enhanced pruning algorithm for improving visual quality in MPEG immersive video” by Hong-Chang Shin and others provides a novel method to improve the visual quality of immersive video using a pruning algorithm. The primary goal of the pruning process is to reduce the amount of data such that the available video must be handled while maintaining image quality. In addition to the traditional pruning algorithm, this paper used two additional approaches: setting the criteria for determining pruning order based on the amount of the overlapping region and a global region-wise color similarity to minimize the matching ambiguity in determining the pruning area. Applying those ideas to pruning algorithms is quite important if someone handles the immersive media instead of two-dimensional (2D) content. As the authors state in their conclusion, this type of algorithms can further be improved using deep learning techniques. The seventh and eighth papers are works about digital hologram manipulation techniques, one of the promising techniques to provide the ultimate form of immersive media. Han-Ju Yeom and others proposed an efficient method to synthesize the mesh-based CGH using an analytic approach in their mesh computation in the seventh paper, “Efficient mesh-based realistic computer-generated hologram synthesis with polygon resolution adjustment.” When calculating the angular spectrum of the 2D plane in mesh-based CGH, a 2D plane with an arbitrary size smaller than the resolution of the final hologram is used to dramatically reduce the computation time. Since one of the major limitations in CGH calculation is its heavy computation complexity and the use of a significant memory source, the proposed work appears to be critical in overcoming such limitations of digital hologram synthesis. Finally, the last paper “A new objective quality metric for phase hologram image processing” by Kwan-Jung Oh and others reports an objective quality assessment for digital phase hologram images. Quality assessment is performed by defining the error value considering the characteristics of phase information. In contrast to conventional 2D images, phase information is often more importantly treated than amplitude information in digital hologram data. Thus, the quality assessment criteria for digital hologram data should differ from those for 2D images. During the compression of CGH data, the signal-to-noise ratio and the effect of additional random noise characteristics may differ significantly from those of conventional media. The proposed work rigorously investigates such characteristics with various object forms, which will provide highly important knowledge to the readers who want to understand the quality metrics that can be effectively applicable to phase hologram images. The guest editors thank all the authors, reviewers, and editorial staff members of the ETRI Journal for making this special issue a success. We are most pleased to have been part of this effort and to ensure the timely publication of high-quality technical articles. Seung-Yeol Lee received the PhD degree from the School of Electrical Engineering, Seoul National University, Seoul, Republic of Korea, in 2014. He participated as a postdoctoral researcher for about 2 years in optical engineering and quantum electronics lab in Seoul National University, and he was a researcher in Electronics and Telecommunications Research Institute (ETRI) for about 1 year until August 2016. He is currently an associate professor in the School of Electronics Engineering, Kyungpook National University, Daegu, Republic of Korea. His research interests include digital holography, diffractive optical elements, nanophotonic devices, and metasurfaces. Donghyeon Cho received the PhD degree from the Electrical and Electronics Engineering Department, KAIST, Daejeon, Republic of Korea, in 2019. He was a full-time student intern with Microsoft Research Asia, and he was a researcher in AI Center, SK-Telecom. He is currently an assistant professor in the Department of Electronics Engineering, Chungnam National University, Daejeon, Republic of Korea. His research interest includes computer vision, image processing, and deep learning. Yong Hyub Won received a PhD in electrical engineering from Cornell University, Ithaca, New York, United States, in 1990. He joined Electronics and Telecommunications Research Institute (ETRI), Republic of Korea, in 1981, where he has conducted his research in optical communication and switching technologies as a section header. He is currently a professor at the School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Republic of Korea. His current researches are focused on 3D display panels using microfluidic device array based on electrowetting technology, bio-inspired robotic visual sensors, and photonic automation using photonic logic circuits based on FP-LD. Munchurl Kim received the BE degree in electronics from Kyungpook National University, Daegu, Republic of Korea, in 1989, and the ME and PhD degrees in Electrical and Computer Engineering from the University of Florida, Gainesville, in 1992 and 1996, respectively. He joined the Electronics and Telecommunications Research Institute, Daejeon, Republic of Korea, as a Senior Research Staff Member, where he led the Realistic Broadcasting Media Research Team. In 2001, he joined the School of Engineering, Information and Communications University, Daejeon, as an assistant professor. Since 2009, he has been with the School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, where he is currently a full professor and an KAIST ICT chair professor. He has published 181 (165) international (domestic) journal and conference papers. He holds several dozens of essential HEVE patents and about 200 registered domestic and international patents in the areas of image restoration and video coding. He received a commendation from the Korea President in the 54th National Innovation Day 2019 and was awarded the Grand Prize for Research Excellence in Commemoration of the 50th Anniversary of Founding for KAIST. He had an invited keynote speech on evolution of conventional and deep video compression technologies in 2020 Multimedia Modeling Conference. His team was awarded the runner-up in Challenge on the Video Temporal Super-Resolution track in AIM (Advances in Image Manipulation workshop and challenges on image and video manipulation) in ECCV 2020 and received the Winner award on “PIRM Challenge on Perceptual Image Enhancement—Track A: Image Super Resolution” in ECCV 2018. His research interest includes image restoration with deep learning, video coding, and image understanding.

Read the paper · More papers on PaperTik