Geometry-Aware Visual Understanding and Stereoscopic Video Generation

This thesis develops a geometry-aware dissolving transformation that selectively removes unnecessary fine-grained visual and geometric detail while preserving coarse structural information, to achieve more robust representation learning, higher-quality stereoscopic video synthesis, and more efficient multi-view geometry processing.

Overview

Visual and geometric signals carry structure at many scales, and not every scale is necessary for every task. The coarse, low-frequency layout of a scene — the arrangement of surfaces, the overall shape of an object — carries most of what a vision task consumes, whereas the fine, high-frequency detail at boundaries and thin structures often adds little. This thesis argues that a geometry-based vision system should control the granularity of its signals. Fully detailed intermediate estimates (e.g., depth, disparity maps) impose fine detail that many tasks may never need, and that surplus detail is where visible artifacts originate. Yet many tasks, we show, do not need overwhelming details.

The thesis develops this idea through a single mechanism, the dissolving transformation: a learned, content-aware operator that removes a signal's fine, instance-specific detail by taking one reverse step of a pretrained diffusion model, while leaving its coarse structure intact. Chapter 2 studies the mechanism on images. Contrasting an image against its dissolved counterpart exposes exactly the detail that was removed, so the dissolving operation amplifies its influence on the learned representation. Such a contrastive learned representation yields improvements in anomaly detection across six medical imaging benchmarks. Chapter 3 applies the technique to stereoscopic video synthesis, where geometry must stay consistent across views. To generate a stereoscopic video, we use only the coarse structure of a depth map in a pretrained diffusion model, leaving the diffusion prior to supply plausible detail. This coarse structure injection improves the video quality and cross-view consistency. Chapter 4 removes the explicit geometric representation altogether by using a layered, implicit disparity relaxation. This relaxation costs nothing in the quality of the generated videos while making synthesis several times faster, since the heavy depth-estimation module is removed entirely.

Together, these contributions trace one idea from images to multi-view geometry: a geometry-aware model needs only as much detail as its task consumes, and often far less than an explicit estimator provides.

Presenters

Brief Biography

Jian Shi is a Ph.D. candidate in Computer Science at King Abdullah University of Science and Technology (KAUST), where he conducts research under the supervision of Prof. Peter Wonka. Before joining KAUST, Shi earned an M.Sc. in Cloud Computing with Distinction (First-Class Honours) from University of Leicester and a bachelor's degree in Information Management and Information Systems from Zhengzhou University of Aeronautics. His professional experience includes research positions at NEC Laboratories China, The Chinese University of Hong Kong, and GE Power, where he worked on a broad range of machine learning and computer vision applications.

Shi has established a strong publication record in leading journals and conferences in artificial intelligence and computer vision. His work has appeared in premier venues including ACM SIGGRAPH, ICML, CVPR, ICCV, ECCV, IEEE TPAMI, and ICRA. His recent research contributions include advances in stereo video generation, depth estimation, anomaly detection, human keypoint estimation from LiDAR data, and the integration of geometric reasoning into generative AI systems.

In addition to his academic research, Shi is a co-founder and core maintainer of Kornia, one of the world's most widely used open-source differentiable computer vision libraries built on PyTorch. He has also served as a mentor and organization administrator for Google Summer of Code, helping foster open-source innovation in computer vision and machine learning.

Among his notable achievements, Shi received the KAUST Dean's List Award in 2025 in recognition of his academic excellence and research accomplishments. He is also an inventor on multiple patent filings and has delivered invited talks on emerging topics such as agentic computer vision and open-source AI infrastructure. Through both his scholarly contributions and open-source leadership, Shi continues to advance the development of next-generation vision systems that enable machines to perceive, understand, and synthesize the three-dimensional world.