This paper demonstrates that our LMM-based approach not only significantly reduces the computational complexity required for sampling based per-title video encoding—by an astounding 13 times—but also maintains the same level of bitrate saving. These findings not only pave the way for more efficient and adaptive video encoding strategies but also highlight the potential of multi-modal models in enhancing multimedia processing tasks.
In the realm of video encoding, achieving the optimal balance between encoding efficiency and computational complexity remains a formidable challenge. This paper introduces a groundbreaking framework that utilizes a Large Multi-modal Model (LMM) to revolutionize the process of per-title video encoding optimization. By harnessing the predictive capabilities of LMMs, our framework estimates the encoding complexity of video content with unprecedented accuracy, enabling the dynamic selection of encoding configurations tailored to each video’s unique characteristics.
The proposed framework marks a significant departure from traditional per-title encoding methods, which often rely on expensive and time-consuming sampling in the rate-distortion space. Through a comprehensive set of experiments, we demonstrate that our LMM-based approach not only significantly reduces the computational complexity required for sampling based per-title video encoding—by an astounding 13 times—but also maintains the same level of bitrate saving.
The implications of this research...
Exclusive Content
This article is available with a Technical Paper Pass
Next generation video compression standards
Tech Papers 2026: This paper presents an overview of the design criteria and development goals for a new video compression standardisation project.
Selective multi-pass encoding for cost-effective video streaming
Tech Papers 2026: This paper presents a content-adaptive strategy, CASE, that predicts whether additional encoding passes would provide meaningful gains using a lightweight mechanism that derives spatial and temporal features from each video segment.
Scalable SSIM estimation from PSNR for per-title and context adaptive encoding workflows
Tech Papers 2026: This paper proposes ApproxSSIMate, a low-complexity method for estimating SSIM from PSNR combined with reference-sequence statistics.
Deep learning super resolution for dense dynamic point Cloud compression
Tech Papers 2026: This paper proposes Video-based Super Sampling Point Cloud Compression (VSS-PCC), a method that uses neural super-resolution to reduce the size of point cloud data before compression.
Client device discovery: An open standard for bridging the first-mile gap in cloud media workflows
Tech Papers 2026: This paper presents the architecture, protocol design, and security model of TR-12, demonstrates its contract-first approach using Smithy models, and discusses early implementation experience with the open source SDK.





