Abstract: This paper introduces a unified edge-cloud system for real-time video vision-language model applications.
Vision-language models are extending video understanding from offline clip analysis to continuous interactive streaming, but most research still emphasises model capability rather than deployable low-latency interaction. This paper presents a unified edge-cloud system for real-time video VLM applications.
You are not signed in
Only registered users can read the rest of this article.
Next generation video compression standards
Tech Papers 2026: This paper presents an overview of the design criteria and development goals for a new video compression standardisation project.
Selective multi-pass encoding for cost-effective video streaming
Tech Papers 2026: This paper presents a content-adaptive strategy, CASE, that predicts whether additional encoding passes would provide meaningful gains using a lightweight mechanism that derives spatial and temporal features from each video segment.
Scalable SSIM estimation from PSNR for per-title and context adaptive encoding workflows
Tech Papers 2026: This paper proposes ApproxSSIMate, a low-complexity method for estimating SSIM from PSNR combined with reference-sequence statistics.
Feasibility and deployment strategies for cloud-based AOIP audio consoles
Tech Papers 2026: This paper investigates the feasibility of cloud-based audio-mixing systems by re-examining existing assumptions about network conditions, multicast transport and synchronisation.
A 100 Hz frame-interleaved approach to live multi-camera switching in led-based virtual production: Experience from a public-broadcaster proof of concept
Tech Papers 2026: This paper reports a Virtual Production (VP) proof of concept carried out by SWR, a German public-service broadcaster, under real broadcast conditions.
.jpg)