Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Self-training for content-aware video quality enhancement in HTTP adaptive streaming

Wei, Yiying, Amirpour, Hadi, Zhou, Wei, Hamidouche, Wassim and Timmerer, Christian 2026. Self-training for content-aware video quality enhancement in HTTP adaptive streaming. IEEE Transactions on Broadcasting 10.1109/tbc.2026.3735588

Full text not available from this repository.

Abstract

Fluctuations in video segment download rates and resolution switching in HTTP Adaptive Streaming (HAS) make it challenging to maintain a consistent Quality of Experience (QoE). However, the impact of such switching is often underestimated, and broadly applicable mitigation strategies remain underexplored. In the past, content-aware approaches have been introduced, using deep neural networks (DNNs) trained on an individual video segment to enhance its quality. These DNNs, transferred as a model stream alongside the video bitstream, allow clients to improve playback quality. However, transferring model streams adds bitrate overhead and additional architectural components, limiting practical use. Furthermore, supporting a wide range of device capabilities with a single DNN is impractical, as it would require device-specific models for each configuration, an approach that becomes unmanageable with increasing device heterogeneity. In this paper, we propose a new self-training method that enables clients to train content-aware video super-resolution (SR) models locally by leveraging previously downloaded high-quality segments. These segments are downscaled and used to train lightweight DNNs, which are then applied to enhance subsequent lower-quality segments. To keep training efficient and real-time, we select only a few predefined frames and extract the most informative patches using a lightweight sampling strategy. Experiments demonstrate that this approach significantly improves visual quality, with average PSNR gains of 1.07 dB ( 2× upscaling), 0.43 dB ( 3× ), and 0.58 dB ( 4× ) using ESPCN, a lightweight SR approach. To further validate the effectiveness of our approach, we conducted a series of ablation studies to analyze the contributions of individual components. Real-device measurements and end-to-end HAS simulations further show that self-training requires only 2.9–13.0% of a 4 s segment interval on CPU and 1.1–5.3% on GPU across tested mobile devices, while improving VMAF/QoE with only marginal additional rebuffering compared with generic SR.

Item Type: Article
Date Type: Published Online
Status: In Press
Schools: Schools > Computational & Mathematical Sciences
Schools > Computer Science & Informatics
Publisher: Institute of Electrical and Electronics Engineers
ISSN: 0018-9316
Last Modified: 05 Oct 2026 10:30
URI: https://orca.cardiff.ac.uk/id/eprint/189985

Actions (repository staff only)

Edit Item Edit Item