|
Liu, Jiang
2026.
Vision-based human action quality assessment.
PhD Thesis,
Cardiff University.
Item availability restricted. |
|
PDF
- Accepted Post-Print Version
Restricted to Repository staff only until 18 September 2027 due to copyright restrictions. Download (16MB) |
|
|
PDF (Cardiff University Electronic Publication Form)
- Supplemental Material
Restricted to Repository staff only Download (214kB) |
Abstract
Action Quality Assessment (AQA) aims to automatically evaluate how well human actions are performed from video, with important applications in sports analysis, medical rehabilitation, and home-based fitness. Although recent advances in deep learning have substantially improved AQA, existing research still faces several important limitations. In particular, current methods often struggle to model long-range spatiotemporal dependencies in long-duration actions, rarely incorporate perception mechanisms aligned with human visual judgement, and remain largely centred on single-score prediction, which is insufficient for the fine-grained and interpretable evaluation required in practical scenarios such as home-based fitness. Motivated by these challenges, this thesis investigates vision-based AQA from three perspectives: long-video spatiotemporal modelling, perception-guided assessment, and expert-perceived multi-dimensional evaluation. First, to address the limitations of existing long-video AQA methods, this thesis proposes an Adaptive Spatiotemporal Graph Transformer Network (ASGTN), which combines adaptive spatial and temporal graph modelling with transformer-based dependency learning to capture both local spatiotemporal relations within clips and global semantic context across clips. Second, inspired by the Human Visual System, this thesis investigates the role of visual saliency in AQA through a Saliency-Guided Action Quality Assessment Network (SAQANet), showing that saliency-guided modelling provides useful complementary information for action assessment. Third, to sup port more practical and fine-grained AQA, this thesis introduces HomeFit-MD, the first expert-perceived multi-dimensional dataset for home-based dumbbell exercises, containing 2,496 dual-view videos across 12 exercise categories with repetition-level annotations over four quality dimensions. Building upon this dataset, the thesis further proposes HomeFit-Net, a unified framework for multi-dimensional action quality assessment that integrates dual-view semantic encoding, a lightweight mo tion stream, and a hierarchical assessment strategy to model dependencies among quality dimensions. Overall, this thesis advances vision-based AQA by improving long-video spatiotemporal modelling, exploring perception-guided feature learning, and extending action quality assessment towards more structured and practically useful multi dimensional evaluation.
| Item Type: | Thesis (PhD) |
|---|---|
| Date Type: | Completion |
| Status: | Unpublished |
| Schools: | Schools > Computational & Mathematical Sciences Schools > Computer Science & Informatics |
| Subjects: | Q Science > QA Mathematics > QA75 Electronic computers. Computer science |
| Date of First Compliant Deposit: | 18 September 2026 |
| Last Modified: | 22 Sep 2026 15:15 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/189711 |
Actions (repository staff only)
![]() |
Edit Item |




Download Statistics
Download Statistics