Alali, Abdulazeez
2026.
Toward partial fake speech detection for voice authentication and speech security.
PhD Thesis,
Cardiff University.
Item availability restricted. |
|
PDF (Cardiff University Electronic Publication Form)
- Supplemental Material
Restricted to Repository staff only Download (446kB) | Request a copy |
|
|
PDF
- Accepted Post-Print Version
Available under License Creative Commons Attribution Non-commercial No Derivatives. Download (4MB) |
Abstract
The proliferation of advanced audio manipulation technologies has led to the emergence of "partial fake speech", where only certain segments of an audio recording are artificially generated or altered, while the rest remains genuine. This poses a significant challenge to audio authenticity verification systems. This thesis addresses the critical issue of partial fake speech detection by proposing novel methodologies that efficiently identify partial fake audio. The research begins with a comprehensive analysis of existing speech synthesis generation and detection methods, highlighting their strengths and limitations, as well as presenting the available datasets that can be used to evaluate the detection models. Building on this analysis, the insights gained from the literature review directly inform the design of the experimental framework adopted in this thesis. The literature review in this thesis led us to create a new dataset called RFP (Real, Fake, and Partial Fake) that includes various synthetic speech generation methods for both fake and partial fake audio. We have made the RFP dataset available online so that other researchers can use it to assess their own detection models. This dataset serves as the foundation for the subsequent experimental evaluations. We evaluated the impact of partial fake speech on machines and humans through realworld attack experiments and a questionnaire to assess humans’ ability to distinguish partial fake speech from genuine audio. Our experimental results highlight the urgent need for dedicated detection methods for partial fake speech. Motivated by these findings, in this thesis, we introduced a detection tool based on deep learning that efficiently identifies partial fake speech. To evaluate the effectiveness of our proposed approach, we utilized three different datasets, and our model outperformed existing publicly available detection models. Our model achieved a substantially lower EER of 0.54% compared to 34.71% obtained by existing baseline methods on partial fake speech, and further demonstrated superior performance on entirely fake speech with an EER of 2.99% versus 9.26% for baseline approaches. This thesis advances fake audio detection by improving our understanding of partial fake speech and offering a scalable, efficient detection solution. The findings impact security, media integrity, and legal domains, where authentic audio evidence is crucial. Subsequent investigations should prioritize the development of real-time detection capabilities and enhanced robustness against increasingly sophisticated manipulation techniques.
| Item Type: | Thesis (PhD) |
|---|---|
| Date Type: | Completion |
| Status: | Unpublished |
| Schools: | Schools > Computer Science & Informatics |
| Subjects: | Q Science > QA Mathematics > QA75 Electronic computers. Computer science |
| Funders: | Kuwait Embassy |
| Date of First Compliant Deposit: | 1 May 2026 |
| Date of Acceptance: | February 2026 |
| Last Modified: | 06 May 2026 08:48 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/186728 |
Actions (repository staff only)
![]() |
Edit Item |




Download Statistics
Download Statistics