Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Lightweight query-adaptive RAG framework for knowledge support in smart manufacturing

Zhou, Tianyu, Liu, Ying ORCID: https://orcid.org/0000-0001-9319-5940, Kumar, Maneesh ORCID: https://orcid.org/0000-0002-2469-1382 and Zou, Lai 2026. Lightweight query-adaptive RAG framework for knowledge support in smart manufacturing. Journal of Manufacturing Systems 86 , pp. 774-789. 10.1016/j.jmsy.2026.04.016

[thumbnail of 1-s2.0-S0278612526001020-main.pdf] PDF - Published Version
Available under License Creative Commons Attribution Non-commercial No Derivatives.

Download (12MB)

Abstract

As smart manufacturing moves beyond automated execution towards knowledge-intensive decision support, efficiently integrating dispersed domain knowledge has become a key challenge. Retrieval-augmented generation (RAG)–based large language models (LLMs) provide a practical approach to incorporating external knowledge. They have been increasingly applied in on-site manufacturing scenarios, where decision processes usually require frequent human–AI interactions and rapid responses. In this context, response latency and the cost of knowledge utilisation impose constraints on deployment. However, conventional RAG methods typically adopt unified, static pipelines for retrieval and generation, which tend to introduce redundant retrieval and contextual overhead when handling varied queries. Therefore, this paper proposes LiteRAG, a lightweight query-adaptive framework that integrates semantic query classification with adaptive retrieval to optimise the necessity and scope of external knowledge. Experimental results show that LiteRAG reduces average token and latency consumption by nearly 50% compared with fixed retrieval RAG, while moderately improving response quality in flexible grinding tasks. The gains are mainly observed in queries requiring retrieval augmentation, reaching about 15–20%. Moreover, the cross-model evaluation indicates that medium-scale models Qwen3_8B and Qwen3_14B achieve around 90% of the performance of a 30B model under this framework, reflecting favourable resource efficiency and scalability. These results demonstrate that LiteRAG provides a deployable solution for balancing knowledge utilisation and deployment efficiency in LLM applications for smart manufacturing.

Item Type: Article
Date Type: Publication
Status: Published
Schools: Schools > Engineering
Schools > Business (Including Economics)
Publisher: Elsevier
ISSN: 0278-6125
Date of First Compliant Deposit: 15 April 2026
Date of Acceptance: 5 April 2026
Last Modified: 15 Apr 2026 09:01
URI: https://orca.cardiff.ac.uk/id/eprint/186394

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics