Gong, Meiyuan
2025.
Robotic object manipulation and language-guided planning.
PhD Thesis,
Cardiff University.
Item availability restricted. |
Preview |
PDF
- Accepted Post-Print Version
Download (3MB) | Preview |
|
PDF (Cardiff University Electronic Publication Form)
- Supplemental Material
Restricted to Repository staff only Download (9MB) |
Abstract
With the rapid development of machine learning and large language models (LLMs), robotic manipulation is undergoing a shift from purely model-based design towards learning-based and data-driven methods. In this context, this thesis studies learning-based methods for robotic grasping and manipulation, with a particular focus on integrating LLMs as high-level advisors while en forcing low-level geometric and physical feasibility through specialised modules. This thesis first studies learning-based grasp prediction in cluttered table top scenes. An autoencoder is used to learn compact latent representations of local RGB-D observations, and a grasp success predictor is trained on top of these representations. The resulting model estimates grasp feasibility di rectly from observations, providing a lightweight prior that can be queried during planning and control. Experiments in simulated grasping tasks show that this learned predictor captures meaningful geometric regularities and can guide the selection of more reliable grasps. Then, this thesis develops an LLM-guided reinforcement learning frame work for goal-object grasping. A pre-trained LLM receives a high-level task description and a textual abstraction of the scene, and generates one-step pro posals in a shared primitive space, selecting push or grasp and outputting continuous pose parameters. During training, an ε-greedy mechanism mixes these proposals with actions from a soft actor–critic policy, and the replay buffer records whether each action originates from the LLM or the learned policy. Experiments demonstrate that such LLM guidance can accelerate early-stage learning and lead to higher success rates compared with purely policy-driven exploration. The LLM is queried only during an early training guidance window and is fully disabled at evaluation time, so improvements arise from reshaping the early replay-buffer distribution rather than from test-time language reasoning. Finally, this thesis investigates LLM-based task and motion planning (TAMP) for box-packing manipulation. Building on recent LLM planners, a lightweight pre-check module is introduced between the LLM and the motion planner, consisting of a fast geometric collision test and a small learned feasi bility model trained from offline motion-planning rollouts. These pre-checks filter out low-value motion-planning attempts and return structured failure signals (e.g., pre-check type and feasibility scores) to the LLM. This makes the closed-loop planner more cost-aware by encouraging cheap local parame ter refinement before expensive backtracking and motion-planning calls. Ex periments on a box-packing benchmark show that this framework matches or improves task success while significantly reducing the number of LLM and motion-planner calls. Overall, the thesis demonstrates that combining learned predictors, LLM guidance, and geometric pre-checks can yield robotic manipulation systems that are both more data-efficient and more physically grounded.
| Item Type: | Thesis (PhD) |
|---|---|
| Date Type: | Completion |
| Status: | Unpublished |
| Schools: | Schools > Engineering |
| Uncontrolled Keywords: | 1. Robotic Manipulation 2. Cluttered Object Manipulation 3. Push Grasp Learning 4. Reinforcement Learning 5. Large Language Models 6. Task and Motion Planning |
| Date of First Compliant Deposit: | 22 July 2026 |
| Last Modified: | 22 Jul 2026 13:40 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/188400 |
Actions (repository staff only)
![]() |
Edit Item |




Download Statistics
Download Statistics