A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies
Khalid Halba, Kylie Cooper, James G. Bellingham
- Published
- Sep 17, 2026 — 16:03 UTC
Problem
Autonomous underwater vehicles (AUVs) are required to autonomously recover from failures without human intervention. This paper addresses the gap in existing literature regarding effective diagnostic strategies for AUVs, particularly in the context of fault recovery. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose a closed-loop simulation architecture named SPAR, which integrates real-time C vehicle software with a higher-level orchestration layer. The methodology includes:
- Fault Injection: A physics-based fault injection mechanism is employed to simulate various failure scenarios.
- Prompting: Structured prompting techniques are utilized to facilitate interaction with language models (LLMs).
- Mission File Generation: The system generates mission files that are used for validation and execution of the AUV's tasks.
- LLM-judge Scoring: A scoring mechanism is implemented to evaluate the performance of the LLMs in diagnosing faults.
- Trials: A total of 480 SPAR trials were conducted specifically for a mass-shift fault scenario.
- Models Evaluated: The study evaluates one frontier model alongside three off-the-shelf locally deployable LLMs to assess their diagnostic capabilities.
Results
The results indicate that the frontier model demonstrates a diagnosis success rate of identifying the center of gravity (CG) shift mechanism within the top three hypotheses in 85-90% of trials. In contrast, the best-performing local model achieves a success rate of 60-78%. This significant difference highlights the effectiveness of the frontier model in diagnosing faults compared to local alternatives.
Limitations
The authors note that the success of local models is contingent upon strict adherence to the complete diagnostic procedure. Additionally, weaker models may exhibit a tendency to prematurely commit to incorrect diagnoses, which could lead to suboptimal recovery strategies. These limitations suggest that while LLMs can enhance fault recovery, their performance is highly dependent on the model's robustness and the fidelity of the diagnostic process.
Why it matters
The implications of this work are substantial for the field of autonomous systems, particularly in enhancing the reliability and autonomy of AUVs. By leveraging LLMs for fault diagnosis, the proposed simulation platform could lead to more resilient AUV operations in complex underwater environments. This research opens avenues for further exploration into the integration of advanced AI techniques in autonomous vehicle systems, potentially influencing future designs and operational protocols.
By Callan Zhang · Sep 17, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
