ABSTRACT This paper presents a soft and novel context‐aware prompt engineering framework to enable adaptive and safe integration of large language models (LLMs) into autonomous vehicle (AV) systems. Unlike prior works, which employed static prompts or offline decision logic, our method dynamically creates prompts based on multimodal context, including speech commands, environmental cues (e.g., traffic, weather), and urgency levels. A hierarchical prioritization model is introduced to classify instructions according to safety sensitivity, enabling fine‐grained control and real‐time response in high‐risk AV scenarios. The system integrates speech‐to‐text (STT) transcription, text inputs, environmental context, and LLMs. Assessment of the Talk2Car dataset and a complementary noisy‐speech testbed indicated consistent improvements in accuracy, precision, recall, and F1 across four backbone LLMs (BERT, GPT‐2, SALMON, and SALMONN). These results demonstrate the effectiveness of prompt‐level adaptation in ensuring robustness and scalability in real‐world AV deployments.
Paper
Full text
Context‐Aware Prompt Engineering for Large Language Models in Autonomous Vehicles
OpenAlex · Autonomous Vehicle Technology and Safety · 2025
Abstract
This paper presents a soft and novel context‐aware prompt engineering framework to enable adaptive and safe integration of large language models (LLMs) into autonomous vehicle (AV) systems. Unlike prior works, which employed static prompts or offline decision logic, our method dynamically creates prompts based on multimodal context, including speech commands, environmental cues (e.g., traffic, weather), and urgency levels. A hierarchical prioritization model is introduced to classify instructions according to safety sensitivity, enabling fine‐grained control and real‐time response in high‐risk AV scenarios. The system integrates speech‐to‐text (STT) transcription, text inputs, environmental context, and LLMs. Assessment of the Talk2Car dataset and a complementary noisy‐speech testbed indicated consistent improvements in accuracy, precision, recall, and F1 across four backbone LLMs (BERT, GPT‐2, SALMON, and SALMONN). These results demonstrate the effectiveness of prompt‐level adaptation in ensuring robustness and scalability in real‐world AV deployments.