Context‐Aware Prompt Engineering for Large Language Models in Autonomous Vehicles

ABSTRACT This paper presents a soft and novel context‐aware prompt engineering framework to enable adaptive and safe integration of large language models (LLMs) into autonomous vehicle (AV) systems. Unlike prior works, which employed static prompts or offline decision logic, our method dynamically creates prompts based on multimodal context, including speech commands, environmental cues (e.g., traffic, weather), and urgency levels. A hierarchical prioritization model is introduced to classify instructions according to safety sensitivity, enabling fine‐grained control and real‐time response in high‐risk AV scenarios. The system integrates speech‐to‐text (STT) transcription, text inputs, environmental context, and LLMs. Assessment of the Talk2Car dataset and a complementary noisy‐speech testbed indicated consistent improvements in accuracy, precision, recall, and F1 across four backbone LLMs (BERT, GPT‐2, SALMON, and SALMONN). These results demonstrate the effectiveness of prompt‐level adaptation in ensuring robustness and scalability in real‐world AV deployments.

Paper

Full text

PDF

Context‐Aware Prompt Engineering for Large Language Models in Autonomous Vehicles

OpenAlex · Autonomous Vehicle Technology and Safety · 2025

Abstract

This paper presents a soft and novel context‐aware prompt engineering framework to enable adaptive and safe integration of large language models (LLMs) into autonomous vehicle (AV) systems. Unlike prior works, which employed static prompts or offline decision logic, our method dynamically creates prompts based on multimodal context, including speech commands, environmental cues (e.g., traffic, weather), and urgency levels. A hierarchical prioritization model is introduced to classify instructions according to safety sensitivity, enabling fine‐grained control and real‐time response in high‐risk AV scenarios. The system integrates speech‐to‐text (STT) transcription, text inputs, environmental context, and LLMs. Assessment of the Talk2Car dataset and a complementary noisy‐speech testbed indicated consistent improvements in accuracy, precision, recall, and F1 across four backbone LLMs (BERT, GPT‐2, SALMON, and SALMONN). These results demonstrate the effectiveness of prompt‐level adaptation in ensuring robustness and scalability in real‐world AV deployments.

Similar papers

© 2026 NYSGPT2525 LLC