Comparative Analysis of Large Language Models for Context-Aware Code Completion using SAFIM Framework
Large language models (LLMs) have transformed code completion and made it a more intelligent, context-aware tool in contemporary integrated development systems. These developments have greatly improved developers' capacity for error-free, effective code writing. Using the Syntax-Aware Fill-in- the- Middle (SAFIM) dataset, this work assesses different chat-based LLMs, including Gemini 1.5 Flash, Gemini 1.5 Pro, GPT-4o, GPT-4o-mini, and GPT-4 Turbo. This benchmark is intended especially to evaluate models' syntactic sensitivity in code creation. Accuracy and efficiency were assessed using performance benchmarks including cosine similarity with ground-truth complements and latency. The results expose significant variations in the code completion skills of the models, therefore providing insightful analysis of their distinct strengths and shortcomings. This paper offers a baseline for next developments in LLM-based code completion by means of a comparative study stressing the trade-offs between correctness and speed.