Reliable and Developer-Aligned Evaluation of Agents for Software Engineering

Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their accelerated adoption, existing evaluation techniques are limited due to their fragmented nature and distorted projection of true model capabilities, often obtained from hypothetical syntactic scenarios. This research aims to bridge this gap by providing a comprehensive evaluation methodology for LLM-powered agents that is grounded in real-world software development practice. Our evaluation approach focuses on contamination-awareness, in-the-wild agentic behavior assessment, and trajectory-aware benchmarks and metrics capturing realistic coding contexts, human-aligned behavior, and model failure modes.

Paper

References (13)

09InvestigatingAutonomousAgentContributionsintheWild:ActivityPatternsandCodeChangeoverTime2026 · arXiv e-prints
10RQ1: Which LLMs have been evaluated in the context of code-related tasks?
11RQ2: How do agent-authored contributions influence the trajectory of code maintenance over time compared to human-authored ones?
12RQ1: What is the difference between agent-authored and human-authored activity in shaping collaboration and development progress?

Scroll for more · 1 remaining

Similar papers

© 2026 NYSGPT2525 LLC