Beyond Benchmarking: A New Paradigm for Evaluation and Assessment of Large Language Models

In current benchmarks for evaluating large language models (LLMs), there are issues such as evaluation content restriction, untimely updates, and lack of optimization guidance. In this paper, we propose a new paradigm for the measurement of LLMs: Benchmarking-Evaluation-Assessment. Our paradigm shifts the "location" of LLM evaluation from the "examination room" to the "hospital". Through conducting a "physical examination" on LLMs, it utilizes specific task-solving as the evaluation content, performs deep attribution of existing problems within LLMs, and provides recommendation for optimization.

Paper

References (9)

042023. Mm-bench: Is your multi-modal model an all-around player?arXiv preprint
05Evaluation: Based on the lacked capabilities, the LLM is evaluated by finishing professional tasks to further explore the specific problems of the LLM in this capability
062023. Safety assessment of chinese large language modelsarXiv preprint
07Qiyue
08the problems of the LLM on a specific capabilitythe assessment of LLMs
092023. Halueval: A large-scale hallucination evaluation benchmark for large language modelsProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

Similar papers

© 2026 NYSGPT2525 LLC