vix.ing · top · new · best · stats

Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts

2025/04/29 by Hanhua Hong, Hong, Hanhua, Chenghao Xiao +9 · 1 voice · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.2504.21117

openalex publication_date 2025/04/29 · arxiv published 2025/04/29 · arxiv updated 2025/09/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardisation, and demographic biases, limiting reproducibility. LLM-based evaluators offer a scalable alternative but are highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation.

Citations

Cited by

Discussions

Related