A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation
Lorenzo Zangari, Davide Picca
- Published
- Sep 22, 2026 — 14:52 UTC
Problem
The paper addresses a significant gap in the evaluation of natural language generation (NLG) systems, specifically the lack of meaningful metrics to assess differences in text framing. Existing evaluation methods often fail to capture the nuances of contextual meaning and discourse references, which are critical for understanding the quality of generated text. This work is presented as a preprint and has not undergone peer review.
Method
The authors propose a semiotic alignment evaluation framework that introduces two novel scores: Semiotic Fidelity and Semiotic Coverage.
- Semiotic Fidelity measures how well the generated text aligns with the intended meaning and context, while
- Semiotic Coverage assesses the breadth of contextual references and discourse elements included in the generated output. The evaluation focuses on the contextual meaning and the references made in discourse, providing a more nuanced understanding of NLG performance compared to traditional metrics.
Results
The findings indicate that Semiotic Coverage is typically lower than Semiotic Fidelity, suggesting that while generated texts may align closely with intended meanings, they often lack comprehensive contextual references. Additionally, the study reveals that the alignment metric shows the highest correlation between large language models (LLMs) and human-curated data at low sampling temperatures. As the sampling temperature increases, the alignment decreases, indicating that higher variability in generation leads to less fidelity in capturing the intended semiotic meaning.
Limitations
The authors do not report any limitations in their study. However, the absence of reported limitations may suggest a need for further exploration of the framework's applicability across diverse NLG tasks and datasets.
Why it matters
This framework has significant implications for downstream work in NLG, as it provides a more robust method for evaluating generated text quality. By focusing on semiotic aspects, researchers and practitioners can better understand the strengths and weaknesses of their models, leading to improved generation techniques that prioritize meaningful communication and contextual relevance.
By Callan Zhang · Sep 22, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
