Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?
AbdulRahman A. Morsy, Aya Zirikly
- Published
- Sep 23, 2026 — 15:08 UTC
Problem
The paper addresses a gap in the capability of large language models (LLMs) to generate classical Arabic maqamat, a form of Arabic prose that combines storytelling with rhetorical techniques. This area remains underexplored in the literature, particularly regarding the effectiveness of various prompting strategies in enhancing the quality of generated maqamat. The work is presented as a preprint and has not undergone peer review.
Method
The authors evaluated five different LLMs using three prompting strategies: zero-shot, few-shot, and rule-based prompting. The evaluation was conducted through a combination of human annotation and an LLM-as-a-judge framework. The evaluation dimensions included rhetorical richness, saj density (a form of rhymed prose), structural coherence, and stylistic authenticity. This multifaceted approach allows for a comprehensive assessment of the models' capabilities in generating maqamat.
Results
The results indicate that few-shot prompting consistently improves saj density compared to zero-shot prompting. However, zero-shot prompting yielded the highest aggregate scores across all five models when evaluated on the specified dimensions. Notably, the strongest models, identified as GPT-4o and GPT-5.4-mini, showed the most significant benefits from rule-based prompting, particularly in terms of rhetorical richness and structural coherence, outperforming other models under the same conditions. The available text does not report quantitative results.
Limitations
The authors did not report any limitations in their study. However, the absence of reported limitations may suggest a lack of comprehensive evaluation across diverse contexts or potential biases in the evaluation methods employed.
Why it matters
This research has implications for the development of LLMs in culturally specific contexts, particularly in generating rich, stylistically authentic content. The findings can inform future work on enhancing LLM capabilities in generating complex literary forms, potentially leading to improved applications in natural language processing tasks that require a deep understanding of cultural and stylistic nuances.
By Callan Zhang · Sep 23, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
