AI struggles with nuance in Classical Chinese Poetry
Summary and headline written by AI from the source article. How we work
Researchers have created a new benchmark, Neo-Classic, to assess how well large language models understand the aesthetic and linguistic rules of Classical Chinese Poetry.
Unlike existing tests that use historical poems, Neo-Classic features new works created by contemporary poets, minimizing the chance models simply recall memorized text. The team evaluated models including Qwen3-Max, Gemini-3-Pro, and DeepSeek-V3.2 using tests that check their ability to follow complex rules of poetry.
The results showed a significant drop in performance, between 20 and 50 percent, when the models moved from analyzing historical verses to contemporary ones. Models also struggled with tasks requiring them to correctly order lines of poetry, achieving only 0 to 13 percent accuracy. While adding reasoning tools improved performance to 36 percent, it still fell short of human expertise.
This suggests current AI can identify local patterns in poetry but lacks the capacity for the larger-scale planning needed for true aesthetic understanding. Further development is needed to bridge the gap between AI and human capabilities in appreciating complex poetic forms.


