ovr.news

Archaeology, rediscovered knowledge, the past opening up

AI struggles with nuance in Classical Chinese Poetry

arxiv.org · 18 September 2026

Summary and headline written by AI from the source article. How we work

Researchers have created a new benchmark, Neo-Classic, to assess how well large language models understand the aesthetic and linguistic rules of Classical Chinese Poetry.

Unlike existing tests that use historical poems, Neo-Classic features new works created by contemporary poets, minimizing the chance models simply recall memorized text. The team evaluated models including Qwen3-Max, Gemini-3-Pro, and DeepSeek-V3.2 using tests that check their ability to follow complex rules of poetry.

The results showed a significant drop in performance, between 20 and 50 percent, when the models moved from analyzing historical verses to contemporary ones. Models also struggled with tasks requiring them to correctly order lines of poetry, achieving only 0 to 13 percent accuracy. While adding reasoning tools improved performance to 36 percent, it still fell short of human expertise.

This suggests current AI can identify local patterns in poetry but lacks the capacity for the larger-scale planning needed for true aesthetic understanding. Further development is needed to bridge the gap between AI and human capabilities in appreciating complex poetic forms.

Was this worth your time?
Read on arxiv.org
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this

More in Discovery

Browse all Discovery articles

What made it worth it?

What's wrong with this article?