AI Research

AI Struggles with Humor: Research Reveals LLMs Misinterpret Puns with 20% Accuracy

Cardiff University research reveals that large language models misinterpret puns with only 20% accuracy, highlighting significant limitations in humor comprehension.

Staff

Published

24 November, 2025

Recent research conducted by teams at Cardiff University in south Wales and Ca’ Foscari University of Venice has provided new insights into the limitations of large language models (LLMs) in understanding humor, specifically puns. This study raises important questions about the capabilities of LLMs in grasping complex linguistic phenomena that often rely on cultural and contextual nuances.

Experimental Setup and Limitations

The research team aimed to explore whether LLMs can comprehend puns by evaluating their performance on a series of pun-based sentences. One of the tested examples was: “I used to be a comedian, but my life became a joke.” When this was altered to “I used to be a comedian, but my life became chaotic,” the models still recognized it as a pun. This indicated that LLMs are sensitive to the structure of puns but lack a deeper understanding of their underlying meanings.

In a similar vein, they tested the sentence, “Long fairy tales have a tendency to dragon.” When “dragon” was replaced with the synonym “prolong” or even a random term, the LLMs continued to identify the presence of a pun. This raises significant concerns regarding the models’ interpretative capabilities: while they can identify patterns from their training sets, they do not seem to genuinely understand the humor involved.

Professor Jose Camacho Collados from Cardiff University’s School of Computer Science and Informatics emphasized that the research highlighted the fragile nature of humor comprehension in LLMs. “In general, LLMs tend to memorize what they have learned in their training,” he stated. “They catch existing puns well, but that doesn’t mean they truly understand them.” The study found that when encountering unfamiliar wordplay, the LLMs’ ability to distinguish between humorous and non-humorous sentences can drop to just 20%.

Results and Findings

Another pun tested was: “Old LLMs never die, they just lose their attention.” When “attention” was substituted with “ukulele,” the LLM still perceived it as a pun, reasoning that “ukulele” phonetically resembled “you-kill-LLM.” This instance further illustrates the models’ reliance on phonetic similarities rather than semantic comprehension.

The findings of this research indicate that LLMs are adept at recognizing established puns from their training data but struggle significantly with newly generated or modified puns, demonstrating a clear limitation in their understanding of humor.

Research Significance and Applications

The implications of these findings are substantial, especially for applications requiring nuanced understanding, such as chatbots, customer service interfaces, and creative writing tools. The researchers caution that developers should exercise restraint when employing LLMs in contexts where humor, empathy, or cultural context is vital. The illusion of humor comprehension exhibited by these models could lead to misinterpretations and miscommunications, underscoring the need for human oversight in such applications.

This research was presented at the 2025 Conference on Empirical Methods in Natural Language Processing, held in Suzhou, China, and is detailed in their paper titled “Pun unintended: LLMs and the illusion of humor understanding.” By shedding light on the limitations of LLMs in one of the more intricate aspects of language, this work contributes to a growing body of literature that seeks to clarify the boundaries of what these models can realistically accomplish.

In summary, while LLMs have demonstrated remarkable prowess in various natural language processing tasks, their grasp of humor remains notably superficial. This study not only emphasizes the necessity for a cautious approach in deploying these models for applications involving humor but also highlights a broader research avenue focusing on understanding and overcoming the limitations of LLMs in interpreting complex linguistic constructs.

1 Experimental Setup and Limitations
2 Results and Findings
3 Research Significance and Applications

AI Cybersecurity

AI’s Cybersecurity Challenges: Setting Data Access Permissions for LLMs and Third-Party Tools

AI integration in corporate workflows demands stringent data access permissions to prevent sensitive information leaks, with shadow AI practices posing significant security risks.

Rachel Torres2 days ago

AI Education

Education System Must Adapt to AI: Teachers Urge Shift from Electronics to Critical Thinking

Educators urge a shift from electronics to critical thinking in classrooms, as AI tools like ChatGPT risk diminishing students' analytical skills.

David Park6 days ago

AI Generative

llama.cpp Achieves 40% VRAM Reduction and 20% Throughput Boost with Speculative Checkpointing

llama.cpp introduces speculative checkpointing, cutting VRAM usage by 40% and boosting throughput by 20%, enhancing local inference for large models.

Staff19 April, 2026

AI Generative

71% of Companies Use AI, Yet Only 11% Achieve Reliable Production Scale

71% of organizations use AI, yet only 11% of AI applications are production-ready, highlighting a critical gap in reliability and accountability

Staff19 April, 2026

AI Tools

AI Content Workflows Transition to Brand-Safe Standards with Enhanced Clarity and Authenticity

AI-assisted writing workflows are evolving to prioritize brand safety and authenticity, shifting focus from speed to clarity and nuanced tone, ensuring higher-quality content outputs.

Staff10 April, 2026

AI Marketing

Cvent Reveals AI-First Strategy as 70% of Event Planners Shift to AI Venue Searches

Cvent reveals a shift as over 70% of event planners now utilize AI for venue searches, emphasizing the critical need for hotels to optimize...

Sofía Méndez8 April, 2026

Anthropic Reveals Claude Sonnet 4.5’s Emotion Signals Impacting AI Behavior and Decision-Making

Anthropic's Claude Sonnet 4.5 reveals 171 emotion-like signals that shape AI decision-making, raising critical implications for educational technology and workforce applications.

Staff6 April, 2026

AI Technology

Intel and Dell Highlight AI PC Upgrades to Cut Cloud Costs and Boost Efficiency

Intel and Dell unveil new AI-capable PCs designed to run smaller language models locally, slashing cloud costs and enhancing operational efficiency for businesses.

Staff2 April, 2026

AIPRESSA.COM

AI Research

AI Struggles with Humor: Research Reveals LLMs Misinterpret Puns with 20% Accuracy

Experimental Setup and Limitations

Results and Findings

Research Significance and Applications

Trending

Top Stories

Albania Appoints AI Bot Minister Diella Amid Corruption Concerns and EU Membership Goals

AI Government

BigBear.ai Launches Biometric Platform at O’Hare, Acquires Generative AI Ask Sage for $250M

AI Cybersecurity

Endpoint Security Market to Reach $23.9B by 2030 with 7.2% CAGR Amid Rising Cyber Threats

AI Business

Enterprise Architecture Shifts to Strategic Enabler in AI-Driven Business Models

AI Technology

AI Hardware Market Grows 30% in 2025, Driven by Generative AI and Edge Computing Demand

You May Also Like

AI Cybersecurity

AI’s Cybersecurity Challenges: Setting Data Access Permissions for LLMs and Third-Party Tools

AI Education

Education System Must Adapt to AI: Teachers Urge Shift from Electronics to Critical Thinking

AI Generative

llama.cpp Achieves 40% VRAM Reduction and 20% Throughput Boost with Speculative Checkpointing

AI Generative

71% of Companies Use AI, Yet Only 11% Achieve Reliable Production Scale

AI Tools

AI Content Workflows Transition to Brand-Safe Standards with Enhanced Clarity and Authenticity

AI Marketing

Cvent Reveals AI-First Strategy as 70% of Event Planners Shift to AI Venue Searches

Top Stories

Anthropic Reveals Claude Sonnet 4.5’s Emotion Signals Impacting AI Behavior and Decision-Making

AI Technology

Intel and Dell Highlight AI PC Upgrades to Cut Cloud Costs and Boost Efficiency