What poetry tells us about AI’s limitations

Intellyx Cortex by Jason Bloomberg

We’ve built the AI tower so tall, so quickly, that now we’re wondering why the whole thing is swaying. Perhaps we should check the foundations.

Agentic AI depends upon generative AI (genAI).

Generative AI depends upon large language models (and in some cases, smaller language models).

Language models depend upon language.

Human language.

And therein lies the problem with our foundation.

Not only are human languages replete with ambiguities and vagaries of meaning, but language models have no clue whatsoever what those meanings are.

All language models do is predict the next set of words in a sentence. They aren’t intelligent. They have no understanding of what the words actually mean.

Language models – and thus, AI generally – also lack creativity. Their response to a prompt – any prompt – is a guess at what words (or images) should come next, based on all the words that went into their training sets plus additional information at the time of the query.

One thing language models are really good at is fooling people – fooling us into believing they have intelligence, understanding, empathy, creativity, and other exclusively human characteristics.

It doesn’t take much critical thought to pierce this charade. Even though AI continues to improve, so too does our ability to identify AI slop vs. human-generated content.

This slop-identifying capability is certainly a skill we could all get better at, however. Let’s take a closer look at how language models actually behave to help us improve.

Our starting point? Poetry.

Why poetry is good for building our slop identification skills

How do you recognize good poetry? I was never an English teacher (math was my jam back in the day), so I’m going to boil the answer down to the essence of creativity: a poet must select words that aren’t predictable.

In other words, good poetry must surprise you.

Given the fact that the only thing genAI can do is provide predictable words, then we should be able to differentiate human-generated poetry from AI slop poetry by how surprising the words are.

Let’s work through an exercise. As the starting point, let’s use Edgar Allan Poe’s “The Raven” – partly because it’s an incontrovertible example of excellent poetry, but also because it’s familiar and in the public domain. (In case you slept through tenth grade English and need to refresh your memory of the poem, click here.)

Now, here’s the exercise: write a witty limerick about Edgar Allen Poe’s “The Raven.”

I chose the limerick form because it’s simple, short, familiar, and follows basic rhyming and meter rules.

Before turning to AI, I tackled the exercise myself, with nothing more than Poe’s poem and a bit of assistance from a rhyming dictionary as inputs. Here’s what I came up with:

“Nevermore!” Poe’s raven doth quoth
A bird and a metaphor both
“Lenore!” Poe bemoans
But deep in his bones
His anger and grief doth both groweth

I’m the first to admit that when it comes to my skills as a poet, I shouldn’t quit my day job. Be that as it may, note that what humor my poem contains is the result of my surprising word choices: ‘doth’ making fun of Poe’s ‘quoth,’ and then repeating the joke with ‘doth both groweth.’

Also note my use of poetic license when I broke from the strict limerick rhyming rules with the two-syllable ‘groweth’ to act as a punch line. In fact, poetic license is one of the poet’s most potent tools, as breaking the rules of poetry lead to surprise – again, the hallmark of creativity.

Can AI beat my skills as a poet?

Limericks are such a simple, rigid form with numerous extant examples that surely AI should have no trouble coming up with some good ones, right?

Furthermore, given the bit of doggerel above, the bar we’re setting for AI to craft its own limericks is quite low. Nevertheless, let’s give it a shot.

Here’s what ChatGPT came up with:

There once was a raven named Poe,
Who answered one question, “No more.”
Though Lenore had gone,
The bird lingered on—
And depressed poor Ed Edgar some more.

If you feed ChatGPT the same prompt, you’ll certainly get a different limerick – but I doubt it will be much better than this one, which is objectively awful.

In fact, it has numerous flaws: the use of ‘there once was’ is cliched; the raven’s name wasn’t Poe; the ‘Poe’/‘more’ rhyme, while an example of poetic license, serves no purpose; and perhaps worst of all the extraneous ‘Ed’ before ‘Edgar,’ presumably for no other reason than to make the meter work out.

Bottom line, there is nothing surprising about this poem (other than just how terrible it is). Comparing it to my poem, it should be clear which one is human-generated and which one is slop.

Let’s give other language models a chance

Here’s a limerick from Microsoft CoPilot:

There once was a raven named Lenore,
Who tapped at a poet’s front door.
“Is this dread?” Poe cried,
“No, snacks,” it replied—
Then croaked, “Nevermore… but encore!”

CoPilot’s limerick is nonsensical, and not in a good way. The AI also mangles the limerick’s meter. And note the use of the em dash, in the same place as the one in the ChatGPT limerick – a surefire signature of AI slop.

Google Gemini can certainly do better, right? Here’s Gemini’s take:

There once was a girl named Lenore,
And a bird and a bust and a door,
And a guy with depression,
And a whole lot of questions,
And the bird always says “Nevermore.”

On first glance, this limerick is surprisingly good. It sticks firmly to the rules of the form, it makes sense, it’s humorous…and it’s surprisingly familiar.

It’s familiar, of course, because it’s not AI slop at all. It’s an old human-generated poem, so old in fact, that apparently nobody knows who wrote it. One thing’s for sure: it predates AI.

What happened here? The other two AI models returned ‘original’ slop, while Gemini acted simply as a search engine, returning the entirety of a result it found online.

On the one hand, Google is first and foremost a search engine, so we shouldn’t be surprised. On the other, it simply plagiarized the poem.

If a middle school English teacher gave this assignment and one of the students turned in Gemini’s answer, they would get a failing grade for simply copying the answer instead of coming up with their own. Why, then, should we give Gemini a pass?

In any case, I wanted to see if Gemini could actually write its own limerick, so I adjusted the prompt to: write an original, witty limerick about Edgar Allen Poe’s “The Raven.”

Here’s Gemini’s response:

A gloomy young man on a bust,
Found a raven perched high on a bust.
When he asked of Lenore,
The bird swore at the door,
With a “Nevermore” starting his rust.

More nonsense, and what poet would ever rhyme a word with itself? We humans find that choice jarring, but from Gemini’s perspective, repeating the word is a logical choice because it follows the rhyming rule.

The Intellyx take

We cannot say what new types of AI will be able to do in the future, but I can confidently guarantee that generative AI lacks the human creativity necessary to produce poetry that people would recognize as a better quality than everyday slop.

True, as genAI improves, I’m sure the various models will be able to get the limerick rules right, create poems that make more sense, and even eschew the use of em dashes. But the poems they produce will always be the result of predicting the next word in a series, rather than any kind of creative process that duplicates human creativity.

The point of any exercise like this one is to generalize the lesson it teaches to broader situations. Any endeavor that requires creativity – whether it be a college essay, a legal brief, marketing copy, application code, or a feature-length film – requires the irreplaceable hand of a human.

All genAI can do is regurgitate word patterns (or image or video elements) it’s seen before. True, it’s very good at such regurgitation, but mimicking some human’s creative output is not the same as creating something original.

I’ll leave you with your own exercise: feed the two prompts above into your AI chat window of choice – not only now, but perhaps weeks or months from now. It will be fascinating to see if any of the word choices I made in my original poem make it into future AI attempts at the limerick.

Then do a simple web search on the phrase “doth both groweth.” This search returns no results for me now before I publish this article, and it will certainly return this article itself once I do.

But over time, will the phrase turn up on other places because AI has been using my poem as fodder for its own designs? If so, what lessons can you learn from how that behavior impacts whatever tasks you’re using AI for?

Only one way to find out!

Copyright © Intellyx BV. Intellyx is the change agent industry analysis and advisory firm focused on enterprise transformation. Covering every angle of enterprise IT from mainframes to artificial intelligence, our broad focus across technologies empowers business executives, IT professionals, and software vendors to leverage disruptive trends to succeed in a dynamic business environment. Microsoft is a former Intellyx customer. None of the other vendors mentioned in this article is an Intellyx customer. This article was written by a human (except for the explicit AI contributions). Image credit: Craiyon.

SHARE THIS: