The growth of audiobooks could well reflect our innate hardwiring.
I concede. It’s a bit clickbait-y. But there is a serious observation here.
Humans are primally oral creatures. We intuitively speak and listen from an early age. Our civilisation began with an oral tradition, passing down stories, connections, learning and wisdom – as well as a prejudice and misapprehension here and there.
On the other hand, we must actively learn to write and read, and the spectrum of ability in doing so is far wider than what we all achieve in our talking and listening.
So, could it be that the growth in audiobooks (and other audio formats) we are seeing simply reflects our innate hardwiring? Are we fundamentally programmed to respond better to expressed thoughts with our ears than our eyes? In a world awash with text – emails, documents, WhatsApp threads, Slack notifications, bite-sized commentaries on doom-scroll feeds – one might ask whether reading has become our default mode of engagement. But was it always meant to be this way? Do we actually prefer aural?
Human language began with sound. As I trace in my book, Shimmer, Don’t Shake – How Publishing Can Embrace AI (available in English, Arabic and very soon Mandarin, Spanish, Greek, Tamil and Uzbek), long before the first alphabets were scratched onto clay or papyrus, we told stories around fires, passed on traditions through chants and remembered our past through oral histories. Spoken language likely emerged between 50,000 and 150,000 years ago. Writing? It’s only a 5,000-year-old invention. Reading obviously followed writing.
That’s not just a matter of chronology – it’s evolutionary. All human cultures have spoken language. Not all have written it. Speech is instinctive. Writing is not.
In his seminal work Orality and Literacy: The Technologizing of the Word, the anthropologist Walter J Ong made the case that literacy fundamentally changed human consciousness – but also that oral culture shaped our minds in deeply social and mnemonic ways. We’re built to speak and to listen. Reading and writing, as Ong puts it, are technologies we had to learn.
The evidence is clear in childhood development. We acquire language through listening and imitation. No one teaches a toddler how to speak in the formal sense – they absorb it, almost like magic. Reading, in contrast, must be explicitly taught. It involves decoding abstract symbols, mapping them onto phonemes and engaging visual and language-processing networks in the brain. It’s hard work. In fact, reading comprehension often lags behind listening comprehension well into adolescence. That’s not a failing. It’s biology. As neuroscientist Maryanne Wolf puts it: “We are not born to read.” The reading brain is a neural workaround – an astonishing one, to be sure, but not an innate one.
Modern neuroscience supports this aural inclination. Studies using Functional Magnetic Resonance Imaging (fMRI) have shown that listening to a story activates not just language centres in the brain but also emotional and sensory areas – we feel more when we hear. Audio delivers prosody, tone, rhythm and nuance. A pause can imply doubt. A sigh, grief. Text can only approximate this with punctuation and formatting.
Studies using fMRI have shown that listening to a story activates not just language centres in the brain, but also emotional and sensory areas — we feel more when we hear
Interestingly, some researchers have found that auditory narratives are often remembered more vividly than written ones. The reason? Oral language evolved in tandem with memory techniques – rhyme, repetition, alliteration – all designed to lodge words in our minds. We’re wired to remember sounds.
If we are indeed aural creatures at heart, the digital era is proving to be very sympathetic. The explosion of audio media – podcasts, audiobooks, voice notes, virtual assistants – suggests that even in our screen-saturated world, we’re returning to our oral roots. Spotify’s podcast listenership now rivals its music base. Audiobooks are growing faster than print books. Millions fall asleep to meditation guides, news roundups, or longform essays read aloud. Far from being obsolete, the human voice is undergoing a renaissance.
There is practicality here, too. Listening is more ambient – you can absorb ideas while walking the dog, commuting or washing the dishes. But it’s also more intimate. A well-told story, spoken directly into your ear, is a different kind of connection. It mimics the evolutionary intimacy of campfire storytelling or whispered confessions.
A personal note here: I love writing. I love reading. I am a fan of curling up with a book and reading, uninterrupted, through the night. But I recognise that our world today is characterised by what I call “multi-modal interactivity”. We just don’t seem to be able to do only one thing at a time. This does pose the question as to whether the printed and written book, off-line and mono-modal, truly synchs up with how our brains work and have nowadays been trained. Is this why, worryingly, reading is in decline? Might more aural be the best way to ensure future generations fall in love with written works that are not delivered by printed words on a page?
There’s one more dimension to consider: trust. We trust voices more than text. Why? Because voice reveals intent. It carries emotion, cadence, vulnerability. A sentence like “I’m fine” can mean radically different things depending on how it’s said. In text, it’s flat. In sound, it’s alive.
That may explain why voice notes are so popular among younger generations – and why remote teams often turn to voice or video when clarity matters. Even in business, tone can do what bullet points cannot.
So, do we need more aural? Yes – not at the expense of reading or writing, but as a recognition of our deeper nature. The shift to aural doesn’t mean a rejection of literacy; it means embracing a more complete understanding of how we absorb, relate to and remember information.
It also means we can guard against anachronism in publishing. We not only need to be more aural, we also need to be more multi-modal.
