Why publishers are rethinking AI audiobook strategies for long-term growth
What to learn from the Frankfurt Book Fair audio whitepaper
The FBM Audio x AI Whitepaper 2026 in cooperation with dosdoce.com is a survey of 85 global publishers, production studios and platforms that reveals that while AI tools are widely adopted in audiobook workflows, strategic integration remains shallow—and the real business opportunity lies not in cost reduction alone, but in catalogue expansion, multilingual reach and operational redesign.

Published: 6.10.2026 | Foto: Dosdoce, AI generated Magnific
Artificial intelligence has moved from experimental curiosity to operational reality in audiobook production. According to a new report published by Dosdoce.com and Frankfurter Buchmesse, 85% of surveyed industry professionals now use AI tools in at least one stage of their audio production workflow. Yet only 17% have integrated AI across more than six production phases, suggesting that adoption remains confined to isolated tasks rather than end-to-end transformation.
The findings emerge from responses by 53 publishers, production studios and distributors, contextualised by insights from 32 additional audio professionals worldwide. Participants span Europe, Latin America, North America, Africa and the Middle East, representing major trade houses, independent publishers, streaming platforms and AI technology providers.
The report, titled Who Is Narrating the Future? AI Adoption, Use and Outlook in Audio Production, identifies a sector at an inflection point: AI is no longer a fringe experiment, but its strategic deployment lags behind the adoption headline. Most organisations have yet to define a coherent AI production strategy, evaluate technical capabilities rigorously, or analyse the full business case for catalogue expansion and multilingual growth.
The three-speed Model: coexistence, not replacement
The dominant response across the industry rejects a binary choice between human and synthetic narration. Instead, a three-tier production model is emerging.
Premium human narration is reserved for flagship titles with high editorial or commercial value—major releases, culturally significant works, or projects where interpretive performance is central to the listener experience. Hybrid production, combining human direction with AI-assisted workflows, is expected to become the most common model, with AI handling mechanical tasks such as text preparation, quality control and draft synthesis while human narrators and editors retain creative control. Fully synthetic narration is deployed for backlist titles, niche academic content, or works in languages and markets where traditional production economics have never been viable.
This segmentation reflects a pragmatic recognition: AI enables catalogue expansion rather than displacing human talent. As Laure Saget, CEO of AudioLib Hachette France, notes, "Human oversight remains essential at every stage of audiobook production. While technology continues to evolve, preserving the integrity of the relationship between publishers, authors, performers and audiences remains paramount."
The economic case is real but nuanced. Respondents report cost reductions in the range of 20% to 50% by the second or third year of AI integration—far below the 85% to 95% savings figures promoted by technology vendors. When complete workflows are accounted for, including text preparation, quality control, post-production and mastering, the actual savings are modest. The value proposition lies elsewhere: making previously unviable titles economically feasible.
WHO IS NARRATING THE FUTURE?
Free whitepaper: AI adoption, use and outlook in audio production.

Barriers to scale: rights, quality and market acceptance
Despite widespread adoption, significant obstacles remain. Half of respondents identify external restrictions as the principal barrier: literary agents who prohibit AI use, publishers and authors who insist on human narration, or platform policies that limit certain AI tools. A further 25% cite quality concerns, including low naturalness of synthetic voices, insufficient availability of voices in local languages or accents, and manuscript security risks.
Technical integration challenges and the absence of industry standards account for another 22% of reported barriers. The implication is clear: AI will scale sustainably only if the sector resolves issues of trust, authorisation, rights frameworks and commercial acceptance.
Transparency and disclosure are non-negotiable. Ninety percent of respondents recognise the need for greater transparency toward both authors and consumers about where and how AI is used. Sixty percent emphasise consent-based voice cloning, ethically licensed voice datasets and fair compensation for voice talent as prerequisites for minimising legal risk and ensuring long-term viability.
Market data underscores the gap between production capability and consumer adoption. In the United States, AI-narrated titles accounted for only 0.03% of total audiobook revenue in 2025. While the volume of AI-narrated audiobooks has increased, willingness to try them remains limited: just 16% of audiobook listeners have listened to an AI-voiced title. Respondents interpret this as evidence that production capacity currently outpaces consumer demand, and that stated sentiment diverges from revealed behaviour such as completion rates and blind-test preference.
Global expansion and the risk of reinforcing inequality
Translation and localisation are repeatedly cited as AI's largest business opportunity. Emerging audio markets—Latin America, the Middle East, Africa—are named specifically as regions where AI-driven cost reduction could meaningfully diversify accent and language coverage and expand catalogues that remain heavily constrained.
Rachel Ghiazza, Audible's Chief Content Officer, frames this as a category growth imperative: "AI production tools will help us close the content gap so that more listeners, across more languages and genres, can find a story that speaks to them. For Audible, this is an 'and' approach. We're expanding our audiobook catalogue through AI, while also continuing to invest more than ever in professionally narrated titles."
Yet this optimism comes with caution. Without deliberate investment in diverse and ethically sourced voice datasets, multilingual models, transparent rights frameworks and fair compensation for local voice talent, AI-driven expansion risks reinforcing existing global inequalities in representation rather than correcting them.
Ama Dadson, Founder and CEO of AkooBooks Audio Ghana and Frankfurt Buchmesse Audio Ambassador, warns: "Without intentional representation of African voices and languages, the next generation of AI audio risks reinforcing existing global inequalities rather than expanding access. For African publishing, AI's greatest opportunity may not be replacing professional narrators, but making it economically viable to produce educational content, local-language titles and backlist works that would otherwise never become audiobooks."

The six-phase framework: from strategy to scale
The report proposes a structured six-phase framework for integrating AI into audio production workflows.
Phase 1: Defining the AI Production Strategy
Before evaluating tools, organisations must analyse business opportunities across markets and languages, assess backlist profitability in audio, identify underserved genres, and determine which production processes can be automated, which require human intervention, and which must remain entirely human. This phase requires senior executive engagement and a solid understanding of AI fundamentals—not merely cost reduction, but growth potential.
Phase 2: Pre-Production and Text Preparation
Text preparation determines final audio quality more than any other step. This phase involves pre-production assessment, extraction and cleanup of source text, editorial treatment of tables, footnotes and images, construction of pronunciation guides, and final normalisation of numbers, abbreviations and punctuation. Human verification at each step is non-negotiable; no AI tool operates without errors.
Phase 3: Performance Tools
Choosing a synthetic voice without testing it with real text from the work is identified as one of the costliest mistakes. The report recommends a four-step evaluation protocol: build a test with real excerpts, generate samples with neutral settings across multiple vendors, conduct blind listening sessions, and document the decision. Directing a synthetic voice requires translating interpretive intuition into parameters the model can process.
Phase 4: Sound Effects and Music
Evaluation criteria for music and sound-effect tools should prioritise clear legal provenance of training data over whichever platform sounds best today. The report recommends verifying licences at the official source and documenting each generated piece with structured prompts.
Phase 5: Quality Control and Post-Production
Strong pre-production is worth more than endless quality-control rounds. The report recommends separating detection from correction, running two complete passes—one for fidelity and interpretive consistency, one for technical artefacts—followed by final sampling. Each complete pass requires approximately 1.2 times the length of the audio.
Phase 6: The Three-Speed Model
Moving from single-title production to catalogue scale requires mapping the entire catalogue across the three production models, building title-by-title allocation tables, designing hybrid workflows, and setting up scale-management systems including stage-by-stage boards, libraries of reusable assets, and per-model checklists.
Outlook: disruptive growth expected, but trust remains fragile
Eighty-six percent of respondents expect disruptive or moderate growth in the audiobook sector attributable to AI tools. Forty-four percent anticipate disruptive growth, 42% moderate growth, and only 14% foresee no significant change.
Several respondents describe a future in which audiobooks cease to be fixed, pre-produced assets and instead become on-demand renderings, generated at the moment of listening in whatever voice and language the listener selects. The binding constraint for this evolution will be licensing negotiations with rights holders rather than technology.
As discoverability becomes harder in an age of abundance, quality rather than price alone is expected to become the primary differentiator. Distinctive human performances and recognisable narrators will function as marketing assets capable of cutting through high-volume AI catalogues.
The recurring thread across nearly every response is that growth is not sustainable without transparency and disclosure toward both authors and listeners, and that trust, once spent, will be far costlier to rebuild than any single title was to produce.
James Long, COO at Pan Macmillan UK, summarises the industry's position: "The main goal for audio publishing right now is to increase the number of titles available for listeners around the world, that they can discover and enjoy wherever they are, in their own language if possible. End-to-end application of AI to enhance the work of our teams will make more titles viable for audio publishing. That only works, though, if it's built on ethically licensed foundations and clear standards around rights and disclosure, so speed and scale never come at the expense of trust."
Expert assessment: building smarter organisations
The real challenge ahead is not whether AI can make publishing more efficient, but whether that efficiency comes at trust's expense. The companies that pull ahead over the next three years will not necessarily be the ones that move fastest with AI. They will be the ones that pair it with openness: keeping humans in the loop, safeguarding creative rights, and never losing sight of the people whose talent made the publishing industry possible in the first place.
The report proposes a four-lane action plan: embrace individual productivity gains from AI tools within existing tasks; redesign entire workflows by coordinating human-AI collaboration; enhance new value-added services based on reader and listener personalisation; and actively search for new technologies that could transform the existing audiobook business model.
Michele Cobb, Partner at Forte Business Consulting USA, observes: "It's fascinating to see how companies are incorporating AI into their workflows while they continue to utilise and support human narration. I do believe the two are not mutually exclusive, and reconciling technology versus art is not unique to audio publishing, but it provides greater challenges and more discussion in an industry built on emotional interpretation and the power of vocal performance."
Conclusion: strategy before tools
The report's central argument is unambiguous: AI in audiobook production is not, above all, a technological revolution, but a business decision. Before evaluating tools or calculating cost savings, publishers must analyse business opportunities, assess catalogue potential, and determine which processes can be automated and which must remain human.
The boundaries between human-created and machine-created audio are blurring rapidly. The industry is at a defining moment. AI tools are enabling creators and producers to scale audiobook production, expand distribution opportunities worldwide, and generate new revenue streams. But the opportunity is contingent on ethical foundations, transparent disclosure, and a recognition that the human voice—interpretive, emotional, culturally grounded—remains irreplaceable for the works that demand it.
As the report concludes: effectively integrating AI into editorial audio production is not about finding the right tool, but about building a management process—documented, verifiable and repeatable—into which any tool can be inserted with confidence at any phase of production. When moving from a few titles to many, the bottleneck stops being production capacity and becomes management and team coordination.

Javier Celaya is a founding partner of Dosdoce.com, where he has been supporting creative and cultural organizations in their digital transformation for more than 20 years. He currently serves as Chairman of the Board of Aniara.One, a Swedish start-up that is revolutionizing the publishing industry through AI-powered translation and multilingual production.
Drawing on his extensive experience advising leading digital platforms such as Bookwire, Storytel, and Podimo on their international expansion strategies, Javier has in-depth expertise in the Spanish-speaking markets of Spain and Latin America. He works as a strategic advisor to publishers, streaming platforms, and media companies, and has taught at institutions including NYU’s Advanced Publishing Institute as well as several Spanish universities.
