
Deepdub Launches Phantom Z 3.4 Conversational: Multilingual Text-to-Speech Built to Survive Real Customers, Not Just Demos
Enterprise-grade real-time text-to-speech delivers 150ms time to first audio at full 48 kHz, with text normalization that gets account numbers, invoice totals and appointment dates right
TEL AVIV, Israel, Sept. 3, 2026 /PRNewswire/ -- Deepdub, a foundational voice AI company pioneering expressive voice technologies, announced today the launch of Phantom Z 3.4 Conversational, a new multilingual text-to-speech model with high-fidelity 48 kHz audio, improved text normalization and extended Hebrew support. The model is available to all Deepdub clients now.
For enterprises running voice agents, a call holds together when four things go right at once. The voice sounds like a person. The response arrives fast enough to feel like a conversation. The agent knows when to speak and when to listen. And every account number, date and amount comes out the way a customer would say it. When one of them slips, the call escalates to a human, and that is where containment and cost are decided. Phantom Z 3.4 Conversational is built for all four.
"Every voice model sounds impressive for two minutes in a demo. Very few survive two weeks with real customers," said Ofir Krakowski, CEO and co-founder of Deepdub. "Deployments don't stall on the 95% a model gets right, they stall on the misread account number, the mangled surname, the one wrong digit on a live call. We built this model for that last few percent, because in production, the last few percent is the whole product."
In English, the work is in text normalization, the step that turns written text into spoken words. A delivery date written 2024-12-31 is read as December thirty first, twenty twenty-four rather than as a run of digits. An invoice total written $1,240 is read as one thousand two hundred forty dollars. An appointment at 14:30 is read as two thirty. A reference written Chapter VII is read as chapter seven rather than as letters. These are the categories where Deepdub's testing puts the model ahead of the other systems it was measured against. An enterprise running more than one language gets one set of behavior to test and one contract to hold rather than two.
Phantom Z 3.4 delivers an end-to-end p95 time-to-first-audio of 150 milliseconds in real-time mode at full-range 48 kHz audio, with cross-language voice transfer from under three seconds of reference audio. Deepdub builds and trains its own speech models from random rather than licensing them, which allows the company to bring a new language into production in two weeks. Deepdub covers more than fifty locales and dialects verified by local voice and language experts, inside a platform supporting more than 50 locales and dialects.
"We run Deepdub in production for live, real-time phone calls, where latency and naturalness aren't nice-to-haves but the key factor in whether a caller stays on the line. 3.4 is the closest we've heard a synthetic voice come to a real person, and our callers show it: they stay longer, talk more, and engage with our agents like we've never seen before," said Adir Haziza, CTO at Voiceman.
The hardest case is Hebrew, which is written without vowels, so the same letters can spell different words. The three letters of שלט are a sign read one way and a remote control read another. A model that reads one word at a time has to guess which the sentence means, and in Hebrew a wrong guess is not an accent, it is a different word that stays invisible until a customer hears it. Phantom Z 3.4 resolves this at the source. Pronunciation is decided from the whole sentence rather than word by word, and every instance of שלט in Deepdub's Hebrew test set was read correctly. Where a brand name or a plan tier has to be said a particular way, marking it in the text is enough. Deepdub ranks first for Hebrew text-to-speech on the public TTS Arena leaderboard hosted by ivrit.ai on Hugging Face.
In Hebrew, national ID numbers, appointment dates and transaction amounts are expanded before speech, so a balance written as 1,240 ₪ is spoken in full rather than read out as digits. In blind listening tests, Phantom Z 3.4 was preferred over Deepdub's previous Hebrew model in 71 percent of decisive comparisons.
"We needed something that would hold up consistently across a large volume of work, so we tested it thoroughly before deciding. What stood out was that the details came out right and the Hebrew was the most natural we'd heard," said Dor Levy, Head of Jeen Talk at Jeen AI.
About Deepdub
Deepdub is the foundational voice AI model company pioneering expressive voice technologies for global enterprises across TV, film, advertising, gaming, e-learning, and AI-agent applications. The company's international team of technology, dubbing, and linguistic experts deliver an end-to-end voice solution that preserves the emotional and cultural integrity of original content in more than 50 locales and dialects. With an advisory board that includes media leaders such as Kevin Reilly, former Chief Content Officer at HBO Max, and Emiliano Calemzuk, former President of Fox Television Studios, Deepdub is eliminating language barriers to enable the global diffusion of media on major streaming platforms like Netflix, Amazon Prime, and Hulu. Visit https://deepdub.ai or follow us on LinkedIn for more information.
Deepdub Media Contact
Zivit Katz
Deepdub
[email protected]
SOURCE Deepdub
Share this article