ElevenLabs, the company whose voice models convert text into human-sounding speech, is reportedly valued at $22 billion by its backers, according to statements from its chief executive. The four-year-old company says it is pacing at $600 million in annual recurring revenue, with more than 55 percent of that total coming from classic enterprise customers. Much of the remaining 45 percent comes from small and medium businesses, developers, builders, and creators. Many U.S. consumers encounter the technology through automated phone support, including first-line service that Klarna runs for 35 million U.S. customers.

Co-founder and CEO Mati Staniszewski discussed the company's position and the voice AI market in an interview at Nrth in Toronto, a local entrepreneurship conference formerly known as Elevate. He said the business does not mind its gross margins narrowing if that helps expand its market share, though he declined to discuss the margins in detail. He also said businesses should disclose to customers when they are speaking with an AI agent rather than a person, a practice he expects to shift as agents become more common.

Staniszewski addressed competition with customers, including Decagon, a conversational AI platform that trained its voice product on ElevenLabs and now competes with it. He said the lines between model companies, platform companies, and application companies have become far more blurred than in the past, pointing to Anthropic as an example of a model company that has expanded into a platform and a broad set of applications. ElevenLabs customers can also select their reasoning layer from a menu of options, and he said the choice between frontier lab models and open-weight models depends on the use case.

In customer experience, Staniszewski said informational calls without executed actions can rely heavily on open source models because a knowledge base defines what counts as a good experience. Financial services calls, by contrast, require authentication and transaction details, leaving no room for error, so he said frontier models will continue to lead there. He said deployments for governments such as Poland and Brazil carry their own requirements and may use open-weight, closed-source, or fine-tuned models while keeping data residency. In a Polish healthcare case, he said agents call patients to remind them about appointments across the public health system, where 18 percent of patients never show up.

On training data, Staniszewski said much of the work has involved annotating data rather than simply amassing volume. He said thousands of internal contractors help annotate not only what was said but when people spoke, how they said it, and which emotions were used, and that voice coaches were brought in to detect accents accurately. He said the company does not train the text models and the intelligence side of models that sit at the center of the broader AI safety debate, and that its technology does not allow agents to create more agents. Every customer goes through KYC, he said, adding that the company has precautions in place for cybersecurity risk.

Staniszewski also said he would like ElevenLabs to be the first to pass the Turing test for conversational AI, which he said requires combining intelligence with emotional intelligence, including understanding the other side's emotions and adjusting pace or volume. He said the quality gap achievable at the model level remains significant, though he expects it to narrow in three to five years. Asked about an initial public offering, he said the company is preparing its foundation to be able to pursue one in the next years but that the decision will depend on time and place.

More company and startup news from TechManNews.