// TECHCRUNCH — INTELLIGENZA ARTIFICIALE
I created an interactive digital avatar of myself — and you can talk to it
When Alexandru Voica, head of corporate affairs at the video-generation startup Synthesia, sent me a link this summer to the newest addition to their PR team, I was surprised. It was an interactive virtual avatar of him, trained to answer common press questions about Synthesia, like what it does and how it works. The day before, I was on a panel where PR people asked me if I minded pitches that used AI-generated text. But Alexandru’s avatar was beyond that — this seemed to me the final boss of using AI in PR.
In September, Synthesia invited me to its new office space in New York. Originally based in the U.K., Synthesia is a hot digital avatar startup alongside others like D-ID, HeyGen, and Colossyan. It hit a $4 billion valuation earlier this year and said last year it had crossed $100 million in ARR.
Synthesia lets enterprises build interactive training videos with AI avatars and recently launched a product called Roleplay Sessions that lets employees practice, for example, sales pitches with an interactive AI avatar that responds and scores their responses.
When I went to the new office opening and they asked me if I would like my own AI avatar, I didn’t even hesitate to say yes. Of course I would like my own digital twin. My outfit was cute that day, and my hair was in place.
Until I met my digital twin, I had been indifferent toward avatars, but I felt they would inevitably become part of everyday online life. I heard of people on Instagram creating them in their likenesses to help them make social content. I find it all to be very interesting, and it’s perhaps why I have no qualms now presenting to you my digital twin. This is the first time Synethsia has made a digital avatar for a journalist (or for anyone, period, outside of Voica). It is trained on my story about why venture-backed startups commit more fraud than non-VC-backed startups and will only answer questions about that story.
Below, simply press “start in a new window” to begin.
To make this, I entered a mini film studio nestled inside Synthesia’s office where they took numerous photos of me and captured a two-minute recording of my voice. I had to consent to these avatars being made and, well, digital Dom was born. They created a personal avatar for me (one that just reads whatever script I give it) — with and without glasses — and they made two interactive avatars for me (ones that can talk back and listen to me), also with and without glasses.
We picked an article to train the interactive avatar on, and then one of the teams built my interactive avatar, which is powered by a combination of voice-to-text, video, language, and text-to-voice models. My avatar’s tech stack includes Synthesia’s own video and voice models, although the company also allows customers to choose alternatives from other labs like Cartesia, ElevenLabs, Google, or OpenAI. Enterprises can also choose to host their avatars on whatever cloud they want or pay Synthesia to host them.
The voice-to-text model turns what people say into text, the agentic language model makes sense of text and can take actions based on it, the text-to-voice model turns a response into audio, and finally a video model (built by Synthesia) animates the avatar as it talks.
Overall, Synthesia builds three types of products — a video-creation and distribution platform with classic avatars, where someone types in a script and the avatar repeats it; an agentic platform called Sessions where people can interact with the avatars in surveys or roleplay; and an API platform where people can take Synthesia video and voice models and combine them with other tech services to build interactive avatars or other types of products.