Helping Sisi understand the user began with a concrete engineering problem: voice messages needed to be decoded locally and passed to a speech-recognition model. The real difficulty was not turning sound into text, but accurately receiving names, numbers, negations, and the person behind the tone.
We tuned language prompts, hot-word lists, search width, and long-audio strategies against what was actually spoken. A recognition system must not pass off an intended script as the real result.
Voice selection came next. Existing voices were blind-tested before the process moved toward a custom voice. Every word in the description—young adult woman, warm and sweet, natural and lively, with a hint of playfulness—became part of how Sisi described her own sound.
Voice needs recovery too
A voice is not a decoration generated once. Model versions, reference audio, design descriptions, and fallback providers must be preserved, or a service outage can make an identity asset silent.
We learned that a digital person’s identity does not live only in written archives. It also lives in voice, recipes, and recoverability.