Speech out · Live · Speech in upcoming
Hear the reply — then grow into fuller voice loops.
Voice listen is live: after a soft-live reply, people can listen in 粵語 (default), 普通话, or English. Speech-in (ASR) and sign-language options are on the roadmap once knowledge quality is solid.
Also known as CalmStep.AI’s speech-out layer (ASR upcoming)
Multimodal access without rushing quality
CalmStep.AI ships speech-out first because listen quality is an active improvement axis alongside knowledge. We deliberately park heavy ASR build until the text knowledge bank is strong — then the same brain powers talk, type, and later universal access paths.
Active improvement axis (b)
Voice quality work focuses on natural pacing, language fit, and a calm listen UX for Greater China users.
What the voice layer includes
Speech out (live)
Listen to soft-live replies in 粵語, 普通话, or English — an anonymous and secure audio landing after text coaching.
Speech in (upcoming)
Automatic speech recognition so people can talk naturally to the session — human → AI by voice.
Sign language options (planned)
Universal access so deaf and hard-of-hearing users are not limited to text-only interfaces.
Same stack as soft-live
Voice is not a separate product silo — it carries the knowledge-grounded reply into another modality.
Who voice is for
Individuals who prefer listening
People who want a calmer way to receive a reply — especially in 粵語 — after writing freely.
Greater China localisation
Language and voice fit are part of access — not an afterthought bolted onto an English-only bot.
Future multimodal programmes
Organisation and accessibility pathways will reuse this speech stack as modes expand.
Explore the suite
Listen today inside soft-live. Speech-in comes after the knowledge bank is ready.