Sesame Launches iOS Voice AI app With 4 Distinct Agents and Sub-300ms Responses
Sesame's iOS app brings four lifelike voice agents to 39 countries, powered by a speech model that searches the web while it talks
- Public iOS launch: Sesame released its Personal Agents app in 39 countries, free during preview with a possible short waitlist.
- Four characters: Maya, Miles, Simone, and Charlie each have distinct personalities, voices, and per-character persistent memory across voice and text.
- Thinks while it talks: Agents run parallel web searches mid-sentence, weaving results into responses in real time using Sesame's custom Conversational Speech Model (CSM).
- Well-funded team: Founded by ex-Oculus CEO Brendan Iribe; backed by $250M Series B from Sequoia, Spark, and others after 1M+ users tried the Research Preview.
- Smart glasses roadmap: The iOS app is a habit-building precursor to intelligent eyewear planned for 2027, with full agentic task execution coming before that.
- Android preview coming: No date confirmed yet; rolling out to more countries progressively.
Sesame, the conversational AI startup co-founded by former Oculus CEO Brendan Iribe and Ankit Kumar, has launched a public iOS preview of its personal agents app. The app puts four named AI characters , Maya, Miles, Simone, and Charlie , in your pocket, each with a distinct voice, personality, and persistent memory. The pitch is simple but technically ambitious: voice-first AI that actually feels like a conversation, not a command prompt.
From a viral demo to a real app
Sesame's story started with a splash. The startup first emerged from stealth offering two demos of its technology , AI voices named Maya and Miles , which were accessed by more than a million people within the first few weeks. That wave of interest helped Sesame close a serious funding round: the company raised a $250 million Series B , with investors including Sequoia, Spark, and other undisclosed backers.
The Research Preview that followed was invite-only and web-based. This iOS launch is the first time the general public can download and use Sesame without waiting for a special invite , though access is free during the preview phase, with potential waitlisting to ensure quality.
The hard problem: thinking while talking
Most voice AI today works in a serial pipeline: listen, think, then speak. That creates an awkward pause that breaks the illusion of real conversation. Sesame's core technical bet is that you have to do all three simultaneously.
There's an inherent tension between replying quickly and taking the time to compose thoughtful responses. A slower response is usually more correct, but it can also feel unnatural if it takes too long. Sesame's answer is a custom architecture called the Conversational Speech Model (CSM). Unlike traditional text-to-speech pipelines that first generate text and then synthesize audio, CSM is end-to-end and multimodal , it processes text and audio context together in a single model, which allows the AI to "think" as it speaks, producing not just words but also the subtle vocal behaviors that convey meaning and emotion.
On top of that voice model, Sesame built a retrieval layer that keeps the conversation grounded in real-world information. Sesame agents can run multiple parallel searches while speaking and seamlessly weave relevant results into their response as they stream in, pivoting mid-sentence if necessary. This matches how humans continue to think as we speak, allowing the agents to carefully balance latency and accuracy , and tap into slower, smarter agentic loops without awkward interruptions.
The underlying CSM architecture is worth understanding at a high level:
- Two-network design: The model taps transformer architecture similar to large language models, adapted for speech generation, and is composed of two neural networks , a "backbone" master model and a decoder.
- Scale: The largest CSM configuration has an 8.3 billion parameter backbone paired with a roughly 300 million parameter decoder.
- Audio tokens: The system represents audio using semantic tokens, which capture linguistic content and high-level speech traits, and acoustic tokens, which capture detailed voice characteristics like timbre, pitch, and timing.
- Latency target: The model achieves sub-300ms first-byte latency on streaming , the rough threshold at which voice responses start to feel natural.
What's actually in the app
The app offers four distinct AI agents called Maya, Miles, Simone, and Charlie, each of which have their own distinct voice, personality, point of view, and memory. Each has a distinct personality and memory that grows over time through conversation, the way you'd actually get to know someone , and that memory is continuous across voice and text, so nothing gets lost.
The feature set is designed around the idea that voice AI needs to work in the real world, not just at a desk:
- Search cards: Visual image results that surface during a conversation to help you understand new concepts
- Notes: Capture takeaways mid-conversation without breaking the flow
- Texting mode: Switch to text input when speaking out loud isn't an option
- Deep Dives: Trigger more thorough research on a topic within the same session
- Incognito Mode: Keeps conversations out of memory and off Sesame servers , while still allowing the agent to draw on prior context within the session
The whole experience is designed to be voice-first but not voice-only. As Sesame's Head of Product Engineering Raven Jiang put it in the launch blog post, the information displayed on screen enriches the agent's responses , but if you go display-free, everything is still available in the app later.
A crowded but still-forming market
Sesame is entering a competitive space. The active voice-agent market places Sesame alongside ElevenLabs, OpenAI Realtime, Hume EVI 4, Vapi, and Deepgram , and competition in that group is already shifting from novelty demos toward systems that can answer quickly, keep context, and support longer back-and-forth use. What differentiates Sesame's approach is the combination of persistent per-character memory, parallel search during speech, and a voice model trained specifically for conversational expressiveness rather than just clean pronunciation.
The iOS launch is also a deliberate habit-formation play. The rollout tests whether Sesame's voice model can build daily habits before its planned 2027 eyewear push. The app is the training ground , both for users who need to learn to interact with voice AI naturally, and for Sesame's models, which improve with real-world conversational data at scale.
The bigger picture: smart glasses by 2027
The iOS app is explicitly a stepping stone. The app is only the first step toward Sesame's bigger plans involving intelligent eyewear, which the team expects to launch in 2027. Before that, the agents will also learn to do more than just think with you , Sesame hints they'll later be able to take action on your behalf, hence why they're called "agents" in the first place.
Working with agentic tools today requires being able to prompt for what you need and have a specific idea of what you want to happen. A conversational agent that you could talk to naturally could help you take the next steps, without you having to perfect the command you're giving it. That's the unlock that makes voice-first agents genuinely different from voice-controlled chatbots.
For now, Sesame has launched its iPhone voice AI app in 39 countries with four agents, testing whether conversational AI can become a daily mobile habit on iOS. An Android preview is also in the works. The full experience is free during the preview period, though a short waitlist may apply at sign-up. You can download it on the App Store now.