Sesame Launches Maya and Miles as Full AI Agents With Their Own Computers
Sesame's four voice agents move from preview to general availability on iOS and Android, each with its own computer, memory, and tool access.
- Sesame moved Maya, Miles, Simone and Charlie from preview to GA on iOS and Android.
- Each agent now has its own computer, persistent memory, and isolated permission scope.
- Agents can connect to Google apps and user-supplied custom MCP servers for tool access.
- Team dogfooded by orchestrating Claude Code, Devin and Codex sessions entirely by voice.
- Release is positioned as the voice OS foundation for Sesame's planned 2027 AI glasses.
- Developer tooling for building third-party Sesame apps is planned for later this year.
Sesame has moved its voice assistants from preview to general availability on iOS and Android. The October release combines new models with a consequential architectural choice: each agent receives a separate computer environment for research, tool use, and long-running tasks.
The company became known for the natural speech in its Maya and Miles demos. Its October release extends that work into hands-free task execution and establishes the software foundation for AI glasses planned for 2027.
Sesame leaves preview
General availability brings four agents, web research, connected apps, scheduled workflows, persistent memory, and background task execution to both mobile platforms. Sesame describes the production stack as a new generation of models, although the available release details do not specify their architecture, context limits, latency, pricing, or benchmark results.
Sesame rolled out the product in several stages:
- Early 2025: Maya and Miles entered research preview.
- May 2026: The native iOS preview opened in 39 countries after a closed beta.
- Later rollout: Sesame released an Android preview.
- October: Both apps reached general availability with Maya, Miles, Simone, and Charlie.
Four voices, separate context
Each agent has an individual voice, personality, viewpoint, memory, and permission scope. Information shared with Maya remains outside Charlie's memory, allowing users to separate work, personal, and project-specific conversations by persona.
- Maya: warm and creative
- Miles: laid-back and sharp
- Simone: curious and intellectual
- Charlie: witty and warm
Sesame develops these personas through scripts, character guides, and recorded scenes created by writers, producers, actors, and researchers. That process treats speech and character as one system, with tone, timing, and conversational behavior shaped alongside the underlying voice.
A computer behind every voice
Sesame says each agent now has its own computer, meaning a separate execution environment that can continue working after a spoken request. A user can ask an agent to research a hobby, plan a trip, or explore a project, then receive an update when the work finishes.
The agents can browse the web, call tools, coordinate other agents, and connect to services, beginning with several Google apps. Scheduled workflows and persistent memory allow tasks to span multiple sessions instead of ending with a single conversation.
Users can also connect custom MCP servers. Model Context Protocol provides a standardized way for an AI agent to access tools and data sources such as files, databases, internal APIs, and third-party services. An MCP server can expose actions with real consequences, including sending messages or changing records, so production deployments require narrow permissions, clear confirmation steps, and audit logs. Separate agent memories do not automatically restrict what a connected server can access.
Voice becomes an agent console
Sesame used Maya, Miles, Simone, and Charlie while developing the release, including to direct coding agents such as Devin, Claude Code, and Codex by voice. The company says employees managed sessions while walking, commuting, and working away from a laptop. Developer tooling for building Sesame apps is planned for later this year.
Sesame shared several examples from internal testing:
- Managing Claude Code and Devin sessions during a walk
- Finding and booking a New York restaurant for a team of 15, then sending the confirmation by text
- Ordering books from an existing task list
- Planning a three-day trip and publishing it as a shareable webpage
Voice handoff and background execution let users delegate these familiar tasks without supervising each intermediate step. The agent interprets the request, uses its assigned tools, handles the workflow in its computer environment, and reports the result.
Mobile rehearses the glasses
Sesame intends to use the assistant as the voice-first operating layer for its planned glasses. Co-founder Brendan Iribe previously led Oculus, and the mobile apps establish the interaction model before Sesame ships dedicated hardware.
A wearable interface also explains the emphasis on spoken delegation, asynchronous work, and concise status updates. Those capabilities reduce the need for a screen while the agent handles web research, service integrations, and other multi-step tasks remotely.
Reliability decides usefulness
Earlier Maya and Miles previews focused on fluid conversation and lacked web access or current information. During extended discussions, they could produce plausible but inaccurate answers to factual questions. Web research, connected apps, and tool use address that limitation by giving agents access to external sources and executable actions.
Production use will depend on how reliably the system handles several operational concerns:
- Task completion: whether multi-step workflows finish without silent failures
- Source quality: whether research results include accurate, traceable information
- Authorization: whether sensitive actions require appropriate confirmation
- Recovery: whether interrupted tasks can retry or resume safely
- Observability: whether users can inspect actions, tool calls, and errors
- Isolation: whether persona-specific memory and permissions remain separated across integrations
The release materials demonstrate the intended workflow, while broader use will show whether the agents can complete purchases, bookings, messages, and coding tasks consistently enough for routine delegation.
Availability and developer access
The iOS and Android apps are generally available now. Sesame is also hosting an AMA thread with members of its engineering and design teams.
Third-party developer tooling is scheduled for later this year. Sesame's open CSM-1B model remains available under the Apache 2.0 license for developers who want to self-host the speech layer. Reproducing the full application also requires the orchestration, memory, permissions, integrations, and computer-use systems surrounding that model.
The launch turns Sesame's natural voice technology into a task-running product with isolated personas and persistent execution environments. Its practical value will depend on reliable handoffs, controlled tool access, and clear reporting when an agent completes or cannot complete a request.