Description
see all eventsBay Area Voice AI Night

Bay Area Voice AI Night
About the Event
Voice AI Night is a Bay Area gathering for founders, builders, product leaders, & investors interested in how voice, multimodal AI, & real-time agents are changing the way people interact with products.
On July 30, StepFun & Seamate bring together people interested in how voice, multimodal AI, & real-time agents are changing the way people interact with products.
Join us for an evening of model insights, product discussion, outdoor dinner, & relaxed conversations with people building & exploring the next wave of AI products.
What's Happening
Keynote : Voice models, from the model side
StepFun is a frontier AI lab building language, multimodal, & audio models. StepFun will share what it is seeing from the model side: where speech models are moving, what StepAudio 2.5 is designed for, & what better realtime voice intelligence could unlock for builders.
Founder panel: What It Takes to Ship Voice AI
We'll bring together founders building in multimodal AI & Voice AI to discuss what it takes to turn new model capabilities into real products. The conversation will focus on product lessons, application-layer opportunities, & what changes when AI can listen, speak, & interact in real time.
The panel will be primarily in Mandarin , & English-speaking attendees are very welcome.
Outdoor buffet dinner + unhurried conversations
After the main session, the conversation continues outdoors over seafood, lamb & beef, with two full hours to meet & exchange notes with other founders & builders.
Who Should Come
People building, investing in, or seriously exploring the next wave of AI products.
You do not need to be deep in Voice AI already. We especially welcome application-layer builders & people curious about where voice & multimodal interaction are going next.
Registration is approval-based so we can keep the room relevant & useful.
Keynote
Yang Yang, Speech Model Researcher, StepFun
Next-Gen Speech Models: End-to-End Voice Intelligence
About StepAudio 2.5
StepAudio 2.5 is StepFun's audio model family for speech understanding, speech synthesis & speech recognition.
It supports speech understanding with vocal cues, instruction-following speech generation, zero-shot voice cloning, & low-latency Chinese & English speech recognition for live captions, voice input, & meeting transcription.
Explore the official StepAudio 2.5 documentation
Agenda
5:00 PM Check-in & Welcome Drinks
5:20 PM Opening
5:30 PM Keynote + Founder Panel
7:00 PM Dinner Mixer, continuing from the main session
Language
The keynote & panel will be primarily in Mandarin. English-speaking founders & builders are more than welcome to join. Networking & dinner will naturally be bilingual.
About the Hosts
Seamate () is a community for founders, builders, & investors exploring AI, startups, & global markets. With 15,000+ members across 10+ countries, we connect entrepreneurs across the U.S., China, & Southeast Asia & beyond through events, conversations, & collaborations.
StepFun is a leading foundation model company committed to the long-term pursuit of AGI, backed by a top-tier technical team. Its Step model family spans language, multimodal, & reasoning capabilities.