ByteDance has launched SeedRealtime, a native audio-video full-duplex model that can continuously process audio, video, and text streams while listening and responding in real time1. The model has been rolled out in the Doubao app, moving the technology from a research demonstration toward a consumer-facing product.
SeedRealtime is designed to support interactions in which it can watch, listen, and speak simultaneously. The headline from a separate report describes the model as enabling Doubao to allow "watching, listening, and speaking without lag"2.
The model is also positioned for automotive integration, according to the title of a report on aibase.com, which describes SeedRealtime as a "full-duplex audio-visual large model for car integration".
The launch gives ByteDance another model for real-time multimodal interaction alongside its existing text and image-generation systems.
ANALYSIS The deployment of SeedRealtime into Doubao — ByteDance's consumer AI app — marks a shift from research-stage multimodal capability to a shipped product, placing ByteDance among the companies actively integrating real-time audio-video processing into consumer applications. The automotive integration angle, referenced in the aibase.com headline, suggests ByteDance is targeting in-vehicle AI assistants as a deployment surface for the model alongside its mobile app.