A generative audio startup run by people who came out of TikTok‘s AI and music divisions has raised $11 million to build models that write music and sound effects to match what happens on screen. San Francisco-based Sonilo secured the round with B Capital leading and Redpoint also taking part. The money will pay for the computing power needed to scale its models.
The company also plans to widen its footprint with creators and developer platforms and build out its music licensing business. Sonilo’s core technology, which it calls the Sound World Model, is a multimodal system that ties video comprehension to audio generation and timing. It doesn’t handle music and effects as standalone pieces.
Instead, it reads a scene’s movement and emotional shifts, then produces audio built for that particular footage. Users can upload a clip, type a description or combine the two to get a synced soundtrack ready for commercial use. You give us the footage, and Sonilo understands what’s happening frame by frame, including the pacing, emotion, and story, and generates music and sound effects that naturally fit the scene,” said Shawn Song, Sonilo’s CEO and co-founder.
Song, who trained at CMU, led multimodal AI technology work at TikTok, where he worked alongside Alex Yin, Sonilo’s CTO. Yin earned a doctorate in computer music at the University of York. COO Keli Li built and ran TikTok’s global music operation and has a background in rights deals and partnerships with the music business.
CMO Trista Taylor, who runs product and the company’s market launch, has helped develop consumer AI apps that made a16z’s Top 50 rankings. The platform is aimed at creators working in short-form video and short dramas, at developer platforms assembling AI video tools, and at enterprise production teams. It also offers AI dubbing with lip-sync, which lets short-drama outfits and content studios localize video into other languages without editing by hand in post.
AI creator aggregator TapNow and ComfyUI have signed on as early distribution partners. Sonilo stands out because its multimodal technology gives creators two ways to generate music and sound effects: describe what they want in text, or provide a video and let the model create audio that fits the scene,” said Daisy Cai, general partner at B Capital. On the rights side, is an early licensing partner.
Extract — continue reading at the source.