BMW i Ventures LLC

08/24/2026 | Press release | Distributed by Public on 08/24/2026 09:06

Voice Interactions That Feel Familiar: Why We Invested in Familiar Labs


Voice AI is growing rapidly, but most interactions still feel like using a walkie-talkie. Interrupt an AI mid-sentence and it keeps talking, finishes its turn, processes what you said, and only then responds. The conversation stalls, and the person on the other end quickly realizes they are talking to software.

The root cause is architectural: systems designed for sequential text exchanges have been retrofitted for real-time spoken conversation. The standard stack is a cascade of ASR, followed by an LLM, followed by TTS. By the time it registers an interruption, it is already too late to respond naturally. Overlapping speech and backchannels break the interaction entirely.

Nowhere is this more obvious than in outbound calls, where even a slight delay or missed interruption can make an agent sound scripted and prompt the person on the other end to hang up. Familiar Labs' own data reflects the stakes: with industry-standard speech models, more than 40% of calls end within the first 30 seconds.

Familiar Labs, formerly MetaVoice Labs, is tackling this problem with a full-duplex speech model purpose-built for production outbound voice agents. We're excited to back Siddharth Sharma, Vatsal Aggarwal, and their team as they build AI that genuinely feels like talking to another person.

Why Now

Three forces are coming together.

Technical Breakthrough. Familiar has built a full-duplex speech architecture around how people actually communicate. It listens while it speaks, allowing for continuous, two-way dialogue. The model can respond to interruptions, overlapping speech, and conversational cues in real time rather than waiting for one turn to end before processing the next. The team has done the applied research required to make this architecture work reliably in real-world production environments.

Market Timing. Outbound voice automation is moving into high-value enterprise workflows such as collections, lead qualification, and appointment scheduling. In these conversations, conversational quality directly affects outcomes. The stakes are especially high in outbound calls, where one mishandled interruption can cost the sale. A voice agent cannot simply follow a script; it needs to handle the unpredictability of a real conversation.

Competitive Gap. Enterprises still face a false choice. Large AI labs have powerful voice capabilities, but they are largely sealed inside closed assistants, making them difficult to customize, self-host, or integrate with external safety and compliance systems. Meanwhile, most speech and TTS incumbents were built for scripted or inbound interactions, not AI-led conversations where timing, interruptions, and overlapping speech determine the outcome.

Familiar is building for this gap: full-duplex speech models that converse more like people and can be deployed within a company's own stack, with the control, flexibility, and reliability required for production.

Built by a Rare, Directly Relevant Team

Our conviction starts with the founders.

Siddharth Sharma (CEO) was a founding engineer at Wayve, building petabyte-scale, real-time ML systems and world models for autonomous vehicles. In that context, timing and edge-case behavior aren't academic. They determine whether a system holds up in the real world. A 100ms latency spike or a missed edge case means the car fails. That experience translates directly: outbound voice agents face the same unforgiving constraints. There is no second take.

Vatsal Aggarwal (CTO) led generative voice work on Alexa and AWS Polly at Amazon. He holds multiple speech patents and publications and created the open-source MetaVoice-1B model. More importantly, he repeatedly took speech research from paper to production, shipping systems that operated reliably at planetary scale and met the demanding latency budgets of real-time voice.

Together they combine frontier systems experience (Wayve: real-time, edge-case robustness) with deep speech research and production discipline (Amazon: shipping at scale). They have translated that into a duplex stack built for real conversations, not proofs of concept. That combination is rare.

Our Conviction

We invested because Familiar Labs is addressing a foundational constraint in today's real-time voice stacks, and doing it with a team we believe can win on both research and production execution.

As Kasper Sage, Managing Partner at BMW i Ventures, put it: "Most voice systems still talk like software. Specifically over the phone. Familiar Labs talks like a person. Full duplex is the difference between a demo and a voice agent you can trust in high-stakes workflows, and this team has the technical depth to make this a reality at scale."

The defensibility is structural. Once companies deploy a voice agent that doesn't drop calls, that reliably handles interruptions, and that sounds human enough to build trust-switching costs become real. And the moat deepens as Familiar Labs accumulates real conversation data and iterates the model against production outcomes, not benchmark datasets.

What's Next

We're backing Familiar Labs as outbound voice automation becomes table-stakes for customer operations. The companies that ship first will own the category.

We're proud to partner with Siddharth, Vatsal, and the Familiar Labs team as they help enterprises win the conversations that matter most.

About Familiar Labs: Familiar Labs is a San Francisco-based conversational AI company. The team is building full-duplex speech models for production outbound voice agents. Learn more at [website].

About BMW i Ventures: BMW i Ventures is a corporate venture fund backed by BMW AG, investing in early-stage companies at the intersection of AI infrastructure, agentic systems, physical AI, and industrial automation.

BMW i Ventures LLC published this content on August 24, 2026, and is solely responsible for the information contained herein. Distributed via Public Technologies (PUBT), unedited and unaltered, on August 24, 2026 at 15:06 UTC. If you believe the information included in the content is inaccurate or outdated and requires editing or removal, please contact us at [email protected]