> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vexa.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Speaker separation on any audio we take in

> Separate speakers in mixed meeting audio inside your own infrastructure — diarization on any audio Vexa takes in, including room recordings and conference bridges.

**Can Vexa tell speakers apart in a single mixed audio stream?**

**Today:** Only where each person joins on their own device. On Meet and Zoom the name is bound at capture; where the audio arrives already mixed, there is no per-speaker split.

**Sponsorable:** yes — Speaker separation on the mixed lane, so mixed audio stops being one undifferentiated speaker.

***

*Tell speakers apart in audio that arrives as one mixed stream, inside your own infrastructure.*

|                 |                                              |
| --------------- | -------------------------------------------- |
| **Status**      | **sponsorable**                              |
| **Size**        | M                                            |
| **North star**  | 2 · Dependable capture                       |
| **Measured by** | served-transcript share and named-word share |

## What it does

* Separates a mixed audio stream into distinct speakers.
* Applies to any audio the system takes in, including a room recording and a bridge.
* Produces Speaker 1..n — separation, not identity.
* Runs inside your tenancy; the audio does not leave it.

## What exists today

On Meet and Zoom, each person joining from their own device arrives on their own audio track, so the name is bound at the moment of capture. Where the stream arrives mixed, there is no per-speaker split today.

## What this adds

Speaker separation on the mixed lane, so mixed audio stops being one undifferentiated speaker.

<Note>
  Speaker separation **separates** speakers. Names come from a roster you supply. Nobody enrols a voice.
</Note>

## How sponsorship works

You fund the item. It moves to the front of the queue. You co-write the acceptance criteria, so "done" means done on your meetings. Your name goes on the release note. Sponsorship buys **the order of the queue** — never exclusivity, never a private build, and never a date the work has not earned. The fee basis is per item or per stage, agreed in writing before work begins.

## Two doors

<Columns cols={2}>
  <Card title="Sponsor this" icon="handshake" color="#0D9373" href="https://cal.com/dmitrygrankin/web?utm_source=docs&utm_medium=roadmap&utm_campaign=sponsor&utm_content=speaker-separation" horizontal arrow cta="Book fifteen minutes">
    Fifteen minutes. Bring the constraint you are under.
  </Card>

  <Card title="Watch this" icon="envelope" color="#0D9373" href="mailto:dmitry@vexa.ai?subject=Roadmap%20update%3A%20Speaker%20separation%20on%20any%20audio%20we%20take%20in&body=watch" horizontal arrow cta="Email me">
    One email when it ships. The mail is already written — send it as it is.
  </Card>
</Columns>

## Questions

<AccordionGroup>
  <Accordion title="Can Vexa tell speakers apart in a single mixed audio stream?">
    Only where each person joins on their own device. On Meet and Zoom the name is bound at capture; where the audio arrives already mixed, there is no per-speaker split.
  </Accordion>

  <Accordion title="What would sponsoring this add?">
    Speaker separation on the mixed lane, so mixed audio stops being one undifferentiated speaker.
  </Accordion>

  <Accordion title="Does this require voice enrolment?">
    No. Speaker separation separates speakers; names come from a roster you supply. Nobody enrols a voice.
  </Accordion>
</AccordionGroup>

***

*Ships upstream Apache-2.0 when sponsored.*
