Skip to main content
Can Vexa tell speakers apart in a single mixed audio stream? Today: Only where each person joins on their own device. On Meet and Zoom the name is bound at capture; where the audio arrives already mixed, there is no per-speaker split. Sponsorable: yes — Speaker separation on the mixed lane, so mixed audio stops being one undifferentiated speaker.
Tell speakers apart in audio that arrives as one mixed stream, inside your own infrastructure.

What it does

  • Separates a mixed audio stream into distinct speakers.
  • Applies to any audio the system takes in, including a room recording and a bridge.
  • Produces Speaker 1..n — separation, not identity.
  • Runs inside your tenancy; the audio does not leave it.

What exists today

On Meet and Zoom, each person joining from their own device arrives on their own audio track, so the name is bound at the moment of capture. Where the stream arrives mixed, there is no per-speaker split today.

What this adds

Speaker separation on the mixed lane, so mixed audio stops being one undifferentiated speaker.
Speaker separation separates speakers. Names come from a roster you supply. Nobody enrols a voice.

How sponsorship works

You fund the item. It moves to the front of the queue. You co-write the acceptance criteria, so “done” means done on your meetings. Your name goes on the release note. Sponsorship buys the order of the queue — never exclusivity, never a private build, and never a date the work has not earned. The fee basis is per item or per stage, agreed in writing before work begins.

Two doors

Sponsor this

Fifteen minutes. Bring the constraint you are under.

Watch this

One email when it ships. The mail is already written — send it as it is.

Questions

Only where each person joins on their own device. On Meet and Zoom the name is bound at capture; where the audio arrives already mixed, there is no per-speaker split.
Speaker separation on the mixed lane, so mixed audio stops being one undifferentiated speaker.
No. Speaker separation separates speakers; names come from a roster you supply. Nobody enrols a voice.

Ships upstream Apache-2.0 when sponsored.