Kara Bembridge
May 28, 2024
Reading time:
The world of AI has made leaps and bounds from what it once was, but there are still some adjustments required for the optimal outcome. In the realm of conversational AI, VoxAI had already developed a platform to capture customer orders. The response time and oratory abilities needed improvement and this is where Collabora stepped in with WhisperLive.
WhisperLive, a real-time transcription service powered by OpenAI's Whisper model departed from traditional speech recognition methods by incorporating voice activity detection (VAD). VAD identifies speech presence, allowing for selective transmission of audio data to enhance transcription accuracy while optimizing data handling.
Simultaneously, Collabora employed a finely tuned Mistral model for the NLP component. Renowned for its efficiency and versatility, Mistral is six times faster and equally or more effective than the Llama 2 70B model across benchmarks. This model supports multiple languages and possesses inherent coding capabilities.
As we look to the WhisperLive's promising possibilities, our Machine Learning Lead, Marcus Edel, puts it best:
"The future of customer interaction lies in the harmonious fusion of sophisticated AI and powerful communication technologies. As we continue our mission and build fully in the open, WhisperLive, and now the award-nominated WhisperFusion, are poised to make an impact in the communication technology landscape."
To learn more about how this project came to life, take a look at out our case study.
If you're eager to implement your own transcription service, please get in touch! Our machine learning team is ready to assist you with your AI needs.
26/05/2026
New upstream BlueZ documentation helps simplify Bluetooth qualification for Linux-based products by mapping supported profiles, test requirements,…
14/05/2026
See how Tyr moves beyond MCU firmware boot to build the group, queue, VM, submission, and completion paths needed to run real Vulkan workloads…
07/05/2026
A complete breakdown of Mesa’s NIR compiler detailing how it optimizes shader memory access with SSA promotion, deref analysis, copy propagation,…
05/05/2026
Collabora brought Bluetooth Auracast broadcasting to MediaTek Genio 700 for Embedded World 2026. Here's the complete, fully Open Source…
22/04/2026
Using our XR expertise, Collabora created a standalone XR experience for our 1% for the Planet partner, SOMAR, to showcase the direct impact…
17/04/2026
BitNet-style ternary brings LLM inference to ExecuTorch via its Vulkan backend, enabling much smaller, bandwidth-efficient models with portable…