WideCodec: 44 kHz speech tokenization from a single codebook at 50 tokens per second
Benchmarked across more than 400 languages, and ready to post-train a language model into a speech model without touching the architecture.
We're excited to share a breakthrough in speech tokenization — designed to deliver maximum compression with minimal compromise.
Our latest tokenizer achieves:
- 44 kHz high-fidelity audio
- Pareto-best quality vs. bitrate
- Benchmarked across 400+ languages
- Single codebook with only 50 tokens per second (TPS)
- LLM-ready — making it significantly easier to post-train language models into speech models without modifying the model architecture.
The benchmark results speak for themselves:
- State-of-the-art naturalness at an exceptionally low bitrate
- Strong CER recovery performance across multilingual evaluation
- Smaller token streams for more efficient training and inference
This is a step toward making high-quality speech AI more efficient, scalable, and accessible — from research to production.
The model is now available on Hugging Face: Scicom-intl/WideCodec.
We look forward to seeing what the community builds with it and will be sharing more technical details soon.
First published on LinkedIn, 5 August 2026. Reproduced as published, with the list markers set as text rather than emoji.