Skip to main content
Scicom AI

WideCodec: 44 kHz speech tokenization from a single codebook at 50 tokens per second

Benchmarked across more than 400 languages, and ready to post-train a language model into a speech model without touching the architecture.

We're excited to share a breakthrough in speech tokenization — designed to deliver maximum compression with minimal compromise.

Our latest tokenizer achieves:

  • 44 kHz high-fidelity audio
  • Pareto-best quality vs. bitrate
  • Benchmarked across 400+ languages
  • Single codebook with only 50 tokens per second (TPS)
  • LLM-ready — making it significantly easier to post-train language models into speech models without modifying the model architecture.

The benchmark results speak for themselves:

  • State-of-the-art naturalness at an exceptionally low bitrate
  • Strong CER recovery performance across multilingual evaluation
  • Smaller token streams for more efficient training and inference

This is a step toward making high-quality speech AI more efficient, scalable, and accessible — from research to production.

The model is now available on Hugging Face: Scicom-intl/WideCodec.

We look forward to seeing what the community builds with it and will be sharing more technical details soon.


First published on LinkedIn, 5 August 2026. Reproduced as published, with the list markers set as text rather than emoji.

More from blog

← Back to the Newsroom