ZICQ
中 Log in / Sign up
ZICQ Info Agentic #Speech Recognition #Open-Source Model #Edge Computing #IoT #Lokutor

Lokutor Releases Oído: Open-Source Speech Recognition Model Outperforming Whisper-tiny

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:The Lokutor team has released Oído, an open-source speech recognition model based on the NVIDIA Conformer-CTC Small architecture (13 million parameters, int8 quantization), capable of running on an ESP32-S3 microcontroller without a GPU or NPU. Oído achieves a Word Error Rate (WER) of 3.7/8.2 on the LibriSpeech dataset, outperforming Whisper-tiny's 6.3/15.9. In real-world noisy environments (e.g., cars, kitchens, cafeterias), Oído's mean WER is 8.4, compared to Whisper-tiny's 12.1. This model de


Technical Highlights

  • Model Architecture and Performance: Oído is based on the NVIDIA Conformer-CTC Small architecture with 13 million parameters and int8 quantization. This allows the model to maintain high performance while significantly reducing computational resource requirements.
  • Hardware Requirements: Oído can run on an ESP32-S3 microcontroller without a GPU or NPU, requiring only 8 MB PSRAM. This demonstrates its strong adaptability in resource-constrained environments.
  • Speech Recognition Performance: On the LibriSpeech dataset, Oído achieves a Word Error Rate (WER) of 3.7/8.2, outperforming Whisper-tiny's 6.3/15.9. In real-world noisy environments (e.g., cars, kitchens, cafeterias), Oído's mean WER is 8.4, compared to Whisper-tiny's 12.1.

Industry Impact and Developer Recommendations

  • IoT and Edge Computing: The release of Oído provides an efficient speech recognition solution for IoT devices (e.g., smart home devices, wearable devices) and edge computing applications. Developers can leverage this model to implement voice interaction features in resource-constrained environments.
  • Open-Source Advantage: As an open-source project, Oído allows developers to customize and optimize the model according to specific needs, further expanding its application scenarios.
  • Real-Time Applications: Oído's low latency and high accuracy make it suitable for real-time speech recognition tasks, such as voice assistants and voice command controls.

Conclusion

The Oído model released by the Lokutor team demonstrates strong performance in the field of speech recognition, particularly on resource-constrained devices. Its open-source nature makes it a powerful tool for developers to build efficient speech recognition applications, driving the development of IoT and edge computing.


Source: Reddit r/LocalLLaMA (2026-09-30)

— END —

Tags: #Speech Recognition #Open-Source Model #Edge Computing #IoT #Lokutor

Community Comments

Loading live comments and annotations…