ithome News score 11

科大讯飞发布全新语音识别大模型 Spark-ASR-2.0,明日上线讯飞输入法

AI Digest

语音识别大模型Spark-ASR-2.0中英文混合识别复杂声学场景讯飞输入法

科大讯飞发布语音识别大模型Spark-ASR-2.0,通过技术创新提升识别准确率与文本流畅性,明日上线讯飞输入法并拓展多场景应用。

Kuaishou releases Spark-ASR-2.0, enhancing speech recognition accuracy and text fluency with new tech, to be integrated into input methods and multiple products.

Key points

  • 采用非自回归与LLM增强自回归协同等技术提升识别效果 Enhanced recognition via non-autoregressive and LLM-enhanced architectures
  • 显著改善中英文混合、方言及复杂声学场景识别能力 Significant improvements in mixed language, dialect, and noisy environments
  • 推理成本仅增加10%,兼顾性能与效率 10% higher efficiency compared to previous model
  • 将应用于讯飞输入法、AI眼镜等多终端产品 Deployment across multiple devices including input methods
  • 在方言识别等维度超越业界最优水平 Outperforms industry benchmarks in dialect recognition

Takeaway: Spark-ASR-2.0代表语音识别技术向自然语言理解迈进一步。 / Spark-ASR-2.0 marks progress toward natural language understanding in speech recognition.

Why it matters 展示语音识别技术在复杂场景下的突破,对智能交互产品有实际指导意义。

View original ↗ Back to hot list

This page is an aggregated digest from ithome; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.