Microsoft’s Speech Recognition Tech Achieves Human Parity--Sort Of

Microsoft researchers from the Speech & Dialog research group include, from back left, Wayne Xiong, Geoffrey Zweig, Xuedong Huang, Dong Yu, Frank Seide, Mike Seltzer, Jasha Droppo and Andreas Stolcke. (Photo by Dan DeLong)

Microsoft announced that its speech recognition technology has achieved a word error rate (WER) of only 5.9%, which the company said was similar to what human transcribers are able to achieve.

Latest Videos FromTom's Hardware
TOPICS
Contributor

Lucian Armasu is a Contributing Writer for Tom's Hardware US. He covers software news and the issues surrounding privacy and security.

  • jaber2
    Just want to know how long until universal translator
    Reply
  • kittle
    I wonder how that thing will work when plugged into siri on a noisy subway?
    Reply
  • Icepilot
    "... we're living in a time when machines are beginning to truly understand humans and the world around us."
    One word too far, truly.
    Reply
  • stuartturner34
    “We’ve reached human parity,” said Xuedong Huang, the company’s chief speech scientist. “This is an historic achievement.”

    A typo in a quote of a scientist talking about word error ratings. So meta.
    Reply
  • alextheblue
    18750343 said:
    I wonder how that thing will work when plugged into siri on a noisy subway?

    That depends on your audio hardware/software more than anything. For example on a PC, it would depend on the type and quality of the microphone / mic array, the sound card, audio drivers, recording software, etc. There's a couple of places where there's opportunities for noise cancellation, depending on the gear and ware used. The result gets handed to this translation software, garbage in garbage out - you have to feed it good audio for it to do it's job. The situation isn't all that different for a smartphone. Unfortunately the iPhone probably wouldn't do the best job compared to a smartphone with a HAAC twin membrane quad-mic array.
    Reply
  • jackt
    using super computers or normal pc ?
    Reply
  • Kafantaris
    Microsoft's AI driven voice recognition has left all rivals in the dust. Great work by Dr. Xuedong Huang's speech team.
    Reply
  • bit_user
    Good job digging into the error rates, Lucian.

    18757879 said:
    using super computers or normal pc ?
    This is the question I had. How much compute does it use? It's not a small detail whether this requires a long time on a big GPU, or whether it can run on a smartphone in realtime. If too much compute is required, then this won't be deployed in most real-world uses cases for years.

    BTW, humans are still way more energy efficient.
    Reply
  • bit_user
    18751434 said:
    “We’ve reached human parity,” said Xuedong Huang, the company’s chief speech scientist. “This is an historic achievement.”

    A typo in a quote of a scientist talking about word error ratings. So meta.
    It would be, but where's the error?
    Reply