Yeah, you are right we should be filtering this sort of stuff out. The algorithms are robust in that they ignore words not in the system's vocabulary (rather than, say, crash) but we did not trap the case in which none of the "words" are familiar.
It is understandable; different people may see different performance depending on what speech service is used. In addition, please keep in mind this is just a beta service; we have only been out there less than a week and know we still have some tuning to do. We also recognize the need for customization and personalization abilities; keep watching our website for continuing developments.
Beta is certainly beta but if you find problems we will try to fix them as quickly as we can. Real users of technologies tend to find issues with the technology much faster than the actual developers...
Sorry, I missed the question at the bottom. We were very proud of ViaVoice at the time but to make an obvious point, the technology has moved on a lot over the past ten or so years...
We really appreciate the comments and will try to fix the problems. To be honest, sometimes developers can't see even the most obvious flaws in documentation. If you can highlight even one incomprehensible point it would help a lot to accelerate the revision process.
We did some work on applying NNs to prosody prediction; see Fernandez, Raul, et al. "Prosody contour prediction with long short-term memory, bi-directional, deep recurrent neural networks." Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH). 2014.
SSML is a speech synthesis markup language that has some degree of popularity in the field. The specific section on markup for emphasis is http://www.w3.org/TR/speech-synthesis11/#S3.2
One thing a machine learning system can do that any one human cannot do is ingest lots of data. For example, for some tasks in which I have tried to compare human vs machine speech recognition performance the machine actually does better because the machine may - for example - know a singer's name that an individual human may not recognize.
We want feedback on all our services. If you are speaking about using data to update the service, I know the speech services do not yet have this capability.
Yeah it was. At the risk of waxing about old history, when we demoed the first large vocabulary speech recognition system back in 1984 it ran on a bunch of IBM mainframes. Within two years it was running on a PC with some special purpose cards. Today much more powerful recognizers run locally on smartphones. We have always found we can shrink something down once we solve the basic problem, and it is important not to let computational limitations prevent you from seeing the best solution.
Well, I don't think the technology is that bad :-). But I agree with you. We have to solve the problem of poorer quality audio input, and the sooner the better! But there are also many scenarios where good audio input is feasible and would like feedback on those sorts of application ideas too.
As a speech technologist, I am amazed and proud about how far long the technology has progressed, especially over the last few years. Even my wife now uses speech input on mobile devices (and may finally think I may be doing something productive...). With that said, speech input is still a surprisingly finicky technology and different people will see different beahviors across systems from different providers.
We know we have strong core speech technology based on various comparisons we have done in the context of competitive evaluations done in conjunction with various government funded speech programs. However, our service is still very new. We could have waited for months to tune it, but our primary goal here is to solicit feedback from the community for how to make our services easier to use, especially in the context of our other platform services. We don't want to wait till the design is so mature that it is impossible to change - so any and all feedback is very welcome!