Significant challenges overcome in speech-to-text include improving accuracy in the presence of background noise, handling diverse accents and speech patterns, and moving from isolated word recognition to continuous, contextual understanding through advancements like Hidden Markov Models and deep learning.