Finding Authentic Video Pronunciations...
Scanning thousands of native English clips with synchronized subtitles and exact timestamps.
right. So how can we take all of these modalities uh into account? And this is really just the idea of you know given images and text can you actually find a
concept that relates these two different sources of uh sources of data. Uh and there are a few challenges and that's why you know models like genative models
uh sometimes proistic models could be useful. In general one of the biggest challenge we've seen is that typically when you're working with images and text
Listen to native speakers pronounce “genative” in real conversational contexts with synchronized timestamps and subtitles.