Visual Speech Recognition
Infers speech content from the temporal motion of lips and surrounding facial structures.
Technical Context
Infers speech content from the temporal motion of lips and surrounding facial structures.
This concept belongs to the speech and audio pipeline from acoustic representation to decision-making.