Visual Speech Recognition and Lip Reading

Visual Speech Recognition and Lip Reading

A multimodal view of lip reading: temporal visual speech representation, viseme ambiguity and fusion with the acoustic channel.

Lip reading, or visual speech recognition, infers speech content from temporal motion of the mouth and surrounding facial structures. It is primarily a temporal and multimodal signal-processing problem rather than a static image-classification task.

Visual speech signal

Lip, jaw and cheek motion correlates with speech production. Different phonemes may nevertheless produce similar mouth shapes, creating viseme-level ambiguity. Camera angle, frame rate, resolution, face coverings and illumination all affect the available information.

Visual speech recognition therefore depends on stable face and mouth-region tracking.

Temporal modelling

3D convolutions, recurrent networks, temporal convolutions and Transformers process motion across frame sequences. CTC or attention-based decoders can map visual features to character or token sequences.

Audio-visual fusion

Audio-visual speech recognition uses the visual channel as additional information when acoustic speech is degraded. Fusion may occur at feature, representation or decision level.

Accurate synchronisation is essential because camera and microphone clocks, buffering and transport delay can introduce temporal offsets between the two modalities.

QR code for this page