Speech recordings can constitute important evidence in digital-forensics investigations.
The project evaluated research and software-development processes for implementing a speaker-recognition system and improving recognition accuracy through speech enhancement. Recordings containing environmental sounds and noise are especially important in forensic investigations. As computer technology and smartphone use became widespread, the nature of offenses involving voice communication also changed, increasing the need to evaluate recorded audio as evidence when establishing material facts.
The study aimed to separate recordings containing environmental sound and noise from non-speech components, improve intelligibility and provide high-accuracy speaker recognition. A substantial portion of the audio-enhancement software used in forensic processes could not perform this operation in real time or did not provide sufficiently accurate results. The intended outcome was clearer recordings for forensic analysis.
Speaker recognition and audio enhancement can be used throughout digital-forensics processes. In some cases, audio may be the only available evidence. Speech- and speaker-recognition applications can contribute to identifying offenders. Analysis techniques use characteristic properties of speech recordings, and these techniques may assist investigations of terrorism, homicide, kidnapping, threats, extortion, sexual assault, organized crime and telephone harassment.
The study assumed a configuration and artificial-neural-network model suitable for forensic use. It anticipated that access to real recordings from forensic cases would support more accurate speech enhancement and speaker recognition.
Work on audio processing requires knowledge from multiple engineering disciplines. It also requires relevant understanding of phonetics, speech physiology and associated health sciences. Limited detailed resources and legal requirements governing the processing of voice data as personal data made the work more difficult. Publicly available speech and audio datasets were considerably more limited than datasets in some other fields.
Speech enhancement is the process of separating a speech signal from noise and environmental components to obtain more intelligible data. Speaker recognition identifies a person by using distinctive characteristics contained in that person’s speech signal.