Audio privacy · Speech removal · WASPAA 2025

Removing speech before domestic audio is shared

A real-time framework for identifying and removing speech from recorded audio while retaining non-speech acoustic information that remains useful for sound event detection.

Real-timespeech identification and removal
CNN + ASTaudio classification models
2 VADsSilero and WebRTC baselines
WASPAA 2025public demonstration

Project overview

Problem

Audio recorded in homes can contain valuable information for machine listening but also highly sensitive personal information in speech. The project treats speech removal as a release-stage privacy mechanism rather than assuming domestic recordings can simply be published as captured.

Systems compared

The framework brings together PANNs and E-PANNs convolutional audio models, the Audio Spectrogram Transformer, Silero VAD and WebRTC VAD in a common workflow for speech versus non-speech decisions.

Usable demonstration

A software GUI exposes the comparison in real time so that model behaviour can be inspected interactively rather than only through offline metrics. The work connects directly to the privacy requirements behind releases such as The Sounds of Home.

Project image

Speech Removal Framework
The portfolio figure is shown complete rather than cropped. Select it to open the original image.

Evidence and links

PANNsASTSilero VADWebRTC VADAudio privacy