Audio Separation

Technical terms
About 1 min read

A technology that uses AI models to analyze and extract individual sound source components (Stems), such as vocals, drums, and bass, from complex audio signals into independent tracks.

Also known as
Sound Source SeparationSource SeparationSource SeparationStem Extraction

Detailed explanation

Audio separation is a technology that analyzes a single-track mixed sound source with machine learning algorithms to identify and isolate the unique frequency and temporal characteristics of each sound. Unlike past simple phase reversal or equalization methods, modern AI models learn complex patterns of waveforms through CNN (Convolutional Neural Network) or hybrid Transformer architectures. This enables precise extraction of vocals and accompaniment (MR), as well as various instrument groups like piano and guitar, while minimizing degradation of the original sound quality. When choosing a tool, users should judge performance based on the degree of distortion (artifacts) occurring after separation and the Signal-to-Distortion Ratio (SDR) metrics. Currently, this technology is utilized as a core technology in various fields, not only in music production but also in noise removal for post-production, dialogue extraction for content localization, and hearing assist devices.

Why It Matters in Tool Selection

The less 'bleeding'—where mechanical noise (artifacts) or sounds of other instruments blend into the extracted individual tracks—the higher the freedom of subsequent editing. In tasks requiring commercial quality, restoration capability without loss of frequency range is a more critical tool selection criterion than simple separation capability.

What to Check

  • Types and number of separable stems (usually in units of 2, 4, 5, or 6)
  • Processing mode options (high-quality slow mode vs. low-quality real-time mode)
  • Ability to preserve and restore data in the high-frequency range (16kHz or higher)
  • Support for batch processing and API integration

Examples

For example, cleanly extracting only an actor's dialogue from an old movie film track where background music and sound effects are mixed, to use it as base data for AI voice cloning or multi-language dubbing.