core

The Sound container, decibels, small numeric helpers, the FFT thread setting, and processing that needs no analysis (padding, mixing, filtering, level changes). Imports nothing else from sonore at module level.

sonore.core.utils

Small numerical helpers: levels and decibels, the ERB and mel scales and the frequency scales built on them, cents and note names, time axes, random generators, and internal array helpers used across sonore.

rms(x: ArrayLike, axis: int | None = None) → ndarray | float[source]

Root-mean-square value.

amp_to_db(x: ArrayLike, ref: float = 1.0, floor_db: float = -300.0) → ndarray[source]

Amplitude (not power) to decibels: 20*log10(|x|/ref), floored.

power_to_db(x: ArrayLike, ref: float = 1.0, floor_db: float = -300.0) → ndarray[source]

Power (not amplitude) to decibels: 10*log10(x/ref), floored at floor_db.

A power is never negative. A negative value is treated as zero and returns floor_db, never the level of its absolute value: a tiny negative from round-off (a power of -1e-17 where the true value is 0) lands where true zeros do, and a negative from a bug upstream shows up at the floor instead of as a plausible level. (amp_to_db takes the absolute value instead, because amplitudes are signed.)

db_to_power(db: ArrayLike) → ndarray[source]

Decibels to power factor: 10**(db/10).

db_to_amp(db: ArrayLike) → ndarray[source]

Decibels to amplitude factor: 10**(db/20).

freq_to_erb(freq: ArrayLike) → ndarray[source]

Frequency [Hz] to ERB-number [Cams] (Glasberg & Moore, 1990).

The exact integral of 1 / erb_bandwidth, 9.265 ln(1 + f / 228.8). The paper prints the rounded form 21.4 log10(1 + 0.00437 f), which is about 0.3% higher (15.62 rather than 15.57 Cams at 1 kHz).

erb_to_freq(erb: ArrayLike) → ndarray[source]

ERB-number [Cams] to frequency [Hz]; inverse of freq_to_erb().

erb_bandwidth(freq: ArrayLike) → ndarray[source]

Equivalent rectangular bandwidth [Hz] of the auditory filter at freq.

freq_to_mel(freq, scale: str = 'htk') → ndarray[source]

Frequency [Hz] to mel.

scale="htk" is 2595 log10(1 + f / 700), the formula HTK uses, as printed in O’Shaughnessy (2000, Eq. 4.2, p. 128), which gives no earlier source for it. "slaney" is the scale of Slaney’s Auditory Toolbox and librosa: linear below 1 kHz (15 mel at 1000 Hz) and logarithmic above it (27 mel per factor of 6.4).

mel_to_freq(mel, scale: str = 'htk') → ndarray[source]

Mel to frequency [Hz]; the inverse of freq_to_mel().

ratio_to_cents(ratio: ArrayLike) → ndarray[source]

The size of a frequency ratio in cents, 1200 log2(ratio): an octave (2) is 1200 cents, an equal-tempered semitone 100, a just fifth (3/2) 702.0.

cents_to_ratio(cents: ArrayLike) → ndarray[source]

The frequency ratio of an interval cents wide, 2 ** (cents / 1200).

note_to_freq(note: str, a4: float = 440.0) → float[source]

The equal-tempered frequency [Hz] of a note name such as "A4", "C#3" or "Bb2" (scientific pitch notation: octave 4 starts at middle C; # or b, or ♯ and ♭, for sharp and flat, repeatable), with A4 at a4 Hz.

class FrequencyScale(name: str, unit: str, to_scale: Callable[[ndarray], ndarray], from_scale: Callable[[ndarray], ndarray])[source]

Bases: object

A frequency axis, such as the one filterbank centers are equally spaced on: its name, the unit of one step on it, and the conversions to and from Hz.

name: str
unit: str
to_scale: Callable[[ndarray], ndarray]
from_scale: Callable[[ndarray], ndarray]
FREQUENCY_SCALES = {'cents': FrequencyScale(name='cents', unit='cent', to_scale=<function cents_scale.<locals>.<lambda>>, from_scale=<function cents_scale.<locals>.<lambda>>), 'erb': FrequencyScale(name='erb', unit='ERB', to_scale=<function freq_to_erb>, from_scale=<function erb_to_freq>), 'linear': FrequencyScale(name='linear', unit='Hz', to_scale=<function _identity>, from_scale=<function _identity>), 'mel': FrequencyScale(name='mel', unit='mel', to_scale=<function freq_to_mel>, from_scale=<function mel_to_freq>), 'octave': FrequencyScale(name='octave', unit='oct', to_scale=<ufunc 'log2'>, from_scale=<ufunc 'exp2'>)}

the ERB-number scale (Glasberg & Moore, 1990), octaves (log2 frequency), cents from A4 at 440 Hz, the HTK mel scale, and Hz.

Type:

The scales a filterbank can be built on by name

cents_scale(reference: float = 440.0) → FrequencyScale[source]

Cents from reference [Hz] (A4 by default): 1200 log2(f / reference). A factor of 2 is 1200 cents and an equal-tempered semitone 100, so a filterbank on this scale with spacing=100 has one band per semitone. The reference only shifts the positions: the bank is the one built on octaves with spacing=1/12.

n_samples(duration: float, fs: float) → int[source]

Number of samples in duration seconds, rounded to the nearest whole number rather than floored (exact halves round to even, as Python’s round does).

time_axis(n: int, fs: float) → ndarray[source]

Sample times k/fs. Use this instead of np.linspace(0, dur, n), whose spacing is dur/(n-1) rather than 1/fs.

as_rng(rng: int | Generator | None) → Generator[source]

Accept a seed, a Generator, or None and return a Generator.

sonore.core.units

A decibel unit, so level changes read like arithmetic: snd + 6*dB.

6*dB is a Decibels value, not a number. Adding it to a Sound changes the level; adding a Sound mixes; adding a bare number is an error. That keeps snd + 0.5 from silently meaning either “+0.5 dB” or “add a DC offset of 0.5”.

class Decibels(value: float)[source]

Bases: object

A level change in decibels (amplitude: gain = 10**(value/20)).

Write the number first (6*dB, 2*(6*dB)); dB*6 is refused. A level is not a plain number either: float(6*dB) raises, and (6*dB).value is 6.

value: float
property gain: float[source]

Linear amplitude factor.

dB = 1 dB

One decibel. Use as snd + 6*dB, snd - 3*dB, or masker + snr*dB.

sonore.core.fft

How many threads sonore’s large FFTs use, and FFT-friendly lengths.

The whole-signal transforms (filterbank analysis and synthesis, Hilbert envelopes of subbands, FFT resampling, modulation filtering) run a batch of 1-D FFTs, one per band and channel. SciPy can split such a batch across threads. Each transform is computed exactly as on one thread, so results are bit-identical whatever the number of threads; only the speed changes.

By default sonore uses every core this process may run on. When you run several sonore jobs in parallel (one per core, say), set it to 1 so they don’t compete:

so.set_fft_workers(1)          # from now on
with so.set_fft_workers(1):    # or only inside the block
    ...

The setting is one value for the whole process, not one per Python thread: two threads that each enter with so.set_fft_workers(...) overwrite each other’s value.

fft_workers() → int[source]

Number of threads sonore’s large FFTs use.

class set_fft_workers(n: int | None = None)[source]

Bases: object

Set the number of threads for sonore’s large FFTs (None: every available core, the default). Takes effect at once; used in a with block, the previous value is restored on exit.

fast_padding(n: int, pad: int) → int[source]

The smallest padding >= pad for which n + 2 * padding has no prime factor above 11. FFTs of such lengths are about four times faster than of a prime length. Padding on both sides keeps the parity of n, so an odd n cannot use the factor 2. For n from 10001 to 60000 the padded length grows by at most about 5% (median 1.1%) for odd n, and by at most 1.6% (median 0.2%) for even n. All of these are measured by tools/measure_docstring_numbers.py.

sonore.core.sound

The Sound container.

A Sound is an immutable (n_samples, n_channels) float array plus a sampling rate. Every operation returns a new Sound; nothing is modified in place.

class Sound(data: ArrayLike, fs: float)[source]

Bases: object

A sampled sound.

Parameters:
  • data – Samples, shape (n_samples,) or (n_samples, n_channels).

  • fs – Sampling rate [Hz].

Notes

Arithmetic works the way you’d expect for signals: a + b mixes, a * b multiplies sample-by-sample (e.g. an envelope), 2 * a scales, and a + 6*dB / a - 3*dB change the level (from sonore import dB). Mono sounds broadcast against multichannel ones. Use pad() or sonore.match_lengths() to match lengths.

Indexing with a slice selects by time in seconds: snd[0.1:0.5]. Times round to the nearest sample, and time slices behave like Python index slicing: a negative time counts from the end (snd[-0.1:] is the last 0.1 s), and a time past either end is clipped to it (snd[0.5:2.0] of a 1 s sound is its last 0.5 s).

classmethod from_channels(*channels: TypeAliasForwardRef('ArrayLike') | Sound, fs: float | None = None) → Sound[source]

Stack 1-D arrays or mono Sounds into a multichannel Sound.

Channels of different lengths are allowed: the shorter ones are zero-padded at the end to the longest.

classmethod load(path: str | PathLike, **kwargs) → Sound[source]

Read an audio file (anything libsndfile supports).

property fs: float[source]

resample() gives the sound at another rate.

Type:

Sampling rate [Hz]. Read-only, like the samples

property data: ndarray[source]

Read-only (n_samples, n_channels) array.

property n_samples: int[source]

Number of samples per channel (the same as len(snd)).

property n_channels: int[source]

Number of channels.

property duration: float[source]

n_samples / fs.

Type:

Length [s]

property t: ndarray[source]

Sample times [s].

property rms: float[source]

RMS over all samples and channels.

property peak: float[source]

Largest absolute sample over all channels (0 for an empty sound).

channel(i: int) → Sound[source]

Channel i as a mono Sound (negative i counts from the last).

property left: Sound[source]

The first channel, as a mono Sound.

property right: Sound[source]

The second channel, as a mono Sound.

mono() → Sound[source]

Average of all channels.

to_channels(n: int = 2) → Sound[source]

Copy a mono sound into n identical channels (diotic).

to_stereo() → Sound[source]

to_channels(2): a mono sound in both ears (diotic).

gain_db(db: float | Decibels) → Sound[source]

Change level by db decibels. Same as snd + db*dB.

normalize(rms: float | None = 1.0, peak: float | None = None) → Sound[source]

Scale to a target RMS (default 1) or, if peak is given, a target peak.

zero_mean() → Sound[source]

Remove each channel’s own mean (DC offset).

ramp(duration: float, shape: str = 'cosine') → Sound[source]

Apply onset and offset ramps of duration seconds.

shape is "cosine" (raised cosine) or "linear".

pad(before: float = 0.0, after: float = 0.0) → Sound[source]

Zero-pad by before/after seconds.

pad_to(n: int, align: str = 'start') → Sound[source]

Zero-pad to n samples. align is start, center, or end.

delay(seconds: float) → Sound[source]

Delay by seconds (may be fractional samples); the result is longer.

Integer-sample delays are exact. Fractional delays use an FFT phase ramp (band-limited interpolation); the output is ceil(delay) samples longer, so sinc ringing past the end is cut off. That’s inaudible for ramped stimuli; for impulse responses use sonore.simple_bir().

resample(fs: float) → Sound[source]

Polyphase resampling to a new rate (scipy.signal.resample_poly() with SciPy’s default anti-aliasing filter, a Kaiser window with beta 5). The ratio of rates is approximated by a fraction with denominator at most 10000, which is exact for the common rates (8, 16, 22.05, 44.1, 48, 96 kHz).

convolve(ir: Sound | TypeAliasForwardRef('ArrayLike')) → Sound[source]

Convolve with an impulse response (full length).

A mono IR is applied to every channel; a multichannel IR convolves a mono sound into that many channels, or channel-by-channel otherwise.

envelope(pad: float | str = 'auto')[source]

Hilbert envelope of each channel, as an Envelope (not a Sound: you apply an envelope to a sound rather than listen to it). snd / snd.envelope() is the fine structure.

The Hilbert transform is FFT-based. By default the sound is zero-padded by its own length on each side first, so a loud start can’t leak into the end of the envelope (or vice versa). pad=0 is circular; a number pads by that many seconds.

play(blocking: bool = False, **kwargs) → None[source]

Play through the default device (requires the sounddevice extra).

save(path: str | PathLike, **kwargs) → None[source]

Write to an audio file (format from the extension; mp3 needs libsndfile>=1.1).

plot(ax=None, **kwargs)[source]

Waveform plot; returns the matplotlib Axes.

load(path: str | PathLike, **kwargs) → Sound[source]

Read an audio file. Alias for Sound.load().

sonore.core.processing

Operations on sounds and lists of sounds.

match_fs(sounds: Sequence[Sound], fs: float | None = None, mode: str = 'down') → list[Sound][source]

Resample so all sounds share a rate: fs if given, else the lowest (mode="down") or highest (mode="up") rate in the list.

match_channels(sounds: Sequence[Sound]) → list[Sound][source]

Upmix mono sounds to the channel count of the others.

match_lengths(sounds: Sequence[Sound], mode: str = 'pad', align: str = 'start') → list[Sound][source]

Bring sounds to one length: zero-pad to the longest (mode="pad") or cut to the shortest (mode="truncate").

align (start, center or end) is where each sound sits in the common length: padding goes after, around or before it, and cutting keeps its start, middle or end. With center, an odd number of extra samples puts the odd one at the end.

normalize(sounds: Sequence[Sound], rms: float | None = 1.0, peak: float | None = None) → list[Sound][source]

Scale each sound on its own to a target RMS (default 1) or, if peak is given, a target peak, as Sound.normalize() does.

Each sound gets its own gain, so their level differences are lost: with rms they all come out equally loud, with peak they all peak at the same value.

concat(sounds: Sequence[Sound]) → Sound[source]

Join end to end (mono sounds are upmixed to match).

mix(sounds: Sequence[Sound], align: str = 'start') → Sound[source]

Sum sounds of possibly different lengths (zero-padded per align).

relative_db(sounds: Sequence[Sound], ref: int = 0) → list[float][source]

Level of each sound in dB relative to sounds[ref].

butter_filter(sound: Sound, cutoff: float | tuple[float, float], btype: str = 'bandpass', order: int = 4, zero_phase: bool = True) → Sound[source]

Butterworth filter along time (each channel separately).

btype is lowpass, highpass, bandpass, or bandstop. zero_phase=True runs it forward and backward: no phase shift, and the magnitude response is squared (twice the attenuation in dB, so the gain at the cutoff is -6 dB instead of -3 dB). This needs a sound longer than SciPy’s padding (15 samples for a 4th-order lowpass).

bandpass(sound: Sound, f_lo: float, f_hi: float, order: int = 4, zero_phase: bool = True) → Sound[source]

Butterworth band-pass from f_lo to f_hi [Hz].

Shorthand for butter_filter(sound, (f_lo, f_hi), "bandpass", order, zero_phase); see butter_filter(). The band edges are where the gain is -3 dB (one way) or -6 dB (zero_phase=True). order is SciPy’s, so the band-pass has 2 * order poles.

amplitude_modulate(sound: Sound, f_mod: float, depth: float = 1.0, phase: float = 0.0) → Sound[source]

Multiply by 1 + depth*sin(2*pi*f_mod*t + phase).

phase is in radians. depth 1 takes the envelope down to zero; above 1 the envelope changes sign (overmodulation).

resonator(sound: Sound, f: float | tuple[ArrayLike, ArrayLike], bw: float | tuple[ArrayLike, ArrayLike]) → Sound[source]

A formant: Klatt’s (1980) second-order digital resonator.

y[n] = A x[n] + B y[n-1] + C y[n-2] with C = -exp(-2 pi bw / fs), B = 2 exp(-pi bw / fs) cos(2 pi f / fs) and A = 1 - B - C.

This is the damped harmonic oscillator x'' + 2 sigma x' + w0^2 x = w0^2 u sampled: its poles are the oscillator’s, mapped by z = exp(s / fs), with f the ringing frequency and bw = sigma / pi the decay rate, and its gain at 0 Hz is exactly 1, as the oscillator’s is (so f = 0 gives a low-pass filter). f and bw are the peak and the -3 dB bandwidth only when bw is much smaller than f, in the oscillator too. Far above the resonance the gain falls more slowly than the oscillator’s, since a digital response repeats every fs.

f and bw are numbers, or (times, values) pairs that are interpolated linearly to every sample (held beyond their ends), so a formant can glide. The filter keeps its state while its coefficients change, which is what keeps a moving formant free of clicks.

antiresonator(sound: Sound, f: float | tuple[ArrayLike, ArrayLike], bw: float | tuple[ArrayLike, ArrayLike]) → Sound[source]

An antiformant: the exact inverse of resonator() with the same f and bw, a spectral zero (Klatt, 1980). Its gain at 0 Hz is 1.

y[n] = (x[n] - B x[n-1] - C x[n-2]) / A; f and bw may change over time as in resonator().