core¶
The Sound container, decibels, small numeric helpers, the FFT thread setting, and processing that needs no analysis (padding, mixing, filtering, level changes). Imports nothing else from sonore at module level.
sonore.core.utils¶
Small numerical helpers: levels and decibels, the ERB and mel scales and the frequency scales built on them, cents and note names, time axes, random generators, and internal array helpers used across sonore.
- amp_to_db(x: ArrayLike, ref: float = 1.0, floor_db: float = -300.0) ndarray[source]¶
Amplitude (not power) to decibels:
20*log10(|x|/ref), floored.
- power_to_db(x: ArrayLike, ref: float = 1.0, floor_db: float = -300.0) ndarray[source]¶
Power (not amplitude) to decibels:
10*log10(x/ref), floored atfloor_db.A power is never negative. A negative value is treated as zero and returns
floor_db, never the level of its absolute value: a tiny negative from round-off (a power of -1e-17 where the true value is 0) lands where true zeros do, and a negative from a bug upstream shows up at the floor instead of as a plausible level. (amp_to_dbtakes the absolute value instead, because amplitudes are signed.)
- freq_to_erb(freq: ArrayLike) ndarray[source]¶
Frequency [Hz] to ERB-number [Cams] (Glasberg & Moore, 1990).
The exact integral of
1 / erb_bandwidth,9.265 ln(1 + f / 228.8). The paper prints the rounded form21.4 log10(1 + 0.00437 f), which is about 0.3% higher (15.62 rather than 15.57 Cams at 1 kHz).
- erb_to_freq(erb: ArrayLike) ndarray[source]¶
ERB-number [Cams] to frequency [Hz]; inverse of
freq_to_erb().
- erb_bandwidth(freq: ArrayLike) ndarray[source]¶
Equivalent rectangular bandwidth [Hz] of the auditory filter at
freq.
- freq_to_mel(freq, scale: str = 'htk') ndarray[source]¶
Frequency [Hz] to mel.
scale="htk"is2595 log10(1 + f / 700), the formula HTK uses, as printed in O’Shaughnessy (2000, Eq. 4.2, p. 128), which gives no earlier source for it."slaney"is the scale of Slaney’s Auditory Toolbox and librosa: linear below 1 kHz (15 mel at 1000 Hz) and logarithmic above it (27 mel per factor of 6.4).
- mel_to_freq(mel, scale: str = 'htk') ndarray[source]¶
Mel to frequency [Hz]; the inverse of
freq_to_mel().
- ratio_to_cents(ratio: ArrayLike) ndarray[source]¶
The size of a frequency ratio in cents,
1200 log2(ratio): an octave (2) is 1200 cents, an equal-tempered semitone 100, a just fifth (3/2) 702.0.
- cents_to_ratio(cents: ArrayLike) ndarray[source]¶
The frequency ratio of an interval
centswide,2 ** (cents / 1200).
- note_to_freq(note: str, a4: float = 440.0) float[source]¶
The equal-tempered frequency [Hz] of a note name such as
"A4","C#3"or"Bb2"(scientific pitch notation: octave 4 starts at middle C;#orb, or ♯ and ♭, for sharp and flat, repeatable), with A4 ata4Hz.
- class FrequencyScale(name: str, unit: str, to_scale: Callable[[ndarray], ndarray], from_scale: Callable[[ndarray], ndarray])[source]¶
Bases:
objectA frequency axis, such as the one filterbank centers are equally spaced on: its
name, theunitof one step on it, and the conversions to and from Hz.- name: str¶
- unit: str¶
- to_scale: Callable[[ndarray], ndarray]¶
- from_scale: Callable[[ndarray], ndarray]¶
- FREQUENCY_SCALES = {'cents': FrequencyScale(name='cents', unit='cent', to_scale=<function cents_scale.<locals>.<lambda>>, from_scale=<function cents_scale.<locals>.<lambda>>), 'erb': FrequencyScale(name='erb', unit='ERB', to_scale=<function freq_to_erb>, from_scale=<function erb_to_freq>), 'linear': FrequencyScale(name='linear', unit='Hz', to_scale=<function _identity>, from_scale=<function _identity>), 'mel': FrequencyScale(name='mel', unit='mel', to_scale=<function freq_to_mel>, from_scale=<function mel_to_freq>), 'octave': FrequencyScale(name='octave', unit='oct', to_scale=<ufunc 'log2'>, from_scale=<ufunc 'exp2'>)}¶
the ERB-number scale (Glasberg & Moore, 1990), octaves (log2 frequency), cents from A4 at 440 Hz, the HTK mel scale, and Hz.
- Type:
The scales a filterbank can be built on by name
- cents_scale(reference: float = 440.0) FrequencyScale[source]¶
Cents from
reference[Hz] (A4 by default):1200 log2(f / reference). A factor of 2 is 1200 cents and an equal-tempered semitone 100, so a filterbank on this scale withspacing=100has one band per semitone. The reference only shifts the positions: the bank is the one built on octaves withspacing=1/12.
- n_samples(duration: float, fs: float) int[source]¶
Number of samples in
durationseconds, rounded to the nearest whole number rather than floored (exact halves round to even, as Python’srounddoes).
sonore.core.units¶
A decibel unit, so level changes read like arithmetic: snd + 6*dB.
6*dB is a Decibels value, not a number. Adding it to a
Sound changes the level; adding a Sound mixes; adding a bare
number is an error. That keeps snd + 0.5 from silently meaning either
“+0.5 dB” or “add a DC offset of 0.5”.
- class Decibels(value: float)[source]¶
Bases:
objectA level change in decibels (amplitude:
gain = 10**(value/20)).Write the number first (
6*dB,2*(6*dB));dB*6is refused. A level is not a plain number either:float(6*dB)raises, and(6*dB).valueis 6.- value: float¶
- dB = 1 dB¶
One decibel. Use as
snd + 6*dB,snd - 3*dB, ormasker + snr*dB.
sonore.core.fft¶
How many threads sonore’s large FFTs use, and FFT-friendly lengths.
The whole-signal transforms (filterbank analysis and synthesis, Hilbert envelopes of subbands, FFT resampling, modulation filtering) run a batch of 1-D FFTs, one per band and channel. SciPy can split such a batch across threads. Each transform is computed exactly as on one thread, so results are bit-identical whatever the number of threads; only the speed changes.
By default sonore uses every core this process may run on. When you run several sonore jobs in parallel (one per core, say), set it to 1 so they don’t compete:
so.set_fft_workers(1) # from now on
with so.set_fft_workers(1): # or only inside the block
...
The setting is one value for the whole process, not one per Python thread:
two threads that each enter with so.set_fft_workers(...) overwrite each
other’s value.
- class set_fft_workers(n: int | None = None)[source]¶
Bases:
objectSet the number of threads for sonore’s large FFTs (
None: every available core, the default). Takes effect at once; used in awithblock, the previous value is restored on exit.
- fast_padding(n: int, pad: int) int[source]¶
The smallest padding
>= padfor whichn + 2 * paddinghas no prime factor above 11. FFTs of such lengths are about four times faster than of a prime length. Padding on both sides keeps the parity ofn, so an oddncannot use the factor 2. Fornfrom 10001 to 60000 the padded length grows by at most about 5% (median 1.1%) for oddn, and by at most 1.6% (median 0.2%) for evenn. All of these are measured by tools/measure_docstring_numbers.py.
sonore.core.sound¶
The Sound container.
A Sound is an immutable (n_samples, n_channels) float array plus a
sampling rate. Every operation returns a new Sound; nothing is modified in
place.
- class Sound(data: ArrayLike, fs: float)[source]¶
Bases:
objectA sampled sound.
- Parameters:
data – Samples, shape
(n_samples,)or(n_samples, n_channels).fs – Sampling rate [Hz].
Notes
Arithmetic works the way you’d expect for signals:
a + bmixes,a * bmultiplies sample-by-sample (e.g. an envelope),2 * ascales, anda + 6*dB/a - 3*dBchange the level (from sonore import dB). Mono sounds broadcast against multichannel ones. Usepad()orsonore.match_lengths()to match lengths.Indexing with a slice selects by time in seconds:
snd[0.1:0.5]. Times round to the nearest sample, and time slices behave like Python index slicing: a negative time counts from the end (snd[-0.1:]is the last 0.1 s), and a time past either end is clipped to it (snd[0.5:2.0]of a 1 s sound is its last 0.5 s).- classmethod from_channels(*channels: TypeAliasForwardRef('ArrayLike') | Sound, fs: float | None = None) Sound[source]¶
Stack 1-D arrays or mono Sounds into a multichannel Sound.
Channels of different lengths are allowed: the shorter ones are zero-padded at the end to the longest.
- classmethod load(path: str | PathLike, **kwargs) Sound[source]¶
Read an audio file (anything libsndfile supports).
- property fs: float[source]¶
resample()gives the sound at another rate.- Type:
Sampling rate [Hz]. Read-only, like the samples
- normalize(rms: float | None = 1.0, peak: float | None = None) Sound[source]¶
Scale to a target RMS (default 1) or, if
peakis given, a target peak.
- ramp(duration: float, shape: str = 'cosine') Sound[source]¶
Apply onset and offset ramps of
durationseconds.shapeis"cosine"(raised cosine) or"linear".
- pad_to(n: int, align: str = 'start') Sound[source]¶
Zero-pad to
nsamples.alignisstart,center, orend.
- delay(seconds: float) Sound[source]¶
Delay by
seconds(may be fractional samples); the result is longer.Integer-sample delays are exact. Fractional delays use an FFT phase ramp (band-limited interpolation); the output is
ceil(delay)samples longer, so sinc ringing past the end is cut off. That’s inaudible for ramped stimuli; for impulse responses usesonore.simple_bir().
- resample(fs: float) Sound[source]¶
Polyphase resampling to a new rate (
scipy.signal.resample_poly()with SciPy’s default anti-aliasing filter, a Kaiser window with beta 5). The ratio of rates is approximated by a fraction with denominator at most 10000, which is exact for the common rates (8, 16, 22.05, 44.1, 48, 96 kHz).
- convolve(ir: Sound | TypeAliasForwardRef('ArrayLike')) Sound[source]¶
Convolve with an impulse response (full length).
A mono IR is applied to every channel; a multichannel IR convolves a mono sound into that many channels, or channel-by-channel otherwise.
- envelope(pad: float | str = 'auto')[source]¶
Hilbert envelope of each channel, as an
Envelope(not a Sound: you apply an envelope to a sound rather than listen to it).snd / snd.envelope()is the fine structure.The Hilbert transform is FFT-based. By default the sound is zero-padded by its own length on each side first, so a loud start can’t leak into the end of the envelope (or vice versa).
pad=0is circular; a number pads by that many seconds.
- play(blocking: bool = False, **kwargs) None[source]¶
Play through the default device (requires the
sounddeviceextra).
- load(path: str | PathLike, **kwargs) Sound[source]¶
Read an audio file. Alias for
Sound.load().
sonore.core.processing¶
Operations on sounds and lists of sounds.
- match_fs(sounds: Sequence[Sound], fs: float | None = None, mode: str = 'down') list[Sound][source]¶
Resample so all sounds share a rate:
fsif given, else the lowest (mode="down") or highest (mode="up") rate in the list.
- match_channels(sounds: Sequence[Sound]) list[Sound][source]¶
Upmix mono sounds to the channel count of the others.
- match_lengths(sounds: Sequence[Sound], mode: str = 'pad', align: str = 'start') list[Sound][source]¶
Bring sounds to one length: zero-pad to the longest (
mode="pad") or cut to the shortest (mode="truncate").align(start,centerorend) is where each sound sits in the common length: padding goes after, around or before it, and cutting keeps its start, middle or end. Withcenter, an odd number of extra samples puts the odd one at the end.
- normalize(sounds: Sequence[Sound], rms: float | None = 1.0, peak: float | None = None) list[Sound][source]¶
Scale each sound on its own to a target RMS (default 1) or, if
peakis given, a target peak, asSound.normalize()does.Each sound gets its own gain, so their level differences are lost: with
rmsthey all come out equally loud, withpeakthey all peak at the same value.
- mix(sounds: Sequence[Sound], align: str = 'start') Sound[source]¶
Sum sounds of possibly different lengths (zero-padded per
align).
- relative_db(sounds: Sequence[Sound], ref: int = 0) list[float][source]¶
Level of each sound in dB relative to
sounds[ref].
- butter_filter(sound: Sound, cutoff: float | tuple[float, float], btype: str = 'bandpass', order: int = 4, zero_phase: bool = True) Sound[source]¶
Butterworth filter along time (each channel separately).
btypeislowpass,highpass,bandpass, orbandstop.zero_phase=Trueruns it forward and backward: no phase shift, and the magnitude response is squared (twice the attenuation in dB, so the gain at the cutoff is -6 dB instead of -3 dB). This needs a sound longer than SciPy’s padding (15 samples for a 4th-order lowpass).
- bandpass(sound: Sound, f_lo: float, f_hi: float, order: int = 4, zero_phase: bool = True) Sound[source]¶
Butterworth band-pass from
f_lotof_hi[Hz].Shorthand for
butter_filter(sound, (f_lo, f_hi), "bandpass", order, zero_phase); seebutter_filter(). The band edges are where the gain is -3 dB (one way) or -6 dB (zero_phase=True).orderis SciPy’s, so the band-pass has2 * orderpoles.
- amplitude_modulate(sound: Sound, f_mod: float, depth: float = 1.0, phase: float = 0.0) Sound[source]¶
Multiply by
1 + depth*sin(2*pi*f_mod*t + phase).phaseis in radians.depth1 takes the envelope down to zero; above 1 the envelope changes sign (overmodulation).
- resonator(sound: Sound, f: float | tuple[ArrayLike, ArrayLike], bw: float | tuple[ArrayLike, ArrayLike]) Sound[source]¶
A formant: Klatt’s (1980) second-order digital resonator.
y[n] = A x[n] + B y[n-1] + C y[n-2]withC = -exp(-2 pi bw / fs),B = 2 exp(-pi bw / fs) cos(2 pi f / fs)andA = 1 - B - C.This is the damped harmonic oscillator
x'' + 2 sigma x' + w0^2 x = w0^2 usampled: its poles are the oscillator’s, mapped byz = exp(s / fs), withfthe ringing frequency andbw = sigma / pithe decay rate, and its gain at 0 Hz is exactly 1, as the oscillator’s is (sof = 0gives a low-pass filter).fandbware the peak and the -3 dB bandwidth only whenbwis much smaller thanf, in the oscillator too. Far above the resonance the gain falls more slowly than the oscillator’s, since a digital response repeats everyfs.fandbware numbers, or(times, values)pairs that are interpolated linearly to every sample (held beyond their ends), so a formant can glide. The filter keeps its state while its coefficients change, which is what keeps a moving formant free of clicks.
- antiresonator(sound: Sound, f: float | tuple[ArrayLike, ArrayLike], bw: float | tuple[ArrayLike, ArrayLike]) Sound[source]¶
An antiformant: the exact inverse of
resonator()with the samefandbw, a spectral zero (Klatt, 1980). Its gain at 0 Hz is 1.y[n] = (x[n] - B x[n-1] - C x[n-2]) / A;fandbwmay change over time as inresonator().