views

One-way analyses, each saying what it drops: spectra and spectrograms (with the reassigned spectrogram), timbre descriptors, envelopes, modulation, the cepstrum and MFCCs, F0 tracking, WORLD’s envelope and aperiodicity, and the routes back to sound they have: the channel vocoder (with the envelopes), WORLD’s synthesis and the phase vocoder.

sonore.views.view

Views: one-way analyses that say what they drop.

A Frame turns a sound into coefficients and back. A view keeps only part of a sound, so many sounds give the same view and none can be read back from it alone. Every view therefore has a View.synthesize() that refuses, raising NotInvertibleError with the mathematical reason (View.discards) and the route that does lead back to a sound (View.back_to_sound), if sonore has one.

That route is to_sound, on the views that have a canonical one, and its arguments are what the view discarded: Spectrum.to_sound takes a carrier for the phase, Cepstrum.to_sound a phase, PVAnalysis.to_sound a time scale and a frequency map. On any other view, View.to_sound() refuses as synthesize does.

Every view has View.plot(). A view with one obvious picture overrides it with a short call into sonore.plotting, where the drawing lives; a view with none keeps the base method, which raises NotImplementedError with View.no_plot, a sentence naming what to plot instead.

exception NotInvertibleError[source]

Bases: NotImplementedError

Raised by View.synthesize(), and by View.to_sound() on a view with no route back: the view keeps too little of a sound for any sound to be recovered from it alone.

class View[source]

Bases: object

Base class of every view.

Each subclass sets two class attributes: discards, one sentence saying what the view drops and so why it cannot be inverted, and back_to_sound, one sentence naming the route to a sound that exists, or an empty string if there is none. A subclass that does not override plot() sets no_plot, one sentence saying why it has no picture and what to plot instead.

discards = ''
back_to_sound = ''
no_plot = ''
synthesize(*args, **kwargs)[source]

Refuse: raise NotInvertibleError with discards and back_to_sound as the message.

to_sound(*args, **kwargs)[source]

Refuse, as synthesize() does, on a view with no canonical route back to a sound; views that have one override it.

plot(*args, **kwargs)[source]

Refuse, on a view with no single picture: raise NotImplementedError with no_plot as the message. Views that have a picture override it.

sonore.views.spectra

Spectra and spectrograms that keep power and drop the phase: the spectrum of a sound, a TANDEM-STRAIGHT-style power spectrogram, and the reassigned spectrogram.

A Spectrum’s levels are a one-sided power spectral density in dB re 1 per Hz, as Welch’s method gives it, so Spectrum.from_sound() and long_term_spectrum() agree for a steady noise (the first scatters, the second is smooth).

class Spectrum(f: ndarray, level: ndarray)[source]

Bases: View

A magnitude spectrum: frequencies [Hz] and levels [dB].

A view: it keeps one level per frequency and drops the phase, and with it all timing, so no sound can be read back from it.

discards = 'Spectrum keeps only the level at each frequency: it discards the phase, and with it all timing, so many sounds share one spectrum.'
back_to_sound = "Spectrum.to_sound makes a sound with this spectrum from a carrier that supplies the phase (a new noise, another sound's phase, or the minimum phase), which is not the analyzed sound."
f: ndarray
level: ndarray
classmethod from_sound(sound: Sound) → Spectrum[source]

The power spectral density of the whole sound from one FFT with no window, |X|**2 / (fs N) with every bin but 0 Hz and Nyquist counted twice (one-sided), in dB re 1 per Hz. The sound is first mixed to mono by averaging its channels, so channels in antiphase cancel. Each bin of a noise scatters about its expected level; see long_term_spectrum() for a smooth estimate on the same scale.

level_at(freqs: ndarray) → ndarray[source]

Levels [dB] linearly interpolated at freqs, held at the end values outside the spectrum’s frequencies.

relative() → Spectrum[source]

Levels re the maximum (0 dB peak).

smooth(fraction: float = 0.3333333333333333) → Spectrum[source]

Fractional-octave smoothing (power average over fraction octave).

to_sound(duration: float, fs: float, carrier: str | Sound = 'noise', rng=None, **noise_kwargs) → Sound[source]

A sound with this magnitude spectrum, its phase supplied by carrier.

"noise" draws Gaussian noise with this spectral shape (extra keyword arguments go to gaussian_noise()). A Sound lends its phase: its first duration seconds, each channel’s spectrum given this magnitude, so Spectrum.from_sound(x).to_sound(x.duration, x.fs, carrier=x) gives back x at RMS 1. "minimum" gives the minimum-phase impulse response with this magnitude, from the folded cepstrum as in Cepstrum.to_stft. The level is read at the FFT bins of duration by linear interpolation, and is silent outside the spectrum’s frequencies (a spectrum measured at 16 kHz gives nothing above 8 kHz at any fs). Every result has RMS 1, as the generators do. None is the analyzed sound: the spectrum keeps no phase to give back.

plot(ax=None, **kwargs)[source]

Level against frequency, on a log frequency axis and re the peak by default (see plot_spectrum()).

long_term_spectrum(sounds: Sound | Sequence[Sound], win_dur: float = 0.1) → Spectrum[source]

Long-term average spectrum via Welch’s method, averaged over sounds (weighted by duration), in dB re 1 per Hz. This is the right input for speech-shaped noise.

win_dur [s] is Welch’s segment length (Hann windows, half overlapping); the frequency spacing is 1 / win_dur, 10 Hz by default. A sound shorter than win_dur is one segment, its own length, zero-padded to win_dur, which puts it on the same frequencies without changing its spectrum. Sounds at another rate are resampled to the first one’s, and each is mixed to mono by averaging its channels.

class TFPower(power: ndarray, t: ndarray, f: ndarray)[source]

Bases: View

A time-frequency power that is not a frame’s coefficients, so it has no synthesis: for example tandem_power(). It keeps the power of each cell and drops the phase.

power has shape (n_channels, n_freqs, n_windows), on frequencies f [Hz] and window times t [s], which need not be uniform.

discards = "TFPower keeps only the power of each cell: it discards the phase, and the cells are not a frame's coefficients, so no synthesis undoes them."
back_to_sound = ''
power: ndarray
t: ndarray
f: ndarray
property db: ndarray[source]

10*log10(power), floored like the other representations.

plot(ax=None, channel: int = 0, **kwargs)[source]

Power in dB (see plot_tf_db()).

tandem_power(sound: Sound, f0_times: Sequence[float], f0: Sequence[float], periods: float = 2.5, overlap: float = 4, window: str | tuple = 'blackman') → TFPower[source]

A TANDEM-STRAIGHT-style power spectrogram: the average of two pitch-adaptive spectrograms whose windows sit a quarter period before and after each window center.

For a periodic sound, the power through a window centered at t fluctuates with period T0 as the window slides across the glottal pulses. Averaging the powers at t - T0/4 and t + T0/4 (half a period apart) cancels the odd harmonics of that fluctuation, including the largest, so a short window shows the spectral envelope steadily. The defaults, a Blackman window 2.5 periods long, are those of Kawahara et al. (2011). Only this averaging is implemented, not TANDEM-STRAIGHT’s smoothing or aperiodicity analysis.

The schedule is pitch_adaptive() over the sound, with F0 bridged across unvoiced stretches. The result is magnitude only: synthesizing from the pair would need a union of two frames, which sonore does not provide.

class ReassignedSpectrogram(t_hat: ndarray, f_hat: ndarray, power: ndarray, keep: ndarray)[source]

Bases: View

Spectrogram cells moved to their reassigned times and frequencies.

t_hat, f_hat and power have shape (n_channels, n_freqs, n_windows), one entry per STFT cell; keep marks the cells within the threshold of the maximum. binned() sums the kept power onto a grid for display. There is no synthesis: reassignment is not linear.

discards = "ReassignedSpectrogram discards the phase and moves each cell's power to a new time and frequency, many cells to one point, so different sounds give the same picture."
back_to_sound = ''
no_plot = 'ReassignedSpectrogram has no plot: its points lie off any grid until one is chosen, so plot binned(t_edges, f_edges).'
t_hat: ndarray
f_hat: ndarray
power: ndarray
keep: ndarray
binned(t_edges: ndarray, f_edges: ndarray) → TFPower[source]

Kept power summed into the cells of t_edges x f_edges [s, Hz]; points outside the edges are dropped.

reassigned_spectrogram(sound: Sound, frame: GaborFrame, threshold_db: float = -60.0) → ReassignedSpectrogram[source]

The reassigned spectrogram (Kodera et al., 1978; Auger & Flandrin, 1995).

Each cell of frame’s spectrogram is moved from its window time t and bin frequency f to

  • t_hat = t + Re(X_tw conj X) / |X|**2,

  • f_hat = f - Im(X_dw conj X) / (2 pi |X|**2),

where X uses the window w, X_tw the time-weighted window tau w(tau) and X_dw its derivative w'(tau), with tau in seconds from the window’s middle sample, which is where SciPy (and so GaborFrame) references each time window’s phase. A tone off the bin grid, an impulse and a linear chirp land on their true frequency, time and instantaneous-frequency line.

The derivative is taken from the window’s formula, so only "hann" and ("gaussian", std) windows are accepted, std in samples as in SciPy. Cells more than threshold_db below the maximum (per channel) are marked not kept: their positions are mostly noise.

sonore.views.envelopes

Envelopes: the slowly varying amplitudes that shape sounds.

An envelope is not a sound. It is non-negative, it usually varies slowly enough to live at a low sampling rate, and you don’t listen to it: you apply it to a sound. The types reflect that:

  • Envelope is one envelope. Envelope * Sound modulates the sound.

  • Envelopes is one envelope per frequency band of a filterbank, i.e. a spectrotemporal envelope. This is conceptually the same thing as a cochleagram (the envelope of each cochlear-filter output over time); here the “cochlea” is whichever Filterbank (by default an ERB cosine_filterbank()) produced it. Envelopes * Subbands modulates each band.

The Hilbert decomposition of a band is then literal:

band == band.envelope() * (band / band.envelope())      # envelope x fine structure
sb == sb.envelopes() * sb.tfs()                          # for every band at once

Envelopes resampled to a lower rate are upsampled automatically (band-limited: polyphase, or the FFT for unusual rate ratios; clipped at zero) when applied to a sound.

to_sound(carrier) is the route back: the envelopes imposed on a carrier. channel_vocode() is a channel vocoder: a sound’s own band envelopes imposed on a carrier. With bands of noise as the carrier (the default) it is the noise vocoder of Shannon et al. (1995), as in simulations of cochlear-implant hearing.

class Envelope(data, fs: float)[source]

Bases: View

A single (possibly multichannel) envelope, shape (n_samples, n_channels). A view: it keeps a magnitude over time and drops the fine structure.

Arithmetic: *, / and + with numbers and other Envelopes (so 1 + 0.5 * env works), and env * snd / snd * env modulate a Sound. snd / env divides the envelope out of a sound. A one-channel envelope applies to every channel of a sound; otherwise the channel counts must match. to_sound() puts it on another sound’s fine structure.

discards = 'Envelope discards the fine structure: only a magnitude over time is kept.'
back_to_sound = "Envelope.to_sound puts it on the fine structure of a carrier sound, which is not the analyzed sound's fine structure."
property data: ndarray[source]

The envelope values, shape (n_samples, n_channels), read-only.

property duration: float[source]

Length in seconds.

property t: ndarray[source]

Sample times in seconds, starting at 0.

property db: ndarray[source]

Envelope in dB (20*log10).

lowpass(cutoff: float, order: int = 4) → Envelope[source]

Zero-phase Butterworth lowpass at cutoff Hz (result clipped at 0).

resample(fs: float) → Envelope[source]

The envelope at rate fs: band-limited, clipped at zero, with the ends extended along a line rather than dragged toward zero.

to_sound(carrier: Sound) → Sound[source]

This envelope on the fine structure of carrier: the carrier is divided by its own Hilbert envelope (carrier / carrier.envelope()) and multiplied by this one. carrier must last as long as the envelope; it sets the sampling rate. Plain envelope * sound keeps the sound’s own envelope as well (amplitude modulation).

plot(ax=None, **kwargs)[source]

The envelope against time (see plot_envelope()).

class Envelopes(data, fs: float, filterbank: Filterbank, pad: int = 0)[source]

Bases: _PaddedBands, View

One envelope per band of a filterbank: a spectrotemporal envelope.

Conceptually this is a cochleagram: the envelope of each filter’s output over time. Shape (n_samples, n_bands, n_channels); the bands include the filterbank’s lowpass and highpass edge filters, as in Subbands.

A view: it keeps each band’s Hilbert magnitude and drops the fine structure (and, with lowpass, the envelope’s own fast detail).

env[i] is an Envelope; env * subbands modulates each band (imposing the envelopes on a carrier, not an inverse); env.modulation_spectrum() gives its 2-D modulation spectrum on the filterbank’s frequency scale (cycles/octave or cycles/ERB).

Envelopes derived from padded Subbands carry the same zero-padding (pad samples at each end, hidden from data), so that modulating and re-synthesizing bands can’t wrap around. Envelopes made from scratch (e.g. a rendered ripple pattern) have pad=0 and are zero outside their own extent when combined with padded bands.

discards = 'Envelopes discard the fine structure: only the Hilbert magnitude of each band is kept, smoothed further when a lowpass was asked for.'
back_to_sound = 'Envelopes.to_sound puts them on the fine structure of a carrier (a new noise, tones at the band centers, or another sound) and synthesizes the bands, as the noise vocoder does.'
property duration: float[source]

Length in seconds (without the padding).

property t: ndarray[source]

Sample times in seconds, starting at 0.

property db: ndarray[source]

Envelopes in dB (20*log10), shape (n_samples, n_bands, n_channels).

lowpass(cutoff: float, order: int = 4) → Envelopes[source]

Zero-phase Butterworth lowpass of every band at cutoff Hz, padding included (result clipped at 0).

resample(fs: float) → Envelopes[source]

The envelopes at rate fs, padding included, as Envelope.resample(). When the rate ratio is a small fraction the inner signal starts exactly on a sample; otherwise within half of one.

without_edges() → Envelopes[source]

Zero the lowpass and highpass edge bands (which lie outside f_lo..f_hi), keeping the band count unchanged. A bank built with edges=False has none, and its envelopes are returned as they are.

modulation_spectrum(scale: str = 'linear', drop_edges: bool = True) → ModulationSpectrum[source]

2-D modulation spectrum (temporal Hz x spectral cycles per scale unit).

scale="db" transforms log envelopes. The edge bands are dropped by default, since they aren’t evenly spaced with the others; a bank built with edges=False has none to drop, and all its bands are kept.

to_sound(carrier: Sound | str = 'noise', fs: float | None = None, rng=None) → Sound[source]

These envelopes imposed on a carrier, band by band, and synthesized with their filterbank.

  • "noise": each band of a Gaussian noise (from rng), scaled to a mean-square envelope of 1. The bands keep their own random envelope fluctuations, as in the classic noise vocoder.

  • "tone": a cosine at each band’s center.

  • a Sound lasting as long as the envelopes, analyzed with the same filterbank. Only its fine structure is used, so this is Envelope.to_sound() in every band.

fs is the output rate for "noise" and "tone" (by default the envelopes’ own); a Sound carrier sets its own. The edge bands are kept (see without_edges()) and the level is the envelopes’ own.

plot(ax=None, **kwargs)[source]

The envelopes as a cochleagram, time by band in dB (see plot_envelopes()).

channel_vocode(sound: Sound, n_bands: int = 16, f_lo: float = 80.0, f_hi: float = 8000.0, carrier: Sound | str = 'noise', env_lowpass: float | None = 50.0, rng=None) → Sound[source]

A channel vocoder with Hilbert envelopes: the sound’s band envelopes on a carrier. The default, carrier="noise", is the noise vocoder of Shannon et al. (1995).

The sound is split by a cosine_filterbank of n_bands bands from f_lo to f_hi (capped at Nyquist), each band’s Hilbert envelope is lowpassed at env_lowpass Hz, and the envelopes are imposed on the carrier with Envelopes.to_sound(): "noise" (bands of noise), "tone" (sinusoids at the band centers), or a Sound at the same rate and at least as long, whose fine structure carries them. The edge bands (outside f_lo..f_hi) are silenced. The output matches the input’s RMS.

sonore.views.mask

Time-frequency masks: gains on a frame’s coefficients.

A mask keeps one gain per coefficient and nothing of any sound, so it is a View. The way back to a sound goes through the coefficients it multiplies: (stft * mask).to_sound() is the frame’s least-squares inverse of the masked coefficients. A mask works with STFT, TVSTFT and Subbands.

The ideal binary mask is the goal Wang (2005) proposed for computational auditory scene analysis: keep the coefficients where the target is stronger than the masker by more than a local criterion lc_db, the parameter whose effect on listeners Brungart, Chang, Simpson & Wang (2006) measured. Ratio masks, soft gains from the target’s share of the power, go back to Srinivasan, Roman & Wang (2006); the ideal ratio mask (T**2 / (T**2 + M**2))**beta with beta = 0.5 is the form Wang, Narayanan & Wang (2014) found best, close to a square-root Wiener filter.

For complex coefficients the power of a cell is |X|**2. Subbands are real and oscillate, so their power is the Hilbert envelope squared: the mask then follows each band’s level, not its waveform.

class Mask(values, like: STFT | TVSTFT | Subbands)[source]

Bases: View

Gains, usually between 0 and 1, one per coefficient of the coefficients like it was made for (same frame, same grid).

mask * coefs and coefs * mask give the masked coefficients; .to_sound() on those is the way back to a sound. values has the layout of the coefficients’ own array: (n_channels, n_freqs, n_windows) for an STFT, (n_samples, n_bands, n_channels) for Subbands (their padding included).

discards = "Mask holds gains for a frame's coefficients, not the coefficients themselves: it keeps no level or phase of any sound."
back_to_sound = 'Multiply coefficients by it and go back from those: (stft * mask).to_sound().'
property t: ndarray[source]

Times [s] of the coefficients’ columns (time windows, or samples for subbands).

property f: ndarray[source]

FFT bins, or the filterbank’s center frequencies for subbands.

Type:

Frequencies [Hz] of the coefficients’ rows

apply(coefs: STFT | TVSTFT | Subbands) → STFT | TVSTFT | Subbands[source]

The coefficients multiplied by this mask (what coefs * mask does).

plot(ax=None, channel: int = 0, **kwargs)[source]

Draw one channel’s gains as an image on the coefficients’ time and frequency axes; see sonore.plotting.plot_mask().

ideal_binary_mask(target: STFT | TVSTFT | Subbands, masker: STFT | TVSTFT | Subbands, lc_db: float = 0.0) → Mask[source]

1 where the target’s level exceeds the masker’s by more than lc_db, the local SNR criterion, else 0 (Wang, 2005; Brungart et al., 2006).

ideal_ratio_mask(target: STFT | TVSTFT | Subbands, masker: STFT | TVSTFT | Subbands, beta: float = 0.5) → Mask[source]

(T**2 / (T**2 + M**2))**beta with T**2 and M**2 the powers of each coefficient (Wang, Narayanan & Wang, 2014; 0 where both are silent).

sonore.views.modulation

Modulation filterbanks: bandpass filters applied to envelopes.

Two banks from McDermott & Simoncelli (2011), usable on their own (e.g. for Dau-style modulation models) and by sonore.texture:

  • ConstantQModulationFilterbank: half-cycle cosines on a linear frequency axis, each spanning cf +/- cf/Q, with centers log-spaced. Used for the modulation power spectrum.

  • OctaveModulationFilterbank: half-cycle cosines on a log2 axis, each two octaves wide (cf/2 .. 2*cf), centers one octave apart. Used for the C1/C2 modulation correlations.

Filtering is done in the frequency domain and is therefore circular: the envelope is treated as one period of a periodic signal. That is exact for texture synthesis (seamless loops); for other uses pad the envelope first.

A third bank, HannModulationFilterbank, is defined in time instead: Hann-windowed complex exponentials of finite length, applied by direct (zero-padded, not circular) correlation, centered or causal. It is the bank behind ModulationSpectrogram, and its finite kernels are what make a causal, block-by-block version possible.

class ModulationFilterbank[source]

Bases: object

Base class. Subclasses define n_bands, cfs and response().

Despite the name, a modulation filterbank is not a Filterbank: it filters envelopes to make views such as ModulationSpectrogram, and has no synthesis.

property cfs: ndarray[source]
response(freqs) → ndarray[source]

Magnitude responses, shape (len(freqs), n_bands) (freqs >= 0).

filter(x, fs: float, axis: int = 0, analytic: bool = False) → ndarray[source]

Filter x (circularly) along axis with every band.

The band index is appended as a new last axis. With analytic=True the output is the complex analytic signal of each band (negative frequencies removed, positive ones doubled), whose real part equals the ordinary filtered output.

class ConstantQModulationFilterbank(n_bands: int = 20, f_lo: float = 0.5, f_hi: float = 200.0, Q: float = 2.0)[source]

Bases: ModulationFilterbank

n_bands constant-Q filters with centers log-spaced from f_lo to f_hi [Hz]. Filter k is a half-cycle cosine on a linear frequency axis, nonzero on cf_k*(1 - 1/Q) .. cf_k*(1 + 1/Q).

Responses are scaled so that the summed squared response, averaged over the middle of the bank (between the 4th and the 4th-from-last center), is 1. That makes band powers sum to about the envelope’s variance.

n_bands: int = 20
f_lo: float = 0.5
f_hi: float = 200.0
Q: float = 2.0
property cfs: ndarray[source]
property scale: float[source]
response(freqs) → ndarray[source]

Magnitude responses, shape (len(freqs), n_bands) (freqs >= 0).

class OctaveModulationFilterbank(n_bands: int = 7, f_hi: float = 100.0)[source]

Bases: ModulationFilterbank

n_bands filters with centers one octave apart, the highest at f_hi [Hz] (default: 1.5625, 3.125, …, 100 Hz). Filter k is a half-cycle cosine on log2 frequency spanning cf_k/2 .. 2*cf_k. Peak gain is 1 (the C1/C2 correlations are scale-invariant, so no further normalization is applied).

n_bands: int = 7
f_hi: float = 100.0
property cfs: ndarray[source]
response(freqs) → ndarray[source]

Magnitude responses, shape (len(freqs), n_bands) (freqs >= 0).

class HannModulationFilterbank(f_lo: float = 0.5, f_hi: float = 64.0, per_octave: float = 2, cycles: int = 3, window: float | None = None)[source]

Bases: ModulationFilterbank

Hann-windowed complex exponentials, defined in time.

Band k correlates the envelope with h_k[j] = w_k[j] exp(i 2 pi f_k (j - (L_k - 1)/2) / fs), where w_k is a Hann window of L_k samples normalized to sum 1. Two ways to set the window:

  • cycles (the default, 3): every window holds that many cycles of its own rate, L_k = cycles * fs / f_k, with centers per_octave to the octave from f_lo to f_hi. A constant-Q bank (Q = cycles / 1.44, about 2.1 for 3 cycles), long windows for slow rates and short ones for fast rates.

  • window [s]: every band uses the same window T, with centers on the linear grid k / T from the first one at or above max(f_lo, 2/T) to f_hi. This is the STFT of the envelope.

Either way each window holds a whole number (at least 2) of cycles of its rate, so the bank ignores the envelope’s mean: the Hann window’s transform is zero at every whole bin from 2 on.

Unlike the other two banks, filter() works in time and is not circular: the envelope is taken to be zero outside its extent. The analytic output 2 y has magnitude A * m for an envelope component A * m * cos(2 pi f_k t), and its real part is the real filtered output, whose frequency response is response().

f_lo: float = 0.5
f_hi: float = 64.0
per_octave: float = 2
cycles: int = 3
window: float | None = None
property n_bands: int[source]
property cfs: ndarray[source]

Center rates [Hz].

durations() → ndarray[source]

Window length [s] of each band.

lengths(fs: float) → ndarray[source]

Window length in samples at fs (rounded, at least 3).

check_fs(fs: float) → None[source]

The envelope rate must be at least 3 * f_hi: the top band reaches about 1.25 * f_hi, and its window must hold enough samples.

kernels(fs: float) → list[tuple[ndarray, ndarray]][source]

(h_k, w_k) for every band at fs: the complex kernel and its normalized Hann window.

response(freqs) → ndarray[source]

Magnitude response of the real filter (the real part of filter()’s analytic output), shape (len(freqs), n_bands), for continuous-time windows. Peak gain is about 1; zero at 0 Hz.

filter(x, fs: float, axis: int = 0, analytic: bool = False, align: str = 'center') → ndarray[source]

Filter x along axis with every band (zero outside x).

The band index is appended as a new last axis. align="center" puts each window’s middle on the output sample; "causal" ends it there, which is the centered output delayed by (L_k - 1) // 2 samples and is what a live, block-by-block analysis would produce.

class ModulationSpectrum(stft: STFT, channel: int = 0)[source]

Bases: View

2-D Fourier transform of a time-frequency envelope (Singh & Theunissen, 2003; Chi et al., 1999).

ModulationSpectrum(stft) uses a dB spectrogram, so spectral modulation is w.r.t. linear frequency (cycles/kHz). octave() uses subband envelopes on a log-frequency axis (cycles/octave), the axis on which ripples (sonore.sources.ripples) are defined; more generally, any Envelopes has .modulation_spectrum().

Sign convention: a ripple sin(2*pi*(rate*t + density*x)) appears at (+rate, +density). Only non-negative spectral modulations are stored (the other half is the complex conjugate).

A view: before the transform the envelope’s mean is removed and a Hann taper applied in time, and only the level (dB, floored) is kept, so the phase of the modulations is dropped and no envelope can be read back.

A spectrum made from Envelopes (by octave() or env.modulation_spectrum()) also keeps, for to_sound(), the magnitude of the untapered transform on the whole plane, the envelope’s mean, and how the envelopes were made. Edit it with with_gain(), and hear it with to_sound(), whose carrier supplies the phases it lacks (see docs/design/views/modulation-targets.md).

discards = 'ModulationSpectrum keeps only the magnitude of the 2-D Fourier transform of an envelope: it discards the phase of the modulations, which holds the timing of every event, and the fine structure under the envelope.'
back_to_sound = "ModulationSpectrum.to_sound(carrier=...) takes both from a carrier: a Sound lends its own modulation phase and fine structure, 'tones' and 'noise' draw a random modulation phase."
classmethod from_array(env: ndarray, dt: float, dx: float, spectral_unit: str) → ModulationSpectrum[source]

From a (frequency, time) envelope array sampled every dt seconds and dx scale units.

classmethod octave(sound: Sound, bands_per_octave: float = 12, f_lo: float = 125.0, f_hi: float = 8000.0, env_fs: float = 1000.0, scale: str = 'linear') → ModulationSpectrum[source]

Modulation spectrum on a log-frequency axis [cycles/octave].

Shorthand for:

fb = cosine_filterbank(f_lo=f_lo, f_hi=f_hi, spacing=1 / bands_per_octave, scale="octave")
fb.analyze(sound.mono()).envelopes(fs=env_fs).modulation_spectrum(scale)
classmethod from_blobs(blobs, duration: float, f_lo: float = 250.0, f_hi: float = 8000.0, bands_per_octave: float = 12, env_fs: float = 1000.0, rms_depth: float = 0.2) → ModulationSpectrum[source]

A target spectrum drawn in code: the sum of ModulationBlob powers, on the grid octave() would measure for a sound of duration seconds, ready for to_sound().

A drawing sets a shape, not a depth, and no long-term spectrum: every band gets the same mean envelope, and rms_depth scales the modulation (the rms of the envelope array about its mean, relative to the mean). Because envelopes cannot go below zero, a random draw reaches only a limited depth: about 0.28 for a one-blob target in docs/design/views/modulation-targets.md (C2), against 0.71 for one full ripple. to_envelopes() refuses a draw that would need clipping and names the largest depth that fits it. level shows the drawn power itself (no taper).

with_gain(gain) → ModulationSpectrum[source]

A copy with every cell’s amplitude multiplied by gain.

gain is a function g(rate, density) of rate [Hz] and spectral modulation, evaluated with rate as a row and density as a column (as ripple patterns are), or a number. A sound’s spectrum is the same at (rate, density) and (-rate, -density), so the gain is averaged over each such pair; g = lambda r, d: abs(r) <= 4 removes every modulation faster than 4 Hz, lambda r, d: r * d <= 0 every downward sweep. The displayed level (Hann-tapered) gets the same gain, which is exact where the gain is smooth across the taper’s rate resolution.

to_envelopes(carrier: Sound | None = None, rng=None) → Envelopes[source]

The envelopes this spectrum describes, given a modulation phase.

The stored magnitudes and mean fix everything but the phase of the 2-D transform, which holds when each event happens and how the bands line up. A Sound carrier lends the phase of its own envelopes (analyzed as this spectrum’s were), so the spectrum of x with carrier=x gives back x’s envelopes; with no carrier the phase is drawn at random from rng at every nonzero rate, while the zero-rate column keeps its own phase, so each band keeps its long-term level (the long-term spectrum). Envelopes rebuilt from a linear-scale spectrum can go below zero; they are clipped, with a warning saying how much, which changes their spectrum. Edge bands that the spectrum dropped are zero.

to_sound(carrier: str | Sound = 'tones', fs: float | None = None, rng=None, iterations: int = 0) → Sound[source]

A sound whose envelopes have this modulation spectrum.

A modulation spectrum lacks two kinds of phase, and the carrier supplies both:

  • the modulation phase (when each event happens, how the bands line up; see to_envelopes()): a Sound lends its own, "tones" and "noise" draw a random one from rng that keeps each band’s long-term level;

  • the fine structure under each band’s envelope: a Sound’s own ((envelopes * subbands.tfs()).to_sound(), the vocoder’s route), narrowband noise ("noise", the same route), or a steady tone at each band’s center ("tones"), added without re-filtering.

to_sound(carrier=x) on the spectrum of x itself rebuilds x’s envelopes, so an edit made with with_gain() keeps the sound’s timing wherever the gain is 1. fs is the audio rate, needed unless the carrier is a Sound, which must be at least as long as the analyzed envelopes. The result has RMS 1.

A fine structure that fluctuates within a band (a sound’s own, or noise) adds modulation of its own when the result is analyzed again, so an edit survives best on "tones": removing every rate above 4 Hz from a sentence leaves about 15 dB less power at 6-40 Hz on tones, but only 3-5 dB less on noise or the sentence’s own fine structure (docs/design/views/modulation-targets.md, C6).

iterations then searches for a sound whose own envelopes come closer to the spectrum, as Griffin & Lim (1984) do for a spectrogram: analyze the sound, keep its fine structure and its modulation phase, impose the stored magnitudes again, and go back. On the same sentence edit, 20 iterations from the sentence’s own fine structure leave 16.5 dB less power at 6-40 Hz (C7 of the design note); each iteration costs one analysis and one synthesis.

peak(exclude_dc: bool = True) → tuple[float, float][source]

(temporal Hz, spectral) coordinates of the largest component.

At zero spectral modulation, +rate and -rate are mirror images (the envelope is the same at every frequency, so it has no direction), and the rate is reported as non-negative.

plot(ax=None, **kwargs)[source]

The modulation spectrum as an image in dB (see plot_modulation_spectrum()).

class ModulationBlob(rate: float, density: float, rate_width: float = 0.5, density_width: float = 0.25, level: float = 0.0)[source]

Bases: object

A patch of modulation power, for drawing a target spectrum in code (ModulationSpectrum.from_blobs()).

A Gaussian bump centered on rate [Hz] and density [cycles/octave]: rate_width is its standard deviation in octaves of rate, density_width in cycles/octave, and level its peak power [dB] relative to the other blobs. The signs follow Ripple: positive rate and density are downward sweeps, a negative rate upward ones, and a blob at density 0 is temporal modulation alone. A ripple is a blob of zero width.

rate: float
density: float
rate_width: float = 0.5
density_width: float = 0.25
level: float = 0.0
power(rate: ndarray, density: ndarray) → ndarray[source]

Relative power at every (rate, density) (broadcast), zero on the other side of the rate axis.

sonore.views.modspectrogram

The modulation spectrogram: how strongly each band’s envelope is modulated at each rate, time window by time window.

An STFT shows how a sound’s power spectrum changes over time; a ModulationSpectrogram shows how its modulation spectrum changes over time. It is built from Envelopes (any filterbank) by passing every band’s envelope through a HannModulationFilterbank and sampling the result every hop seconds. docs/design/views/modulation-spectrogram.md has the design and the numbers behind it.

class ModulationSpectrogram(envelopes: Envelopes, f_lo: float = 0.5, f_hi: float = 64.0, per_octave: float = 2, cycles: int = 3, window: float | None = None, hop: float = 0.01, align: str = 'center', bank: HannModulationFilterbank | None = None)[source]

Bases: View

Modulation power, local mean and depth for every time window, acoustic band and modulation band.

For band envelope e_b and modulation band k with complex kernel h_k and window w_k (see HannModulationFilterbank), at each time window:

y    = sum_j e_b[n + j - c] conj(h_k[j])     (c: centered or causal)
mean = sum_j e_b[n + j - c] w_k[j]

and the stored arrays are power = |2 y|**2 and mean, both of shape (n_channels, n_bands, n_mod, n_windows). depth is 2 |y| / mean: a sinusoidal AM of depth m at the band’s rate reads m, so 100% modulation is 0 dB. Power says how much of the sound a modulation is; depth says how modulated the band is.

The acoustic axis f is the envelopes’ filterbank without its edge bands, fm holds the modulation rates and t the window times. valid ((n_bands, n_mod, n_windows)) is False where a number can’t be trusted: where the modulation rate exceeds the band’s -3 dB width (a band’s envelope can’t move faster than the band is wide), and where the window runs past either end of the envelopes (outside them the envelope is taken to be zero, so an abrupt start reads as modulation).

The representation is not invertible: the phase of y is dropped, as a magnitude spectrogram drops an STFT’s phase (the local mean is kept, and depth divides by it). For modulation filtering, filter the envelopes with a modulation bank directly.

Parameters:
  • envelopes – Band envelopes, for example fb.analyze(snd).envelopes(fs=1000). Their rate must be at least 3 times the highest modulation rate.

  • f_lo – The modulation bank (see HannModulationFilterbank): cycles sets a window of that many cycles of each band’s rate (constant Q); window [s] one window for every rate (the STFT of each envelope).

  • f_hi – The modulation bank (see HannModulationFilterbank): cycles sets a window of that many cycles of each band’s rate (constant Q); window [s] one window for every rate (the STFT of each envelope).

  • per_octave – The modulation bank (see HannModulationFilterbank): cycles sets a window of that many cycles of each band’s rate (constant Q); window [s] one window for every rate (the STFT of each envelope).

  • cycles – The modulation bank (see HannModulationFilterbank): cycles sets a window of that many cycles of each band’s rate (constant Q); window [s] one window for every rate (the STFT of each envelope).

  • window – The modulation bank (see HannModulationFilterbank): cycles sets a window of that many cycles of each band’s rate (constant Q); window [s] one window for every rate (the STFT of each envelope).

  • hop – Spacing between time windows [s], independent of the window, as in an STFT.

  • align – "center" (windows centered on the window times) or "causal" (windows ending at them: what a live analysis would see).

  • bank – A ready-made bank, instead of f_lo .. window.

discards = "ModulationSpectrogram discards the fine structure under the envelopes and the phase of every modulation band, as a magnitude spectrogram drops an STFT's phase."
back_to_sound = ''
property depth: ndarray[source]

Modulation depth 2 |y| / mean, same shape as power; NaN where the window holds only digital silence.

at(t: float, kind: str = 'depth') → ndarray[source]

The acoustic band x modulation rate image at the time window nearest t [s], shape (n_channels, n_bands, n_mod): Atlas and Shamma’s joint acoustic and modulation frequency display at one moment. kind is "depth" or "power".

average(kind: str = 'power') → ndarray[source]

Average over time windows, shape (n_channels, n_bands, n_mod). The average power is the per-band modulation power spectrum, the same kind of quantity as the texture statistics’ mod_power.

pooled_depth() → ndarray[source]

Depth pooled over acoustic bands, shape (n_channels, n_mod, n_windows): sqrt(sum_b power) / sqrt(sum_b mean**2) over the bands whose cell is valid, which weights bands by their level. NaN where no band is valid.

plot(ax=None, band: float | None = None, rate: float | None = None, **kwargs)[source]

Depth in dB as an image, invalid cells in gray. By default, modulation rate against time, pooled over bands (pooled_depth()); band= [Hz] shows the acoustic band nearest that frequency instead; rate= [Hz] shows acoustic band against time at the modulation band nearest that rate.

slices(t: float, rate: float = 4.0, **kwargs)[source]

Three linked cuts through the time x band x rate cube with a cursor at t [s]: rate against time (pooled), band against time at rate, and band against rate at t. Returns the figure.

animate(path=None, sound=None, fps: float = 25.0, **kwargs)[source]

The band x rate image (depth, dB) time window by time window, as a matplotlib animation. With path, it is written to a video file (needs ffmpeg), and with sound as well, the sound becomes its audio track. Returns the animation.

sonore.views.cepstrum

The real cepstrum of a short-time Fourier transform: liftering, resynthesis with the original or minimum phase, and cepstral F0.

class Cepstrum(coefs: STFT | TVSTFT, floor_db: float = -200.0)[source]

Bases: View

The real cepstrum of each time window of an STFT or TVSTFT.

For a time window with spectrum X[k] on n_fft bins, the cepstrum is c[n] = IDFT(ln |X[k]|) at quefrency n / fs seconds. It is real and even in n because the sound is real, so only n = 0 .. n_fft // 2 is stored: data has shape (n_channels, n_fft // 2 + 1, n_windows) on quefrencies q [s] and window times t [s].

The natural log is used, so exp undoes it exactly: without liftering, to_sound() gives back the analyzed sound. Scaling the sound changes only c[0]. Before the log, magnitudes are floored at floor_db below the channel’s largest magnitude over all time windows, so a time window of digital silence has a flat log spectrum rather than -inf; real recordings stay far above the default. Only the real cepstrum is provided: the complex cepstrum needs phase unwrapping, and the minimum phase comes from the real one.

Parameters:
  • coefs – The coefficients to take the cepstrum of. They are kept (as source) for their phase, frame and times.

  • floor_db – The floor, in dB below each channel’s maximum.

discards = 'Cepstrum discards the phase: its coefficients are the transform of the log magnitude alone, and only the source STFT it keeps holds the phase.'
back_to_sound = "Cepstrum.to_sound borrows the phase of the STFT it was computed from, and gives the sound back exactly only for an unliftered cepstrum with phase='original'."
property q: ndarray[source]

Quefrencies [s].

property t: ndarray[source]

Window center times [s], those of source.

lifter(cutoff: float | Sequence[float], keep: str = 'low') → Cepstrum[source]

A rectangular lifter. keep="low" keeps the quefrencies below cutoff [s], the smooth spectral envelope; "high" keeps the rest, the fine structure such as the harmonics. cutoff is one value or one per time window, so it can follow an F0 track (half a period separates the envelope from the harmonics).

envelope() → ndarray[source]

exp(DFT(c)): the magnitude spectrum this cepstrum stands for, shape (n_channels, n_freqs, n_windows) on the source’s frequencies. After a low lifter it is the cepstral spectral envelope, which follows the shape of the true envelope but sits a few dB below the harmonic peaks, because the lifter averages the peaks with the dips between them.

envelope_view() → GridEnvelope[source]

envelope() as power on this cepstrum’s time windows and frequencies, read as env(t, f), so it goes wherever a spectral envelope is taken (warp_frequency(), world_synthesize(), harmonic_complex()). Lifter first: cep.lifter(0.5 / f0).envelope_view().

to_stft(phase: str = 'original') → STFT | TVSTFT[source]

Coefficients of the source’s type and frame with envelope() as magnitude.

phase="original" takes the phase from source, so an unliftered cepstrum returns the source’s coefficients. "minimum" uses the minimum phase for that magnitude, from the folded cepstrum (c[0] kept, 2 c[n] up to n_fft / 2, zero beyond). Each time window’s response then starts at its phase reference, the middle of its window. The fold is exact up to the time aliasing of the cepstrum on n_fft bins, which is negligible once n_fft is several times the response’s length.

to_sound(phase: str = 'original') → Sound[source]

to_stft() synthesized by the source’s frame: exact for an unliftered cepstrum with the original phase, as long as no magnitude fell below floor_db (those are raised to the floor), and otherwise the least-squares signal for those coefficients. It is not an inverse of the cepstrum alone: the phase comes from the source STFT.

f0(f_lo: float = 75.0, f_hi: float = 400.0, threshold: float = 0.1) → tuple[ndarray, ndarray, ndarray][source]

Classic cepstral F0 (Noll, 1967): the largest cepstral peak between quefrencies 1/f_hi and 1/f_lo, refined by a parabola through it and its neighbors.

Returns (t, f0, peak): window times [s], and F0 [Hz] and the peak’s height for each channel and time window, shape (n_channels, n_windows). F0 is 0 where the peak is below threshold, a crude voicing rule.

A periodic sound puts ripples in the log spectrum, one per harmonic, and they are resolved only if the window holds about three periods: at two periods, many time windows come out an octave off. So every window must be at least 3 / f_lo long, or this raises. On a male spoken sentence (CMU ARCTIC bdl, 40 ms Hann time windows), the result agrees with WORLD’s Harvest within 5% on about 79% of the time windows Harvest calls voiced, and on 95% of those whose peak also exceeds 0.1. It is a baseline, not an F0 tracker: each time window is judged alone.

plot(ax=None, channel: int = 0, **kwargs)[source]

Cepstrum against time and quefrency in ms (see plot_cepstrum()).

sonore.views.mfcc

Mel-frequency cepstral coefficients (MFCCs): the log power in triangular bands on a mel scale, summarized by a discrete cosine transform.

mel_filterbank(n_mels: int, freqs: ndarray, f_lo: float, f_hi: float, scale: str = 'htk', triangles: str = 'height', triangle_axis: str = 'mel') → tuple[ndarray, ndarray][source]

Triangular mel weights on the frequencies freqs [Hz].

The n_mels + 2 band edges are equally spaced in mel from f_lo to f_hi; band m rises linearly from edge m to edge m + 1 and falls to edge m + 2. “Linearly” is in mel with triangle_axis="mel", as HTK and Kaldi build them, and in Hz with "hz", as librosa does; the two differ inside each band, most in the wide high bands. With triangles="height" each peak is 1, and between the first and last centers the weights sum to 1. With "area" band m is scaled by 2 / (edge[m+2] - edge[m]) (Slaney’s normalization), which after the log only adds a constant to each band. The triangles are evaluated at the exact frequencies, not rounded to FFT bins.

Returns the weights, shape (n_mels, len(freqs)), and the edges [Hz].

symmetric_hamming(n_samples: int) → ndarray[source]

The Hamming window as HTK and Kaldi define it, 0.54 - 0.46 cos(2 pi i / (n - 1)): symmetric, with both ends at 0.08 (SciPy’s "hamming" spec, as a GaborFrame samples it, is the periodic one).

delta_features(values: ndarray, order: int = 1, width: int = 5, axis: int = -1) → ndarray[source]

Time derivatives of a feature track, computed as librosa’s feature.delta does.

Along axis, a polynomial of degree order is fitted by least squares to each run of width (odd) values and its order-th derivative taken: SciPy’s savgol_filter with mode="interp", so the first and last width // 2 values use the fit to the first or last width. For order=1 and width=5 this is HTK’s regression sum(n * c[t+n]) / sum(n**2) over n = -2..2 away from the ends. order=2 is the second derivative of a local quadratic, not the delta of the delta. Units are per time window.

class MFCC(source: Sound | STFT | TVSTFT, *, n_mfcc: int = 13, n_mels: int = 26, f_lo: float = 0.0, f_hi: float | None = None, mel_scale: str = 'htk', triangles: str = 'height', triangle_axis: str = 'mel', floor_db: float = -200.0, lifter: float = 0.0, win_dur: float | None = None, hop_dur: float | None = None)[source]

Bases: View

Mel-frequency cepstral coefficients of each time window (Davis & Mermelstein, 1980).

For each time window with power spectrum P[k]: band powers E[m] = sum_k W[m, k] P[k] through n_mels triangular mel bands (mel_filterbank()), their natural log, and the orthonormal DCT-II of that, of which the first n_mfcc coefficients are kept. The DCT of a log spectrum is a cepstrum, so the low coefficients are the slow ripples of the log mel spectrum: its envelope, without the pitch.

source is a Sound or the coefficients of an analysis:

  • a Sound is analyzed with the usual speech settings: a symmetric Hamming window win_dur long (25 ms), as HTK and Kaldi use, a hop of hop_dur (10 ms), and an FFT length of the next power of two;

  • an STFT or TVSTFT is used as it is, so any window, hop or pitch-adaptive analysis can be summarized.

This is a view that discards information: the phase, everything inside each mel band, the faster ripples of the log mel spectrum beyond n_mfcc, and the level apart from c0. Many spectra have the same coefficients, and no sound is recovered from them; envelope() shows the smoothed spectrum they stand for.

The mel power is floored at floor_db below each channel’s largest band power over all time windows before the log, so scaling a sound changes only c0 and digital silence gives finite coefficients. Because the floor is set by the loudest moment, a deep floor (the default) keeps each time window independent of the rest of the sound.

The defaults follow Kaldi and HTK: HTK’s mel scale, height-1 triangles straight in mel, and, for a sound, a symmetric Hamming window. With no pre-emphasis, no DC removal and use_energy=false, Kaldi’s MFCCs (compute-mfcc-feats, here through kaldi-native-fbank) are reproduced to float32 precision; Kaldi’s own lifter, low-freq of 20 Hz, 23 bins and “povey” window (a Hann window to the power 0.85, as a callable window of the GaborFrame) are all available. Kaldi’s time windows start at the first sample (snip-edges) where sonore’s are centered on multiples of the hop, so the grids line up only after trimming the sound by half a window modulo the hop.

To reproduce librosa’s feature.mfcc (Slaney mel, area-normalized triangles straight in Hz, power in dB floored 80 dB below the loudest cell), analyze with so.GaborFrame(n_fft / fs, hop / fs, window="hann", n_fft=n_fft) and use mel_scale="slaney", triangles="area", triangle_axis="hz", n_mels=128, n_mfcc=20, floor_db=-80 and db. sonore’s grid has one more time window, centered one hop before the first sample, and one or two more at the end; the others are librosa’s. librosa also floors the power at an absolute 1e-10, which matters only for sounds far quieter than any recording, and takes the loudest cell over all channels rather than per channel.

Pre-emphasis (y[n] = x[n] - 0.97 x[n-1], as in HTK and python_speech_features) is a change to the sound, so it is applied to the sound first, for example so.Sound(scipy.signal.lfilter([1, -0.97], 1, snd.data, axis=0), snd.fs). python_speech_features also rounds the triangle feet to FFT bins and replaces c0 with the log energy of the time window; neither is done here.

Parameters:
  • source – A sound, or the STFT or TVSTFT to summarize.

  • n_mfcc – Coefficients kept, c0 included.

  • n_mels – Mel bands.

  • f_lo – The outer feet of the first and last bands [Hz]; f_hi defaults to half the sampling rate.

  • f_hi – The outer feet of the first and last bands [Hz]; f_hi defaults to half the sampling rate.

  • mel_scale – "htk" or "slaney" (see freq_to_mel()).

  • triangles – "height" (peaks of 1, as HTK and Kaldi) or "area" (Slaney’s and librosa’s area normalization).

  • triangle_axis – "mel": the triangles are straight lines in mel (HTK, Kaldi); "hz": straight in Hz (librosa).

  • floor_db – Floor of the band powers, in dB below each channel’s largest.

  • lifter – L of HTK’s sinusoidal lifter, which multiplies c[n] by 1 + (L / 2) sin(pi n / L); 0 for none. It is a fixed gain per coefficient: it changes Euclidean distances and nothing else.

  • win_dur – Window length and hop [s] used when source is a sound.

  • hop_dur – Window length and hop [s] used when source is a sound.

data

The coefficients, shape (n_channels, n_mfcc, n_windows), in natural log units (db for dB).

mel_power

Band powers before the floor, shape (n_channels, n_mels, n_windows): the mel spectrogram.

edges

Band edges [Hz], n_mels + 2 of them; cfs are the centers.

source

The STFT or TVSTFT analyzed.

discards = 'MFCC discards the phase, the detail inside each mel band, and every coefficient past the last kept, so no sound has these MFCCs alone.'
back_to_sound = 'For an approximate voice, read the envelope with mfcc.envelope_view() and pass it to so.world_synthesize with an F0 track and an aperiodicity.'
property t: ndarray[source]

Window center times [s], those of source.

property cfs: ndarray[source]

Band centers [Hz].

property db: ndarray[source]

data times 10 / ln 10, as if the band powers had been taken in dB (10 log10) before the DCT.

Type:

The coefficients in dB units

property mel_db: ndarray[source]

The floored mel spectrogram in dB, shape (n_channels, n_mels, n_windows).

deltas(order: int = 1, width: int = 5) → ndarray[source]

Time derivatives of the coefficients, per time window, computed as librosa’s feature.delta (see delta_features()); width=9 is librosa’s default. Same shape as data.

With a 10 ms hop, the first-order deltas of width 5 act as a band-pass filter on each coefficient’s track, strongest near 14 Hz: they pick out changes at the rate of syllables and phonemes.

envelope(f) → ndarray[source]

The smoothed power spectrum the kept coefficients stand for, at frequencies f [Hz], shape (n_channels, len(f), n_windows).

The lifter is undone, the missing coefficients are taken as zero, and the inverse DCT gives a smoothed log band power at each band center; between centers it is interpolated linearly in mel, and held constant beyond the first and last. It is a display of what the coefficients keep, for plotting against a spectrum or envelope, not an estimate of the spectrum: with as many coefficients as bands it returns the floored band powers at the centers, which are sums over bands, not spectral densities.

envelope_view(f=None) → GridEnvelope[source]

envelope() on this analysis’s time windows, at frequencies f (by default its FFT bins), read as env(t, f), so it goes wherever a spectral envelope is taken (warp_frequency(), world_synthesize(), harmonic_complex()). It holds band powers, sums over triangles that widen with frequency, so with triangles="height" it tilts upward against a spectral density; triangles="area" removes most of that tilt.

plot(ax=None, channel: int = 0, kind: str = 'mfcc', **kwargs)[source]

The coefficients against time (kind="mfcc") or the mel spectrogram in dB (kind="mel"); see plot_mfcc().

sonore.views.f0

F0 tracking: candidates from a difference function, refinement by the instantaneous frequency of harmonics, a periodicity score, and a Viterbi pass that decides voicing.

class F0Track(t: ndarray, f0: ndarray, score: ndarray, candidates: ndarray, candidate_scores: ndarray, fs: float, threshold: float)[source]

Bases: View

An F0 track: one estimate per channel every hop seconds. A view: it keeps only the pitch and the voicing.

f0 and score have shape (n_channels, n_windows); f0 is 0 where the time window is unvoiced, and score is the periodicity score of the chosen candidate (0 where unvoiced). candidates and candidate_scores have shape (n_channels, n_windows, 4) and hold every refined candidate the tracker weighed, NaN where a time window had fewer. t and a row of f0 go straight into pitch_adaptive() and lifter().

discards = 'F0Track keeps only the pitch and the voicing of each time window.'
back_to_sound = 'so.world_synthesize rebuilds an approximation of the voice from it together with a spectral envelope and an aperiodicity.'
t: ndarray
f0: ndarray
score: ndarray
candidates: ndarray
candidate_scores: ndarray
fs: float
threshold: float
property voiced: ndarray[source]

True where the time window is voiced, shape (n_channels, n_windows).

plot(ax=None, channel: int = 0, candidates: bool = False, **kwargs)[source]

F0 against time (see plot_f0_track()).

f0_track(sound: Sound, f_lo: float = 60.0, f_hi: float = 500.0, hop: float = 0.005, threshold: float = 0.5, *, octave_cost: float = 2.0, switch_cost: float = 0.5, subharmonic_margin: float | None = 0.05) → F0Track[source]

Track the F0 of each channel of a sound.

Four stages, every hop seconds from time 0:

  1. Candidates. Up to eight local minima of the cumulative-mean-normalized difference function (de Cheveigné & Kawahara, 2002) over a 25 ms window, at lags between 1/f_hi and 1/f_lo, each refined by a parabola. This is the autocorrelation idea with YIN’s corrections: it finds the period from all the harmonics, so it works when the fundamental is weak or filtered out. Every minimum below 1 is offered, so that the score alone decides voicing. A very regular voice has a minimum at every multiple of its period, so eight leave room for the period itself (with four, a steady 300 Hz vowel could lose it).

  2. Refinement. For a candidate f, a Blackman window three periods long; the instantaneous frequency of each of the first six harmonics is its phase advance over one sample, and the new F0 is the power-weighted mean of each harmonic’s frequency divided by its number. Done twice. This is the idea of WORLD’s StoneMask, written from its description.

  3. Score. The normalized correlation between a stretch three periods long and the same stretch one period later, the fractional part of the period done with a windowed sinc. For a periodic sound in independent noise it is about s / (1 + s), s the ratio of periodic to noise power, so 0.5 means the two are equally strong.

  4. Tracking. A Viterbi pass over (unvoiced, candidates) per time window. A candidate costs 1 - score and the unvoiced state 1 - threshold, so a time window leans voiced when its best score exceeds threshold; moving between candidates costs octave_cost per octave, and a voicing change costs switch_cost. A candidate is never chosen if another one at a whole multiple of its frequency (an octave, a twelfth, …) scores at least as well, less subharmonic_margin: anything periodic with period T is also periodic with period 2T, 3T, …, and without this rule a female voice is tracked an octave low on about 1% of time windows, and a steady synthetic vowel at 300 Hz at a third of its F0. No smoothing follows.

On synthetic vowels the refined F0 is within 0.12% of the truth on a one octave per second glide and a 5.5 Hz vibrato. Against laryngograph reference F0 (a male and a female speaker), it gets the voicing of 5.6% and 1.5% of time windows wrong, where WORLD’s Harvest, which leans toward calling time windows voiced, gets about 21% wrong; where both it and the reference say voiced, 98% and 95% of time windows are within 5%. Lower threshold to voice more weak or creaky stretches, for example before resynthesis.

Parameters:
  • sound – The sound; each channel is tracked separately.

  • f_lo – The search range [Hz]. The defaults cover speaking voices; raise f_hi for singing.

  • f_hi – The search range [Hz]. The defaults cover speaking voices; raise f_hi for singing.

  • hop – Spacing between time windows [s].

  • threshold – The periodicity score above which a time window leans voiced.

  • octave_cost – Tracking costs, as above.

  • switch_cost – Tracking costs, as above.

  • subharmonic_margin – How much lower a candidate at a whole multiple may score and still rule out a candidate. None turns the rule off, to see the raw choice.

scale_f0(contour, ratio: float, *, range: float = 1.0)[source]

An F0 contour with its pitch changed: every voiced value multiplied by ratio and, if range is not 1, spread around its median on a log scale, median * ratio * (f0 / median) ** range (range=0 is a monotone at ratio times the median, 2 doubles every interval from it). Unvoiced time windows (0) stay unvoiced.

contour is an F0Track, which comes back as an F0Track (its candidates changed the same way), a (times, f0) pair, which comes back as a pair, or anything with .t and .f0, which comes back as a (times, f0) pair. A ratio in semitones st is 2 ** (st / 12). With ratio=1 and range=1 the contour itself is returned.

The envelope is not touched, so a voice resynthesized on the new contour keeps its formants (unlike pitch_shift(), which moves them with the pitch).

sonore.views.spectral_envelope

Spectral envelopes, and moving them along frequency.

cheaptrick() is WORLD’s CheapTrick (Morise, 2015), ported exactly: every step reproduces WORLD’s C++ code (github.com/mmorise/World, commit d625e76) to floating-point precision, including the tiny noise WORLD adds to keep logarithms and divisions finite, drawn from WORLD’s own generator in WORLD’s order (sonore.views.world, where DIFFERENCES_FROM_WORLD lists every option that departs from WORLD).

An envelope here is anything read as env(t, f), giving power at times t and frequencies f with shape (n_channels, len(f), len(t)): a SpectralEnvelope (CheapTrick), a GridEnvelope (from Cepstrum.envelope_view or MFCC.envelope_view), or a function. Aperiodicity is read the same way. warp_frequency() moves an envelope (or an aperiodicity) along frequency without looking at how it was made, so any envelope and any synthesizer combine: world_synthesize() and harmonic_complex() both take any envelope.

class SpectralEnvelope(data: ndarray, t: ndarray, fs: float, q1: float)[source]

Bases: _FrequencyView

A spectral envelope from cheaptrick(): power on n_fft // 2 + 1 frequencies f from 0 to fs / 2, at the window times t of the F0 track. data has shape (n_channels, n_freqs, n_windows); to_world() gives one channel in WORLD’s layout.

It is a view (it keeps the envelope, not the sound): calling it, env(t, f), reads the power at any times and frequencies, linearly in time and in dB over frequency.

discards = "SpectralEnvelope discards the harmonics, the phase, and everything finer than the envelope's smoothing."
back_to_sound = 'so.world_synthesize rebuilds an approximation of the voice from it together with an F0 track and an aperiodicity.'
property db: ndarray[source]

10 log10 of the power.

amplitude(t, f, channel: int = 0) → ndarray[source]

The amplitude (square root of the power) at the points (t[i], f[i]), t and f of the same shape: the form harmonic_complex() reads, one value per sample and harmonic. Linear in time and in dB over frequency, held beyond the first and last time windows.

plot(ax=None, channel: int = 0, db_range: float = 70.0, fmax: float | None = None)[source]

The envelope in dB as a time-frequency image.

cheaptrick(sound: Sound, f0, *, q1: float = -0.15, f0_floor: float = 71.0) → SpectralEnvelope[source]

WORLD’s CheapTrick spectral envelope (Morise, 2015), ported exactly.

At each time window of the F0 track: the sound under a Hann window three periods long, scaled to unit energy, less its weighted mean; its power spectrum on world_fft_size() bins; the power below F0 folded back about F0 / 2; a moving average over 2 F0 / 3; then the log, liftered by sinc(F0 q) (smoothing over F0) times the recovery lifter (1 - 2 q1) + 2 q1 cos(2 pi F0 q), and back. The result follows the shape of the harmonic peaks a few dB below them and, because of the smoothing over 2 F0 / 3, barely changes with the time window’s position within a period. Time windows with F0 at or below the floor 3 fs / (n_fft - 3) (unvoiced ones included) are analyzed at 500 Hz.

Parameters:
  • sound – The sound. Its sampling rate must be a whole number of Hz.

  • f0 – An F0Track (one row per channel) or a (times, f0) pair; 0 marks unvoiced time windows. The envelope is computed at these times.

  • q1 – The recovery lifter’s parameter. -0.15 is WORLD’s code; the paper gives -0.09; 0 is no recovery.

  • f0_floor – The lowest F0 the FFT size must hold (WORLD’s 71 Hz).

class GridEnvelope(data: ndarray, t, f)[source]

Bases: View

A spectral envelope held as power on any grid of times and frequencies: data has shape (n_channels, len(f), len(t)). A view: it keeps the envelope and drops the harmonics and the phase.

Any envelope estimate can be put in this form, so that it reads like a SpectralEnvelope and goes wherever an envelope is taken: warp_frequency(), world_synthesize(), harmonic_complex(). Cepstrum.envelope_view and MFCC.envelope_view return one.

Calling it, env(t, f), reads the power at any times and frequencies, linear in time and in dB over frequency, as SpectralEnvelope does; beyond the ends of the grid the end values are held.

discards = "GridEnvelope discards the harmonics, the phase, and everything finer than the envelope's smoothing."
back_to_sound = 'so.world_synthesize rebuilds an approximation of the voice from it together with an F0 track and an aperiodicity.'
property db: ndarray[source]

10 log10 of the power.

amplitude(t, f, channel: int = 0) → ndarray[source]

The amplitude (square root of the power) at the points (t[i], f[i]), t and f of the same shape: the form harmonic_complex() reads, one value per sample and harmonic.

plot(ax=None, channel: int = 0, db_range: float = 70.0, fmax: float | None = None)[source]

The envelope in dB as a time-frequency image.

warp_frequency(view, ratio)[source]

A spectral envelope (or aperiodicity) moved along frequency: new(t, f) = view(t, f / ratio). A ratio above 1 moves every formant up by that factor, as a shorter vocal tract would; below 1, down. The F0 is not touched, so the pitch stays.

view is anything read as view(t, f): a SpectralEnvelope, an Aperiodicity, a GridEnvelope, or a function. A view held on a grid comes back as the same type on the same grid (so world_synthesize takes it as before); a function comes back as a function.

ratio is:

  • a number: one ratio for the whole sound;

  • a (times, ratios) pair: a ratio that changes over time, read at each time window (linearly between the given times, held beyond them); only for a view held on a grid;

  • a function f -> source frequency: any frequency map, giving for each new frequency the one it is read from (a piecewise or bilinear warp, for example).

Values above the view’s top frequency (when lowering) hold the top value. The level is not adjusted: a warp stretches or squeezes the envelope’s area, which changes the output level by a little (about 1.5 dB for a ratio of 0.8 on the speech example), which normalize removes. A ratio of exactly 1 returns the view itself, so a voice resynthesized with it is the same, sample for sample.

sonore.views.aperiodicity

Aperiodicity, the share of noise in a voice at each frequency: WORLD’s D4C (Morise, 2016), ported exactly, and a harmonic-residual aperiodicity that measures the share of noise directly.

Every step named after WORLD reproduces WORLD’s C++ code (github.com/mmorise/World, commit d625e76) to floating-point precision, with WORLD’s own noise in WORLD’s order; options that depart from WORLD are listed in DIFFERENCES_FROM_WORLD.

class Aperiodicity(data: ndarray, t: ndarray, fs: float, method: str)[source]

Bases: _FrequencyView

An aperiodicity, from d4c() or harmonic_aperiodicity(): per frequency and time window, how much of the power is noise rather than harmonics. Stored as WORLD stores it, an amplitude ratio between 0 and 1 (data, shape (n_channels, n_freqs, n_windows)): its square share is the share of the power that is noise, which is what world_synthesize() uses. method names the measure that made it (“D4C” or “harmonic residual”).

discards = 'Aperiodicity keeps only the share of noise in each frequency and time window: it discards the spectrum, the pitch and the phase.'
back_to_sound = 'so.world_synthesize rebuilds an approximation of the voice from it together with an F0 track and a spectral envelope.'
property share: ndarray[source]

The noise share of the power, data ** 2.

property db: ndarray[source]

The noise share in dB, 20 log10 data.

bands(edges: Sequence[float], envelope: SpectralEnvelope | None = None) → ndarray[source]

The noise share averaged over each band [edges[i], edges[i+1]), shape (n_channels, n_bands, n_windows). With an envelope, the average is weighted by its power, so the result is the band’s noise power over its total power.

plot(ax=None, channel: int = 0, db_range: float = 60.0, fmax: float | None = None)[source]

The noise share in dB as a time-frequency image (0 dB is all noise).

d4c(sound: Sound, f0, *, threshold: float = 0.85, f0_floor: float = 71.0) → Aperiodicity[source]

WORLD’s D4C aperiodicity (Morise, 2016), ported exactly.

At each voiced time window: a “static group delay” from two Blackman-windowed spectra four periods long, a quarter period either side of the window center, divided by a smoothed power spectrum, smoothed over F0 / 2, less itself smoothed over F0. Around each multiple of 3 kHz up to min(15 kHz, fs / 2 - 3 kHz), a Nuttall window over 3 kHz of that group delay is transformed, its power sorted, and the band’s aperiodicity is the share outside the largest bins, in dB, plus (F0 - 100) / 50 dB, capped at 0. The curve is then linear in dB from -60 dB at 0 Hz through those values to 0 dB at Nyquist. At 16 kHz that is one measured value per time window, at 3 kHz.

A time window is left fully aperiodic (amplitude ratio 1 - 1e-12) where F0 is 0, or where less than threshold of its power between 100 Hz and 7.9 kHz lies below 4 kHz (WORLD’s “LoveTrain” test; threshold=0 keeps every voiced time window voiced).

D4C was tuned so that WORLD’s resynthesis sounds natural; it is robust to F0 errors but does not report the share of noise below 3 kHz, which the -60 dB anchor sets. harmonic_aperiodicity() measures that share directly.

Parameters:
  • sound – As cheaptrick().

  • f0 – As cheaptrick().

  • threshold – The LoveTrain voicing threshold.

  • f0_floor – Sets the output’s frequency grid, which must match the envelope’s for synthesis (WORLD’s 71 Hz, as cheaptrick()).

harmonic_aperiodicity(sound: Sound, f0, *, periods: float = 4.0, f0_floor: float = 71.0, cell_harmonics: float = 2.0) → Aperiodicity[source]

The share of noise, measured by fitting the harmonics and keeping what is left. Not part of WORLD; d4c() is WORLD’s measure.

The F0 track is interpolated to every sample and integrated to a running phase Phi(t). At each voiced time window, under a Hann window periods periods long, a weighted least-squares fit of the harmonics cos(k Phi), sin(k Phi) below Nyquist, each with a linear change of amplitude across the window, is subtracted. The residual’s power, over the windowed sound’s power, is the noise share. The fit also absorbs a little of the noise near every harmonic; how much is known exactly from the fit (the share of white noise the residual keeps at each frequency), and is divided out. Both are summed over cells cell_harmonics harmonics wide, and the cells’ shares are interpolated in dB onto the grid of cheaptrick(). Unvoiced time windows are all noise.

The measure is what the word means, so it reads a known share of noise within a few tenths of a dB on a steady vowel; but it needs F0 to about 0.1%, since a harmonic that drifts out of phase with the fit over the window reads as noise, and so do jitter and shimmer.

Parameters:
  • sound – As cheaptrick().

  • f0 – As cheaptrick().

  • periods – The window’s length in periods of the time window’s F0.

  • f0_floor – Sets the output’s frequency grid, as cheaptrick().

  • cell_harmonics – The width of the cells, in harmonics.

sonore.views.world

WORLD’s synthesis (Morise, Yokomori & Ozawa, 2016), ported exactly: a sound from an F0 track, a spectral envelope and an aperiodicity, as made by cheaptrick() and d4c() (or harmonic_aperiodicity()).

Also here are the pieces WORLD’s analysis shares with its synthesis: WORLD’s own noise generator (world_randn()), its FFT size, MATLAB’s rounding, the reading of an F0 track onto time windows, and DIFFERENCES_FROM_WORLD. The synthesis reads its inputs by what they provide, not by their type: an F0 track as .t and .f0 or a (times, f0) pair, an envelope as env(t, f), and an aperiodicity as its grid .t, .f and .data.

world_randn(n: int) → ndarray[source]

The first n values of WORLD’s randn after randn_reseed: each is the sum of twelve xorshift128 draws (Marsaglia, 2003), shifted to 28 bits, scaled to [0, 12) and less 6, so approximately standard normal. WORLD restarts this stream at the start of every CheapTrick, D4C and Synthesis call, so the same values come back each time; they are computed once, many lanes at a time, and kept. Read-only.

world_fft_size(fs: float, f0_floor: float = 71.0) → int[source]

CheapTrick’s FFT size: 2 ** (1 + floor(log2(3 fs / f0_floor + 1))), long enough for a window three periods of f0_floor long. 1024 at 16 kHz and 2048 at 44.1 or 48 kHz with WORLD’s floor of 71 Hz.

world_synthesize(f0, envelope, aperiodicity, *, rng=None) → Sound[source]

WORLD’s synthesis, ported exactly.

Pulses are placed where the running phase of the F0 track (linearly interpolated to every sample) crosses a multiple of 2 pi, the fraction of a sample kept as a linear phase shift; unvoiced stretches get a pulse every 1/500 s. At each pulse the envelope S and the amplitude ratio A are interpolated between time windows. The periodic part is the minimum-phase response of S (1 - A^2), scaled by the square root of the interval to the next pulse, with its DC removed; the aperiodic part is noise as long as that interval through the minimum-phase response of S A^2 (of S where unvoiced). The responses are overlap-added.

Minimum phase is an approximation: the waveform within each period is not the original’s, which WORLD’s authors note is audible at low F0.

Parameters:
  • f0 – An F0 track (an F0Track or a (times, f0) pair), usually the one the envelope and aperiodicity were measured with, from any tracker. It is used as it is when its time windows are the aperiodicity’s; otherwise it is read onto them (voiced where the nearest time window is voiced, linear between voiced values). Those time windows must be evenly spaced from time 0, as WORLD assumes; the spacing is the hop (WORLD’s frame period). F0 can be changed before synthesis (scale_f0(), a pitch change); the envelope stays, so the formants stay.

  • envelope – A SpectralEnvelope on the aperiodicity’s time windows and frequencies (any envelope whose .fs, .t, .f and .data match the aperiodicity’s grid), used as it is; or any other envelope read as envelope(t, f) (power, shape (n_channels, len(f), len(t))), such as a GridEnvelope from the cepstrum or MFCCs, or one moved by warp_frequency(), which is read at the aperiodicity’s time windows and frequencies first. WORLD’s synthesis computes a minimum-phase response on its own FFT length, which suits a smooth envelope like CheapTrick’s; an envelope with deep, narrow valleys (tens of dB) comes out a few dB off at the harmonics, so smooth such an envelope first, or give it to harmonic_complex() as its amplitudes, which reads the envelope at each harmonic exactly.

  • aperiodicity – An Aperiodicity (or anything with its grid: .t, .f, .fs and .data, the amplitude ratio), which sets the time windows and frequencies.

  • rng – None (default): WORLD’s own noise stream, restarted at every call, so the same inputs give WORLD’s output, sample for sample. A seed or numpy.random.Generator: fresh Gaussian noise instead, for independent tokens (listed in DIFFERENCES_FROM_WORLD).

Returns:

int(n_windows * hop * fs) samples, one channel per channel of the envelope.

Return type:

Sound

sonore.views.phasevocoder

Phase vocoder (Gordon & Strawn, 1985; Dolson, 1986).

The phase vocoder treats each STFT channel as a slowly varying sinusoid with an amplitude and an instantaneous frequency, estimated from how much the channel’s phase advances between time windows beyond what its center frequency predicts. With amplitudes and frequencies in hand you can:

  • change duration without changing pitch (time_stretch()), by re-accumulating phase at a different hop and overlap-adding;

  • change pitch without changing duration (pitch_shift()), by stretching and then resampling;

  • go back to a sound through an oscillator bank with any frequency remapping (PVAnalysis.to_sound()), e.g. to shift or stretch partials.

Time stretching uses identity phase locking (Laroche & Dolson, 1999) by default: bins around each spectral peak keep their original phase relationship to the peak, which removes most of the classic “phasiness”.

class PVAnalysis(magnitude: ndarray, phase: ndarray, freq: ndarray, t: ndarray, fs: float, n_win: int, hop: int, n_samples: int)[source]

Bases: View

Phase-vocoder analysis: per-channel magnitude, phase and instantaneous frequency, all shaped (n_channels, n_bins, n_windows). A view: it reads each bin as one sinusoid, and to_sound() rebuilds a sound from those sinusoids, closely for tonal sounds but not exactly.

discards = 'PVAnalysis reads each STFT bin as a single sinusoid, which a sound with more than one component per bin (noise, close partials) is not, so its oscillator resynthesis is not an inverse.'
back_to_sound = 'PVAnalysis.to_sound rebuilds an approximation through an oscillator bank; so.GaborFrame gives an exact STFT.'
no_plot = 'PVAnalysis has no plot: its magnitudes are those of an STFT, so plot so.GaborFrame(...).analyze(sound).'
magnitude: ndarray
phase: ndarray
freq: ndarray
t: ndarray
fs: float
n_win: int
hop: int
n_samples: int
to_sound(time_scale: float = 1.0, freq_map: float | Callable[[ndarray], ndarray] | None = None, floor_db: float = -80.0, chunk: int = 64) → Sound[source]

Oscillator-bank resynthesis.

Each bin drives a sinusoid whose amplitude and frequency are interpolated between time windows and whose phase is the running integral of frequency. time_scale stretches the time axis; freq_map is a ratio (e.g. 1.5) or a function mapping frequencies in Hz to new frequencies. For example lambda f: f + 70 makes a 220 Hz harmonic complex inharmonic; note that shifting by a multiple of half the f0 keeps it harmonic (f + 110 gives odd harmonics of 110 Hz). Bins that never come within floor_db of the loudest bin are skipped, and partials mapped above Nyquist are dropped.

This is designed for tonal sounds, which it reconstructs closely (r > 0.999 for harmonic complexes). For noise, neighboring channels drift out of phase and partially cancel (about 2 dB low, r about 0.93); use time_stretch() for noisy material. Both figures are measured by tools/measure_docstring_numbers.py.

pv_analyze(sound: Sound, win_dur: float = 0.046, hop_dur: float | None = None) → PVAnalysis[source]

Phase-vocoder analysis with a Hann window of win_dur seconds and a hop of hop_dur (default: a quarter window, the largest hop at which instantaneous frequency is unambiguous across a Hann main lobe).

time_stretch(sound: Sound, factor: float, win_dur: float = 0.046, phase_lock: bool = True) → Sound[source]

Change duration by factor (2 = twice as long) without changing pitch.

The synthesis hop is fixed at a quarter window; the analysis hop is synthesis_hop / factor (rounded), so the achieved factor can differ slightly from the requested one for extreme values. The output length is round(len(sound) * factor).

pitch_shift(sound: Sound, semitones: float, win_dur: float = 0.046, phase_lock: bool = True) → Sound[source]

Shift pitch by semitones without changing duration (time-stretch, then resample). Formants shift along with the pitch.

sonore.views.timbre

Timbre descriptors after Peeters et al. (2011), “The Timbre Toolbox”: the log attack time, the spectral centroid and the spectral flux.

Written from the paper’s equations, not from the Timbre Toolbox’s code, whose license forbids redistribution. Where the defaults differ from the paper’s, the reason is measured in tools/check_timbre_claims.py and set out in docs/design/music/timbre-page.md.

class DescriptorTrack(t: ndarray, values: ndarray, name: str, unit: str)[source]

Bases: View

One descriptor’s value in every time window, for each channel. A view: it keeps one number per time window.

values has shape (n_channels, n_windows) and is NaN where a time window is silent. median and iqr are the summaries Peeters et al. (2011) use, since silent and near-silent time windows make a mean or a standard deviation meaningless; both skip NaN.

discards = 'DescriptorTrack keeps one number per time window: it discards the rest of the spectrum or envelope, so many sounds share one track.'
t: ndarray
values: ndarray
name: str
unit: str
property median: ndarray[source]

Median over time, one per channel.

property iqr: ndarray[source]

Interquartile range over time, one per channel.

plot(ax=None, channel: int = 0, **kwargs)[source]

The descriptor against time (see plot_descriptor_track()).

attack_segment(sound: Sound, channel: int = 0, cutoff: float = 20.0, zero_phase: bool = True) → tuple[float, float][source]

Start and end of the attack [s], by the weakest-effort method of Peeters et al. (2011), III A 2 a.

The energy envelope (the amplitude of the analytic signal, low-passed by a third-order Butterworth filter at cutoff Hz) is cut at 0.1, 0.2, …, 1 times its maximum, and the “efforts” are the times it takes to climb from one level to the next. The attack starts within the first effort at most three times the mean effort and ends within the last one, at the envelope’s minimum and maximum inside those two efforts. It measures one channel, channel, where the spectral descriptors measure every channel. The envelope filter sets how short an attack it can resolve, as the next paragraph measures.

The paper computes its descriptors with a 5 Hz filter applied once (cutoff=5, zero_phase=False), which measures a 5 ms attack as about 80 ms. The default here is the paper’s onset setting (20 Hz, forward and backward), which measures it as about 16 ms; see tools/check_timbre_claims.py.

log_attack_time(sound: Sound, channel: int = 0, cutoff: float = 20.0, zero_phase: bool = True) → float[source]

The base-10 logarithm of the attack’s duration in seconds, log10(end - start) with the start and end of attack_segment() (Peeters et al., 2011, eq. 4). One sample is the shortest duration.

spectral_centroid(sound: Sound, scale: str = 'power', win_dur: float = 0.0232, hop_dur: float = 0.0058) → DescriptorTrack[source]

The center of gravity of the spectrum in each time window, sum(f_k a_k) / sum(a_k) (Peeters et al., 2011, eq. 7), on a Hamming STFT of 23.2 ms windows every 5.8 ms by default.

scale is "power" (the default) or "magnitude", the two spectra the paper offers; they differ a lot (2.5 times for a tone with harmonics falling as 1/n), so say which one you report. The power centroid of a steady harmonic tone matches the one computed from its partials. The magnitude centroid also counts the window’s sidelobes in every bin up to Nyquist, which for a dull tone can outweigh the weak high partials: for harmonics falling as n^-3 at E-flat 4 it measures 810 Hz at 44.1 kHz and 592 Hz at 22.05 kHz, against 414 Hz from the partials (tools/check_timbre_claims.py).

spectral_flux(sound: Sound, spacing: float | None = 0.1, scale: str = 'magnitude', win_dur: float = 0.0232, hop_dur: float = 0.0058) → DescriptorTrack[source]

How much the spectrum changes: 1 minus the normalized correlation of two spectra (Peeters et al., 2011, eq. 24, “spectral variation”), 0 for spectra of the same shape and up to 1.

The paper correlates successive time windows, one hop (5.8 ms) apart: spacing=None. At that spacing a steady tone and a tone whose spectral slope glides over a second measure almost alike; spacing correlates spectra that far apart instead (100 ms by default), at which the two differ 25 times (tools/check_timbre_claims.py). Each value is placed at the later of its two time windows.