views¶
One-way analyses, each saying what it drops: spectra and spectrograms (with the reassigned spectrogram), timbre descriptors, envelopes, modulation, the cepstrum and MFCCs, F0 tracking, WORLD’s envelope and aperiodicity, and the routes back to sound they have: the channel vocoder (with the envelopes), WORLD’s synthesis and the phase vocoder.
sonore.views.view¶
Views: one-way analyses that say what they drop.
A Frame turns a sound into coefficients and
back. A view keeps only part of a sound, so many sounds give the same view
and none can be read back from it alone. Every view therefore has a
View.synthesize() that refuses, raising NotInvertibleError
with the mathematical reason (View.discards) and the route that does
lead back to a sound (View.back_to_sound), if sonore has one.
That route is to_sound, on the views that have a canonical one, and its
arguments are what the view discarded: Spectrum.to_sound takes a carrier
for the phase, Cepstrum.to_sound a phase, PVAnalysis.to_sound a time
scale and a frequency map. On any other view, View.to_sound() refuses
as synthesize does.
Every view has View.plot(). A view with one obvious picture overrides
it with a short call into sonore.plotting, where the drawing lives; a
view with none keeps the base method, which raises NotImplementedError
with View.no_plot, a sentence naming what to plot instead.
- exception NotInvertibleError[source]¶
Bases:
NotImplementedErrorRaised by
View.synthesize(), and byView.to_sound()on a view with no route back: the view keeps too little of a sound for any sound to be recovered from it alone.
- class View[source]¶
Bases:
objectBase class of every view.
Each subclass sets two class attributes:
discards, one sentence saying what the view drops and so why it cannot be inverted, andback_to_sound, one sentence naming the route to a sound that exists, or an empty string if there is none. A subclass that does not overrideplot()setsno_plot, one sentence saying why it has no picture and what to plot instead.- discards = ''¶
- back_to_sound = ''¶
- no_plot = ''¶
- synthesize(*args, **kwargs)[source]¶
Refuse: raise
NotInvertibleErrorwithdiscardsandback_to_soundas the message.
- to_sound(*args, **kwargs)[source]¶
Refuse, as
synthesize()does, on a view with no canonical route back to a sound; views that have one override it.
sonore.views.spectra¶
Spectra and spectrograms that keep power and drop the phase: the spectrum of a sound, a TANDEM-STRAIGHT-style power spectrogram, and the reassigned spectrogram.
A Spectrum’s levels are a one-sided power spectral density in dB re
1 per Hz, as Welch’s method gives it, so Spectrum.from_sound() and
long_term_spectrum() agree for a steady noise (the first scatters, the
second is smooth).
- class Spectrum(f: ndarray, level: ndarray)[source]¶
Bases:
ViewA magnitude spectrum: frequencies [Hz] and levels [dB].
A view: it keeps one level per frequency and drops the phase, and with it all timing, so no sound can be read back from it.
- discards = 'Spectrum keeps only the level at each frequency: it discards the phase, and with it all timing, so many sounds share one spectrum.'¶
- back_to_sound = "Spectrum.to_sound makes a sound with this spectrum from a carrier that supplies the phase (a new noise, another sound's phase, or the minimum phase), which is not the analyzed sound."¶
- f: ndarray¶
- level: ndarray¶
- classmethod from_sound(sound: Sound) Spectrum[source]¶
The power spectral density of the whole sound from one FFT with no window,
|X|**2 / (fs N)with every bin but 0 Hz and Nyquist counted twice (one-sided), in dB re 1 per Hz. The sound is first mixed to mono by averaging its channels, so channels in antiphase cancel. Each bin of a noise scatters about its expected level; seelong_term_spectrum()for a smooth estimate on the same scale.
- level_at(freqs: ndarray) ndarray[source]¶
Levels [dB] linearly interpolated at
freqs, held at the end values outside the spectrum’s frequencies.
- smooth(fraction: float = 0.3333333333333333) Spectrum[source]¶
Fractional-octave smoothing (power average over
fractionoctave).
- to_sound(duration: float, fs: float, carrier: str | Sound = 'noise', rng=None, **noise_kwargs) Sound[source]¶
A sound with this magnitude spectrum, its phase supplied by
carrier."noise"draws Gaussian noise with this spectral shape (extra keyword arguments go togaussian_noise()). ASoundlends its phase: its firstdurationseconds, each channel’s spectrum given this magnitude, soSpectrum.from_sound(x).to_sound(x.duration, x.fs, carrier=x)gives backxat RMS 1."minimum"gives the minimum-phase impulse response with this magnitude, from the folded cepstrum as inCepstrum.to_stft. The level is read at the FFT bins ofdurationby linear interpolation, and is silent outside the spectrum’s frequencies (a spectrum measured at 16 kHz gives nothing above 8 kHz at anyfs). Every result has RMS 1, as the generators do. None is the analyzed sound: the spectrum keeps no phase to give back.
- plot(ax=None, **kwargs)[source]¶
Level against frequency, on a log frequency axis and re the peak by default (see
plot_spectrum()).
- long_term_spectrum(sounds: Sound | Sequence[Sound], win_dur: float = 0.1) Spectrum[source]¶
Long-term average spectrum via Welch’s method, averaged over sounds (weighted by duration), in dB re 1 per Hz. This is the right input for speech-shaped noise.
win_dur[s] is Welch’s segment length (Hann windows, half overlapping); the frequency spacing is1 / win_dur, 10 Hz by default. A sound shorter thanwin_duris one segment, its own length, zero-padded towin_dur, which puts it on the same frequencies without changing its spectrum. Sounds at another rate are resampled to the first one’s, and each is mixed to mono by averaging its channels.
- class TFPower(power: ndarray, t: ndarray, f: ndarray)[source]¶
Bases:
ViewA time-frequency power that is not a frame’s coefficients, so it has no synthesis: for example
tandem_power(). It keeps the power of each cell and drops the phase.powerhas shape(n_channels, n_freqs, n_windows), on frequenciesf[Hz] and window timest[s], which need not be uniform.- discards = "TFPower keeps only the power of each cell: it discards the phase, and the cells are not a frame's coefficients, so no synthesis undoes them."¶
- back_to_sound = ''¶
- power: ndarray¶
- t: ndarray¶
- f: ndarray¶
- plot(ax=None, channel: int = 0, **kwargs)[source]¶
Power in dB (see
plot_tf_db()).
- tandem_power(sound: Sound, f0_times: Sequence[float], f0: Sequence[float], periods: float = 2.5, overlap: float = 4, window: str | tuple = 'blackman') TFPower[source]¶
A TANDEM-STRAIGHT-style power spectrogram: the average of two pitch-adaptive spectrograms whose windows sit a quarter period before and after each window center.
For a periodic sound, the power through a window centered at
tfluctuates with period T0 as the window slides across the glottal pulses. Averaging the powers att - T0/4andt + T0/4(half a period apart) cancels the odd harmonics of that fluctuation, including the largest, so a short window shows the spectral envelope steadily. The defaults, a Blackman window 2.5 periods long, are those of Kawahara et al. (2011). Only this averaging is implemented, not TANDEM-STRAIGHT’s smoothing or aperiodicity analysis.The schedule is
pitch_adaptive()over the sound, with F0 bridged across unvoiced stretches. The result is magnitude only: synthesizing from the pair would need a union of two frames, which sonore does not provide.
- class ReassignedSpectrogram(t_hat: ndarray, f_hat: ndarray, power: ndarray, keep: ndarray)[source]¶
Bases:
ViewSpectrogram cells moved to their reassigned times and frequencies.
t_hat,f_hatandpowerhave shape(n_channels, n_freqs, n_windows), one entry per STFT cell;keepmarks the cells within the threshold of the maximum.binned()sums the kept power onto a grid for display. There is no synthesis: reassignment is not linear.- discards = "ReassignedSpectrogram discards the phase and moves each cell's power to a new time and frequency, many cells to one point, so different sounds give the same picture."¶
- back_to_sound = ''¶
- no_plot = 'ReassignedSpectrogram has no plot: its points lie off any grid until one is chosen, so plot binned(t_edges, f_edges).'¶
- t_hat: ndarray¶
- f_hat: ndarray¶
- power: ndarray¶
- keep: ndarray¶
- reassigned_spectrogram(sound: Sound, frame: GaborFrame, threshold_db: float = -60.0) ReassignedSpectrogram[source]¶
The reassigned spectrogram (Kodera et al., 1978; Auger & Flandrin, 1995).
Each cell of
frame’s spectrogram is moved from its window timetand bin frequencyftot_hat = t + Re(X_tw conj X) / |X|**2,f_hat = f - Im(X_dw conj X) / (2 pi |X|**2),
where
Xuses the windoww,X_twthe time-weighted windowtau w(tau)andX_dwits derivativew'(tau), withtauin seconds from the window’s middle sample, which is where SciPy (and soGaborFrame) references each time window’s phase. A tone off the bin grid, an impulse and a linear chirp land on their true frequency, time and instantaneous-frequency line.The derivative is taken from the window’s formula, so only
"hann"and("gaussian", std)windows are accepted,stdin samples as in SciPy. Cells more thanthreshold_dbbelow the maximum (per channel) are marked not kept: their positions are mostly noise.
sonore.views.envelopes¶
Envelopes: the slowly varying amplitudes that shape sounds.
An envelope is not a sound. It is non-negative, it usually varies slowly enough to live at a low sampling rate, and you don’t listen to it: you apply it to a sound. The types reflect that:
Envelopeis one envelope.Envelope * Soundmodulates the sound.Envelopesis one envelope per frequency band of a filterbank, i.e. a spectrotemporal envelope. This is conceptually the same thing as a cochleagram (the envelope of each cochlear-filter output over time); here the “cochlea” is whicheverFilterbank(by default an ERBcosine_filterbank()) produced it.Envelopes * Subbandsmodulates each band.
The Hilbert decomposition of a band is then literal:
band == band.envelope() * (band / band.envelope()) # envelope x fine structure
sb == sb.envelopes() * sb.tfs() # for every band at once
Envelopes resampled to a lower rate are upsampled automatically (band-limited: polyphase, or the FFT for unusual rate ratios; clipped at zero) when applied to a sound.
to_sound(carrier) is the route back: the envelopes imposed on a carrier.
channel_vocode() is a channel vocoder: a sound’s own band envelopes imposed
on a carrier. With bands of noise as the carrier (the default) it is the
noise vocoder of Shannon et al. (1995), as in simulations of cochlear-implant
hearing.
- class Envelope(data, fs: float)[source]¶
Bases:
ViewA single (possibly multichannel) envelope, shape
(n_samples, n_channels). A view: it keeps a magnitude over time and drops the fine structure.Arithmetic:
*,/and+with numbers and other Envelopes (so1 + 0.5 * envworks), andenv * snd/snd * envmodulate aSound.snd / envdivides the envelope out of a sound. A one-channel envelope applies to every channel of a sound; otherwise the channel counts must match.to_sound()puts it on another sound’s fine structure.- discards = 'Envelope discards the fine structure: only a magnitude over time is kept.'¶
- back_to_sound = "Envelope.to_sound puts it on the fine structure of a carrier sound, which is not the analyzed sound's fine structure."¶
- lowpass(cutoff: float, order: int = 4) Envelope[source]¶
Zero-phase Butterworth lowpass at
cutoffHz (result clipped at 0).
- resample(fs: float) Envelope[source]¶
The envelope at rate
fs: band-limited, clipped at zero, with the ends extended along a line rather than dragged toward zero.
- to_sound(carrier: Sound) Sound[source]¶
This envelope on the fine structure of
carrier: the carrier is divided by its own Hilbert envelope (carrier / carrier.envelope()) and multiplied by this one.carriermust last as long as the envelope; it sets the sampling rate. Plainenvelope * soundkeeps the sound’s own envelope as well (amplitude modulation).
- plot(ax=None, **kwargs)[source]¶
The envelope against time (see
plot_envelope()).
- class Envelopes(data, fs: float, filterbank: Filterbank, pad: int = 0)[source]¶
Bases:
_PaddedBands,ViewOne envelope per band of a filterbank: a spectrotemporal envelope.
Conceptually this is a cochleagram: the envelope of each filter’s output over time. Shape
(n_samples, n_bands, n_channels); the bands include the filterbank’s lowpass and highpass edge filters, as inSubbands.A view: it keeps each band’s Hilbert magnitude and drops the fine structure (and, with
lowpass, the envelope’s own fast detail).env[i]is anEnvelope;env * subbandsmodulates each band (imposing the envelopes on a carrier, not an inverse);env.modulation_spectrum()gives its 2-D modulation spectrum on the filterbank’s frequency scale (cycles/octave or cycles/ERB).Envelopes derived from padded Subbands carry the same zero-padding (
padsamples at each end, hidden fromdata), so that modulating and re-synthesizing bands can’t wrap around. Envelopes made from scratch (e.g. a rendered ripple pattern) havepad=0and are zero outside their own extent when combined with padded bands.- discards = 'Envelopes discard the fine structure: only the Hilbert magnitude of each band is kept, smoothed further when a lowpass was asked for.'¶
- back_to_sound = 'Envelopes.to_sound puts them on the fine structure of a carrier (a new noise, tones at the band centers, or another sound) and synthesizes the bands, as the noise vocoder does.'¶
- lowpass(cutoff: float, order: int = 4) Envelopes[source]¶
Zero-phase Butterworth lowpass of every band at
cutoffHz, padding included (result clipped at 0).
- resample(fs: float) Envelopes[source]¶
The envelopes at rate
fs, padding included, asEnvelope.resample(). When the rate ratio is a small fraction the inner signal starts exactly on a sample; otherwise within half of one.
- without_edges() Envelopes[source]¶
Zero the lowpass and highpass edge bands (which lie outside
f_lo..f_hi), keeping the band count unchanged. A bank built withedges=Falsehas none, and its envelopes are returned as they are.
- modulation_spectrum(scale: str = 'linear', drop_edges: bool = True) ModulationSpectrum[source]¶
2-D modulation spectrum (temporal Hz x spectral cycles per scale unit).
scale="db"transforms log envelopes. The edge bands are dropped by default, since they aren’t evenly spaced with the others; a bank built withedges=Falsehas none to drop, and all its bands are kept.
- to_sound(carrier: Sound | str = 'noise', fs: float | None = None, rng=None) Sound[source]¶
These envelopes imposed on a carrier, band by band, and synthesized with their filterbank.
"noise": each band of a Gaussian noise (fromrng), scaled to a mean-square envelope of 1. The bands keep their own random envelope fluctuations, as in the classic noise vocoder."tone": a cosine at each band’s center.a Sound lasting as long as the envelopes, analyzed with the same filterbank. Only its fine structure is used, so this is
Envelope.to_sound()in every band.
fsis the output rate for"noise"and"tone"(by default the envelopes’ own); a Sound carrier sets its own. The edge bands are kept (seewithout_edges()) and the level is the envelopes’ own.
- plot(ax=None, **kwargs)[source]¶
The envelopes as a cochleagram, time by band in dB (see
plot_envelopes()).
- channel_vocode(sound: Sound, n_bands: int = 16, f_lo: float = 80.0, f_hi: float = 8000.0, carrier: Sound | str = 'noise', env_lowpass: float | None = 50.0, rng=None) Sound[source]¶
A channel vocoder with Hilbert envelopes: the sound’s band envelopes on a
carrier. The default,carrier="noise", is the noise vocoder of Shannon et al. (1995).The sound is split by a
cosine_filterbankofn_bandsbands fromf_lotof_hi(capped at Nyquist), each band’s Hilbert envelope is lowpassed atenv_lowpassHz, and the envelopes are imposed on thecarrierwithEnvelopes.to_sound():"noise"(bands of noise),"tone"(sinusoids at the band centers), or a Sound at the same rate and at least as long, whose fine structure carries them. The edge bands (outsidef_lo..f_hi) are silenced. The output matches the input’s RMS.
sonore.views.mask¶
Time-frequency masks: gains on a frame’s coefficients.
A mask keeps one gain per coefficient and nothing of any sound, so it is a
View. The way back to a sound goes through the
coefficients it multiplies: (stft * mask).to_sound() is the frame’s
least-squares inverse of the masked coefficients. A mask works with
STFT, TVSTFT and
Subbands.
The ideal binary mask is the goal Wang (2005) proposed for computational
auditory scene analysis: keep the coefficients where the target is stronger
than the masker by more than a local criterion lc_db, the parameter
whose effect on listeners Brungart, Chang, Simpson & Wang (2006) measured.
Ratio masks, soft gains from the target’s share of the power, go back to
Srinivasan, Roman & Wang (2006); the ideal ratio mask
(T**2 / (T**2 + M**2))**beta with beta = 0.5 is the form Wang,
Narayanan & Wang (2014) found best, close to a square-root Wiener filter.
For complex coefficients the power of a cell is |X|**2. Subbands are
real and oscillate, so their power is the Hilbert envelope squared: the mask
then follows each band’s level, not its waveform.
- class Mask(values, like: STFT | TVSTFT | Subbands)[source]¶
Bases:
ViewGains, usually between 0 and 1, one per coefficient of the coefficients
likeit was made for (same frame, same grid).mask * coefsandcoefs * maskgive the masked coefficients;.to_sound()on those is the way back to a sound.valueshas the layout of the coefficients’ own array:(n_channels, n_freqs, n_windows)for an STFT,(n_samples, n_bands, n_channels)for Subbands (their padding included).- discards = "Mask holds gains for a frame's coefficients, not the coefficients themselves: it keeps no level or phase of any sound."¶
- back_to_sound = 'Multiply coefficients by it and go back from those: (stft * mask).to_sound().'¶
- property t: ndarray[source]¶
Times [s] of the coefficients’ columns (time windows, or samples for subbands).
- property f: ndarray[source]¶
FFT bins, or the filterbank’s center frequencies for subbands.
- Type:
Frequencies [Hz] of the coefficients’ rows
- apply(coefs: STFT | TVSTFT | Subbands) STFT | TVSTFT | Subbands[source]¶
The coefficients multiplied by this mask (what
coefs * maskdoes).
- plot(ax=None, channel: int = 0, **kwargs)[source]¶
Draw one channel’s gains as an image on the coefficients’ time and frequency axes; see
sonore.plotting.plot_mask().
sonore.views.modulation¶
Modulation filterbanks: bandpass filters applied to envelopes.
Two banks from McDermott & Simoncelli (2011), usable on their own (e.g. for
Dau-style modulation models) and by sonore.texture:
ConstantQModulationFilterbank: half-cycle cosines on a linear frequency axis, each spanningcf +/- cf/Q, with centers log-spaced. Used for the modulation power spectrum.OctaveModulationFilterbank: half-cycle cosines on a log2 axis, each two octaves wide (cf/2 .. 2*cf), centers one octave apart. Used for the C1/C2 modulation correlations.
Filtering is done in the frequency domain and is therefore circular: the envelope is treated as one period of a periodic signal. That is exact for texture synthesis (seamless loops); for other uses pad the envelope first.
A third bank, HannModulationFilterbank, is defined in time instead:
Hann-windowed complex exponentials of finite length, applied by direct
(zero-padded, not circular) correlation, centered or causal. It is the bank
behind ModulationSpectrogram, and
its finite kernels are what make a causal, block-by-block version possible.
- class ModulationFilterbank[source]¶
Bases:
objectBase class. Subclasses define
n_bands,cfsandresponse().Despite the name, a modulation filterbank is not a
Filterbank: it filters envelopes to make views such asModulationSpectrogram, and has no synthesis.- filter(x, fs: float, axis: int = 0, analytic: bool = False) ndarray[source]¶
Filter
x(circularly) alongaxiswith every band.The band index is appended as a new last axis. With
analytic=Truethe output is the complex analytic signal of each band (negative frequencies removed, positive ones doubled), whose real part equals the ordinary filtered output.
- class ConstantQModulationFilterbank(n_bands: int = 20, f_lo: float = 0.5, f_hi: float = 200.0, Q: float = 2.0)[source]¶
Bases:
ModulationFilterbankn_bandsconstant-Q filters with centers log-spaced fromf_lotof_hi[Hz]. Filterkis a half-cycle cosine on a linear frequency axis, nonzero oncf_k*(1 - 1/Q) .. cf_k*(1 + 1/Q).Responses are scaled so that the summed squared response, averaged over the middle of the bank (between the 4th and the 4th-from-last center), is 1. That makes band powers sum to about the envelope’s variance.
- n_bands: int = 20¶
- f_lo: float = 0.5¶
- f_hi: float = 200.0¶
- Q: float = 2.0¶
- class OctaveModulationFilterbank(n_bands: int = 7, f_hi: float = 100.0)[source]¶
Bases:
ModulationFilterbankn_bandsfilters with centers one octave apart, the highest atf_hi[Hz] (default: 1.5625, 3.125, …, 100 Hz). Filterkis a half-cycle cosine on log2 frequency spanningcf_k/2 .. 2*cf_k. Peak gain is 1 (the C1/C2 correlations are scale-invariant, so no further normalization is applied).- n_bands: int = 7¶
- f_hi: float = 100.0¶
- class HannModulationFilterbank(f_lo: float = 0.5, f_hi: float = 64.0, per_octave: float = 2, cycles: int = 3, window: float | None = None)[source]¶
Bases:
ModulationFilterbankHann-windowed complex exponentials, defined in time.
Band
kcorrelates the envelope withh_k[j] = w_k[j] exp(i 2 pi f_k (j - (L_k - 1)/2) / fs), wherew_kis a Hann window ofL_ksamples normalized to sum 1. Two ways to set the window:cycles(the default, 3): every window holds that many cycles of its own rate,L_k = cycles * fs / f_k, with centersper_octaveto the octave fromf_lotof_hi. A constant-Q bank (Q = cycles / 1.44, about 2.1 for 3 cycles), long windows for slow rates and short ones for fast rates.window[s]: every band uses the same window T, with centers on the linear gridk / Tfrom the first one at or abovemax(f_lo, 2/T)tof_hi. This is the STFT of the envelope.
Either way each window holds a whole number (at least 2) of cycles of its rate, so the bank ignores the envelope’s mean: the Hann window’s transform is zero at every whole bin from 2 on.
Unlike the other two banks,
filter()works in time and is not circular: the envelope is taken to be zero outside its extent. The analytic output2 yhas magnitudeA * mfor an envelope componentA * m * cos(2 pi f_k t), and its real part is the real filtered output, whose frequency response isresponse().- f_lo: float = 0.5¶
- f_hi: float = 64.0¶
- per_octave: float = 2¶
- cycles: int = 3¶
- window: float | None = None¶
- check_fs(fs: float) None[source]¶
The envelope rate must be at least 3 * f_hi: the top band reaches about 1.25 * f_hi, and its window must hold enough samples.
- kernels(fs: float) list[tuple[ndarray, ndarray]][source]¶
(h_k, w_k)for every band atfs: the complex kernel and its normalized Hann window.
- response(freqs) ndarray[source]¶
Magnitude response of the real filter (the real part of
filter()’s analytic output), shape(len(freqs), n_bands), for continuous-time windows. Peak gain is about 1; zero at 0 Hz.
- filter(x, fs: float, axis: int = 0, analytic: bool = False, align: str = 'center') ndarray[source]¶
Filter
xalongaxiswith every band (zero outsidex).The band index is appended as a new last axis.
align="center"puts each window’s middle on the output sample;"causal"ends it there, which is the centered output delayed by(L_k - 1) // 2samples and is what a live, block-by-block analysis would produce.
- class ModulationSpectrum(stft: STFT, channel: int = 0)[source]¶
Bases:
View2-D Fourier transform of a time-frequency envelope (Singh & Theunissen, 2003; Chi et al., 1999).
ModulationSpectrum(stft)uses a dB spectrogram, so spectral modulation is w.r.t. linear frequency (cycles/kHz).octave()uses subband envelopes on a log-frequency axis (cycles/octave), the axis on which ripples (sonore.sources.ripples) are defined; more generally, anyEnvelopeshas.modulation_spectrum().Sign convention: a ripple
sin(2*pi*(rate*t + density*x))appears at(+rate, +density). Only non-negative spectral modulations are stored (the other half is the complex conjugate).A view: before the transform the envelope’s mean is removed and a Hann taper applied in time, and only the level (dB, floored) is kept, so the phase of the modulations is dropped and no envelope can be read back.
A spectrum made from
Envelopes(byoctave()orenv.modulation_spectrum()) also keeps, forto_sound(), the magnitude of the untapered transform on the whole plane, the envelope’s mean, and how the envelopes were made. Edit it withwith_gain(), and hear it withto_sound(), whose carrier supplies the phases it lacks (seedocs/design/views/modulation-targets.md).- discards = 'ModulationSpectrum keeps only the magnitude of the 2-D Fourier transform of an envelope: it discards the phase of the modulations, which holds the timing of every event, and the fine structure under the envelope.'¶
- back_to_sound = "ModulationSpectrum.to_sound(carrier=...) takes both from a carrier: a Sound lends its own modulation phase and fine structure, 'tones' and 'noise' draw a random modulation phase."¶
- classmethod from_array(env: ndarray, dt: float, dx: float, spectral_unit: str) ModulationSpectrum[source]¶
From a (frequency, time) envelope array sampled every
dtseconds anddxscale units.
- classmethod octave(sound: Sound, bands_per_octave: float = 12, f_lo: float = 125.0, f_hi: float = 8000.0, env_fs: float = 1000.0, scale: str = 'linear') ModulationSpectrum[source]¶
Modulation spectrum on a log-frequency axis [cycles/octave].
Shorthand for:
fb = cosine_filterbank(f_lo=f_lo, f_hi=f_hi, spacing=1 / bands_per_octave, scale="octave") fb.analyze(sound.mono()).envelopes(fs=env_fs).modulation_spectrum(scale)
- classmethod from_blobs(blobs, duration: float, f_lo: float = 250.0, f_hi: float = 8000.0, bands_per_octave: float = 12, env_fs: float = 1000.0, rms_depth: float = 0.2) ModulationSpectrum[source]¶
A target spectrum drawn in code: the sum of
ModulationBlobpowers, on the gridoctave()would measure for a sound ofdurationseconds, ready forto_sound().A drawing sets a shape, not a depth, and no long-term spectrum: every band gets the same mean envelope, and
rms_depthscales the modulation (the rms of the envelope array about its mean, relative to the mean). Because envelopes cannot go below zero, a random draw reaches only a limited depth: about 0.28 for a one-blob target indocs/design/views/modulation-targets.md(C2), against 0.71 for one full ripple.to_envelopes()refuses a draw that would need clipping and names the largest depth that fits it.levelshows the drawn power itself (no taper).
- with_gain(gain) ModulationSpectrum[source]¶
A copy with every cell’s amplitude multiplied by
gain.gainis a functiong(rate, density)of rate [Hz] and spectral modulation, evaluated withrateas a row anddensityas a column (as ripple patterns are), or a number. A sound’s spectrum is the same at(rate, density)and(-rate, -density), so the gain is averaged over each such pair;g = lambda r, d: abs(r) <= 4removes every modulation faster than 4 Hz,lambda r, d: r * d <= 0every downward sweep. The displayedlevel(Hann-tapered) gets the same gain, which is exact where the gain is smooth across the taper’s rate resolution.
- to_envelopes(carrier: Sound | None = None, rng=None) Envelopes[source]¶
The envelopes this spectrum describes, given a modulation phase.
The stored magnitudes and mean fix everything but the phase of the 2-D transform, which holds when each event happens and how the bands line up. A
Soundcarrierlends the phase of its own envelopes (analyzed as this spectrum’s were), so the spectrum ofxwithcarrier=xgives backx’s envelopes; with no carrier the phase is drawn at random fromrngat every nonzero rate, while the zero-rate column keeps its own phase, so each band keeps its long-term level (the long-term spectrum). Envelopes rebuilt from a linear-scale spectrum can go below zero; they are clipped, with a warning saying how much, which changes their spectrum. Edge bands that the spectrum dropped are zero.
- to_sound(carrier: str | Sound = 'tones', fs: float | None = None, rng=None, iterations: int = 0) Sound[source]¶
A sound whose envelopes have this modulation spectrum.
A modulation spectrum lacks two kinds of phase, and the
carriersupplies both:the modulation phase (when each event happens, how the bands line up; see
to_envelopes()): aSoundlends its own,"tones"and"noise"draw a random one fromrngthat keeps each band’s long-term level;the fine structure under each band’s envelope: a Sound’s own (
(envelopes * subbands.tfs()).to_sound(), the vocoder’s route), narrowband noise ("noise", the same route), or a steady tone at each band’s center ("tones"), added without re-filtering.
to_sound(carrier=x)on the spectrum ofxitself rebuildsx’s envelopes, so an edit made withwith_gain()keeps the sound’s timing wherever the gain is 1.fsis the audio rate, needed unless the carrier is a Sound, which must be at least as long as the analyzed envelopes. The result has RMS 1.A fine structure that fluctuates within a band (a sound’s own, or noise) adds modulation of its own when the result is analyzed again, so an edit survives best on
"tones": removing every rate above 4 Hz from a sentence leaves about 15 dB less power at 6-40 Hz on tones, but only 3-5 dB less on noise or the sentence’s own fine structure (docs/design/views/modulation-targets.md, C6).iterationsthen searches for a sound whose own envelopes come closer to the spectrum, as Griffin & Lim (1984) do for a spectrogram: analyze the sound, keep its fine structure and its modulation phase, impose the stored magnitudes again, and go back. On the same sentence edit, 20 iterations from the sentence’s own fine structure leave 16.5 dB less power at 6-40 Hz (C7 of the design note); each iteration costs one analysis and one synthesis.
- peak(exclude_dc: bool = True) tuple[float, float][source]¶
(temporal Hz, spectral)coordinates of the largest component.At zero spectral modulation,
+rateand-rateare mirror images (the envelope is the same at every frequency, so it has no direction), and the rate is reported as non-negative.
- plot(ax=None, **kwargs)[source]¶
The modulation spectrum as an image in dB (see
plot_modulation_spectrum()).
- class ModulationBlob(rate: float, density: float, rate_width: float = 0.5, density_width: float = 0.25, level: float = 0.0)[source]¶
Bases:
objectA patch of modulation power, for drawing a target spectrum in code (
ModulationSpectrum.from_blobs()).A Gaussian bump centered on
rate[Hz] anddensity[cycles/octave]:rate_widthis its standard deviation in octaves of rate,density_widthin cycles/octave, andlevelits peak power [dB] relative to the other blobs. The signs followRipple: positive rate and density are downward sweeps, a negative rate upward ones, and a blob at density 0 is temporal modulation alone. A ripple is a blob of zero width.- rate: float¶
- density: float¶
- rate_width: float = 0.5¶
- density_width: float = 0.25¶
- level: float = 0.0¶
sonore.views.modspectrogram¶
The modulation spectrogram: how strongly each band’s envelope is modulated at each rate, time window by time window.
An STFT shows how a sound’s power spectrum changes over time; a
ModulationSpectrogram shows how its modulation spectrum changes over
time. It is built from Envelopes (any
filterbank) by passing every band’s envelope through a
HannModulationFilterbank and sampling the
result every hop seconds. docs/design/views/modulation-spectrogram.md has the
design and the numbers behind it.
- class ModulationSpectrogram(envelopes: Envelopes, f_lo: float = 0.5, f_hi: float = 64.0, per_octave: float = 2, cycles: int = 3, window: float | None = None, hop: float = 0.01, align: str = 'center', bank: HannModulationFilterbank | None = None)[source]¶
Bases:
ViewModulation power, local mean and depth for every time window, acoustic band and modulation band.
For band envelope
e_band modulation bandkwith complex kernelh_kand windoww_k(seeHannModulationFilterbank), at each time window:y = sum_j e_b[n + j - c] conj(h_k[j]) (c: centered or causal) mean = sum_j e_b[n + j - c] w_k[j]
and the stored arrays are
power = |2 y|**2andmean, both of shape(n_channels, n_bands, n_mod, n_windows).depthis2 |y| / mean: a sinusoidal AM of depthmat the band’s rate readsm, so 100% modulation is 0 dB. Power says how much of the sound a modulation is; depth says how modulated the band is.The acoustic axis
fis the envelopes’ filterbank without its edge bands,fmholds the modulation rates andtthe window times.valid((n_bands, n_mod, n_windows)) is False where a number can’t be trusted: where the modulation rate exceeds the band’s -3 dB width (a band’s envelope can’t move faster than the band is wide), and where the window runs past either end of the envelopes (outside them the envelope is taken to be zero, so an abrupt start reads as modulation).The representation is not invertible: the phase of
yis dropped, as a magnitude spectrogram drops an STFT’s phase (the local mean is kept, anddepthdivides by it). For modulation filtering, filter the envelopes with a modulation bank directly.- Parameters:
envelopes – Band envelopes, for example
fb.analyze(snd).envelopes(fs=1000). Their rate must be at least 3 times the highest modulation rate.f_lo – The modulation bank (see
HannModulationFilterbank):cyclessets a window of that many cycles of each band’s rate (constant Q);window[s] one window for every rate (the STFT of each envelope).f_hi – The modulation bank (see
HannModulationFilterbank):cyclessets a window of that many cycles of each band’s rate (constant Q);window[s] one window for every rate (the STFT of each envelope).per_octave – The modulation bank (see
HannModulationFilterbank):cyclessets a window of that many cycles of each band’s rate (constant Q);window[s] one window for every rate (the STFT of each envelope).cycles – The modulation bank (see
HannModulationFilterbank):cyclessets a window of that many cycles of each band’s rate (constant Q);window[s] one window for every rate (the STFT of each envelope).window – The modulation bank (see
HannModulationFilterbank):cyclessets a window of that many cycles of each band’s rate (constant Q);window[s] one window for every rate (the STFT of each envelope).hop – Spacing between time windows [s], independent of the window, as in an STFT.
align –
"center"(windows centered on the window times) or"causal"(windows ending at them: what a live analysis would see).bank – A ready-made bank, instead of
f_lo..window.
- discards = "ModulationSpectrogram discards the fine structure under the envelopes and the phase of every modulation band, as a magnitude spectrogram drops an STFT's phase."¶
- back_to_sound = ''¶
- property depth: ndarray[source]¶
Modulation depth
2 |y| / mean, same shape aspower; NaN where the window holds only digital silence.
- at(t: float, kind: str = 'depth') ndarray[source]¶
The acoustic band x modulation rate image at the time window nearest
t[s], shape(n_channels, n_bands, n_mod): Atlas and Shamma’s joint acoustic and modulation frequency display at one moment.kindis"depth"or"power".
- average(kind: str = 'power') ndarray[source]¶
Average over time windows, shape
(n_channels, n_bands, n_mod). The average power is the per-band modulation power spectrum, the same kind of quantity as the texture statistics’mod_power.
- pooled_depth() ndarray[source]¶
Depth pooled over acoustic bands, shape
(n_channels, n_mod, n_windows):sqrt(sum_b power) / sqrt(sum_b mean**2)over the bands whose cell isvalid, which weights bands by their level. NaN where no band is valid.
- plot(ax=None, band: float | None = None, rate: float | None = None, **kwargs)[source]¶
Depth in dB as an image, invalid cells in gray. By default, modulation rate against time, pooled over bands (
pooled_depth());band=[Hz] shows the acoustic band nearest that frequency instead;rate=[Hz] shows acoustic band against time at the modulation band nearest that rate.
sonore.views.cepstrum¶
The real cepstrum of a short-time Fourier transform: liftering, resynthesis with the original or minimum phase, and cepstral F0.
- class Cepstrum(coefs: STFT | TVSTFT, floor_db: float = -200.0)[source]¶
Bases:
ViewThe real cepstrum of each time window of an
STFTorTVSTFT.For a time window with spectrum
X[k]onn_fftbins, the cepstrum isc[n] = IDFT(ln |X[k]|)at quefrencyn / fsseconds. It is real and even innbecause the sound is real, so onlyn = 0 .. n_fft // 2is stored:datahas shape(n_channels, n_fft // 2 + 1, n_windows)on quefrenciesq[s] and window timest[s].The natural log is used, so
expundoes it exactly: without liftering,to_sound()gives back the analyzed sound. Scaling the sound changes onlyc[0]. Before the log, magnitudes are floored atfloor_dbbelow the channel’s largest magnitude over all time windows, so a time window of digital silence has a flat log spectrum rather than-inf; real recordings stay far above the default. Only the real cepstrum is provided: the complex cepstrum needs phase unwrapping, and the minimum phase comes from the real one.- Parameters:
coefs – The coefficients to take the cepstrum of. They are kept (as
source) for their phase, frame and times.floor_db – The floor, in dB below each channel’s maximum.
- discards = 'Cepstrum discards the phase: its coefficients are the transform of the log magnitude alone, and only the source STFT it keeps holds the phase.'¶
- back_to_sound = "Cepstrum.to_sound borrows the phase of the STFT it was computed from, and gives the sound back exactly only for an unliftered cepstrum with phase='original'."¶
- lifter(cutoff: float | Sequence[float], keep: str = 'low') Cepstrum[source]¶
A rectangular lifter.
keep="low"keeps the quefrencies belowcutoff[s], the smooth spectral envelope;"high"keeps the rest, the fine structure such as the harmonics.cutoffis one value or one per time window, so it can follow an F0 track (half a period separates the envelope from the harmonics).
- envelope() ndarray[source]¶
exp(DFT(c)): the magnitude spectrum this cepstrum stands for, shape(n_channels, n_freqs, n_windows)on the source’s frequencies. After a low lifter it is the cepstral spectral envelope, which follows the shape of the true envelope but sits a few dB below the harmonic peaks, because the lifter averages the peaks with the dips between them.
- envelope_view() GridEnvelope[source]¶
envelope()as power on this cepstrum’s time windows and frequencies, read asenv(t, f), so it goes wherever a spectral envelope is taken (warp_frequency(),world_synthesize(),harmonic_complex()). Lifter first:cep.lifter(0.5 / f0).envelope_view().
- to_stft(phase: str = 'original') STFT | TVSTFT[source]¶
Coefficients of the source’s type and frame with
envelope()as magnitude.phase="original"takes the phase fromsource, so an unliftered cepstrum returns the source’s coefficients."minimum"uses the minimum phase for that magnitude, from the folded cepstrum (c[0]kept,2 c[n]up ton_fft / 2, zero beyond). Each time window’s response then starts at its phase reference, the middle of its window. The fold is exact up to the time aliasing of the cepstrum onn_fftbins, which is negligible oncen_fftis several times the response’s length.
- to_sound(phase: str = 'original') Sound[source]¶
to_stft()synthesized by the source’s frame: exact for an unliftered cepstrum with the original phase, as long as no magnitude fell belowfloor_db(those are raised to the floor), and otherwise the least-squares signal for those coefficients. It is not an inverse of the cepstrum alone: the phase comes from the source STFT.
- f0(f_lo: float = 75.0, f_hi: float = 400.0, threshold: float = 0.1) tuple[ndarray, ndarray, ndarray][source]¶
Classic cepstral F0 (Noll, 1967): the largest cepstral peak between quefrencies
1/f_hiand1/f_lo, refined by a parabola through it and its neighbors.Returns
(t, f0, peak): window times [s], and F0 [Hz] and the peak’s height for each channel and time window, shape(n_channels, n_windows). F0 is 0 where the peak is belowthreshold, a crude voicing rule.A periodic sound puts ripples in the log spectrum, one per harmonic, and they are resolved only if the window holds about three periods: at two periods, many time windows come out an octave off. So every window must be at least
3 / f_lolong, or this raises. On a male spoken sentence (CMU ARCTICbdl, 40 ms Hann time windows), the result agrees with WORLD’s Harvest within 5% on about 79% of the time windows Harvest calls voiced, and on 95% of those whose peak also exceeds 0.1. It is a baseline, not an F0 tracker: each time window is judged alone.
- plot(ax=None, channel: int = 0, **kwargs)[source]¶
Cepstrum against time and quefrency in ms (see
plot_cepstrum()).
sonore.views.mfcc¶
Mel-frequency cepstral coefficients (MFCCs): the log power in triangular bands on a mel scale, summarized by a discrete cosine transform.
- mel_filterbank(n_mels: int, freqs: ndarray, f_lo: float, f_hi: float, scale: str = 'htk', triangles: str = 'height', triangle_axis: str = 'mel') tuple[ndarray, ndarray][source]¶
Triangular mel weights on the frequencies
freqs[Hz].The
n_mels + 2band edges are equally spaced in mel fromf_lotof_hi; bandmrises linearly from edgemto edgem + 1and falls to edgem + 2. “Linearly” is in mel withtriangle_axis="mel", as HTK and Kaldi build them, and in Hz with"hz", as librosa does; the two differ inside each band, most in the wide high bands. Withtriangles="height"each peak is 1, and between the first and last centers the weights sum to 1. With"area"bandmis scaled by2 / (edge[m+2] - edge[m])(Slaney’s normalization), which after the log only adds a constant to each band. The triangles are evaluated at the exact frequencies, not rounded to FFT bins.Returns the weights, shape
(n_mels, len(freqs)), and the edges [Hz].
- symmetric_hamming(n_samples: int) ndarray[source]¶
The Hamming window as HTK and Kaldi define it,
0.54 - 0.46 cos(2 pi i / (n - 1)): symmetric, with both ends at 0.08 (SciPy’s"hamming"spec, as aGaborFramesamples it, is the periodic one).
- delta_features(values: ndarray, order: int = 1, width: int = 5, axis: int = -1) ndarray[source]¶
Time derivatives of a feature track, computed as librosa’s
feature.deltadoes.Along
axis, a polynomial of degreeorderis fitted by least squares to each run ofwidth(odd) values and itsorder-th derivative taken: SciPy’ssavgol_filterwithmode="interp", so the first and lastwidth // 2values use the fit to the first or lastwidth. Fororder=1andwidth=5this is HTK’s regressionsum(n * c[t+n]) / sum(n**2)overn = -2..2away from the ends.order=2is the second derivative of a local quadratic, not the delta of the delta. Units are per time window.
- class MFCC(source: Sound | STFT | TVSTFT, *, n_mfcc: int = 13, n_mels: int = 26, f_lo: float = 0.0, f_hi: float | None = None, mel_scale: str = 'htk', triangles: str = 'height', triangle_axis: str = 'mel', floor_db: float = -200.0, lifter: float = 0.0, win_dur: float | None = None, hop_dur: float | None = None)[source]¶
Bases:
ViewMel-frequency cepstral coefficients of each time window (Davis & Mermelstein, 1980).
For each time window with power spectrum
P[k]: band powersE[m] = sum_k W[m, k] P[k]throughn_melstriangular mel bands (mel_filterbank()), their natural log, and the orthonormal DCT-II of that, of which the firstn_mfcccoefficients are kept. The DCT of a log spectrum is a cepstrum, so the low coefficients are the slow ripples of the log mel spectrum: its envelope, without the pitch.sourceis aSoundor the coefficients of an analysis:a
Soundis analyzed with the usual speech settings: a symmetric Hamming windowwin_durlong (25 ms), as HTK and Kaldi use, a hop ofhop_dur(10 ms), and an FFT length of the next power of two;an
STFTorTVSTFTis used as it is, so any window, hop or pitch-adaptive analysis can be summarized.
This is a view that discards information: the phase, everything inside each mel band, the faster ripples of the log mel spectrum beyond
n_mfcc, and the level apart fromc0. Many spectra have the same coefficients, and no sound is recovered from them;envelope()shows the smoothed spectrum they stand for.The mel power is floored at
floor_dbbelow each channel’s largest band power over all time windows before the log, so scaling a sound changes onlyc0and digital silence gives finite coefficients. Because the floor is set by the loudest moment, a deep floor (the default) keeps each time window independent of the rest of the sound.The defaults follow Kaldi and HTK: HTK’s mel scale, height-1 triangles straight in mel, and, for a sound, a symmetric Hamming window. With no pre-emphasis, no DC removal and
use_energy=false, Kaldi’s MFCCs (compute-mfcc-feats, here through kaldi-native-fbank) are reproduced to float32 precision; Kaldi’s own lifter,low-freqof 20 Hz, 23 bins and “povey” window (a Hann window to the power 0.85, as a callablewindowof theGaborFrame) are all available. Kaldi’s time windows start at the first sample (snip-edges) where sonore’s are centered on multiples of the hop, so the grids line up only after trimming the sound by half a window modulo the hop.To reproduce librosa’s
feature.mfcc(Slaney mel, area-normalized triangles straight in Hz, power in dB floored 80 dB below the loudest cell), analyze withso.GaborFrame(n_fft / fs, hop / fs, window="hann", n_fft=n_fft)and usemel_scale="slaney",triangles="area",triangle_axis="hz",n_mels=128,n_mfcc=20,floor_db=-80anddb. sonore’s grid has one more time window, centered one hop before the first sample, and one or two more at the end; the others are librosa’s. librosa also floors the power at an absolute 1e-10, which matters only for sounds far quieter than any recording, and takes the loudest cell over all channels rather than per channel.Pre-emphasis (
y[n] = x[n] - 0.97 x[n-1], as in HTK and python_speech_features) is a change to the sound, so it is applied to the sound first, for exampleso.Sound(scipy.signal.lfilter([1, -0.97], 1, snd.data, axis=0), snd.fs). python_speech_features also rounds the triangle feet to FFT bins and replacesc0with the log energy of the time window; neither is done here.- Parameters:
source – A sound, or the STFT or TVSTFT to summarize.
n_mfcc – Coefficients kept,
c0included.n_mels – Mel bands.
f_lo – The outer feet of the first and last bands [Hz];
f_hidefaults to half the sampling rate.f_hi – The outer feet of the first and last bands [Hz];
f_hidefaults to half the sampling rate.mel_scale –
"htk"or"slaney"(seefreq_to_mel()).triangles –
"height"(peaks of 1, as HTK and Kaldi) or"area"(Slaney’s and librosa’s area normalization).triangle_axis –
"mel": the triangles are straight lines in mel (HTK, Kaldi);"hz": straight in Hz (librosa).floor_db – Floor of the band powers, in dB below each channel’s largest.
lifter –
Lof HTK’s sinusoidal lifter, which multipliesc[n]by1 + (L / 2) sin(pi n / L); 0 for none. It is a fixed gain per coefficient: it changes Euclidean distances and nothing else.win_dur – Window length and hop [s] used when
sourceis a sound.hop_dur – Window length and hop [s] used when
sourceis a sound.
- mel_power¶
Band powers before the floor, shape
(n_channels, n_mels, n_windows): the mel spectrogram.
- source¶
The STFT or TVSTFT analyzed.
- discards = 'MFCC discards the phase, the detail inside each mel band, and every coefficient past the last kept, so no sound has these MFCCs alone.'¶
- back_to_sound = 'For an approximate voice, read the envelope with mfcc.envelope_view() and pass it to so.world_synthesize with an F0 track and an aperiodicity.'¶
- property db: ndarray[source]¶
datatimes10 / ln 10, as if the band powers had been taken in dB (10 log10) before the DCT.- Type:
The coefficients in dB units
- property mel_db: ndarray[source]¶
The floored mel spectrogram in dB, shape
(n_channels, n_mels, n_windows).
- deltas(order: int = 1, width: int = 5) ndarray[source]¶
Time derivatives of the coefficients, per time window, computed as librosa’s
feature.delta(seedelta_features());width=9is librosa’s default. Same shape asdata.With a 10 ms hop, the first-order deltas of width 5 act as a band-pass filter on each coefficient’s track, strongest near 14 Hz: they pick out changes at the rate of syllables and phonemes.
- envelope(f) ndarray[source]¶
The smoothed power spectrum the kept coefficients stand for, at frequencies
f[Hz], shape(n_channels, len(f), n_windows).The lifter is undone, the missing coefficients are taken as zero, and the inverse DCT gives a smoothed log band power at each band center; between centers it is interpolated linearly in mel, and held constant beyond the first and last. It is a display of what the coefficients keep, for plotting against a spectrum or envelope, not an estimate of the spectrum: with as many coefficients as bands it returns the floored band powers at the centers, which are sums over bands, not spectral densities.
- envelope_view(f=None) GridEnvelope[source]¶
envelope()on this analysis’s time windows, at frequenciesf(by default its FFT bins), read asenv(t, f), so it goes wherever a spectral envelope is taken (warp_frequency(),world_synthesize(),harmonic_complex()). It holds band powers, sums over triangles that widen with frequency, so withtriangles="height"it tilts upward against a spectral density;triangles="area"removes most of that tilt.
- plot(ax=None, channel: int = 0, kind: str = 'mfcc', **kwargs)[source]¶
The coefficients against time (
kind="mfcc") or the mel spectrogram in dB (kind="mel"); seeplot_mfcc().
sonore.views.f0¶
F0 tracking: candidates from a difference function, refinement by the instantaneous frequency of harmonics, a periodicity score, and a Viterbi pass that decides voicing.
- class F0Track(t: ndarray, f0: ndarray, score: ndarray, candidates: ndarray, candidate_scores: ndarray, fs: float, threshold: float)[source]¶
Bases:
ViewAn F0 track: one estimate per channel every
hopseconds. A view: it keeps only the pitch and the voicing.f0andscorehave shape(n_channels, n_windows);f0is 0 where the time window is unvoiced, andscoreis the periodicity score of the chosen candidate (0 where unvoiced).candidatesandcandidate_scoreshave shape(n_channels, n_windows, 4)and hold every refined candidate the tracker weighed, NaN where a time window had fewer.tand a row off0go straight intopitch_adaptive()andlifter().- discards = 'F0Track keeps only the pitch and the voicing of each time window.'¶
- back_to_sound = 'so.world_synthesize rebuilds an approximation of the voice from it together with a spectral envelope and an aperiodicity.'¶
- t: ndarray¶
- f0: ndarray¶
- score: ndarray¶
- candidates: ndarray¶
- candidate_scores: ndarray¶
- fs: float¶
- threshold: float¶
- property voiced: ndarray[source]¶
True where the time window is voiced, shape
(n_channels, n_windows).
- plot(ax=None, channel: int = 0, candidates: bool = False, **kwargs)[source]¶
F0 against time (see
plot_f0_track()).
- f0_track(sound: Sound, f_lo: float = 60.0, f_hi: float = 500.0, hop: float = 0.005, threshold: float = 0.5, *, octave_cost: float = 2.0, switch_cost: float = 0.5, subharmonic_margin: float | None = 0.05) F0Track[source]¶
Track the F0 of each channel of a sound.
Four stages, every
hopseconds from time 0:Candidates. Up to eight local minima of the cumulative-mean-normalized difference function (de Cheveigné & Kawahara, 2002) over a 25 ms window, at lags between
1/f_hiand1/f_lo, each refined by a parabola. This is the autocorrelation idea with YIN’s corrections: it finds the period from all the harmonics, so it works when the fundamental is weak or filtered out. Every minimum below 1 is offered, so that the score alone decides voicing. A very regular voice has a minimum at every multiple of its period, so eight leave room for the period itself (with four, a steady 300 Hz vowel could lose it).Refinement. For a candidate f, a Blackman window three periods long; the instantaneous frequency of each of the first six harmonics is its phase advance over one sample, and the new F0 is the power-weighted mean of each harmonic’s frequency divided by its number. Done twice. This is the idea of WORLD’s StoneMask, written from its description.
Score. The normalized correlation between a stretch three periods long and the same stretch one period later, the fractional part of the period done with a windowed sinc. For a periodic sound in independent noise it is about
s / (1 + s),sthe ratio of periodic to noise power, so 0.5 means the two are equally strong.Tracking. A Viterbi pass over (unvoiced, candidates) per time window. A candidate costs
1 - scoreand the unvoiced state1 - threshold, so a time window leans voiced when its best score exceedsthreshold; moving between candidates costsoctave_costper octave, and a voicing change costsswitch_cost. A candidate is never chosen if another one at a whole multiple of its frequency (an octave, a twelfth, …) scores at least as well, lesssubharmonic_margin: anything periodic with period T is also periodic with period 2T, 3T, …, and without this rule a female voice is tracked an octave low on about 1% of time windows, and a steady synthetic vowel at 300 Hz at a third of its F0. No smoothing follows.
On synthetic vowels the refined F0 is within 0.12% of the truth on a one octave per second glide and a 5.5 Hz vibrato. Against laryngograph reference F0 (a male and a female speaker), it gets the voicing of 5.6% and 1.5% of time windows wrong, where WORLD’s Harvest, which leans toward calling time windows voiced, gets about 21% wrong; where both it and the reference say voiced, 98% and 95% of time windows are within 5%. Lower
thresholdto voice more weak or creaky stretches, for example before resynthesis.- Parameters:
sound – The sound; each channel is tracked separately.
f_lo – The search range [Hz]. The defaults cover speaking voices; raise
f_hifor singing.f_hi – The search range [Hz]. The defaults cover speaking voices; raise
f_hifor singing.hop – Spacing between time windows [s].
threshold – The periodicity score above which a time window leans voiced.
octave_cost – Tracking costs, as above.
switch_cost – Tracking costs, as above.
subharmonic_margin – How much lower a candidate at a whole multiple may score and still rule out a candidate.
Noneturns the rule off, to see the raw choice.
- scale_f0(contour, ratio: float, *, range: float = 1.0)[source]¶
An F0 contour with its pitch changed: every voiced value multiplied by
ratioand, ifrangeis not 1, spread around its median on a log scale,median * ratio * (f0 / median) ** range(range=0is a monotone atratiotimes the median, 2 doubles every interval from it). Unvoiced time windows (0) stay unvoiced.contouris anF0Track, which comes back as anF0Track(its candidates changed the same way), a(times, f0)pair, which comes back as a pair, or anything with.tand.f0, which comes back as a(times, f0)pair. A ratio in semitonesstis2 ** (st / 12). Withratio=1andrange=1the contour itself is returned.The envelope is not touched, so a voice resynthesized on the new contour keeps its formants (unlike
pitch_shift(), which moves them with the pitch).
sonore.views.spectral_envelope¶
Spectral envelopes, and moving them along frequency.
cheaptrick() is WORLD’s CheapTrick (Morise, 2015), ported exactly: every
step reproduces WORLD’s C++ code (github.com/mmorise/World, commit d625e76) to
floating-point precision, including the tiny noise WORLD adds to keep
logarithms and divisions finite, drawn from WORLD’s own generator in WORLD’s
order (sonore.views.world, where DIFFERENCES_FROM_WORLD lists every option that departs
from WORLD).
An envelope here is anything read as env(t, f), giving power at times
t and frequencies f with shape (n_channels, len(f), len(t)):
a SpectralEnvelope (CheapTrick), a GridEnvelope (from
Cepstrum.envelope_view
or MFCC.envelope_view), or a
function. Aperiodicity is read the same
way. warp_frequency() moves an envelope (or an aperiodicity) along
frequency without looking at how it was made, so any envelope and any
synthesizer combine: world_synthesize() and
harmonic_complex() both take any envelope.
- class SpectralEnvelope(data: ndarray, t: ndarray, fs: float, q1: float)[source]¶
Bases:
_FrequencyViewA spectral envelope from
cheaptrick(): power onn_fft // 2 + 1frequenciesffrom 0 tofs / 2, at the window timestof the F0 track.datahas shape(n_channels, n_freqs, n_windows);to_world()gives one channel in WORLD’s layout.It is a view (it keeps the envelope, not the sound): calling it,
env(t, f), reads the power at any times and frequencies, linearly in time and in dB over frequency.- discards = "SpectralEnvelope discards the harmonics, the phase, and everything finer than the envelope's smoothing."¶
- back_to_sound = 'so.world_synthesize rebuilds an approximation of the voice from it together with an F0 track and an aperiodicity.'¶
- amplitude(t, f, channel: int = 0) ndarray[source]¶
The amplitude (square root of the power) at the points
(t[i], f[i]),tandfof the same shape: the formharmonic_complex()reads, one value per sample and harmonic. Linear in time and in dB over frequency, held beyond the first and last time windows.
- cheaptrick(sound: Sound, f0, *, q1: float = -0.15, f0_floor: float = 71.0) SpectralEnvelope[source]¶
WORLD’s CheapTrick spectral envelope (Morise, 2015), ported exactly.
At each time window of the F0 track: the sound under a Hann window three periods long, scaled to unit energy, less its weighted mean; its power spectrum on
world_fft_size()bins; the power below F0 folded back about F0 / 2; a moving average over 2 F0 / 3; then the log, liftered bysinc(F0 q)(smoothing over F0) times the recovery lifter(1 - 2 q1) + 2 q1 cos(2 pi F0 q), and back. The result follows the shape of the harmonic peaks a few dB below them and, because of the smoothing over 2 F0 / 3, barely changes with the time window’s position within a period. Time windows with F0 at or below the floor3 fs / (n_fft - 3)(unvoiced ones included) are analyzed at 500 Hz.- Parameters:
sound – The sound. Its sampling rate must be a whole number of Hz.
f0 – An
F0Track(one row per channel) or a(times, f0)pair; 0 marks unvoiced time windows. The envelope is computed at these times.q1 – The recovery lifter’s parameter.
-0.15is WORLD’s code; the paper gives-0.09;0is no recovery.f0_floor – The lowest F0 the FFT size must hold (WORLD’s 71 Hz).
- class GridEnvelope(data: ndarray, t, f)[source]¶
Bases:
ViewA spectral envelope held as power on any grid of times and frequencies:
datahas shape(n_channels, len(f), len(t)). A view: it keeps the envelope and drops the harmonics and the phase.Any envelope estimate can be put in this form, so that it reads like a
SpectralEnvelopeand goes wherever an envelope is taken:warp_frequency(),world_synthesize(),harmonic_complex().Cepstrum.envelope_viewandMFCC.envelope_viewreturn one.Calling it,
env(t, f), reads the power at any times and frequencies, linear in time and in dB over frequency, asSpectralEnvelopedoes; beyond the ends of the grid the end values are held.- discards = "GridEnvelope discards the harmonics, the phase, and everything finer than the envelope's smoothing."¶
- back_to_sound = 'so.world_synthesize rebuilds an approximation of the voice from it together with an F0 track and an aperiodicity.'¶
- amplitude(t, f, channel: int = 0) ndarray[source]¶
The amplitude (square root of the power) at the points
(t[i], f[i]),tandfof the same shape: the formharmonic_complex()reads, one value per sample and harmonic.
- warp_frequency(view, ratio)[source]¶
A spectral envelope (or aperiodicity) moved along frequency:
new(t, f) = view(t, f / ratio). A ratio above 1 moves every formant up by that factor, as a shorter vocal tract would; below 1, down. The F0 is not touched, so the pitch stays.viewis anything read asview(t, f): aSpectralEnvelope, anAperiodicity, aGridEnvelope, or a function. A view held on a grid comes back as the same type on the same grid (soworld_synthesizetakes it as before); a function comes back as a function.ratiois:a number: one ratio for the whole sound;
a
(times, ratios)pair: a ratio that changes over time, read at each time window (linearly between the given times, held beyond them); only for a view held on a grid;a function
f -> source frequency: any frequency map, giving for each new frequency the one it is read from (a piecewise or bilinear warp, for example).
Values above the view’s top frequency (when lowering) hold the top value. The level is not adjusted: a warp stretches or squeezes the envelope’s area, which changes the output level by a little (about 1.5 dB for a ratio of 0.8 on the speech example), which
normalizeremoves. A ratio of exactly 1 returns the view itself, so a voice resynthesized with it is the same, sample for sample.
sonore.views.aperiodicity¶
Aperiodicity, the share of noise in a voice at each frequency: WORLD’s D4C (Morise, 2016), ported exactly, and a harmonic-residual aperiodicity that measures the share of noise directly.
Every step named after WORLD reproduces WORLD’s C++ code
(github.com/mmorise/World, commit d625e76) to floating-point precision, with
WORLD’s own noise in WORLD’s order; options that depart from WORLD are listed
in DIFFERENCES_FROM_WORLD.
- class Aperiodicity(data: ndarray, t: ndarray, fs: float, method: str)[source]¶
Bases:
_FrequencyViewAn aperiodicity, from
d4c()orharmonic_aperiodicity(): per frequency and time window, how much of the power is noise rather than harmonics. Stored as WORLD stores it, an amplitude ratio between 0 and 1 (data, shape(n_channels, n_freqs, n_windows)): its squareshareis the share of the power that is noise, which is whatworld_synthesize()uses.methodnames the measure that made it (“D4C” or “harmonic residual”).- discards = 'Aperiodicity keeps only the share of noise in each frequency and time window: it discards the spectrum, the pitch and the phase.'¶
- back_to_sound = 'so.world_synthesize rebuilds an approximation of the voice from it together with an F0 track and a spectral envelope.'¶
The noise share of the power,
data ** 2.
- bands(edges: Sequence[float], envelope: SpectralEnvelope | None = None) ndarray[source]¶
The noise share averaged over each band
[edges[i], edges[i+1]), shape(n_channels, n_bands, n_windows). With anenvelope, the average is weighted by its power, so the result is the band’s noise power over its total power.
- d4c(sound: Sound, f0, *, threshold: float = 0.85, f0_floor: float = 71.0) Aperiodicity[source]¶
WORLD’s D4C aperiodicity (Morise, 2016), ported exactly.
At each voiced time window: a “static group delay” from two Blackman-windowed spectra four periods long, a quarter period either side of the window center, divided by a smoothed power spectrum, smoothed over F0 / 2, less itself smoothed over F0. Around each multiple of 3 kHz up to
min(15 kHz, fs / 2 - 3 kHz), a Nuttall window over 3 kHz of that group delay is transformed, its power sorted, and the band’s aperiodicity is the share outside the largest bins, in dB, plus(F0 - 100) / 50dB, capped at 0. The curve is then linear in dB from -60 dB at 0 Hz through those values to 0 dB at Nyquist. At 16 kHz that is one measured value per time window, at 3 kHz.A time window is left fully aperiodic (amplitude ratio
1 - 1e-12) where F0 is 0, or where less thanthresholdof its power between 100 Hz and 7.9 kHz lies below 4 kHz (WORLD’s “LoveTrain” test;threshold=0keeps every voiced time window voiced).D4C was tuned so that WORLD’s resynthesis sounds natural; it is robust to F0 errors but does not report the share of noise below 3 kHz, which the -60 dB anchor sets.
harmonic_aperiodicity()measures that share directly.- Parameters:
sound – As
cheaptrick().f0 – As
cheaptrick().threshold – The LoveTrain voicing threshold.
f0_floor – Sets the output’s frequency grid, which must match the envelope’s for synthesis (WORLD’s 71 Hz, as
cheaptrick()).
- harmonic_aperiodicity(sound: Sound, f0, *, periods: float = 4.0, f0_floor: float = 71.0, cell_harmonics: float = 2.0) Aperiodicity[source]¶
The share of noise, measured by fitting the harmonics and keeping what is left. Not part of WORLD;
d4c()is WORLD’s measure.The F0 track is interpolated to every sample and integrated to a running phase
Phi(t). At each voiced time window, under a Hann windowperiodsperiods long, a weighted least-squares fit of the harmonicscos(k Phi),sin(k Phi)below Nyquist, each with a linear change of amplitude across the window, is subtracted. The residual’s power, over the windowed sound’s power, is the noise share. The fit also absorbs a little of the noise near every harmonic; how much is known exactly from the fit (the share of white noise the residual keeps at each frequency), and is divided out. Both are summed over cellscell_harmonicsharmonics wide, and the cells’ shares are interpolated in dB onto the grid ofcheaptrick(). Unvoiced time windows are all noise.The measure is what the word means, so it reads a known share of noise within a few tenths of a dB on a steady vowel; but it needs F0 to about 0.1%, since a harmonic that drifts out of phase with the fit over the window reads as noise, and so do jitter and shimmer.
- Parameters:
sound – As
cheaptrick().f0 – As
cheaptrick().periods – The window’s length in periods of the time window’s F0.
f0_floor – Sets the output’s frequency grid, as
cheaptrick().cell_harmonics – The width of the cells, in harmonics.
sonore.views.world¶
WORLD’s synthesis (Morise, Yokomori & Ozawa, 2016), ported exactly: a
sound from an F0 track, a spectral envelope and an aperiodicity, as made by
cheaptrick() and
d4c() (or
harmonic_aperiodicity()).
Also here are the pieces WORLD’s analysis shares with its synthesis: WORLD’s
own noise generator (world_randn()), its FFT size, MATLAB’s rounding,
the reading of an F0 track onto time windows, and
DIFFERENCES_FROM_WORLD. The synthesis reads its inputs by what they
provide, not by their type: an F0 track as .t and .f0 or a (times,
f0) pair, an envelope as env(t, f), and an aperiodicity as its grid
.t, .f and .data.
- world_randn(n: int) ndarray[source]¶
The first
nvalues of WORLD’srandnafterrandn_reseed: each is the sum of twelve xorshift128 draws (Marsaglia, 2003), shifted to 28 bits, scaled to [0, 12) and less 6, so approximately standard normal. WORLD restarts this stream at the start of every CheapTrick, D4C and Synthesis call, so the same values come back each time; they are computed once, many lanes at a time, and kept. Read-only.
- world_fft_size(fs: float, f0_floor: float = 71.0) int[source]¶
CheapTrick’s FFT size:
2 ** (1 + floor(log2(3 fs / f0_floor + 1))), long enough for a window three periods off0_floorlong. 1024 at 16 kHz and 2048 at 44.1 or 48 kHz with WORLD’s floor of 71 Hz.
- world_synthesize(f0, envelope, aperiodicity, *, rng=None) Sound[source]¶
WORLD’s synthesis, ported exactly.
Pulses are placed where the running phase of the F0 track (linearly interpolated to every sample) crosses a multiple of 2 pi, the fraction of a sample kept as a linear phase shift; unvoiced stretches get a pulse every 1/500 s. At each pulse the envelope
Sand the amplitude ratioAare interpolated between time windows. The periodic part is the minimum-phase response ofS (1 - A^2), scaled by the square root of the interval to the next pulse, with its DC removed; the aperiodic part is noise as long as that interval through the minimum-phase response ofS A^2(ofSwhere unvoiced). The responses are overlap-added.Minimum phase is an approximation: the waveform within each period is not the original’s, which WORLD’s authors note is audible at low F0.
- Parameters:
f0 – An F0 track (an
F0Trackor a(times, f0)pair), usually the one the envelope and aperiodicity were measured with, from any tracker. It is used as it is when its time windows are the aperiodicity’s; otherwise it is read onto them (voiced where the nearest time window is voiced, linear between voiced values). Those time windows must be evenly spaced from time 0, as WORLD assumes; the spacing is the hop (WORLD’s frame period). F0 can be changed before synthesis (scale_f0(), a pitch change); the envelope stays, so the formants stay.envelope – A
SpectralEnvelopeon the aperiodicity’s time windows and frequencies (any envelope whose.fs,.t,.fand.datamatch the aperiodicity’s grid), used as it is; or any other envelope read asenvelope(t, f)(power, shape(n_channels, len(f), len(t))), such as aGridEnvelopefrom the cepstrum or MFCCs, or one moved bywarp_frequency(), which is read at the aperiodicity’s time windows and frequencies first. WORLD’s synthesis computes a minimum-phase response on its own FFT length, which suits a smooth envelope like CheapTrick’s; an envelope with deep, narrow valleys (tens of dB) comes out a few dB off at the harmonics, so smooth such an envelope first, or give it toharmonic_complex()as itsamplitudes, which reads the envelope at each harmonic exactly.aperiodicity – An
Aperiodicity(or anything with its grid:.t,.f,.fsand.data, the amplitude ratio), which sets the time windows and frequencies.rng –
None(default): WORLD’s own noise stream, restarted at every call, so the same inputs give WORLD’s output, sample for sample. A seed ornumpy.random.Generator: fresh Gaussian noise instead, for independent tokens (listed inDIFFERENCES_FROM_WORLD).
- Returns:
int(n_windows * hop * fs)samples, one channel per channel of the envelope.- Return type:
sonore.views.phasevocoder¶
Phase vocoder (Gordon & Strawn, 1985; Dolson, 1986).
The phase vocoder treats each STFT channel as a slowly varying sinusoid with an amplitude and an instantaneous frequency, estimated from how much the channel’s phase advances between time windows beyond what its center frequency predicts. With amplitudes and frequencies in hand you can:
change duration without changing pitch (
time_stretch()), by re-accumulating phase at a different hop and overlap-adding;change pitch without changing duration (
pitch_shift()), by stretching and then resampling;go back to a sound through an oscillator bank with any frequency remapping (
PVAnalysis.to_sound()), e.g. to shift or stretch partials.
Time stretching uses identity phase locking (Laroche & Dolson, 1999) by default: bins around each spectral peak keep their original phase relationship to the peak, which removes most of the classic “phasiness”.
- class PVAnalysis(magnitude: ndarray, phase: ndarray, freq: ndarray, t: ndarray, fs: float, n_win: int, hop: int, n_samples: int)[source]¶
Bases:
ViewPhase-vocoder analysis: per-channel magnitude, phase and instantaneous frequency, all shaped
(n_channels, n_bins, n_windows). A view: it reads each bin as one sinusoid, andto_sound()rebuilds a sound from those sinusoids, closely for tonal sounds but not exactly.- discards = 'PVAnalysis reads each STFT bin as a single sinusoid, which a sound with more than one component per bin (noise, close partials) is not, so its oscillator resynthesis is not an inverse.'¶
- back_to_sound = 'PVAnalysis.to_sound rebuilds an approximation through an oscillator bank; so.GaborFrame gives an exact STFT.'¶
- no_plot = 'PVAnalysis has no plot: its magnitudes are those of an STFT, so plot so.GaborFrame(...).analyze(sound).'¶
- magnitude: ndarray¶
- phase: ndarray¶
- freq: ndarray¶
- t: ndarray¶
- fs: float¶
- n_win: int¶
- hop: int¶
- n_samples: int¶
- to_sound(time_scale: float = 1.0, freq_map: float | Callable[[ndarray], ndarray] | None = None, floor_db: float = -80.0, chunk: int = 64) Sound[source]¶
Oscillator-bank resynthesis.
Each bin drives a sinusoid whose amplitude and frequency are interpolated between time windows and whose phase is the running integral of frequency.
time_scalestretches the time axis;freq_mapis a ratio (e.g.1.5) or a function mapping frequencies in Hz to new frequencies. For examplelambda f: f + 70makes a 220 Hz harmonic complex inharmonic; note that shifting by a multiple of half the f0 keeps it harmonic (f + 110gives odd harmonics of 110 Hz). Bins that never come withinfloor_dbof the loudest bin are skipped, and partials mapped above Nyquist are dropped.This is designed for tonal sounds, which it reconstructs closely (r > 0.999 for harmonic complexes). For noise, neighboring channels drift out of phase and partially cancel (about 2 dB low, r about 0.93); use
time_stretch()for noisy material. Both figures are measured by tools/measure_docstring_numbers.py.
- pv_analyze(sound: Sound, win_dur: float = 0.046, hop_dur: float | None = None) PVAnalysis[source]¶
Phase-vocoder analysis with a Hann window of
win_durseconds and a hop ofhop_dur(default: a quarter window, the largest hop at which instantaneous frequency is unambiguous across a Hann main lobe).
- time_stretch(sound: Sound, factor: float, win_dur: float = 0.046, phase_lock: bool = True) Sound[source]¶
Change duration by
factor(2 = twice as long) without changing pitch.The synthesis hop is fixed at a quarter window; the analysis hop is
synthesis_hop / factor(rounded), so the achieved factor can differ slightly from the requested one for extreme values. The output length isround(len(sound) * factor).
sonore.views.timbre¶
Timbre descriptors after Peeters et al. (2011), “The Timbre Toolbox”: the log attack time, the spectral centroid and the spectral flux.
Written from the paper’s equations, not from the Timbre Toolbox’s code,
whose license forbids redistribution. Where the defaults differ from the
paper’s, the reason is measured in tools/check_timbre_claims.py and set
out in docs/design/music/timbre-page.md.
- class DescriptorTrack(t: ndarray, values: ndarray, name: str, unit: str)[source]¶
Bases:
ViewOne descriptor’s value in every time window, for each channel. A view: it keeps one number per time window.
valueshas shape(n_channels, n_windows)and is NaN where a time window is silent.medianandiqrare the summaries Peeters et al. (2011) use, since silent and near-silent time windows make a mean or a standard deviation meaningless; both skip NaN.- discards = 'DescriptorTrack keeps one number per time window: it discards the rest of the spectrum or envelope, so many sounds share one track.'¶
- t: ndarray¶
- values: ndarray¶
- name: str¶
- unit: str¶
- plot(ax=None, channel: int = 0, **kwargs)[source]¶
The descriptor against time (see
plot_descriptor_track()).
- attack_segment(sound: Sound, channel: int = 0, cutoff: float = 20.0, zero_phase: bool = True) tuple[float, float][source]¶
Start and end of the attack [s], by the weakest-effort method of Peeters et al. (2011), III A 2 a.
The energy envelope (the amplitude of the analytic signal, low-passed by a third-order Butterworth filter at
cutoffHz) is cut at 0.1, 0.2, …, 1 times its maximum, and the “efforts” are the times it takes to climb from one level to the next. The attack starts within the first effort at most three times the mean effort and ends within the last one, at the envelope’s minimum and maximum inside those two efforts. It measures one channel,channel, where the spectral descriptors measure every channel. The envelope filter sets how short an attack it can resolve, as the next paragraph measures.The paper computes its descriptors with a 5 Hz filter applied once (
cutoff=5, zero_phase=False), which measures a 5 ms attack as about 80 ms. The default here is the paper’s onset setting (20 Hz, forward and backward), which measures it as about 16 ms; seetools/check_timbre_claims.py.
- log_attack_time(sound: Sound, channel: int = 0, cutoff: float = 20.0, zero_phase: bool = True) float[source]¶
The base-10 logarithm of the attack’s duration in seconds,
log10(end - start)with the start and end ofattack_segment()(Peeters et al., 2011, eq. 4). One sample is the shortest duration.
- spectral_centroid(sound: Sound, scale: str = 'power', win_dur: float = 0.0232, hop_dur: float = 0.0058) DescriptorTrack[source]¶
The center of gravity of the spectrum in each time window,
sum(f_k a_k) / sum(a_k)(Peeters et al., 2011, eq. 7), on a Hamming STFT of 23.2 ms windows every 5.8 ms by default.scaleis"power"(the default) or"magnitude", the two spectra the paper offers; they differ a lot (2.5 times for a tone with harmonics falling as 1/n), so say which one you report. The power centroid of a steady harmonic tone matches the one computed from its partials. The magnitude centroid also counts the window’s sidelobes in every bin up to Nyquist, which for a dull tone can outweigh the weak high partials: for harmonics falling as n^-3 at E-flat 4 it measures 810 Hz at 44.1 kHz and 592 Hz at 22.05 kHz, against 414 Hz from the partials (tools/check_timbre_claims.py).
- spectral_flux(sound: Sound, spacing: float | None = 0.1, scale: str = 'magnitude', win_dur: float = 0.0232, hop_dur: float = 0.0058) DescriptorTrack[source]¶
How much the spectrum changes: 1 minus the normalized correlation of two spectra (Peeters et al., 2011, eq. 24, “spectral variation”), 0 for spectra of the same shape and up to 1.
The paper correlates successive time windows, one hop (5.8 ms) apart:
spacing=None. At that spacing a steady tone and a tone whose spectral slope glides over a second measure almost alike;spacingcorrelates spectra that far apart instead (100 ms by default), at which the two differ 25 times (tools/check_timbre_claims.py). Each value is placed at the later of its two time windows.