frames

Invertible analyses: frames and their coefficients, whose resynthesis after a change is the frame’s least-squares inverse.

sonore.frames.frame

Frames: invertible analyses with a common contract.

A Frame turns a Sound into coefficients and back. The contract (docs/design/frames/frames.md) is:

  • synthesize(analyze(x)) == x to floating-point precision.

  • For modified coefficients, synthesize returns the least-squares signal: the one whose coefficients are nearest to the given ones, in the norm of Frame.energy().

  • frame_bounds(n_samples, fs) -> (A, B) are the extreme eigenvalues of the frame operator. A == B means tight. A == 0 means not a frame, and synthesize raises (analyze still works).

Every frame also has Frame.adjoint(), the adjoint of analyze for the coefficient inner product that energy uses: synthesis without the division by s. It is the gradient of a coefficient-domain loss with respect to the signal.

The frames themselves are in sonore.frames.filterbank (filters on the DFT grid) and sonore.frames.gabor (the STFT and its time-varying version).

class Frame[source]

Bases: ABC

An invertible analysis. See the module docstring for the contract.

abstractmethod analyze(sound: Sound, **kwargs)[source]

Coefficients of sound.

abstractmethod synthesize(coefs) → Sound[source]

The least-squares signal for coefs (exact for unmodified ones).

abstractmethod frame_bounds(n_samples: int, fs: float) → tuple[float, float][source]

Extreme eigenvalues (A, B) of the frame operator for signals of n_samples samples at fs Hz.

abstractmethod energy(coefs) → ndarray[source]

Coefficient energy per channel, in the norm the bounds and the least squares are defined in, so that bounds, SNRs and tests all agree.

abstractmethod adjoint(coefs) → Sound[source]

The adjoint of analyze(): the signal y with <analyze(x), coefs> == <x, y> for every x, the coefficient inner product being the one energy() uses. It is synthesize() without the division by the frame operator, and the gradient of a loss on the coefficients with respect to the signal.

sonore.frames.filterbank

Filterbanks: frames whose filters are given by their responses on the DFT grid, and the subband representation they produce.

One class, Filterbank, holds any undecimated bank: a FrequencyScale, the filter centers as positions on that scale, and a FilterType that gives every filter’s response. Functions return the banks in common use, as pure_tone() returns a Sound:

so.cosine_filterbank(30, 50, 8000)                    # ERB-spaced cosines
so.cosine_filterbank(f_lo=125, f_hi=8000, spacing=1/12, scale="octave")
so.gammatone_filterbank(30, 50, 8000)
so.morlet_filterbank(30, 50, 8000, cycles=6)

The frame operator of such a bank is diagonal in frequency, s(f) = sum_k |H_k(f)|**2, and its canonical dual filters are H_k / s (Balazs et al., 2011). Whether s is constant (the bank is tight) is measured on the grid the coefficients live on, never declared (docs/design/frames/filterbanks.md).

n_bands bandpass filters sit at equally spaced points strictly between f_lo and f_hi on the scale; with edges=True (the default) a lowpass and a highpass take over below f_lo and above f_hi, which makes the bank a frame on the whole band.

The cosine filter type is a half-cycle cosine on the scale, spanning two band spacings. Its squared responses sum to 1 (they are power complementary), so filtering on analysis and synthesis reconstructs the input. McDermott & Simoncelli (2011) used these filters on the ERB scale for sound texture, with a lowpass and a highpass at the ends so that “the summed squared frequency response of the filter bank was constant across frequency”. The construction comes from image processing: a squared response summing to one is the “flat system response” of the steerable pyramid (Simoncelli & Freeman, 1995), and its radial filters in Portilla & Simoncelli (2000) are these cosines on a log2 scale. The gammatone and Morlet types are not tight; their synthesis is the canonical dual.

class FilterType[source]

Bases: object

The kind of filter a bank is made of: it gives every filter’s response. A filter type defines band_response(), the bandpass filters at the bank’s centers; the edge filters default to the raised-cosine lowpass and highpass described under responses(). Filter types must be hashable (a frozen dataclass).

edge_width = 1.0
band_response(bank: Filterbank, freqs: ndarray) → ndarray[source]

Responses of the bandpass filters at freqs [Hz] (all >= 0), shape (len(freqs), bank.n_bands).

edge_cfs(bank: Filterbank) → tuple[float, float][source]

Nominal centers [Hz] of the lowpass and highpass: where their transitions start, edge_width gaps outside the outermost bandpass centers (0 Hz at the lowest).

responses(bank: Filterbank, freqs: ndarray) → ndarray[source]

All of the bank’s responses at freqs [Hz], shape (len(freqs), bank.n_filters), edges first and last.

The default edge filters have magnitude sqrt(s_floor), where s_floor is the smallest s of the bandpass filters between the lowest and highest center, and a raised-cosine transition from the nominal center (edge_cfs()) to the outermost bandpass center. One gap trades frame bounds (A/B about 0.5-0.6 for typical banks) against ringing (2-3 times the bank’s own); filling the gap exactly gives better bounds but rings for about half a second (docs/design/frames/frames.md, step 2).

class Cosine(width: float = 1.0)[source]

Bases: FilterType

Half-cycle cosines on the bank’s scale, each width gaps to either side of its center (width=1: zero at the neighboring centers). The lowpass and highpass are the cosines beyond the edges summed in power, so they are flat outside f_lo..f_hi.

With width=1 the squared responses sum to 1 on any increasing centers. Wider cosines on equally spaced centers sum to width when 2 * width is a whole number, and ripple otherwise; the bank measures which (Filterbank.is_tight()).

width: float = 1.0
edge_cfs(bank: Filterbank) → tuple[float, float][source]

The outermost knots, f_lo and f_hi, where the flat parts begin.

band_response(bank: Filterbank, freqs: ndarray) → ndarray[source]

Responses of the bandpass filters at freqs [Hz] (all >= 0), shape (len(freqs), bank.n_bands).

responses(bank: Filterbank, freqs: ndarray, edges: bool | None = None) → ndarray[source]

All of the bank’s responses at freqs [Hz], shape (len(freqs), bank.n_filters), edges first and last.

The default edge filters have magnitude sqrt(s_floor), where s_floor is the smallest s of the bandpass filters between the lowest and highest center, and a raised-cosine transition from the nominal center (edge_cfs()) to the outermost bandpass center. One gap trades frame bounds (A/B about 0.5-0.6 for typical banks) against ringing (2-3 times the bank’s own); filling the gap exactly gives better bounds but rings for about half a second (docs/design/frames/frames.md, step 2).

class Gammatone(order: int = 4, bandwidth_factor: float = 1.019, phase: str = 'causal', edge_width: float = 1.0)[source]

Bases: FilterType

Gammatone filters of the given order with bandwidth parameter b = bandwidth_factor * ERB(cf) (Patterson et al., 1992). ERB is Glasberg & Moore’s (1990); the factor 1.019 for order 4 is the usual convention: Slaney (1993) gives it as Patterson’s recommendation for the fourth-order filter.

The responses are the exact Fourier transform of the impulse response t**(order - 1) exp(-2 pi b t) cos(2 pi cf t), t >= 0, not an IIR approximation, normalized to unit gain at cf. phase="causal" is that filter, with its CF-dependent delay (group delay order / (2 pi b) at cf); phase="zero" keeps the magnitude and aligns onsets across bands. Synthesis is exact either way: the dual filters with conj(H), which undoes the delay.

order: int = 4
bandwidth_factor: float = 1.019
phase: str = 'causal'
edge_width: float = 1.0
band_response(bank: Filterbank, freqs: ndarray) → ndarray[source]

Responses of the bandpass filters at freqs [Hz] (all >= 0), shape (len(freqs), bank.n_bands).

class Morlet(cycles: float = 6.0, edge_width: float = 1.0)[source]

Bases: FilterType

Morlet wavelets: a Gaussian in frequency with standard deviation cf / cycles, minus the standard DC correction so that the response is exactly 0 at 0 Hz, normalized to unit gain at cf; zero-phase. Because of the correction a bank without edge filters is not a frame: every response is 0 at DC, so A = 0.

cycles: float = 6.0
edge_width: float = 1.0
band_response(bank: Filterbank, freqs: ndarray) → ndarray[source]

Responses of the bandpass filters at freqs [Hz] (all >= 0), shape (len(freqs), bank.n_bands).

class Filterbank(scale: FrequencyScale, knots: tuple[float, ...], filter_type: FilterType, *, edges: bool = True)[source]

Bases: Frame

Undecimated filters applied by multiplication on the DFT grid.

knots are positions on scale: the bandpass centers, with one more point at each end (f_lo and f_hi) marking where the edge filters take over. filter_type gives the responses; edges says whether the bank has its lowpass and highpass. Build one with cosine_filterbank(), gammatone_filterbank() or morlet_filterbank(), or directly from any FilterType.

Responses are zero-phase or, more generally, satisfy H(-f) = conj(H(f)); only f >= 0 is ever evaluated, and the subbands are real.

scale: FrequencyScale
knots: tuple[float, ...]
filter_type: FilterType
edges: bool = True
property n_bands: int[source]

Number of bandpass filters.

property n_filters: int[source]

Number of filters (the band axis of Subbands).

property f_lo: float[source]

below it the lowpass takes over.

Type:

Lowest knot [Hz]

property f_hi: float[source]

above it the highpass takes over.

Type:

Highest knot [Hz]

property unit: str[source]

The scale’s unit, e.g. "ERB" or "oct".

property spacing: float | None[source]

Distance between adjacent knots in scale units, or None if they are not equally spaced.

property band_cfs: ndarray[source]

Center frequencies [Hz] of the bandpass filters.

property cfs: ndarray[source]

Center frequency [Hz] of each filter, shape (n_filters,); the edge filters’ nominal centers first and last.

property s_floor: float[source]

The smallest s of the bandpass filters between the lowest and highest center, on a grid of 64 points per gap.

property envelope_peak_delay: ndarray[source]

Time [s] from an impulse to the peak of each filter’s Hilbert envelope, one per filter (as cfs), measured from the impulse responses and refined between samples; 0 for zero-phase filters. Drawing each band this much earlier shows a click as a vertical line (see plot_envelopes(align="peak")). For a causal gammatone it is close to (order - 1) / (2 pi b), the peak of the envelope t**(order - 1) exp(-2 pi b t), not the group delay order / (2 pi b).

response(freqs: ndarray) → ndarray[source]

Responses at freqs [Hz] (all >= 0), shape (len(freqs), n_filters).

rfft_response(n: int, fs: float) → ndarray[source]

Responses on the rfft grid of an n-sample signal.

On an even grid, a complex response is replaced by its real part at Nyquist: irfft keeps only the real part of that bin, so this is the filter actually applied, and analysis, frame_power() and synthesize() must all see the same one. Without this, the dual divides by |H(fs/2)|**2 while analysis applied Re H(fs/2), and reconstruction is off by up to ~1e-5. Real responses are returned unchanged.

frame_power(n: int, fs: float) → ndarray[source]

s(f) = sum_k |H_k(f)|**2 on the rfft grid of n samples: the eigenvalues of the (circular) frame operator.

is_tight(n_samples: int, fs: float, pad: float | str = 'auto') → bool[source]

Whether s is constant (to TIGHT_TOLERANCE) on the grid analyze() would use for n_samples at fs, padding included. Synthesis is then re-filtering, divided by that constant.

ringing(fs: float, level_db: float = -60.0) → int[source]

How long [samples] the filters ring: the longest time, over all filters, before the zero-phase impulse response stays below level_db re its peak (capped at 2 s). The minimum padding of pad="auto".

analyze(sound: Sound, pad: float | str = 'auto') → Subbands[source]

Split sound into subbands (via FFT).

By default the sound is zero-padded by at least the filters’ ringing time (ringing()), so filter ringing near one end can’t wrap around to the other. The padding is rounded up (by a median 0.3% of the length, at most about 4%, for a 30-band ERB bank on 1 to 10 s) so the FFT length has no large prime factor, which makes the transforms about four times faster than at a prime length (measured by tools/measure_docstring_numbers.py). The padding travels with the Subbands (and any Envelopes derived from them) and is removed on output, so .data and Subbands.to_sound() have the sound’s own length, and analysis followed by synthesis is exact.

pad=0 makes the analysis circular, which is what you want for periodic signals and for texture synthesis (seamless loops). A number pads by that many seconds.

synthesize(coefs: Subbands) → Sound[source]

Canonical dual synthesis: filter each band with conj(H) / s and sum, then remove the padding. When s is constant on the grid (is_tight()) this is re-filtering with H and dividing by that constant.

The frame is the circular operator on the (padded) grid the coefficients live on. Unmodified coefficients reconstruct exactly. For modified ones the result is the least-squares fit on that grid, cropped; with pad=0 this is the canonical least squares on the signal itself. With padding on a non-tight bank it differs slightly (about 1% on masked coefficients) from the canonical least squares on the unpadded signal, because masked energy that lands in the padding is discarded rather than refit; this is a deliberate choice, exact and non-iterative. Raises if the bank leaves a frequency uncovered (A = 0).

frame_bounds(n_samples: int, fs: float, pad: float | str = 'auto') → tuple[float, float][source]

(min s, max s) on the DFT grid that analyze() would use for a signal of n_samples at fs, including its padding: the bounds of the grid the coefficients actually live on, which depends on the signal length and rate (hence a method, not a property). A tight bank returns its constant twice, (1.0, 1.0) for the cosine banks.

energy(coefs: Subbands) → ndarray[source]

Sum of squares of the subbands, padding included, per channel. Subbands are real, so every sample has weight 1.

adjoint(coefs: Subbands) → Sound[source]

Filter each band with conj(H) and sum, with no division by s, then remove the padding (the adjoint of zero-padding is cropping). For a bank with s = 1 this equals synthesize().

cosine_filterbank(n_bands: int | None = None, f_lo: float = 50.0, f_hi: float = 8000.0, *, scale: str | FrequencyScale = 'erb', spacing: float | None = None, centers=None, width: float = 1.0, edges: bool = True) → Filterbank[source]

Half-cycle cosine filters (Cosine) on scale.

n_bands bandpass filters (30 if neither it nor spacing is given) sit at equally spaced points strictly between f_lo and f_hi [Hz]; spacing (in scale units, e.g. 1/12 octave) chooses n_bands instead. centers [Hz] gives every knot directly, increasing, the first and last being where the lowpass and highpass take over; the cosines then follow the gaps and stay tight. With edges=True the bank has n_bands + 2 filters and its squared responses sum to 1 (width=1).

gammatone_filterbank(n_bands: int | None = None, f_lo: float = 50.0, f_hi: float = 8000.0, *, scale: str | FrequencyScale = 'erb', spacing: float | None = None, centers=None, phase: str = 'causal', order: int = 4, bandwidth_factor: float = 1.019, edges: bool = True, edge_width: float = 1.0) → Filterbank[source]

Gammatone filters (Gammatone), equally spaced on scale (the ERB-number scale by default), with the raised-cosine edge filters of FilterType.responses(). n_bands, f_lo, f_hi, spacing and centers place the filters as in cosine_filterbank(). edges=False gives the bare bank (for cochleagrams); it may be badly conditioned or not a frame at all, in which case analyze still works but synthesize refuses.

morlet_filterbank(n_bands: int | None = None, f_lo: float = 50.0, f_hi: float = 8000.0, *, scale: str | FrequencyScale = 'octave', spacing: float | None = None, centers=None, cycles: float = 6.0, edges: bool = True, edge_width: float = 1.0) → Filterbank[source]

Morlet wavelets (Morlet), equally spaced on scale (log2 frequency by default), with the raised-cosine edge filters of FilterType.responses(). n_bands, f_lo, f_hi, spacing and centers place the filters as in cosine_filterbank(). Without edges the bank is not a frame (A = 0 at DC).

class Subbands(data, fs: float, filterbank: Filterbank, pad: int = 0)[source]

Bases: _PaddedBands

The output of a Filterbank: one band-limited Sound per filter.

sb[i] is a Sound, iterating yields Sounds, and sb.cfs labels them. The first and last bands are the filterbank’s lowpass and highpass edges. Internally the bands are stored as one (n_samples, n_bands, n_channels) array (data).

The Hilbert decomposition of every band at once:

sb.envelopes()      # Envelopes: one Envelope per band (a cochleagram)
sb.tfs()            # Subbands: the fine structure, itself audible
sb.envelopes() * sb.tfs()   # == sb

so a vocoder is (speech.envelopes() * carrier.tfs()).to_sound().

envelopes(lowpass: float | None = None, fs: float | None = None)[source]

Hilbert envelope of every band, as Envelopes. Optionally lowpass-filtered [Hz] and then resampled to fs.

tfs() → Subbands[source]

Temporal fine structure of every band: cos of the instantaneous phase (unit amplitude).

to_sound() → Sound[source]

Back to a Sound with the filterbank’s canonical dual (synthesize()): the exact inverse of analyze(), and the least-squares signal after the bands are modified. For the cosine banks this is re-filtering each band and summing; with the default padding, re-filtering can’t wrap around either.

sum() → Sound[source]

Add the bands without re-filtering. to_sound() is almost always what you want; sum is for bands that were already shaped to add up correctly.

plot(axes=None, channel: int = 0, **kwargs)[source]

Stacked band waveforms (see sonore.plotting.plot_subbands()). For an image, plot the envelopes: sb.envelopes().plot().

sonore.frames.gabor

Gabor frames: the STFT as a frame, and its coefficients.

GaborFrame is the one-sided STFT (wrapping scipy.signal.ShortTimeFFT). Its frame operator is diagonal in time, s(t) = K sum_q |w(t - q hop)|**2 with K = n_fft (this holds because SciPy requires the window to fit in n_fft). Only the non-negative frequencies are stored, so the coefficient norm weights each stored bin by 2, except DC and, for even K, Nyquist, which have no mirror image: this recovers the energy of the full two-sided STFT, and it is the norm in which SciPy’s istft is the least-squares inverse.

TVGaborFrame is the Gabor frame with one window per position (nonstationary Gabor, painless case): every window fits in one FFT length M, so the frame operator is again diagonal in time, s(t) = M sum_q |w_q(t - a_q)|**2.

STFT and TVSTFT are their coefficients, which synthesize back exactly. Levels are in dB, 20 log10 |X|.

class GaborFrame(win_dur: float, hop_dur: float | None = None, window: str | tuple | Callable[[int], ndarray] = 'hann', n_fft: int | None = None)[source]

Bases: _ShortTimeFourierFrame

The one-sided short-time Fourier transform as a frame.

Durations are rounded to samples per sampling rate, and the ShortTimeFFT is built once per rate (sft()). analyze returns an STFT.

Analysis always works, even for a window and hop that leave gaps (A = 0, including a hop longer than the window): the STFT is still a picture of the sound. synthesize() (and Griffin-Lim) refuse such a pair, and the error gives the frame bounds.

Parameters:
  • win_dur (float) – Window length [s].

  • hop_dur (float | None) – Hop [s]; defaults to a quarter window (75% overlap).

  • window (str | tuple | collections.abc.Callable[[int], numpy.ndarray]) – A scipy.signal.get_window() spec (name or tuple), sampled periodically, or a callable n -> array of length n.

  • n_fft (int | None) – FFT length K (>= the window length); defaults to the window length.

win_dur: float
hop_dur: float | None = None
window: str | tuple | Callable[[int], ndarray] = 'hann'
n_fft: int | None = None
lengths(fs: float) → tuple[int, int, int][source]

(window, hop, n_fft) in samples at fs.

window_samples(fs: float) → ndarray[source]

The analysis window at fs.

sft(fs: float) → ShortTimeFFT[source]

The ShortTimeFFT at fs (cached), with the canonical dual window for synthesis. Raises, with the frame bounds, when the window and hop leave gaps.

frame_power(n_samples: int, fs: float) → ndarray[source]

s(t) = K sum_q |w(t - q hop)|**2 for t in 0..n_samples-1: the diagonal of the frame operator. Sums over every time window that overlaps the signal; time windows SciPy leaves out overlap it only where the window is zero, so they add nothing.

bin_weights(fs: float) → ndarray[source]

Weight of each stored frequency bin: 1 for DC and, for even n_fft, Nyquist; 2 for the rest, which stand for their negative- frequency mirror images too.

analyze(sound: Sound) → STFT[source]

The STFT of sound, data shape (n_channels, n_freqs, n_windows).

synthesize(coefs: STFT) → Sound[source]

SciPy’s istft: the weighted real least-squares signal for coefs, exact for unmodified coefficients. Padding does not change this, since the frame operator is diagonal in time.

adjoint(coefs: STFT) → Sound[source]

n_fft times SciPy’s istft with dual_win = win: the adjoint for the bin-weighted inner product of energy() (verified against a dense matrix, including odd and zero-padded FFTs).

class TVGaborFrame(times: Sequence[float], win_durs: Sequence[float], n_fft: int | None = None, window: str | tuple | Callable[[int], ndarray] = 'hann')[source]

Bases: _ShortTimeFourierFrame

A Gabor frame with a time-varying window.

Window q is win_durs[q] long and centered at times[q] [s]; each windowed segment is zero-padded to one FFT length n_fft and transformed, with the phase referenced to the window’s center, as GaborFrame (SciPy) does. Every window must fit in n_fft samples, the painless condition that keeps the frame operator diagonal (Balazs et al., 2011); n_fft defaults to the longest window. Durations and times are rounded to samples at each sampling rate, and a window longer than n_fft is refused when the frame is first used at a rate.

analyze returns a TVSTFT, data shape (n_channels, n_fft // 2 + 1, len(times)). The signal is zero outside its own extent, as in GaborFrame. A constant schedule over the time windows SciPy uses reproduces GaborFrame exactly.

Parameters:
  • times (collections.abc.Sequence[float]) – Window centers [s], strictly increasing.

  • win_durs (collections.abc.Sequence[float]) – Window lengths [s], one per center.

  • n_fft (int | None) – FFT length [samples], at least the longest window.

  • window (str | tuple | collections.abc.Callable[[int], numpy.ndarray]) – A scipy.signal.get_window() spec, sampled periodically at each window’s length, or a callable n -> array.

times: Sequence[float]
win_durs: Sequence[float]
n_fft: int | None = None
window: str | tuple | Callable[[int], ndarray] = 'hann'
classmethod from_function(win_dur_of_t: Callable[[float], float], t_end: float, overlap: float = 4, t_start: float = 0.0, **kwargs) → TVGaborFrame[source]

A schedule that steps from t_start to t_end [s] with hop win_dur_of_t(t) / overlap at each center t. Use t_end at least the signal’s duration so the windows cover it. Other arguments go to the constructor.

classmethod pitch_adaptive(f0_times: Sequence[float], f0: Sequence[float], t_end: float, periods: float = 3.0, overlap: float = 4, t_start: float = 0.0, **kwargs) → TVGaborFrame[source]

Windows periods fundamental periods long, following an F0 track.

f0 [Hz] at f0_times [s] is 0 or NaN where unvoiced. F0 is carried through unvoiced stretches, log-linearly between the voiced neighbors and held constant before the first and after the last voiced point, so the window length never jumps: jumps are what cost frame bounds, not the adaptation itself. The hop is window / overlap, as in from_function(), from t_start to t_end; to cover a sound evenly up to its last sample, t_end should be its duration plus half the longest window.

Three periods is the shortest whole number of periods at which neighboring harmonics separate clearly (a peak-to-dip ratio of about 12 dB with a Hann window, the same at every F0) while the flicker at the period rate cancels; it is also the window length of WORLD’s CheapTrick (Morise, 2015). Other arguments go to the constructor.

layout(fs: float) → _TVLayout[source]

Window lengths, centers and FFT length in samples at fs (cached).

frame_power(n_samples: int, fs: float) → ndarray[source]

s(t) = M sum_q |w_q(t - a_q)|**2 for t in 0..n_samples-1: the diagonal of the frame operator.

bin_weights(fs: float) → ndarray[source]

Weight of each stored frequency bin, as for GaborFrame.

analyze(sound: Sound) → TVSTFT[source]

The time-varying STFT of sound, data shape (n_channels, n_freqs, n_windows).

adjoint(coefs: TVSTFT) → Sound[source]

Conjugate-window overlap-add of n_fft * irfft of each time window, without dividing by s.

synthesize(coefs: TVSTFT) → Sound[source]

The adjoint divided by s(t): the canonical dual, so the weighted least-squares signal for coefs, exact for unmodified ones. Raises if the windows leave part of the signal uncovered.

class STFT(sound: Sound, win_dur: float = 0.02, hop_dur: float | None = None, *, frame: GaborFrame | None = None)[source]

Bases: _ShortTimeCoefficients

Short-time Fourier transform (wraps scipy.signal.ShortTimeFFT).

data has shape (n_channels, n_freqs, n_windows). Resynthesis with to_sound() is exact for an unmodified STFT, and the least-squares signal for a modified one.

STFT(sound, win_dur, hop_dur) is GaborFrame(win_dur, hop_dur).analyze(sound); the frame is kept as frame and SciPy’s transform as sft.

Parameters:
  • win_dur – Window length [s] (periodic Hann).

  • hop_dur – Hop [s]; defaults to a quarter window (75% overlap).

  • frame – A GaborFrame to use instead (other windows, zero-padded FFTs); win_dur and hop_dur are then ignored.

property f: ndarray[source]
property t: ndarray[source]

Window center times [s].

property n_fft: int[source]

FFT length [samples].

property shortest_window: int[source]

every window has it.

Type:

Window length [samples]

griffin_lim(n_iter: int = 100, momentum: float = 0.99, rng=None) → Sound[source]

Reconstruct a signal from the magnitude only (fast Griffin-Lim, Perraudin et al., 2013). momentum=0 gives classic Griffin-Lim.

plot(ax=None, channel: int = 0, **kwargs)[source]
class TVSTFT(data: ndarray, fs: float, n_samples: int, frame: TVGaborFrame)[source]

Bases: _ShortTimeCoefficients

Coefficients of a TVGaborFrame: a short-time Fourier transform whose window changes over time.

data has shape (n_channels, n_freqs, n_windows) like STFT, on one frequency grid f (every window is zero-padded to the same FFT length) and at non-uniform window centers t. to_sound() is exact for unmodified coefficients and the least-squares signal for modified ones. Multiply by an array or a Mask to mask.

property f: ndarray[source]
property t: ndarray[source]

Window center times [s], rounded to samples.

property n_fft: int[source]

FFT length [samples], shared by every window.

property shortest_window: int[source]

Length of the shortest window [samples].

plot(ax=None, channel: int = 0, **kwargs)[source]

Spectrogram in dB at the window centers (see plot_tf_db()).