sources

Sounds made from parameters alone: tones, noises, chirps and the LF glottal source, Klatt’s formant synthesizer, and spectrotemporal ripples.

sonore.sources.waveforms

Waveforms made from parameters: silence, tones, harmonic complexes, chirps, noises, and the LF glottal source.

Every generator returns a Sound normalized to RMS = 1. Anything random takes an rng argument (a seed or np.random.Generator) so stimuli can be regenerated exactly.

silence(duration: float, fs: float, n_channels: int = 1) → Sound[source]

Zeros, duration seconds long, in n_channels channels (RMS 0, the one generator not scaled to 1).

pure_tone(duration: float, fs: float, freq: float, phase: float = 0.0) → Sound[source]

cos(2*pi*freq*t + phase).

harmonic_complex(duration: float, fs: float, f0: float, harmonics: ArrayLike | None = None, amplitudes: ArrayLike | Callable[[ndarray, ndarray], ArrayLike] | Callable[[ndarray, ndarray, int], ArrayLike] | None = None, phases: ArrayLike | str = 'cosine', rng: int | Generator | None = None, *, f_max: float | None = None) → Sound[source]
harmonic_complex(duration: float, fs: float, f0: tuple[ArrayLike, ArrayLike] | F0Contour, harmonics: ArrayLike | None = None, amplitudes: ArrayLike | Callable[[ndarray, ndarray], ArrayLike] | Callable[[ndarray, ndarray, int], ArrayLike] | None = None, phases: ArrayLike | str = 'cosine', rng: int | Generator | None = None, *, f_max: float | None = None, ramp: float = 0.005, unvoiced: str = 'silence') → Sound

Sum of harmonics a_n cos(n * Phi(t) + phi_n) of a fixed or moving F0.

f0 is either a number, for Phi = 2*pi*f0*t, or an F0 contour: a (times, values) pair or anything with .t and .f0, such as an F0Track, with values in Hz and 0 where unvoiced (one row per channel for a multichannel sound).

harmonics are the harmonic numbers; by default every one below Nyquist (fixed F0) or below f_max (contour). amplitudes is one value per harmonic, or a function amplitudes(t, f) giving the gain at times t and frequencies f (a spectral envelope, evaluated at every harmonic’s frequency). The function may also take the harmonic number as a third argument, amplitudes(t, f, n), and may return complex values: harmonic n is then Re(a_n exp(i(n Phi + phi_n))), so the angle of a_n adds to its phase and can change over time (an LF glottal pulse whose shape changes, as in glottal_source()). amplitudes may also be a spectral envelope, such as SpectralEnvelope or GridEnvelope (power on a grid of times and frequencies), which is read point by point as amplitude, from its first channel. phases are starting phases: an array, one per harmonic, or one of "cosine", "sine", "alternating", "random", "schroeder+", "schroeder-".

A fixed F0 drops harmonics at or above Nyquist, with a warning. Given f_max, it also takes harmonics only below f_max by default and fades them out as a contour does (below); without it, nothing is faded.

A contour (docs/design/sources/harmonic-source.md) is filled across its unvoiced gaps and interpolated linearly to the sample rate, and the phase is its exact running integral, so every harmonic follows n times the contour with no jumps. Each harmonic fades out (cos^2) between 0.9 * f_max and f_max (default 0.45 * fs), so nothing aliases as the pitch rises. Harmonics are switched off where the contour is unvoiced, with Hann ramps ramp seconds long; unvoiced="noise" fills those stretches with white noise of the harmonics’ power instead of silence. Values are held beyond the first and last times. With a contour, phases has one value per harmonic number up to the highest that fits below f_max at the lowest voiced F0, unless harmonics is given.

schroeder_complex(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, n_harmonics: int | None = None, sign: int = 1, **contour) → Sound[source]

Equal-amplitude harmonic complex with Schroeder phases sign * pi * n(n+1)/N. n_harmonics defaults to all below Nyquist (or below f_max for an F0 contour). f0 and the keyword arguments for a contour are as in harmonic_complex().

square_wave(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, phase: float = 0.0, bandlimited: bool = True, **contour) → Sound[source]

Square wave. Band-limited (odd harmonics below Nyquist) unless bandlimited=False, which aliases badly at high f0. A band-limited square wave also takes an F0 contour, as harmonic_complex() does.

sawtooth_wave(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, phase: float = 0.0, bandlimited: bool = True, **contour) → Sound[source]

Rising sawtooth. Band-limited unless bandlimited=False. A band-limited sawtooth also takes an F0 contour, as harmonic_complex() does.

pulse_train(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, phase: float = 0.0, bandlimited: bool = True, **contour) → Sound[source]

Periodic pulse (click) train.

Band-limited: equal-amplitude cosine-phase harmonics up to Nyquist; this also takes an F0 contour, as harmonic_complex() does. Otherwise single-sample impulses at the nearest sample to each period. phase (radians of the fundamental) shifts the pulses in time.

linear_chirp(duration: float, fs: float, f0: float, f1: float, phase: float = 0.0) → Sound[source]

Linear frequency sweep from f0 to f1.

exponential_chirp(duration: float, fs: float, f0: float, f1: float, phase: float = 0.0) → Sound[source]

Exponential (log-frequency) sweep from f0 to f1.

gaussian_noise(duration: float, fs: float, band: tuple[float, float] | None = None, tilt: float = 0.0, spectrum: Callable[[ndarray], ndarray] | tuple[ArrayLike, ArrayLike] | None = None, n_channels: int = 1, rng: int | Generator | None = None) → Sound[source]

Gaussian noise, optionally spectrally shaped.

White Gaussian noise for the whole duration is shaped in the frequency domain: each bin of its FFT is multiplied by 10^(L(f)/20), the target level L(f) in dB at that bin from band, tilt and spectrum together, then transformed back, its mean removed, and scaled to RMS 1. The shape is the expected spectrum: the bins of any one draw scatter around it, as they do for any Gaussian noise.

Parameters:
  • band – (f_lo, f_hi) passband in Hz (brick-wall).

  • tilt – Spectral slope in dB/octave (-3 = pink, power 1/f; -6 = brown, power 1/f^2).

  • spectrum – Target spectrum level in dB (its shape only; the output is scaled to RMS 1): a Spectrum, a function f -> dB, or a (freqs, dB) pair, interpolated linearly in Hz. A Spectrum is silent outside its frequencies, since it was measured only there; a pair is held flat beyond its ends.

  • n_channels – Independent noise in each channel.

correlated_noise(duration: float, fs: float, corr: float = 1.0, rng: int | Generator | None = None, **noise_kwargs) → Sound[source]

Two-channel noise whose interaural correlation is exactly corr (in [-1, 1]), each channel at RMS 1. Extra keyword arguments go to gaussian_noise().

The symmetric-generator method, made exact as Hartmann and Cho (2011) describe: two independent noises x1 and x2 are made orthogonal (Gram-Schmidt) and equal in power, then mixed as a x1 + b x2 (left) and a x1 - b x2 (right) with a^2 - b^2 = corr. Without those two steps each draw would miss corr by chance, more so for few spectral components (narrow bands, short sounds). The correlation is the zero-lag one, over the whole sound.

iterated_ripple_noise(duration: float, fs: float, delay: float, gain: float = 1.0, iterations: int = 16, network: str = 'add-same', rng: int | Generator | None = None, **noise_kwargs) → Sound[source]

Iterated rippled noise (Yost, 1996). Pitch is at 1/delay.

network="add-same" (IRNS) iterates y <- y + g*delay(y); "add-original" (IRNO) iterates y <- x + g*delay(y). Implemented exactly in the frequency domain (fractional delays are fine), with a warm-up segment discarded so the output is stationary.

lf_harmonics(harmonics: ArrayLike, rd: float = 0.7, *, ra: float | None = None, rg: float | None = None, rk: float | None = None, flow: bool = False) → ndarray[source]

Complex Fourier coefficients of the LF pulse, one per harmonic number.

Harmonic k of a train of LF pulses is 2 |c_k| cos(k Phi + angle(c_k)) where Phi is the running phase of F0, so c_k depends only on k and the pulse’s shape, not on F0. The coefficients are those of the flow derivative scaled to -1 at the main excitation (Ee = 1); with flow=True, those of the flow itself, c_k / (2 pi i k), in units of Ee * T0. They are exact (a closed form), so a source built from them does not alias.

The shape is set by Fant’s (1995) rd (0.3-2.7, his main range, the one his predictions were fitted to; 0.7 is close to his typical adult male values), through his prediction of ra, rg and rk from it; or directly by ra = ta/T0, rg = T0/(2 tp) and rk = (te - tp)/tp (below 1), given together.

lf_pulse(x: ArrayLike, rd: float = 0.7, *, ra: float | None = None, rg: float | None = None, rk: float | None = None, flow: bool = False) → ndarray[source]

The LF flow derivative (or with flow=True, the flow) at times x in fractions of a period (wrapped into [0, 1)), scaled to -1 at the main excitation te (Ee = 1; for the laxest voices, near Rd 2.7, the open phase dips slightly below that just before it).

Evaluating the formula at sample times aliases, most for tense voices with an abrupt closure; use it to draw a period, and glottal_source() to make a sound. The shape is set as in lf_harmonics().

glottal_source(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, rd: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] = 0.7, *, flow: bool = False, f_max: float | None = None, ramp: float = 0.005) → Sound[source]

A voiced source of LF glottal pulses on a fixed F0 or an F0 contour.

The LF model (Fant, Liljencrants & Lin, 1985) describes one period of the glottal flow derivative: an exponentially growing sinusoid while the glottis opens, a sharp negative peak where it closes, and an exponential return phase. Fant’s (1995) single shape parameter Rd sets the whole pulse, from tense, pressed voice (small Rd: short open phase, abrupt closure, strong high harmonics) to lax, breathy voice (large Rd: long open phase, gradual closure, a dominant fundamental). The design, and the checks behind it, are in docs/design/sources/glottal-source.md.

The pulses are built from their harmonics (lf_harmonics()) by harmonic_complex(), so they do not alias, each period takes its own length on a moving F0, and f0, f_max and ramp work as there (unvoiced time windows are silent, and a contour that is never voiced gives silence). The result is the flow derivative, the source as it excites the vocal tract with radiation folded in, or with flow=True the flow itself; normalized to RMS 1.

rd is Fant’s (1995) shape parameter (0.3-2.7, default 0.7, close to his typical adult male values): a number, or a (times, values) track for a voice quality that changes, interpolated to every sample. LF models the periodic pulse only: breathy voice also needs aspiration noise, as in klatt_synthesize().

sonore.sources.klatt

A Klatt-style cascade/parallel formant synthesizer.

Speech from named acoustic parameters, after Klatt (1980): a voiced source and two noise sources, formant resonators in cascade (voicing and aspiration) and in parallel (frication), and radiation from the lips. Every control is something visible on a spectrogram, so the source-filter model can be taken apart one parameter at a time, and stimuli such as vowel and /ba/-/da/-/ga/ continua are set exactly. The design, and the checks behind it, are in docs/design/sources/klatt.md.

KLATT_DEFAULTS = mappingproxy({'F0': 100.0, 'AV': 60.0, 'SS': 1, 'RD': 0.7, 'AH': 0.0, 'AF': 0.0, 'AB': 0.0, 'F1': 500.0, 'B1': 60.0, 'F2': 1500.0, 'B2': 90.0, 'F3': 2500.0, 'B3': 150.0, 'F4': 3500.0, 'B4': 200.0, 'F5': 4500.0, 'B5': 250.0, 'F6': 4900.0, 'B6': 1000.0, 'A1': 0.0, 'A2': 0.0, 'A3': 0.0, 'A4': 0.0, 'A5': 0.0, 'A6': 0.0, 'FNP': 250.0, 'BNP': 100.0, 'FNZ': 250.0, 'BNZ': 100.0})

a neutral adult male vowel, voiced at 100 Hz, with no noise. Frequencies and bandwidths are in Hz, amplitudes in dB (60 = the source at unit level, 0 or less = off).

Type:

Every parameter klatt_synthesize() takes, with its default

klatt_synthesize(duration: float, fs: float, params: Mapping[str, float | tuple[ArrayLike, ArrayLike]] | None = None, *, rng: int | Generator | None = None, **kwargs: float | tuple[ArrayLike, ArrayLike]) → Sound[source]

Speech from Klatt’s (1980) parameters, as a cascade/parallel formant synthesizer.

Parameters are given by Klatt’s names, in params (a mapping, such as one made by klatt_continuum()) or as keyword arguments, which take precedence; anything not given takes its value in KLATT_DEFAULTS. Each is a number, or a (times, values) pair: times in seconds, interpolated linearly to every sample and held beyond the ends. A table of values every 5 ms, as Klatt used, is the pair (0.005 * np.arange(len(values)), values). F0 may also be an F0Track, to resynthesize a measured pitch contour.

  • Sources. Voicing: harmonics of F0 (harmonic_complex(), so the pitch can glide with no phase jumps or aliasing) at level AV, with the spectrum set by the source switch SS, numbered as in KLSYN88 (Klatt & Klatt, 1990). SS = 1 (the default): Klatt’s (1980) impulses through his glottal low-pass, falling about 12 dB per octave. SS = 3: Liljencrants-Fant glottal pulses (glottal_source()), their shape set by Fant’s RD, from tense (0.3, strong high harmonics) to lax (2.7, a dominant fundamental); at the default 0.7 its 3 kHz harmonic is about 12 dB weaker, relative to the first, than with SS = 1. KLSYN88’s SS = 2 (KLGLOTT88) is not included. Where F0 is 0 the harmonics are switched off. Aspiration (AH) and frication (AF, AB): white Gaussian noise. While voicing is on, both noises are amplitude modulated by a square wave at F0, 50% deep, as in Klatt’s synthesizer.

  • Cascade branch. Voicing and aspiration pass through the nasal pole and zero (FNP, BNP, FNZ, BNZ) and formants 1-5 (F1-F5, bandwidths B1-B5) in series: formant levels follow from the frequencies, as in a vowel. A formant default at or above Nyquist is left out (F5 at fs below 9 kHz).

  • Parallel branch. Frication passes through formants 1-6 side by side, each at its own level A1-A6 (dB at the formant’s peak), added with alternating signs so that they sum like the cascade between peaks; formants 2-6 get a differenced input, as in Klatt’s, to keep low frequencies out. AB adds the frication with no formants.

  • Radiation. The voiced source is differenced (+6 dB per octave), the radiation from the lips; the noises are white as they leave the lips.

Amplitudes are in dB: at 60 a source has RMS 1 as it leaves the lips (before the formants), so equal values mean equal source levels; each 20 dB is a factor of 10, and 0 or less is off. The result is normalized to RMS 1, like every generator; relative levels within one call are kept. Klatt’s quasi-sinusoidal voicing (AVS) and KLSYN88’s other voice-quality controls (open quotient, spectral tilt, flutter, double pulsing) are not included.

rng seeds the noise, so a call can be repeated exactly.

klatt_continuum(start: Mapping[str, float | tuple[ArrayLike, ArrayLike]], end: Mapping[str, float | tuple[ArrayLike, ArrayLike]], steps: int) → list[dict[str, float | tuple[ArrayLike, ArrayLike]]][source]

steps parameter sets for klatt_synthesize(), evenly spaced from start to end (both included), such as a /ba/-/da/ continuum.

A parameter given in only one of them takes its KLATT_DEFAULTS value in the other. The source switch SS is not interpolated: it must be the same at both ends. Numbers are interpolated directly; (times, values) pairs are interpolated value by value, so they must share their times (a number paired with a track is held at every one of its times).

sonore.sources.ripples

Spectrotemporal ripples: sounds defined by their modulation content.

A pattern is an envelope over time and log-frequency, E(t, x), where x is octaves above f_lo. ripple_sound() imposes a pattern on a carrier, which supplies the fine structure.

Patterns

Ripple

A moving ripple (Kowalski, Depireux & Shamma, 1996; Chi et al., 1999): 1 + depth * sin(2*pi*(rate*t + density*x) + phase). rate is in Hz and density in cycles/octave. With positive rate and density the ripple drifts downward in frequency; a negative rate drifts upward. Ripples add: Ripple(4, 1) + Ripple(-8, 2) is a RippleSum.

DynamicRipple

A dynamic moving ripple (Escabí & Schreiner, 2002) whose rate and density wander slowly and randomly within given ranges, for STRF estimation.

Any callable f(t, x)

Evaluated with t as a row and x as a column, so ordinary numpy broadcasting works, e.g. lambda t, x: 1 + 0.5 * np.sin(2*np.pi*(3*t + x**2)).

Carriers

"tones": log-spaced tones with random phases, the classic ripple carrier. "harmonic": harmonics of f0. "noise": narrowband Gaussian noise per channel. "low-noise": the fine structure of that noise with its envelope flattened, so it adds no envelope fluctuations of its own. Or any Sound, whose fine structure is used the same way.

All carriers are scaled to equal energy per octave, so switching carriers changes the fine structure but not the long-term spectrum.

class Ripple(rate: float, density: float, depth: float = 0.9, phase: float = 0.0, scale: str = 'linear')[source]

Bases: _Pattern

A moving ripple 1 + depth*sin(2*pi*(rate*t + density*x) + phase).

Parameters:
  • rate (float) – Temporal modulation [Hz]. Positive: drifts down in frequency.

  • density (float) – Spectral modulation [cycles/octave].

  • depth (float) – scale="linear": modulation depth in [0, 1]. scale="db": peak-to-peak depth in dB (the envelope is 10**((depth/2)*sin(...)/20)).

  • phase (float) – Starting phase [radians].

rate: float
density: float
depth: float = 0.9
phase: float = 0.0
scale: str = 'linear'
property direction: str[source]
envelope(t: ndarray, x: ndarray) → ndarray[source]
class RippleSum(components: tuple[Ripple, ...])[source]

Bases: _Pattern

A sum of ripples sharing one depth scale. Linear depths must total <= 1 so the envelope stays non-negative.

components: tuple[Ripple, ...]
property scale: str[source]
envelope(t: ndarray, x: ndarray) → ndarray[source]
class DynamicRipple(rate_range: tuple[float, float] = (-350.0, 350.0), density_range: tuple[float, float] = (0.0, 4.0), rate_change: float = 3.0, density_change: float = 6.0, depth: float = 45.0, seed: int | None = None, grid_fs: float = 1000.0)[source]

Bases: _Pattern

Dynamic moving ripple (Escabí & Schreiner, 2002).

The envelope is 10**((depth/2) * sin(2*pi*density(t)*x + Phi(t)) / 20) with Phi(t) = 2*pi * integral of rate(t). rate(t) and density(t) are independent, slowly varying random processes, uniformly distributed over their ranges, whose fastest changes are limited to rate_change and density_change Hz.

Note that rate(t) is the temporal modulation at x = 0 (f_lo). Elsewhere the local rate is rate(t) + x * d(density)/dt, because a changing density fans the ripple out across frequency, so some of the modulation power lies outside rate_range: about 7% for the defaults over 5 octaves (14% in the top octave), under 1% with density_change=0.25 (measured by tools/measure_docstring_numbers.py). Analyze with ModulationSpectrum.octave(..., scale="db"), since the pattern is defined in dB.

The defaults follow the ranges commonly used after Escabí & Schreiner (2002); check them against the study you are matching.

The pattern is fully determined by seed (drawn at random if not given, and stored), and it doesn’t depend on the sampling rate.

rate_range: tuple[float, float] = (-350.0, 350.0)
density_range: tuple[float, float] = (0.0, 4.0)
rate_change: float = 3.0
density_change: float = 6.0
depth: float = 45.0
seed: int | None = None
grid_fs: float = 1000.0
property scale: str[source]
trajectories(t: ndarray) → tuple[ndarray, ndarray][source]

rate(t) [Hz] and density(t) [cycles/octave] at times t.

envelope(t: ndarray, x: ndarray) → ndarray[source]
ripple_sound(pattern: Ripple | RippleSum | DynamicRipple | Callable[[ndarray, ndarray], ndarray], duration: float, fs: float, f_lo: float = 250.0, f_hi: float = 8000.0, carrier: str | Sound = 'tones', tones_per_octave: float = 20.0, f0: float = 100.0, bands_per_octave: float = 24.0, rng=None, chunk: int = 32) → Sound[source]

Synthesize a sound whose spectrotemporal envelope is pattern.

Parameters:
  • pattern – A Ripple, RippleSum, DynamicRipple, or a function f(t, x) of time [s] and octaves above f_lo.

  • f_lo – Frequency range [Hz]; x runs from 0 to log2(f_hi/f_lo).

  • f_hi – Frequency range [Hz]; x runs from 0 to log2(f_hi/f_lo).

  • carrier – "tones", "harmonic", "noise", "low-noise", or a Sound.

  • tones_per_octave – Density of the tone carrier. Spectral modulation up to half this (cycles/octave) is representable.

  • f0 – Fundamental of the harmonic carrier. Harmonic spacing in octaves is coarse at low harmonic numbers, which limits the representable ripple density there; a warning is issued when it’s exceeded.

  • bands_per_octave – Channel density for noise and Sound carriers.

  • (tones (The result has RMS = 1. Phases)

  • from (harmonics) and noise are drawn)

  • rng.

render(pattern: Ripple | RippleSum | DynamicRipple | Callable[[ndarray, ndarray], ndarray], filterbank: Filterbank, duration: float, fs: float) → Envelopes[source]

Evaluate any pattern (including a plain f(t, x)) on the band centers of an octave filterbank, giving Envelopes. Compare it with a sound’s measured envelopes on the same filterbank.