sources¶
Sounds made from parameters alone: tones, noises, chirps and the LF glottal source, Klatt’s formant synthesizer, and spectrotemporal ripples.
sonore.sources.waveforms¶
Waveforms made from parameters: silence, tones, harmonic complexes, chirps, noises, and the LF glottal source.
Every generator returns a Sound normalized to RMS = 1.
Anything random takes an rng argument (a seed or np.random.Generator)
so stimuli can be regenerated exactly.
- silence(duration: float, fs: float, n_channels: int = 1) Sound[source]¶
Zeros,
durationseconds long, inn_channelschannels (RMS 0, the one generator not scaled to 1).
- pure_tone(duration: float, fs: float, freq: float, phase: float = 0.0) Sound[source]¶
cos(2*pi*freq*t + phase).
- harmonic_complex(duration: float, fs: float, f0: float, harmonics: ArrayLike | None = None, amplitudes: ArrayLike | Callable[[ndarray, ndarray], ArrayLike] | Callable[[ndarray, ndarray, int], ArrayLike] | None = None, phases: ArrayLike | str = 'cosine', rng: int | Generator | None = None, *, f_max: float | None = None) Sound[source]¶
- harmonic_complex(duration: float, fs: float, f0: tuple[ArrayLike, ArrayLike] | F0Contour, harmonics: ArrayLike | None = None, amplitudes: ArrayLike | Callable[[ndarray, ndarray], ArrayLike] | Callable[[ndarray, ndarray, int], ArrayLike] | None = None, phases: ArrayLike | str = 'cosine', rng: int | Generator | None = None, *, f_max: float | None = None, ramp: float = 0.005, unvoiced: str = 'silence') Sound
Sum of harmonics
a_n cos(n * Phi(t) + phi_n)of a fixed or moving F0.f0is either a number, forPhi = 2*pi*f0*t, or an F0 contour: a(times, values)pair or anything with.tand.f0, such as anF0Track, with values in Hz and 0 where unvoiced (one row per channel for a multichannel sound).harmonicsare the harmonic numbers; by default every one below Nyquist (fixed F0) or belowf_max(contour).amplitudesis one value per harmonic, or a functionamplitudes(t, f)giving the gain at timestand frequenciesf(a spectral envelope, evaluated at every harmonic’s frequency). The function may also take the harmonic number as a third argument,amplitudes(t, f, n), and may return complex values: harmonicnis thenRe(a_n exp(i(n Phi + phi_n))), so the angle ofa_nadds to its phase and can change over time (an LF glottal pulse whose shape changes, as inglottal_source()).amplitudesmay also be a spectral envelope, such asSpectralEnvelopeorGridEnvelope(power on a grid of times and frequencies), which is read point by point as amplitude, from its first channel.phasesare starting phases: an array, one per harmonic, or one of"cosine","sine","alternating","random","schroeder+","schroeder-".A fixed F0 drops harmonics at or above Nyquist, with a warning. Given
f_max, it also takes harmonics only belowf_maxby default and fades them out as a contour does (below); without it, nothing is faded.A contour (
docs/design/sources/harmonic-source.md) is filled across its unvoiced gaps and interpolated linearly to the sample rate, and the phase is its exact running integral, so every harmonic followsntimes the contour with no jumps. Each harmonic fades out (cos^2) between0.9 * f_maxandf_max(default0.45 * fs), so nothing aliases as the pitch rises. Harmonics are switched off where the contour is unvoiced, with Hann rampsrampseconds long;unvoiced="noise"fills those stretches with white noise of the harmonics’ power instead of silence. Values are held beyond the first and last times. With a contour,phaseshas one value per harmonic number up to the highest that fits belowf_maxat the lowest voiced F0, unlessharmonicsis given.
- schroeder_complex(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, n_harmonics: int | None = None, sign: int = 1, **contour) Sound[source]¶
Equal-amplitude harmonic complex with Schroeder phases
sign * pi * n(n+1)/N.n_harmonicsdefaults to all below Nyquist (or belowf_maxfor an F0 contour).f0and the keyword arguments for a contour are as inharmonic_complex().
- square_wave(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, phase: float = 0.0, bandlimited: bool = True, **contour) Sound[source]¶
Square wave. Band-limited (odd harmonics below Nyquist) unless
bandlimited=False, which aliases badly at high f0. A band-limited square wave also takes an F0 contour, asharmonic_complex()does.
- sawtooth_wave(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, phase: float = 0.0, bandlimited: bool = True, **contour) Sound[source]¶
Rising sawtooth. Band-limited unless
bandlimited=False. A band-limited sawtooth also takes an F0 contour, asharmonic_complex()does.
- pulse_train(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, phase: float = 0.0, bandlimited: bool = True, **contour) Sound[source]¶
Periodic pulse (click) train.
Band-limited: equal-amplitude cosine-phase harmonics up to Nyquist; this also takes an F0 contour, as
harmonic_complex()does. Otherwise single-sample impulses at the nearest sample to each period.phase(radians of the fundamental) shifts the pulses in time.
- linear_chirp(duration: float, fs: float, f0: float, f1: float, phase: float = 0.0) Sound[source]¶
Linear frequency sweep from
f0tof1.
- exponential_chirp(duration: float, fs: float, f0: float, f1: float, phase: float = 0.0) Sound[source]¶
Exponential (log-frequency) sweep from
f0tof1.
- gaussian_noise(duration: float, fs: float, band: tuple[float, float] | None = None, tilt: float = 0.0, spectrum: Callable[[ndarray], ndarray] | tuple[ArrayLike, ArrayLike] | None = None, n_channels: int = 1, rng: int | Generator | None = None) Sound[source]¶
Gaussian noise, optionally spectrally shaped.
White Gaussian noise for the whole duration is shaped in the frequency domain: each bin of its FFT is multiplied by
10^(L(f)/20), the target levelL(f)in dB at that bin fromband,tiltandspectrumtogether, then transformed back, its mean removed, and scaled to RMS 1. The shape is the expected spectrum: the bins of any one draw scatter around it, as they do for any Gaussian noise.- Parameters:
band –
(f_lo, f_hi)passband in Hz (brick-wall).tilt – Spectral slope in dB/octave (
-3= pink, power1/f;-6= brown, power1/f^2).spectrum – Target spectrum level in dB (its shape only; the output is scaled to RMS 1): a
Spectrum, a functionf -> dB, or a(freqs, dB)pair, interpolated linearly in Hz. ASpectrumis silent outside its frequencies, since it was measured only there; a pair is held flat beyond its ends.n_channels – Independent noise in each channel.
Two-channel noise whose interaural correlation is exactly
corr(in [-1, 1]), each channel at RMS 1. Extra keyword arguments go togaussian_noise().The symmetric-generator method, made exact as Hartmann and Cho (2011) describe: two independent noises
x1andx2are made orthogonal (Gram-Schmidt) and equal in power, then mixed asa x1 + b x2(left) anda x1 - b x2(right) witha^2 - b^2 = corr. Without those two steps each draw would misscorrby chance, more so for few spectral components (narrow bands, short sounds). The correlation is the zero-lag one, over the whole sound.
- iterated_ripple_noise(duration: float, fs: float, delay: float, gain: float = 1.0, iterations: int = 16, network: str = 'add-same', rng: int | Generator | None = None, **noise_kwargs) Sound[source]¶
Iterated rippled noise (Yost, 1996). Pitch is at
1/delay.network="add-same"(IRNS) iteratesy <- y + g*delay(y);"add-original"(IRNO) iteratesy <- x + g*delay(y). Implemented exactly in the frequency domain (fractional delays are fine), with a warm-up segment discarded so the output is stationary.
- lf_harmonics(harmonics: ArrayLike, rd: float = 0.7, *, ra: float | None = None, rg: float | None = None, rk: float | None = None, flow: bool = False) ndarray[source]¶
Complex Fourier coefficients of the LF pulse, one per harmonic number.
Harmonic
kof a train of LF pulses is2 |c_k| cos(k Phi + angle(c_k))wherePhiis the running phase of F0, soc_kdepends only onkand the pulse’s shape, not on F0. The coefficients are those of the flow derivative scaled to -1 at the main excitation (Ee = 1); withflow=True, those of the flow itself,c_k / (2 pi i k), in units ofEe * T0. They are exact (a closed form), so a source built from them does not alias.The shape is set by Fant’s (1995)
rd(0.3-2.7, his main range, the one his predictions were fitted to; 0.7 is close to his typical adult male values), through his prediction ofra,rgandrkfrom it; or directly byra = ta/T0,rg = T0/(2 tp)andrk = (te - tp)/tp(below 1), given together.
- lf_pulse(x: ArrayLike, rd: float = 0.7, *, ra: float | None = None, rg: float | None = None, rk: float | None = None, flow: bool = False) ndarray[source]¶
The LF flow derivative (or with
flow=True, the flow) at timesxin fractions of a period (wrapped into [0, 1)), scaled to -1 at the main excitationte(Ee = 1; for the laxest voices, near Rd 2.7, the open phase dips slightly below that just before it).Evaluating the formula at sample times aliases, most for tense voices with an abrupt closure; use it to draw a period, and
glottal_source()to make a sound. The shape is set as inlf_harmonics().
- glottal_source(duration: float, fs: float, f0: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] | F0Contour, rd: float | tuple[TypeAliasForwardRef('ArrayLike'), TypeAliasForwardRef('ArrayLike')] = 0.7, *, flow: bool = False, f_max: float | None = None, ramp: float = 0.005) Sound[source]¶
A voiced source of LF glottal pulses on a fixed F0 or an F0 contour.
The LF model (Fant, Liljencrants & Lin, 1985) describes one period of the glottal flow derivative: an exponentially growing sinusoid while the glottis opens, a sharp negative peak where it closes, and an exponential return phase. Fant’s (1995) single shape parameter Rd sets the whole pulse, from tense, pressed voice (small Rd: short open phase, abrupt closure, strong high harmonics) to lax, breathy voice (large Rd: long open phase, gradual closure, a dominant fundamental). The design, and the checks behind it, are in
docs/design/sources/glottal-source.md.The pulses are built from their harmonics (
lf_harmonics()) byharmonic_complex(), so they do not alias, each period takes its own length on a moving F0, andf0,f_maxandrampwork as there (unvoiced time windows are silent, and a contour that is never voiced gives silence). The result is the flow derivative, the source as it excites the vocal tract with radiation folded in, or withflow=Truethe flow itself; normalized to RMS 1.rdis Fant’s (1995) shape parameter (0.3-2.7, default 0.7, close to his typical adult male values): a number, or a(times, values)track for a voice quality that changes, interpolated to every sample. LF models the periodic pulse only: breathy voice also needs aspiration noise, as inklatt_synthesize().
sonore.sources.klatt¶
A Klatt-style cascade/parallel formant synthesizer.
Speech from named acoustic parameters, after Klatt (1980): a voiced source
and two noise sources, formant resonators in cascade (voicing and
aspiration) and in parallel (frication), and radiation from the lips. Every
control is something visible on a spectrogram, so the source-filter model
can be taken apart one parameter at a time, and stimuli such as vowel and
/ba/-/da/-/ga/ continua are set exactly. The design, and the checks behind
it, are in docs/design/sources/klatt.md.
- KLATT_DEFAULTS = mappingproxy({'F0': 100.0, 'AV': 60.0, 'SS': 1, 'RD': 0.7, 'AH': 0.0, 'AF': 0.0, 'AB': 0.0, 'F1': 500.0, 'B1': 60.0, 'F2': 1500.0, 'B2': 90.0, 'F3': 2500.0, 'B3': 150.0, 'F4': 3500.0, 'B4': 200.0, 'F5': 4500.0, 'B5': 250.0, 'F6': 4900.0, 'B6': 1000.0, 'A1': 0.0, 'A2': 0.0, 'A3': 0.0, 'A4': 0.0, 'A5': 0.0, 'A6': 0.0, 'FNP': 250.0, 'BNP': 100.0, 'FNZ': 250.0, 'BNZ': 100.0})¶
a neutral adult male vowel, voiced at 100 Hz, with no noise. Frequencies and bandwidths are in Hz, amplitudes in dB (60 = the source at unit level, 0 or less = off).
- Type:
Every parameter
klatt_synthesize()takes, with its default
- klatt_synthesize(duration: float, fs: float, params: Mapping[str, float | tuple[ArrayLike, ArrayLike]] | None = None, *, rng: int | Generator | None = None, **kwargs: float | tuple[ArrayLike, ArrayLike]) Sound[source]¶
Speech from Klatt’s (1980) parameters, as a cascade/parallel formant synthesizer.
Parameters are given by Klatt’s names, in
params(a mapping, such as one made byklatt_continuum()) or as keyword arguments, which take precedence; anything not given takes its value inKLATT_DEFAULTS. Each is a number, or a(times, values)pair: times in seconds, interpolated linearly to every sample and held beyond the ends. A table of values every 5 ms, as Klatt used, is the pair(0.005 * np.arange(len(values)), values).F0may also be anF0Track, to resynthesize a measured pitch contour.Sources. Voicing: harmonics of
F0(harmonic_complex(), so the pitch can glide with no phase jumps or aliasing) at levelAV, with the spectrum set by the source switchSS, numbered as in KLSYN88 (Klatt & Klatt, 1990).SS = 1(the default): Klatt’s (1980) impulses through his glottal low-pass, falling about 12 dB per octave.SS = 3: Liljencrants-Fant glottal pulses (glottal_source()), their shape set by Fant’sRD, from tense (0.3, strong high harmonics) to lax (2.7, a dominant fundamental); at the default 0.7 its 3 kHz harmonic is about 12 dB weaker, relative to the first, than withSS = 1. KLSYN88’sSS = 2(KLGLOTT88) is not included. WhereF0is 0 the harmonics are switched off. Aspiration (AH) and frication (AF,AB): white Gaussian noise. While voicing is on, both noises are amplitude modulated by a square wave atF0, 50% deep, as in Klatt’s synthesizer.Cascade branch. Voicing and aspiration pass through the nasal pole and zero (
FNP,BNP,FNZ,BNZ) and formants 1-5 (F1-F5, bandwidthsB1-B5) in series: formant levels follow from the frequencies, as in a vowel. A formant default at or above Nyquist is left out (F5atfsbelow 9 kHz).Parallel branch. Frication passes through formants 1-6 side by side, each at its own level
A1-A6(dB at the formant’s peak), added with alternating signs so that they sum like the cascade between peaks; formants 2-6 get a differenced input, as in Klatt’s, to keep low frequencies out.ABadds the frication with no formants.Radiation. The voiced source is differenced (+6 dB per octave), the radiation from the lips; the noises are white as they leave the lips.
Amplitudes are in dB: at 60 a source has RMS 1 as it leaves the lips (before the formants), so equal values mean equal source levels; each 20 dB is a factor of 10, and 0 or less is off. The result is normalized to RMS 1, like every generator; relative levels within one call are kept. Klatt’s quasi-sinusoidal voicing (
AVS) and KLSYN88’s other voice-quality controls (open quotient, spectral tilt, flutter, double pulsing) are not included.rngseeds the noise, so a call can be repeated exactly.
- klatt_continuum(start: Mapping[str, float | tuple[ArrayLike, ArrayLike]], end: Mapping[str, float | tuple[ArrayLike, ArrayLike]], steps: int) list[dict[str, float | tuple[ArrayLike, ArrayLike]]][source]¶
stepsparameter sets forklatt_synthesize(), evenly spaced fromstarttoend(both included), such as a /ba/-/da/ continuum.A parameter given in only one of them takes its
KLATT_DEFAULTSvalue in the other. The source switchSSis not interpolated: it must be the same at both ends. Numbers are interpolated directly;(times, values)pairs are interpolated value by value, so they must share their times (a number paired with a track is held at every one of its times).
sonore.sources.ripples¶
Spectrotemporal ripples: sounds defined by their modulation content.
A pattern is an envelope over time and log-frequency, E(t, x), where
x is octaves above f_lo. ripple_sound() imposes a pattern on a
carrier, which supplies the fine structure.
Patterns¶
RippleA moving ripple (Kowalski, Depireux & Shamma, 1996; Chi et al., 1999):
1 + depth * sin(2*pi*(rate*t + density*x) + phase).rateis in Hz anddensityin cycles/octave. With positive rate and density the ripple drifts downward in frequency; a negative rate drifts upward. Ripples add:Ripple(4, 1) + Ripple(-8, 2)is aRippleSum.DynamicRippleA dynamic moving ripple (Escabí & Schreiner, 2002) whose rate and density wander slowly and randomly within given ranges, for STRF estimation.
- Any callable
f(t, x) Evaluated with
tas a row andxas a column, so ordinary numpy broadcasting works, e.g.lambda t, x: 1 + 0.5 * np.sin(2*np.pi*(3*t + x**2)).
Carriers¶
"tones": log-spaced tones with random phases, the classic ripple carrier.
"harmonic": harmonics of f0. "noise": narrowband Gaussian noise
per channel. "low-noise": the fine structure of that noise with its
envelope flattened, so it adds no envelope fluctuations of its own. Or any
Sound, whose fine structure is used the same way.
All carriers are scaled to equal energy per octave, so switching carriers changes the fine structure but not the long-term spectrum.
- class Ripple(rate: float, density: float, depth: float = 0.9, phase: float = 0.0, scale: str = 'linear')[source]¶
Bases:
_PatternA moving ripple
1 + depth*sin(2*pi*(rate*t + density*x) + phase).- Parameters:
rate (float) – Temporal modulation [Hz]. Positive: drifts down in frequency.
density (float) – Spectral modulation [cycles/octave].
depth (float) –
scale="linear": modulation depth in [0, 1].scale="db": peak-to-peak depth in dB (the envelope is10**((depth/2)*sin(...)/20)).phase (float) – Starting phase [radians].
- rate: float¶
- density: float¶
- depth: float = 0.9¶
- phase: float = 0.0¶
- scale: str = 'linear'¶
- class RippleSum(components: tuple[Ripple, ...])[source]¶
Bases:
_PatternA sum of ripples sharing one depth scale. Linear depths must total <= 1 so the envelope stays non-negative.
- class DynamicRipple(rate_range: tuple[float, float] = (-350.0, 350.0), density_range: tuple[float, float] = (0.0, 4.0), rate_change: float = 3.0, density_change: float = 6.0, depth: float = 45.0, seed: int | None = None, grid_fs: float = 1000.0)[source]¶
Bases:
_PatternDynamic moving ripple (Escabí & Schreiner, 2002).
The envelope is
10**((depth/2) * sin(2*pi*density(t)*x + Phi(t)) / 20)withPhi(t) = 2*pi * integral of rate(t).rate(t)anddensity(t)are independent, slowly varying random processes, uniformly distributed over their ranges, whose fastest changes are limited torate_changeanddensity_changeHz.Note that
rate(t)is the temporal modulation atx = 0(f_lo). Elsewhere the local rate israte(t) + x * d(density)/dt, because a changing density fans the ripple out across frequency, so some of the modulation power lies outsiderate_range: about 7% for the defaults over 5 octaves (14% in the top octave), under 1% withdensity_change=0.25(measured by tools/measure_docstring_numbers.py). Analyze withModulationSpectrum.octave(..., scale="db"), since the pattern is defined in dB.The defaults follow the ranges commonly used after Escabí & Schreiner (2002); check them against the study you are matching.
The pattern is fully determined by
seed(drawn at random if not given, and stored), and it doesn’t depend on the sampling rate.- rate_range: tuple[float, float] = (-350.0, 350.0)¶
- density_range: tuple[float, float] = (0.0, 4.0)¶
- rate_change: float = 3.0¶
- density_change: float = 6.0¶
- depth: float = 45.0¶
- seed: int | None = None¶
- grid_fs: float = 1000.0¶
- ripple_sound(pattern: Ripple | RippleSum | DynamicRipple | Callable[[ndarray, ndarray], ndarray], duration: float, fs: float, f_lo: float = 250.0, f_hi: float = 8000.0, carrier: str | Sound = 'tones', tones_per_octave: float = 20.0, f0: float = 100.0, bands_per_octave: float = 24.0, rng=None, chunk: int = 32) Sound[source]¶
Synthesize a sound whose spectrotemporal envelope is
pattern.- Parameters:
pattern – A
Ripple,RippleSum,DynamicRipple, or a functionf(t, x)of time [s] and octaves abovef_lo.f_lo – Frequency range [Hz];
xruns from 0 tolog2(f_hi/f_lo).f_hi – Frequency range [Hz];
xruns from 0 tolog2(f_hi/f_lo).carrier –
"tones","harmonic","noise","low-noise", or a Sound.tones_per_octave – Density of the tone carrier. Spectral modulation up to half this (cycles/octave) is representable.
f0 – Fundamental of the harmonic carrier. Harmonic spacing in octaves is coarse at low harmonic numbers, which limits the representable ripple density there; a warning is issued when it’s exceeded.
bands_per_octave – Channel density for noise and Sound carriers.
(tones (The result has RMS = 1. Phases)
from (harmonics) and noise are drawn)
rng.
- render(pattern: Ripple | RippleSum | DynamicRipple | Callable[[ndarray, ndarray], ndarray], filterbank: Filterbank, duration: float, fs: float) Envelopes[source]¶
Evaluate any pattern (including a plain
f(t, x)) on the band centers of an octave filterbank, givingEnvelopes. Compare it with a sound’s measured envelopes on the same filterbank.