spatial¶
Two ears, heads and rooms: spatialization through HRIRs, binaural cues and reverberation.
sonore.spatial.spatialization¶
Head-related spatialization.
Coordinates¶
Cartesian positions are in meters, head-centered: x = right, y = front,
z = up. Head-centered “hcc” coordinates are (dist [cm], elev [deg],
azim [deg]) with azimuth measured clockwise from the front (90° = right),
as in the PKU-IOA database. (SOFA files use counter-clockwise azimuth;
HRIRSet.from_sofa() converts.)
Moving sources¶
A path is where the source is as a function of the time it emits, in
seconds. A sample emitted at time t_e reaches each ear after that ear’s
delay d(t_e), the onset of the HRIR at the source’s position. Measured
HRIRs such as PKU-IOA’s keep the travel time from the source, so this
delay is the propagation delay plus the interaural delay. move_sound()
solves t = t_e + d(t_e) for every output sample t and reads the sound at
t_e between samples, so the delay changes smoothly, and the Doppler shift
of a source that changes distance comes out of the same read. The rest of
the HRIR, its shape with the onset removed, is switched every hop.
Beyond the measured distances, distance acts only through travel time and
level: the HRIR of the nearest measured distance in the same direction,
delayed by the extra distance over the speed of sound and scaled by the
ratio of distances (1/r).
- linear_trajectory(start: ArrayLike, end: ArrayLike, n: int) ndarray[source]¶
Straight line between two Cartesian points; shape
(n, 3).
- circular_trajectory(start_hcc: ArrayLike, end_hcc: ArrayLike, n: int) ndarray[source]¶
Interpolate linearly in
(dist, elev, azim)and return Cartesian points, shape(n, 3). E.g.(100, 0, 90) -> (100, 0, -90)sweeps from right to left through the front at 1 m.
- hcc_trajectory(dist, elev, azim) Callable[[ndarray], ndarray][source]¶
A path in head-centered coordinates, as a function of time.
Each coordinate is a number, a
(times, values)pair (interpolated linearly, held before the first and after the last time), or a function of time in seconds.distis in cm,elevandazimin degrees (azimuth clockwise from the front). Interpolating in these coordinates, rather than between Cartesian points, keeps an arc an arc.Returns a function
t -> positions, Cartesian [m], shape(n, 3), whichmove_sound()takes as a trajectory. For example, the azimuth swing of Cho & Kidd (2022), 30 degrees to either side at 2 Hz, 1 m away:so.hcc_trajectory(100, 0, lambda t: 30 * np.sin(2 * np.pi * 2 * t))
- distance_gain_db(distance: ArrayLike, ref: float = 1.0) ndarray[source]¶
Inverse-square-law level change [dB] for a point source re
refmeters:-20*log10(d/ref), i.e. -6 dB per doubling of distance.
- class HRIRSet(irs: ndarray, positions: ndarray, fs: float, align: bool = True, onset_threshold_db: float = -20.0, _pre: int = 4)[source]¶
Bases:
objectA set of measured head-related impulse responses.
- Parameters:
irs (numpy.ndarray) – Shape
(n_positions, 2, n_taps)(left, right).positions (numpy.ndarray) – Cartesian source positions [m], shape
(n_positions, 3).fs (float) – Sampling rate of the IRs.
align (bool) – Interpolate onset-aligned IRs and interpolate the onset delays separately. Avoids the comb filtering you get from averaging HRIRs with different delays.
- irs: ndarray¶
- positions: ndarray¶
- fs: float¶
- align: bool = True¶
- onset_threshold_db: float = -20.0¶
- classmethod from_pku_ioa(directory: str | PathLike, fs: float = 65536, **kwargs) HRIRSet[source]¶
Load the PKU-IOA database from its original
.datfiles (azi{A}_elev{E}_dist{D}.dat, float64, left then right), found indirectoryor any folder below it, such as the distributeddist{D}/elev{E}/layout. Empty files are skipped with a warning.load_hrirs()downloads the SOFA copy instead.
- classmethod from_sofa(path: str | PathLike, **kwargs) HRIRSet[source]¶
Load a SOFA
SimpleFreeFieldHRIRfile (needsh5py).
- spatialize(sound: Sound, position: ArrayLike, hrirs: HRIRSet, speed_of_sound: float = 343.0) Sound[source]¶
Render a mono sound at a fixed Cartesian position [m].
Beyond the measured distances (or off the one distance of a single-distance set) the HRIR of the nearest measured distance in the same direction is delayed by the extra distance over
speed_of_soundand scaled by the ratio of distances (1/r).
- move_sound(sound: Sound, trajectory, hrirs: HRIRSet, *, room: Sound | None = None, drr_db: float = 0.0, hop: float = 0.005, speed_of_sound: float = 343.0) Sound[source]¶
Render a mono sound moving along a path, through measured HRIRs.
- Parameters:
trajectory –
Where the source is, as a function of the time it emits:
a function
t -> positions, Cartesian [m], shape(n, 3), such ashcc_trajectory()returns;a
(times, points)pair:(N,)times [s] and(N, 3)Cartesian points, interpolated linearly in between and held before and after;an
(N, 3)array of points, passed at a constant rate over the sound’s duration (linear_trajectory(),circular_trajectory()).
hrirs – The measured HRIRs. Between measured distances they are interpolated; beyond them, distance acts through travel time and 1/r only (see the module notes). A set with one distance renders every distance that way.
room – A reverberant tail, one or two channels, such as
so.synth_ir(rt60, fs, n_channels=2)(withoutdrr_db, so it has no direct sound). It is driven by the source as it arrives, without the 1/r, so its level stays put while the direct sound falls with distance, and each ear hears it through the HRIRs averaged over directions, since reverberation arrives from everywhere.drr_db – Direct-to-reverberant energy ratio at 1 m, averaged over directions.
hop – Spacing [s] at which the HRIR shapes are interpolated and switched with raised-cosine cross-fades.
speed_of_sound – In m/s, for distances beyond the measured ones.
delay (Each ear reads the sound through its own)
the (the HRIR onset at)
position (source's)
sample (which changes smoothly from sample to)
a (so)
being (source changing distance glides in pitch (Doppler) instead of)
delays (cross-faded between fixed)
output (which would comb-filter. The)
later. (starts when the source starts emitting; each ear hears it a delay)
sonore.spatial.hrir_data¶
Public HRIR databases, downloaded on demand and cached.
sonore bundles no HRIRs. load_hrirs() downloads the files a database
needs the first time they are asked for, checks each against a pinned SHA-256,
and keeps them in a cache directory: $SONORE_DATA_DIR if set, else the
platform’s user cache directory.
- data_dir() Path[source]¶
Where downloaded data is cached:
$SONORE_DATA_DIR, else the user cache directory.
- load_hrirs(name: str = 'pku-ioa', distances=None, cache_dir: str | PathLike | None = None, **kwargs) HRIRSet[source]¶
Load a public HRIR database, downloading its files on first use.
- Parameters:
name – Database name, a key of
HRIR_DATABASES."pku-ioa": KEMAR at 20, 30, 40, 50, 75, 100, 130 and 160 cm, about 13 MB per distance. Cite Qu et al. (2009) when you use it.distances – Source distances [cm] to load. Several distances give a set that interpolates across distance as well as direction.
Noneloads the database’s default (1 m for PKU-IOA),"all"every distance.cache_dir – Overrides
data_dir().**kwargs – Passed to
HRIRSet, e.g.align.
sonore.spatial.binaural¶
Binaural manipulation and analysis.
Sign convention throughout: positive ITD = right ear leads and positive ILD = right ear louder, i.e. positive values point to the right.
- apply_itd_ild(sound: Sound, itd: float = 0.0, ild: float = 0.0) Sound[source]¶
Impose an ITD [s] and ILD [dB] on a mono (or diotic stereo) sound.
The lagging ear is delayed by
|itd|(fractional delays are exact up to band-limiting); the ILD is split symmetrically,±ild/2per ear.
- simple_bir(fs: float, itd: float = 0.0, ild: float = 0.0, half_width: int = 32) Sound[source]¶
A binaural impulse response with only an ITD [s] and ILD [dB].
Each ear is a (possibly fractionally delayed) unit impulse scaled by
±ild/2dB. Fractional delays use a windowed sinc, so the IR startshalf_widthsamples late whenever the ITD isn’t a whole number of samples.
- class InterauralCues(t: ndarray, itd: ndarray, ild: ndarray, iac: ndarray, corr0: ndarray, cfs: ndarray | None = None)[source]¶
Bases:
ViewShort-time interaural cues. Arrays are
(n_windows,)for broadband analysis or(n_windows, n_bands)per band. Silent time windows are NaN. A view: it keeps a few numbers per time window and drops the sounds.- discards = 'InterauralCues keeps only the time and level differences and the coherence between the ears in each time window: it discards the sounds themselves.'¶
- t: ndarray¶
- itd: ndarray¶
- ild: ndarray¶
- iac: ndarray¶
- corr0: ndarray¶
- cfs: ndarray | None = None¶
- plot(ax=None, **kwargs)[source]¶
ITD and ILD on twin axes, with IAC underneath (see
plot_interaural_cues()).
- interaural_cues(sound: Sound, win_dur: float = 0.02, hop_dur: float | None = None, max_itd: float = 0.001, filterbank: Filterbank | None = None, silence_db: float = -40.0) InterauralCues[source]¶
Windowed ITD, ILD and interaural coherence.
ITD is the lag (within
±max_itd) of the peak of the normalized cross-correlation, refined by parabolic interpolation.iacis that peak’s height (interaural coherence);corr0is the zero-lag correlation, which can be negative (use it for Oscor/Phasewarp). Time windows more thansilence_dbbelow the loudest one are NaN. If afilterbankis given, cues are computed per band.
- oscor(duration: float, fs: float, f_mod: float, rng=None, **noise_kwargs) Sound[source]¶
Oscillating-correlation noise (Siveke et al., 2008): the interaural correlation follows
sin(2*pi*f_mod*t).
- phasewarp(duration: float, fs: float, f_mod: float, rng=None, **noise_kwargs) Sound[source]¶
Phasewarp noise (Siveke et al., 2008): every component’s interaural phase difference rotates at
f_modHz, so the IPD sweeps through 360° and the zero-lag interaural correlation followscos(2*pi*f_mod*t).Implemented as a single-sideband frequency shift of the left-ear noise.
sonore.spatial.reverb¶
Synthetic room impulse responses from the statistics of natural reverberation (Traer & McDermott, 2016).
- synth_ir(rt60: float, fs: float, drr_db: float | None = None, decay_db: float = 60.0, n_bands: int = 32, f_lo: float = 20.0, f_hi: float = 16000.0, n_channels: int = 1, decay_shape: str = 'exponential', rt60_profile: str = 'ecological', drr_profile: str = 'ecological', rng=None) Sound[source]¶
Synthesize a room impulse response (Traer & McDermott, 2016).
Gaussian noise is split into ERB bands, each band is multiplied by a decay envelope, and the bands are re-filtered and summed. With the defaults the decay is exponential with frequency-dependent RT60s and onset levels taken from the paper’s regressions on 271 real-world rooms, scaled to the requested broadband
rt60(the median across bands). Listeners can’t tell such IRs from real ones.The paper’s “atypical” IRs, which listeners readily hear as unnatural, are available through
decay_shape,rt60_profileanddrr_profile. Note that in the paper each atypical IR was further adjusted to distort sounds as much as the ecological one (equating cochleagram error); that psychophysical control isn’t done here, so variants can differ somewhat in loudness and audible length.- Parameters:
rt60 – Median reverberation time [s].
drr_db – Direct-to-reverberant energy ratio [dB]. If given, a unit impulse is placed at t=0 and the tail is scaled to this DRR. If None, only the (unit-RMS) reverberant tail is returned.
decay_db – For exponential decays, truncate once the slowest band has decayed this many dB.
n_channels – Independent tails per channel (e.g. 2 for a decorrelated binaural tail); the direct impulse is identical in all channels.
decay_shape –
"exponential"(natural),"time_reversed","linear_matched_start"(linear from the natural starting level, same energy per band), or"linear_matched_end"(linear, reaching zero where the natural decay is 60 dB down, same energy per band).rt60_profile – Frequency dependence of decay; see
band_rt60s().drr_profile –
"ecological"onset levels per band, or"constant"(their mean).
- band_rt60s(rt60: float, freqs: ndarray, profile: str = 'ecological') ndarray[source]¶
Per-band RT60s [s] at
freqsfor a broadband (median) RT60, from the regression fits of Traer & McDermott (2016, Eq. S11), or one of the paper’s atypical variants (Eq. S15 and following):"ecological": mid frequencies decay slowest, as in real rooms."inverted":max + min - ecological; slow where real rooms are fast."exaggerated": the profile of a room twice as reverberant, scaled down to this length (more sharply peaked than real rooms)."reduced": the profile of a room half as reverberant, scaled up (flatter).
- measure_rt60(ir: Sound, n_bands: int = 30, f_lo: float = 50.0, f_hi: float = 8000.0, fit_range_db=(-5.0, -25.0)) tuple[ndarray, ndarray][source]¶
Per-band RT60s of an impulse response:
(center freqs [Hz], RT60s [s]).Each ERB band’s energy decay curve (Schroeder backward integration) is fit with a line over
fit_range_dband extrapolated to -60 dB (so the default measures T20 x 3). The direct sound, if any, should be removed first; bands that never decay through the fit range give NaN.