spatial

Two ears, heads and rooms: spatialization through HRIRs, binaural cues and reverberation.

sonore.spatial.spatialization

Head-related spatialization.

Coordinates

Cartesian positions are in meters, head-centered: x = right, y = front, z = up. Head-centered “hcc” coordinates are (dist [cm], elev [deg], azim [deg]) with azimuth measured clockwise from the front (90° = right), as in the PKU-IOA database. (SOFA files use counter-clockwise azimuth; HRIRSet.from_sofa() converts.)

Moving sources

A path is where the source is as a function of the time it emits, in seconds. A sample emitted at time t_e reaches each ear after that ear’s delay d(t_e), the onset of the HRIR at the source’s position. Measured HRIRs such as PKU-IOA’s keep the travel time from the source, so this delay is the propagation delay plus the interaural delay. move_sound() solves t = t_e + d(t_e) for every output sample t and reads the sound at t_e between samples, so the delay changes smoothly, and the Doppler shift of a source that changes distance comes out of the same read. The rest of the HRIR, its shape with the onset removed, is switched every hop. Beyond the measured distances, distance acts only through travel time and level: the HRIR of the nearest measured distance in the same direction, delayed by the extra distance over the speed of sound and scaled by the ratio of distances (1/r).

rect_to_sph(x, y, z)[source]

Cartesian -> (rho, theta=elevation, phi=angle from +x), radians.

sph_to_rect(rho, theta, phi)[source]
rect_to_hcc(x, y, z)[source]
hcc_to_rect(dist, elev, azim)[source]
linear_trajectory(start: ArrayLike, end: ArrayLike, n: int) → ndarray[source]

Straight line between two Cartesian points; shape (n, 3).

circular_trajectory(start_hcc: ArrayLike, end_hcc: ArrayLike, n: int) → ndarray[source]

Interpolate linearly in (dist, elev, azim) and return Cartesian points, shape (n, 3). E.g. (100, 0, 90) -> (100, 0, -90) sweeps from right to left through the front at 1 m.

hcc_trajectory(dist, elev, azim) → Callable[[ndarray], ndarray][source]

A path in head-centered coordinates, as a function of time.

Each coordinate is a number, a (times, values) pair (interpolated linearly, held before the first and after the last time), or a function of time in seconds. dist is in cm, elev and azim in degrees (azimuth clockwise from the front). Interpolating in these coordinates, rather than between Cartesian points, keeps an arc an arc.

Returns a function t -> positions, Cartesian [m], shape (n, 3), which move_sound() takes as a trajectory. For example, the azimuth swing of Cho & Kidd (2022), 30 degrees to either side at 2 Hz, 1 m away:

so.hcc_trajectory(100, 0, lambda t: 30 * np.sin(2 * np.pi * 2 * t))
distance_gain_db(distance: ArrayLike, ref: float = 1.0) → ndarray[source]

Inverse-square-law level change [dB] for a point source re ref meters: -20*log10(d/ref), i.e. -6 dB per doubling of distance.

class HRIRSet(irs: ndarray, positions: ndarray, fs: float, align: bool = True, onset_threshold_db: float = -20.0, _pre: int = 4)[source]

Bases: object

A set of measured head-related impulse responses.

Parameters:
  • irs (numpy.ndarray) – Shape (n_positions, 2, n_taps) (left, right).

  • positions (numpy.ndarray) – Cartesian source positions [m], shape (n_positions, 3).

  • fs (float) – Sampling rate of the IRs.

  • align (bool) – Interpolate onset-aligned IRs and interpolate the onset delays separately. Avoids the comb filtering you get from averaging HRIRs with different delays.

irs: ndarray
positions: ndarray
fs: float
align: bool = True
onset_threshold_db: float = -20.0
classmethod from_pku_ioa(directory: str | PathLike, fs: float = 65536, **kwargs) → HRIRSet[source]

Load the PKU-IOA database from its original .dat files (azi{A}_elev{E}_dist{D}.dat, float64, left then right), found in directory or any folder below it, such as the distributed dist{D}/elev{E}/ layout. Empty files are skipped with a warning. load_hrirs() downloads the SOFA copy instead.

classmethod from_sofa(path: str | PathLike, **kwargs) → HRIRSet[source]

Load a SOFA SimpleFreeFieldHRIR file (needs h5py).

classmethod concat(sets, **kwargs) → HRIRSet[source]

Merge sets measured at the same sampling rate and length, e.g. one per distance, into a single set. kwargs go to HRIRSet.

property distances: ndarray[source]

The measured distances [m], nearest first.

at(positions: ArrayLike, fs: float | None = None) → ndarray[source]

Interpolated HRIRs at Cartesian positions (N, 3), resampled to fs. Returns shape (N, 2, n_taps).

spatialize(sound: Sound, position: ArrayLike, hrirs: HRIRSet, speed_of_sound: float = 343.0) → Sound[source]

Render a mono sound at a fixed Cartesian position [m].

Beyond the measured distances (or off the one distance of a single-distance set) the HRIR of the nearest measured distance in the same direction is delayed by the extra distance over speed_of_sound and scaled by the ratio of distances (1/r).

move_sound(sound: Sound, trajectory, hrirs: HRIRSet, *, room: Sound | None = None, drr_db: float = 0.0, hop: float = 0.005, speed_of_sound: float = 343.0) → Sound[source]

Render a mono sound moving along a path, through measured HRIRs.

Parameters:
  • trajectory –

    Where the source is, as a function of the time it emits:

    • a function t -> positions, Cartesian [m], shape (n, 3), such as hcc_trajectory() returns;

    • a (times, points) pair: (N,) times [s] and (N, 3) Cartesian points, interpolated linearly in between and held before and after;

    • an (N, 3) array of points, passed at a constant rate over the sound’s duration (linear_trajectory(), circular_trajectory()).

  • hrirs – The measured HRIRs. Between measured distances they are interpolated; beyond them, distance acts through travel time and 1/r only (see the module notes). A set with one distance renders every distance that way.

  • room – A reverberant tail, one or two channels, such as so.synth_ir(rt60, fs, n_channels=2) (without drr_db, so it has no direct sound). It is driven by the source as it arrives, without the 1/r, so its level stays put while the direct sound falls with distance, and each ear hears it through the HRIRs averaged over directions, since reverberation arrives from everywhere.

  • drr_db – Direct-to-reverberant energy ratio at 1 m, averaged over directions.

  • hop – Spacing [s] at which the HRIR shapes are interpolated and switched with raised-cosine cross-fades.

  • speed_of_sound – In m/s, for distances beyond the measured ones.

  • delay (Each ear reads the sound through its own)

  • the (the HRIR onset at)

  • position (source's)

  • sample (which changes smoothly from sample to)

  • a (so)

  • being (source changing distance glides in pitch (Doppler) instead of)

  • delays (cross-faded between fixed)

  • output (which would comb-filter. The)

  • later. (starts when the source starts emitting; each ear hears it a delay)

sonore.spatial.hrir_data

Public HRIR databases, downloaded on demand and cached.

sonore bundles no HRIRs. load_hrirs() downloads the files a database needs the first time they are asked for, checks each against a pinned SHA-256, and keeps them in a cache directory: $SONORE_DATA_DIR if set, else the platform’s user cache directory.

data_dir() → Path[source]

Where downloaded data is cached: $SONORE_DATA_DIR, else the user cache directory.

load_hrirs(name: str = 'pku-ioa', distances=None, cache_dir: str | PathLike | None = None, **kwargs) → HRIRSet[source]

Load a public HRIR database, downloading its files on first use.

Parameters:
  • name – Database name, a key of HRIR_DATABASES. "pku-ioa": KEMAR at 20, 30, 40, 50, 75, 100, 130 and 160 cm, about 13 MB per distance. Cite Qu et al. (2009) when you use it.

  • distances – Source distances [cm] to load. Several distances give a set that interpolates across distance as well as direction. None loads the database’s default (1 m for PKU-IOA), "all" every distance.

  • cache_dir – Overrides data_dir().

  • **kwargs – Passed to HRIRSet, e.g. align.

sonore.spatial.binaural

Binaural manipulation and analysis.

Sign convention throughout: positive ITD = right ear leads and positive ILD = right ear louder, i.e. positive values point to the right.

apply_itd_ild(sound: Sound, itd: float = 0.0, ild: float = 0.0) → Sound[source]

Impose an ITD [s] and ILD [dB] on a mono (or diotic stereo) sound.

The lagging ear is delayed by |itd| (fractional delays are exact up to band-limiting); the ILD is split symmetrically, ±ild/2 per ear.

simple_bir(fs: float, itd: float = 0.0, ild: float = 0.0, half_width: int = 32) → Sound[source]

A binaural impulse response with only an ITD [s] and ILD [dB].

Each ear is a (possibly fractionally delayed) unit impulse scaled by ±ild/2 dB. Fractional delays use a windowed sinc, so the IR starts half_width samples late whenever the ITD isn’t a whole number of samples.

class InterauralCues(t: ndarray, itd: ndarray, ild: ndarray, iac: ndarray, corr0: ndarray, cfs: ndarray | None = None)[source]

Bases: View

Short-time interaural cues. Arrays are (n_windows,) for broadband analysis or (n_windows, n_bands) per band. Silent time windows are NaN. A view: it keeps a few numbers per time window and drops the sounds.

discards = 'InterauralCues keeps only the time and level differences and the coherence between the ears in each time window: it discards the sounds themselves.'
t: ndarray
itd: ndarray
ild: ndarray
iac: ndarray
corr0: ndarray
cfs: ndarray | None = None
plot(ax=None, **kwargs)[source]

ITD and ILD on twin axes, with IAC underneath (see plot_interaural_cues()).

interaural_cues(sound: Sound, win_dur: float = 0.02, hop_dur: float | None = None, max_itd: float = 0.001, filterbank: Filterbank | None = None, silence_db: float = -40.0) → InterauralCues[source]

Windowed ITD, ILD and interaural coherence.

ITD is the lag (within ±max_itd) of the peak of the normalized cross-correlation, refined by parabolic interpolation. iac is that peak’s height (interaural coherence); corr0 is the zero-lag correlation, which can be negative (use it for Oscor/Phasewarp). Time windows more than silence_db below the loudest one are NaN. If a filterbank is given, cues are computed per band.

oscor(duration: float, fs: float, f_mod: float, rng=None, **noise_kwargs) → Sound[source]

Oscillating-correlation noise (Siveke et al., 2008): the interaural correlation follows sin(2*pi*f_mod*t).

phasewarp(duration: float, fs: float, f_mod: float, rng=None, **noise_kwargs) → Sound[source]

Phasewarp noise (Siveke et al., 2008): every component’s interaural phase difference rotates at f_mod Hz, so the IPD sweeps through 360° and the zero-lag interaural correlation follows cos(2*pi*f_mod*t).

Implemented as a single-sideband frequency shift of the left-ear noise.

sonore.spatial.reverb

Synthetic room impulse responses from the statistics of natural reverberation (Traer & McDermott, 2016).

synth_ir(rt60: float, fs: float, drr_db: float | None = None, decay_db: float = 60.0, n_bands: int = 32, f_lo: float = 20.0, f_hi: float = 16000.0, n_channels: int = 1, decay_shape: str = 'exponential', rt60_profile: str = 'ecological', drr_profile: str = 'ecological', rng=None) → Sound[source]

Synthesize a room impulse response (Traer & McDermott, 2016).

Gaussian noise is split into ERB bands, each band is multiplied by a decay envelope, and the bands are re-filtered and summed. With the defaults the decay is exponential with frequency-dependent RT60s and onset levels taken from the paper’s regressions on 271 real-world rooms, scaled to the requested broadband rt60 (the median across bands). Listeners can’t tell such IRs from real ones.

The paper’s “atypical” IRs, which listeners readily hear as unnatural, are available through decay_shape, rt60_profile and drr_profile. Note that in the paper each atypical IR was further adjusted to distort sounds as much as the ecological one (equating cochleagram error); that psychophysical control isn’t done here, so variants can differ somewhat in loudness and audible length.

Parameters:
  • rt60 – Median reverberation time [s].

  • drr_db – Direct-to-reverberant energy ratio [dB]. If given, a unit impulse is placed at t=0 and the tail is scaled to this DRR. If None, only the (unit-RMS) reverberant tail is returned.

  • decay_db – For exponential decays, truncate once the slowest band has decayed this many dB.

  • n_channels – Independent tails per channel (e.g. 2 for a decorrelated binaural tail); the direct impulse is identical in all channels.

  • decay_shape – "exponential" (natural), "time_reversed", "linear_matched_start" (linear from the natural starting level, same energy per band), or "linear_matched_end" (linear, reaching zero where the natural decay is 60 dB down, same energy per band).

  • rt60_profile – Frequency dependence of decay; see band_rt60s().

  • drr_profile – "ecological" onset levels per band, or "constant" (their mean).

band_rt60s(rt60: float, freqs: ndarray, profile: str = 'ecological') → ndarray[source]

Per-band RT60s [s] at freqs for a broadband (median) RT60, from the regression fits of Traer & McDermott (2016, Eq. S11), or one of the paper’s atypical variants (Eq. S15 and following):

"ecological": mid frequencies decay slowest, as in real rooms. "inverted": max + min - ecological; slow where real rooms are fast. "exaggerated": the profile of a room twice as reverberant, scaled down to this length (more sharply peaked than real rooms). "reduced": the profile of a room half as reverberant, scaled up (flatter).

measure_rt60(ir: Sound, n_bands: int = 30, f_lo: float = 50.0, f_hi: float = 8000.0, fit_range_db=(-5.0, -25.0)) → tuple[ndarray, ndarray][source]

Per-band RT60s of an impulse response: (center freqs [Hz], RT60s [s]).

Each ERB band’s energy decay curve (Schroeder backward integration) is fit with a line over fit_range_db and extrapolated to -60 dB (so the default measures T20 x 3). The direct sound, if any, should be removed first; bands that never decay through the fit range give NaN.