Project Resonant · paper 02, interactive

Does the physics do the work?

An untrained network of 1,024 coupled oscillators, a linear readout, and spoken digits. How the second paper tests what the oscillators actually contribute, with every model it compares running live in your browser on real recordings.

Twenty held-out clips from AudioMNIST · every model the paper compares · ~25 min read

01 · the questionWhat does the physics add?

In How a machine hears a number I walked from a pressure wave to a network of coupled oscillators that could tell “three” from “eight”. The oscillators were never trained. Their coupling was drawn at random once and never updated, and only a linear readout on top was fitted, and it still read spoken digits at 96%.

That number invites an obvious question, and the first post ended on it: how much of it comes from the oscillators? A good accuracy can come from at least three places that have nothing to do with synchronization physics. The input may already carry most of the answer. The readout may be wide enough to find the answer in almost anything. Or the dynamics may genuinely help. The first guide’s own fine print showed that the readout alone was worth a lot: squeeze its features through a narrow projection and 96.2% fell to 67.3%.

The second paper in this research programme, Spoken-Digit Recognition Without Training, is built to separate those three explanations. It takes a small untrained oscillator network, reads it the same way it reads every alternative, and compares it with models that each remove one ingredient: the dynamics entirely, the oscillation, the coupling. Then it changes the physics one factor at a time, across eight experiments and 6,353 runs.

This page explains that study: what each model is, how they were compared, and why the comparisons are the ones they are, with what the paper found at the end. Every model at the heart of the paper is also running here, in your browser, on real recordings, so you can do the experiment yourself one clip at a time.

FIXED, SHAREDVARIES BY MODELFIXED, SHAREDfront end16 kHz audio→ 16 mel bands62.5 frames/s× gainband r → row rone armnothing (spectrogram-only baseline)coupled oscillators · uncoupled oscillatorsleaky integrators (1,024 or 2,048)GRU · TCN · CNN · transformer · S4Deach exposes its signals over timethe readframes 16–614 windowsmean · SD · |Δ|readout→ 192 featuresridge, 1,930weights · 0–9every model is read by the same function and scored by a readout of the same size
Figure 1 · the shared pipeline. The paper’s design in one picture. The front end, the way signals are summarized and the readout are identical for every model; only the box in the middle changes. The untrained arms have no fitted parameters at all: only the readout’s 1,930 weights are fitted, so any difference between two arms is a difference in what they did to the same input.

02 · oscillatorsOscillators, coupled and uncoupled

An oscillator is anything that cycles: a pendulum, a firefly’s flash, a neuron that fires rhythmically. The simplest mathematical version (Kuramoto) keeps a single number, its phase θ, an angle that advances around a circle. Left alone, it advances at its own natural frequency ω. The interesting part is what happens when many of them can feel each other.

Below are 24 oscillators with different natural frequencies. With the coupling at zero they are an uncoupled population: each turns at its own rate, and their phases smear evenly around the circle. Turn the coupling up and each one starts to be pulled toward the others. Past a threshold, a cluster forms and grows until most of them move together. The white arrow is the order parameter R, the average of all their phases taken as unit vectors: near 0 when they are scattered, near 1 when they are locked.

24 oscillators, all coupled to allR = —
Figure 2 · coupled and uncoupled. Left, each oscillator as a clock hand; right, all 24 phases on one circle. Drag the coupling to zero and the population falls apart into independent clocks; raise it and they synchronize, faster when their natural frequencies are close together. The restoring strength λ pulls every oscillator toward phase 0: one whose natural frequency is below λ stops turning altogether and is held in place, and a faster one keeps turning but unevenly, slowed where the pull opposes it. This toy couples every oscillator to every other with the same positive weight; the network in the paper does neither, as below.

Written out, each oscillator i in the paper’s network advances by

dθi/dt = ωi + Σj Kij · sin(θj − θi) − λ · sin θi + g · ur(i)(t)
ωNatural frequency. Drawn at random from N(1, 0.1²) and then fixed. One step per 16 ms frame with a time step of 0.1 means a free oscillator turns about once a second.
K, sin(Δθ)Coupling. Each oscillator is pulled toward the phases of the others, in proportion to a weight that depends only on their offset on the lattice. This is the Kuramoto coupling function; section 07 swaps it for others.
λRestoring strength. A pull toward phase 0, λ = 0.3. It gives the network a fading memory: what it heard a second ago matters less than what it hears now.
g · uThe input. The energy in mel band r, scaled by the input gain g, added to the turning rate of every oscillator in row r. A loud band speeds its row up.
Figure 3 · the whole network, in one line. Integrated with one Euler step per frame. The readout never sees θ itself, since 0.01 and 6.27 are neighbours on the circle but far apart as numbers; it sees sin θ and cos θ, two signals per oscillator.

The network under test

The coupled oscillator network is 1,024 of these: four channels, each a 16 × 16 lattice of 256 oscillators. Row r of every channel is driven by mel band r, lowest at the bottom, so the lattice inherits the ear’s layout. Within a channel, every oscillator acts on every other through a coupling kernel: a 16 × 16 table of weights, one per offset, drawn from N(0, 0.05²). A positive weight pulls a pair toward the same phase, a negative one pushes them apart. The same table applies at every site, which makes the coupling a convolution. Channels do not act on each other: they are four differently drawn copies read side by side, the way a convolutional layer’s channels are.

One more constant bounds the whole thing. A random kernel can amplify some spatial patterns of phase much more than others; the coupling ceiling rescales each channel’s kernel so that its largest amplification is 1. Random kernels here peak between 1.4 and 2.9, so the ceiling always applies, and it sets every channel’s overall coupling strength.

That is the entire model: 1,024 natural frequencies and 1,024 kernel weights, 2,048 numbers, all drawn once from a seed and never updated. The uncoupled oscillator network is the same network with every kernel weight set to zero. Each oscillator keeps its natural frequency, its restoring pull and its input, and none acts on another, so any difference between the two is what the coupling contributes.

03 · the control with memoryLeaky integrators: memory without oscillation

An oscillator network driven by sound is, among other things, a bank of filters with memory. Each oscillator’s state at any moment depends on what it heard recently, weighted toward the recent past. So is a much simpler object, and if the oscillators cannot beat it, their accuracy comes from being filters with memory, not from being oscillators.

That simpler object is the leaky integrator. It holds one number and, every frame, moves a fixed fraction of the way toward its input:

x ← (1 − a) · x + a · tanh(gin · g · u)
(1 − a) xWhat it keeps. Each frame the unit keeps a share 1 − a of its old value; the rest leaks away, so old input fades exponentially.
aLeak rate. Set by the unit’s time constant τ as a = 1 − e−1/(62.5 τ): 0.63 at τ = 16 ms, a unit that forgets within a few frames, down to 0.016 at τ = 1 s, one that averages the whole clip.
gin · gGains. The unit’s own input weight, drawn from N(1, 0.1²), times the input gain g that every reservoir shares.
uThe input. The energy in the unit’s mel band, the same row value that drives the oscillators. The tanh keeps a loud band from pushing the state past 1.
Figure 4 · a leaky integrator, in one line. In signal-processing terms it is a first-order low-pass filter: a running average of its recent input. In reservoir computing it is the unit of a leaky echo state network.
one band’s energy through five leaky integrators
band 7 · 1,004 to 1,591 Hz
input gain
Figure 5 · fast and slow memories. Time along the bottom, over the selected clip’s one second (pick a clip in section 04); each unit’s state x up the side. Grey: the input every unit moves toward, tanh(g · u) of one mel band’s energy. Coloured: the states of five leaky integrators fed that band, with time constants of 16 ms (orange), 45 ms, 125 ms, 350 ms and 1 s (violet). The slider picks the band, and the line under it gives the frequencies that band covers. The fast ones track every syllable; the slow ones only know roughly how loud the band has been. Raise the gain and the tanh flattens the peaks. None of them can ever do what an oscillator does: turn.

The leaky-integrator bank is 1,024 of these, routed exactly like the oscillator network: band r drives every unit in row r of every channel. Its time constants run on one log-spaced schedule from one frame (16 ms) to one clip (1 s), laid out so that every band is read at every time scale, and its input weights gin are drawn from N(1, 0.1²), the distribution the network draws its natural frequencies from. Nothing about the schedule was tuned: it is a falsification control, not a claim that this is a good design.

It comes in two sizes, because the two models expose different numbers of signals. The state-matched bank has the network’s 1,024 states and 2,048 parameters, but each unit exposes one signal where an oscillator exposes two (sin θ and cos θ). The width-matched bank has 2,048 units, matching the 2,048 signals, at twice the parameters.

the state-matched bank’s 1,024 time constantsorange = fast (16 ms) · violet = slow (1 s)
Figure 6 · every band at every time scale. The bank stored as the network is: four channels of 16 × 16, row r fed by mel band r. Within a row, the 64 units run from the fastest (channel 1, left) to the slowest (channel 4, right), so each band is integrated over 64 time scales. The colours are figure 5’s: orange for 16 ms through violet for 1 s.

04 · the front endSixteen numbers every 16 milliseconds

Every model starts from the same fixed front end, which turns one second of 16 kHz audio into 61 frames of 16 numbers. The first post builds it up piece by piece; here it is with the constants the paper fixed, and why.

choose a recording—
loading…
waveform · 16,000 samples
STFT · 512-point, every 256 samples · 257 bins × 61 frames
16 mel bands, log energy
drive rows: (log + 10) / 10, clamped at 0 · what every arm receives
now
Figure 7 · the front end, live. Every clip on this page comes from AudioMNIST, 30,000 recordings of 60 speakers saying the ten digits. These twenty come from the test speakers (49 to 60), whom no readout on this page was ever fitted on. Each is stored as the paper’s bank stores it: resampled to 16 kHz, its peak set to 0.5, trimmed to the word, and zero-padded to one second.
constantvaluewhy
sample rate16 kHzSpeech carries little above 8 kHz; one second is 16,000 samples.
window, hop512, 256 samplesA 32 ms window hopping 16 ms: 62.5 frames a second, 61 per clip. No centre padding, so the frame count is exact.
mel bands16One per lattice row, so band r can drive row r directly.
loglog(energy + 10⁻⁵)Loudness is heard on a log scale; the offset keeps silence finite.
rescale(x + 10) / 10, clamped at 0One fixed map into a drive of about 0 (silence) to 1.5 (loud). Nothing is normalized per clip, because a clip’s own statistics are unknown until it ends: a model that has to listen as the sound arrives cannot use them.
warm-up16 frames (256 ms)Skipped by the read, so every model is read after its state has settled from the same initial condition. The spectrogram-only baseline, which has no state, is read from frame 0 (section 05).

Nothing in the front end is trained, and nothing depends on the clip. The same arithmetic runs here in your browser: the rows you see are identical, to rounding, to the rows the paper’s harness computed. The paper measures the effect of one of these constants: in figure 9, read the spectrogram-only baseline from frame 0 and then with the four windows, which skip the warm-up, to see what the first 256 ms of a word are worth on their own.

05 · the readoutOne readout for every model

If two models are read differently, a difference in their accuracy can come from the readers rather than the models. A pilot of this study read its models three different ways, so the paper reads every model with exactly one function, in four steps.

  1. Signals. An arm only says which signals it exposes over time: sin θ and cos θ of each oscillator, each leaky integrator’s state, a trained network’s hidden units, or, for the spectrogram-only baseline, the 16 band energies themselves.
  2. Statistics. Over frames 16 to 61, cut into four equal windows, each signal is summarized by its mean, its standard deviation and its mean absolute change from frame to frame. None of them depends on an endpoint, so no arm can hand the readout its last state.
  3. Projection. The features are standardized with the training set’s own statistics and multiplied by one fixed random Gaussian matrix down to 192, the spectrogram-only baseline’s own width. An arm already at or below 192 is read as it is. The matrix is drawn once, from a fixed seed, and serves every run of the same native width.
  4. Ridge. One linear layer from 192 features to ten digit scores, fitted in closed form by least squares with an L2 penalty. The penalty is chosen from four values on the last eighth of the 2,048 training clips, then the layer is refit on all of them. The highest score wins.
reading the selected clip
1 · six of the arm’s signals over the clip, the warm-up dimmed, the four windows marked
2 · every statistic of every signal: four windows × mean, SD, change (blue low, orange high)
3 · projected to 192
4 · ten digit scores, 0 at the bottom

Figure 8 · the read, step by step. The selected clip (section 04) through each arm at gain 1 on clean audio, and through that arm’s fitted readout. The coupled network exposes 2,048 signals, so its read is 24,576 numbers before the projection brings it to 192; the spectrogram-only baseline exposes 16 and reaches exactly 192 on its own.

Why these choices

The same width for every arm. A ridge’s capacity grows with the number of features it is handed, and the reservoirs expose 12,288 to 24,576 against the input’s 192. So that readout capacity cannot pass for dynamics, every arm is read at 192 features: exactly the spectrogram-only baseline’s own count (16 bands × 3 statistics × 4 windows), so the input is read without compression, and the native width of four of the five trained baselines. The paper also reads every model wider: at width 4,096 the reservoirs gain up to about 6 points, but their readout then fits 40,970 weights against the 1,930 that read the input. Figure 9 shows every width.

One fixed window. The natural shortcut is to read each clip over its own length. But an oscillator keeps turning whether or not anything drives it, so statistics over a span encode how long the span was, and in speech, how long a word lasts says something about which word it is. The paper’s leak check measures it: an undriven network, with no input at all, read over each clip’s own length recognizes digits about 18% of the time, against 10% for chance. Read over the same frames for every clip, it reads exactly 10%. So every arm is read over frames 16 to 61, and everything a read carries arrives through the arm’s response to the sound.

Why the read starts at frame 16. Every reservoir starts each clip from the same fixed state, and for the first 16 frames (256 ms) its state reflects that starting point more than the sound. So every model, the trained baselines included, is read over frames 16 to 61. Nothing is thrown away: the input drives each model from the first frame, and what a model remembers of the first 256 ms reaches the read through its state. The spectrogram-only baseline has no memory, so read over the same frames it would miss the start of the word, which the other models hear. It is therefore read over the whole clip, frames 0 to 61, seeing everything the others were driven with, and that is the comparison the paper makes. Read from frame 16 instead, it loses 10 to 17 points, which is what the first 256 ms of a word are worth on their own (the from frame 0 read below).

Below are spoken-digit classification accuracies for the models in the paper, at different readout widths and training sizes: the mean over three seeds on all 6,000 test clips, copied from the paper’s record of the controls experiment. The primary cell, the one every comparison is made at, is width 192, 2,048 training clips and the four-window read.

the record: what each readout choice doescontrols experiment · test speakers 49–60
noise
input gain
read
readout width
training clips
—
accuracy against readout width
accuracy against training clips
Figure 9 · accuracy by readout. Pick a model and a condition and toggle the readout’s settings. Width is how many features the ridge sees (native is the arm’s own count, unprojected, fitted at 2,048 clips only). Read: the four windows, the whole span as one window, the oscillator networks with their rotation rates added (a read that favours them, since no other arm has an analogue), and the baseline from frame 0. Noise is the signal-to-noise ratio: at 0 dB the noise is as loud as the speech, at −5 dB louder. Chance is 10%.

06 · the comparisonThe ten models compared

The controls experiment compares ten models. Each removes one candidate explanation for the network’s accuracy.

modelbetween the front end and the readoutstatesparametersfeatureswhat it isolates
spectrogram-only baselinenothing: the readout reads the band energies00192what the input alone supports
coupled oscillator network1,024 untrained coupled oscillators1,0242,04824,576the system under test
uncoupled oscillator networkthe same, coupling set to zero1,0242,048 (1,024 in effect)24,576the coupling
leaky-integrator bank, state-matched1,024 independent leaky integrators1,0242,04812,288oscillation, at equal states
leaky-integrator bank, width-matched2,048 independent leaky integrators2,0484,09624,576oscillation, at equal signals
GRU, TCN, CNN, transformer, S4Dsmall conventional networks, trained end to end·1,840–2,109, all trained192–216what learning buys at the same budget

The four untrained dynamical arms (both oscillator networks and both banks) are the paper’s reservoirs, the term reservoir computing uses for a fixed dynamical system read by a trained linear layer. The five trained baselines play a different role. Their input layers are themselves trained, so they cannot isolate any physics; they answer whether a conventional network of the same size, trained normally, does better or worse. Each is trained with AdamW for 30 epochs through a learned linear head on the very statistics the ridge reads, standardized, so it is trained for the read it is judged by, and is then read by the same ridge as everyone else. The TCN is a dilated residual one, as Bai, Kolter and Koltun define it; the CNN is two plain causal convolutions that see 9 frames, the local, memoryless baseline. The controls experiment also runs the Stuart–Landau network of section 07 on clean audio, so that network stands beside every arm in all three conditions.

The conditions

Noise. Clean audio nearly saturates this task, so white noise is added at signal-to-noise ratios of 0 dB (the noise as loud as the speech) and −5 dB (5 dB louder), and clean audio is reported for reference. Each clip’s noise is drawn from a generator seeded by the clip itself, so a clip sounds the same to every arm. The design experiments of section 07 read only the noisy conditions.

Input gain. The reservoirs are nonlinear, so how hard the input pushes relative to their own dynamics changes how they respond: for an oscillator, the input competes with its natural frequency, its coupling and the restoring pull; for a leaky integrator, it sets how far into the tanh the input reaches. Every reservoir runs at gains 1 and 2, and the sweep experiment takes each coupling function’s reference network from gain 0.25 to 12. Gain does not apply to the spectrogram-only baseline, whose standardized statistics would divide any fixed scale out exactly, nor to the trained baselines, whose first layer learns its own.

Seeds, training sets and reporting. Every condition runs at three seeds. A seed sets an arm’s random draws and which 2,048 of the 24,000 training clips (speakers 1 to 48) the readout is fitted on; the test set is always all 6,000 clips of speakers 49 to 60. Accuracies are reported as the mean and standard deviation over seeds, and every comparison between two arms is paired on the same test clips, with a 95% interval from resampling them. No threshold decides a result. For comparison with published numbers, the controls arms also run once on each of Becker et al.’s five speaker folds, on clean audio.

Temporal order. The controls experiment also runs a second task, built so that an order-free read of the input cannot solve it: two digits spoken one after the other, and the question is which came first. The mean and spread of a signal do not depend on the order of its frames, so the spectrogram-only baseline, read over the whole span, sits at 50% by construction. Anything above that has to come from an arm’s memory of what happened when.

Eight experiments

The study runs these comparisons, and the ablations of the next section, as eight experiments.

experimentrunswhat it asks
leak check13Does the pipeline leak? Every reservoir with no input must read exactly chance, and a per-clip read window is measured for how much it would give away.
controls852Does the network add anything beyond its input, a leaky-integrator bank of its size, or its own oscillators uncoupled, and how do trained baselines compare? On recognition and on temporal order.
design3,744Does the network’s design matter: its coupling function, lattice geometry, natural frequencies, restoring strength and coupling ceiling?
cochlea432Do a coil and a cochlea, lattices built to follow the ear, do better than the torus?
sweep684What happens beyond the design’s levels: restoring strengths up to 1, ceilings of 1.5 and 2, and input gains from 0.25 to 12?
quadrature126Can the networks use a drive that tells them when in a band’s cycle to push (section 08)?
projection432Does the readout’s fixed random projection matter? The reservoirs are read again through one drawn from each run’s seed.
Becker folds70Where do the arms sit against published AudioMNIST results, on the corpus’s own speaker folds?

07 · the ablationsChanging the physics one factor at a time

The design experiment asks whether the design of the network matters. It crosses five factors of the coupled network, 312 configurations at 0 and −5 dB, both gains and three seeds, and measures every effect on matched pairs: two runs identical in every factor but the one compared, at the same noise level, gain and seed.

The coupling function

How the phases of two coupled oscillators turn into a push on one of them. The paper tests six, each a standard object in the synchronization literature, each changing one property of the pull.

functionthe coupling term for oscillator iwhat it changes
KuramotoΣ Kij sin(θj − θi)The reference: each oscillator pulled toward the others’ phases.
Kuramoto–SakaguchiΣ Kij sin(θj − θi − π/4)A phase lag that breaks the pull’s symmetry and admits travelling waves.
second harmonicKuramoto + ½ Σ Kij sin 2(θj − θi)Favours two-cluster states: pairs in phase or in opposition.
Winfree−sin θi Σ Kij (1 + cos θj)How strongly an oscillator responds, and acts, depends on its own phase.
Stuart–LandauΣ Kij (zj − zi), z = x + iyAmplitude joins phase as state: each oscillator relaxes to a cycle of radius 1.
Stuart–Landau, fixed amplitudethe same, |z| held at 1The phase-only limit, which reduces to Kuramoto: it separates what amplitude adds.

The lattice geometry

Every geometry stores the same 256 oscillators per channel in the same 16 × 16 grid, with row r driven by band r. A geometry only changes how the grid’s edges are glued, and so which oscillators are neighbours. The kernel is the same table of weights in every case; the gluing decides where each weight lands. The design experiment runs six geometries. The cochlea experiment adds two that follow the ear itself: a coil, the 256 oscillators on one open spiral from the apex (the lowest band) to the base (the highest), an octave per turn, and a cochlea, the coil with two features of the cochlea’s mechanics. Influence runs three times as strongly from base to apex as back, as the travelling wave does, and the coupling into each site grows with the spiral’s curvature, from a quarter at the base to full strength at the apex. Because that weighting halves the cochlea’s average coupling, a control keeps its shape at the coil’s average. Click any oscillator on the grid to see who acts on it.

who acts on whom · one channel’s kernel under each geometry
the flat 16 × 16 grid · row 1 (lowest band) at the bottom · click to choose an oscillator to focus on
the shape the gluing makes · drag to turn

Figure 10 · one kernel, every gluing. Blue: the chosen oscillator. Every other oscillator is lit by how strongly it acts on the chosen one, from dim fuchsia for a small weight to white for the largest, whether it pulls toward its phase or pushes away. The darkest, a near-black fuchsia, cannot act on it at all. The kernel has 16 offsets on each axis, reaching 8 rows or columns one way and 7 the other, so from the starting point, row 8 and column 8, it reaches every oscillator on every geometry but the coils. Click one near an edge to see where each geometry cuts the reach off. On the torus both axes wrap, so every oscillator reaches every other from anywhere, and the top band couples to the bottom. The cylinder opens the frequency axis, as in the cochlea, so the reach stops at the top and bottom rows: from the top band, only the 7 bands below it can act. The sheet opens both axes. The helix reads all 256 as one closed coil, 64 to a turn, so a turn away is an octave away. The cube folds each row into a 4 × 4 slab, a 16 × 4 × 4 lattice that wraps on all three axes, with shorter paths between the same oscillators. The sphere makes rows latitudes and weights each oscillator’s influence by the cosine of its latitude, an approximation to a sphere rather than exact spherical coupling. The coil is the helix opened, so its ends never meet and each oscillator reaches two turns either way. The cochlea adds a direction (arrowheads, pointing to the apex), which shows on the grid as brighter oscillators in the rows above the chosen one than in the rows below, and a curvature weighting, which scales everything arriving at an oscillator from 1 at the apex to 1/4 at the base: choose higher rows and the largest weight, below the grid, falls. The channel buttons switch between the four channels. Each is a complete copy of the lattice with its own kernel, not another part of the shape.

The other three factors

  • Natural frequencies. Random, drawn from N(1, 0.1²); tonotopic, each row set to the centre frequency of the mel band that drives it, with small jitter, the arrangement of the cochlea; identical, all 1.
  • Restoring strength λ. 0.3 and 0.1: a longer or shorter memory.
  • Coupling ceiling. 1 and 0.5: stronger or weaker coupling overall.

The two Stuart–Landau functions were run on the torus only, a limit of the study’s scope: their core was built for the torus, and the other geometries were wired into the phase oscillators alone. The sweep experiment then pushes each coupling function’s reference network past these levels, to restoring strengths of 0.5 to 1, ceilings of 1.5 and 2, and gains from 0.25 to 12. The console below runs the random-frequency slice of the design and cochlea experiments: every coupling function on every geometry it was run on, at restoring strength 0.3 and coupling ceiling 1.

08 · the drive signalTwo input pathways

Everything so far drives the oscillators with the spectrogram: how loud each band is, frame by frame. That throws away something an oscillator could use. Sound is itself oscillation, and a spectrogram drive tells an oscillator how hard to push, never when in the sound’s own cycle to push. An oscillator nudged at the right moment of every cycle can lock to a rhythm; one pushed at random moments cannot. So the paper also drives the network through a pathway that keeps timing, with its own spectrogram-only baseline.

  • Spectrogram. The front end of section 04: each band’s energy adds to the turning rate of its row, the same push whatever the oscillator’s phase. Everything above runs on it.
  • Quadrature. Each band’s energy and phase, 62.5 times a second: the phase of the band’s centre frequency, demodulated so what remains is how the band drifts around that centre (within ±31 Hz). The push becomes g · A · sin(φ − θ): it depends on where the oscillator is relative to the band’s own cycle, the phase-referenced drive Adler analysed in 1946, which pulls an oscillator into step with the band.
the selected clip, two ways16 bands, lowest at the bottom
spectrogram · 61 frames · brightness = energy
quadrature · 61 frames · brightness = energy, hue = the band’s phase
Figure 11 · what each pathway hands the network. The same recording (chosen in section 04). The spectrogram keeps 0 to 8 kHz at 16 ms resolution and discards phase. Quadrature keeps the same energies and adds each band’s drifting phase.

The paper runs the quadrature pathway on coupled networks only, the four phase coupling functions on the torus: a leaky integrator has no phase for the push to act on, and the question is whether coupled oscillators can use a phase-referenced drive at all. The console below adds the uncoupled network on the quadrature pathway, which the paper did not run, as the obvious reference.

09 · run it yourselfThe experiment, one clip at a time

Everything below runs in your browser: the front end, the model, the read and the readout, with the physics ported line for line from the paper’s harness and the readout for each condition fitted exactly the way the paper fits it. Pick a model, a pathway and a condition, pick a recording or record yourself saying a digit, and press play. The clip is heard and the network is driven at the same moment; the scores appear when the read’s last window closes.

Gain and noise snap to the levels the paper ran (gains 1 and 2; clean, 0 dB and −5 dB). Where the paper did not run a combination you pick, the console moves the other settings to the nearest one it did and says so under the controls. Each prediction comes from a readout fitted at exactly that condition at seed 0, and the accuracy beside it is the paper’s, over all three seeds.

the network on its toruschannel 1 of 4

choose a modelin the paper
input pathway
input gain
noise

or record yourself saying a digit
input0.00 s
order parameter R · drive rows now—
ten scores · highest wins—
source—
detected—
accuracy in the paper—
run time—

all four channels, flattenedhue = phase
Figure 12 · the model explorer. The coupled network starts at the paper’s reference configuration: Kuramoto coupling on a torus, random natural frequencies, restoring strength 0.3, ceiling 1. Change its coupling function or geometry and you are in the design or cochlea experiment; switch the pathway and you are in the quadrature experiment, where the rows panel shows each band’s phase as hue and its energy as brightness. Noise is added the way the paper adds it, with the same noise samples the paper drew for each clip; your own recording gets Gaussian white noise at the same signal-to-noise ratio. The ten scores are the readout’s outputs, fitted to 1 for the spoken digit and 0 for the others, so they are not probabilities and not accuracy: the highest wins. Accuracy is the share of the 6,000 test clips on which the highest score is the right digit, and one clip says little about it.

10 · what we foundMostly the input, and a memory

Against its controls. Read at the common width, the spectrogram alone reads 93.6% on clean audio, 78.0% at 0 dB and 71.6% at −5 dB. The Kuramoto network at gain 1 reads 91.1%, 77.8% and 70.7%: within a point of its own input with noise, and 2.5 points below it on clean audio. It reads 0.4 to 1.1 points above the same network uncoupled, and about 4 points above the state-matched leaky-integrator bank with noise, though 2.35 below it on clean audio. With noise every trained baseline reads above it, from 79.2% (CNN) to 82.8% (GRU) at 0 dB.

Recognition accuracy for every arm at clean, 0 dB and minus 5 dB, from the paper
Figure 13 · every arm at the primary cell, from the paper: mean ± one standard deviation over three seeds, the reservoirs at gain 1 (filled) and 2 (open). Dashed line: the whole-clip spectrogram-only baseline.

What the dynamics add is memory. Read only after its 16-frame warm-up, the network reads 7.5 to 15.9 points above its input read over the same frames: it carries the onset, which the input loses without its first 16 frames. On the order task, which the input alone cannot solve (49.8% to 50.6%), it reads 95.5% to 97.7%. The leaky-integrator banks do both, and better on order, 98.7% to 99.9%. The reason is the same in both: apart from the restoring pull and the coupling, an oscillator’s phase is the running sum of its band’s energy wrapped around a circle. Oscillation adds no memory the bank lacks; the phase is an integrator whose sum wraps. Coupling adds about a point on noisy recognition and 2 to 19 points on order.

Temporal-order accuracy for every arm, from the paper
Figure 14 · temporal order, from the paper: which of two digits came first, averaged over five digit pairs. Dotted line: chance. The CNN sees 9 frames, less than a digit, so with noise it reads chance.

Design. Among phase oscillators, the coupling function moved accuracy by at most half a point, and every lattice geometry read within 0.65 points of the torus. The coil read within 0.22 points of it and the cochlea 0.41 to 0.97 below: built to follow the ear, they read no better. The one design choice that mattered was a free amplitude. The Stuart–Landau network reads 2.1 to 3.8 points above its Kuramoto match and, at the reference configuration, above its own input in every noisy condition, by 1.1 to 2.3 points, which puts it in the trained baselines’ range. With its amplitude fixed it reads like Kuramoto, so the effect is the amplitude, not the coupling form.

Each lattice geometry minus the torus, from the paper
Figure 15 · lattice geometry, from the paper: each geometry minus the torus, paired, with 95% intervals. Left of the grey line, the design experiment; right, the cochlea experiment.

Gain and drive. Gain separates the Stuart–Landau network from the rest: from gain 1 to gain 8 it lost 3 to 4 points while every phase-oscillator network lost 15 to 19, as if a free amplitude lets an oscillator absorb a strong drive in its radius where a phase can only be pushed faster. Below gain 1 the order reverses, and at their better low gain the phase networks read above their input at 0 dB. Driven in quadrature, every network read near chance, 11.7% to 15.1%: in no run did more than 0.3% of its oscillators lock to their band.

Accuracy of each coupling function against restoring strength, coupling ceiling and input gain, from the paper
Figure 16 · restoring strength, ceiling and gain, from the paper, for each coupling function at the reference configuration, at 0 dB (top) and −5 dB (bottom).

The readout. Widening the readout to 4,096 features lifts the reservoirs above their input, but through a readout with 21 times as many fitted weights, so the fair comparison stays at the common width. More training data helps the trained baselines three to six times as much as the network, since in a reservoir only the readout learns.

In short. Untrained and read at a common width, a coupled oscillator network classifies noisy, held-out spoken digits about as well as its own input, and carries the order of events by integrating it, as a leaky-integrator bank does. Coupling adds a little to both; a free amplitude lifts it above its input and into the trained baselines’ range; the geometry and the phase coupling function each move it by less than a point. Those numbers are the baseline a trained oscillator network of this size, on this task, should beat.

11 · go deeperThe paper, the code and the prior art

The paper, the harness that ran every model on this page, and the per-run record behind every number quoted here are public. The browser code here is a port of that harness, checked against it: the scores your browser computes match the harness’s on every demo clip, and every untrained arm’s readout, rescored on the full test set, reproduces the paper’s seed-0 accuracy.

Read the paper (PDF) →Its code and record on GitHub →

Further reading