Skip to content

feat: Gate worklet output at startup to prevent underrun clicks - #87

Open
bhj wants to merge 1 commit into
cutterbl:masterfrom
bhj:feat/startup-output-gate
Open

bhj wants to merge 1 commit into
cutterbl:masterfrom
bhj:feat/startup-output-gate

Conversation

@bhj

@bhj bhj commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Hey again :)

This came up while trying to switch my bespoke phase vocoder implementation over to v2.1.1. Can't say I noticed the transients myself, but it's plausible. It does add ~3ms of latency which seemed like a reasonable tradeoff for correctness.

Model used: Anthropic Fable 5

A bursty stretch stage refills the output buffer one hop at a time while the render thread drains one block per quantum, so the buffer's trough rides at exactly one render block. Wherever the stage's production schedule comes up a few frames short, the block is zero-filled — an audible click of up to full signal amplitude (measured: 1-8 sample dropouts).

SoundTouchProcessorBase now supports holding extraction for startupHoldBlocks render blocks after the output buffer first fills a block. Held blocks are emitted before any real output, so the added silence is inaudible, and the unconsumed frames become a permanent cushion under the trough. The phase vocoder worklet holds one block: a 128-frame cushion (16x the worst measured shortfall) for one render block (~2.9 ms) of added latency. Warmup blocks no longer count as underruns, so the metric now reports only real faults.

Measured A/B (real pipeline, gate vs none) across the full fftSize/overlapFactor grid at 44.1/48/96 kHz: group delay shifts by exactly 128 frames with output otherwise bit-identical, and hop-aligned pitch ratios drop from 9-21 underruns per 20 s to zero at every functional configuration. Arbitrary (unaligned) ratios are a sustained stretch/transposer rate mismatch that outruns any fixed cushion — the gate only delays onset there; matching the transposer to the stage's realized tempo is the actual fix for that case. Configs with hop <= 128 starve with or without the gate because PhaseVocoder.process() handles at most one hop per call (unlike Stretch, which loops until input is exhausted) — a pre-existing defect to fix separately. The default of 0 leaves the other worklets' timing unchanged.

Highlighting the two other issues the model called out:

Arbitrary (unaligned) ratios are a sustained stretch/transposer rate mismatch that outruns any fixed cushion — the gate only delays onset there; matching the transposer to the stage's realized tempo is the actual fix for that case

Already a PR for this with #85

Configs with hop <= 128 starve with or without the gate because PhaseVocoder.process() handles at most one hop per call (unlike Stretch, which loops until input is exhausted) — a pre-existing defect to fix separately.

I'll look at creating a separate PR for this - not sure how invasive it may end up being.

A bursty stretch stage refills the output buffer one hop at a time
while the render thread drains one block per quantum, so the buffer's
trough rides at exactly one render block. Wherever the stage's
production schedule comes up a few frames short, the block is
zero-filled — an audible click of up to full signal amplitude
(measured: 1-8 sample dropouts).

SoundTouchProcessorBase now supports holding extraction for
startupHoldBlocks render blocks after the output buffer first fills a
block. Held blocks are emitted before any real output, so the added
silence is inaudible, and the unconsumed frames become a permanent
cushion under the trough. The phase vocoder worklet holds one block: a
128-frame cushion (16x the worst measured shortfall) for one render
block (~2.9 ms) of added latency. Warmup blocks no longer count as
underruns, so the metric now reports only real faults.

Measured A/B (real pipeline, gate vs none) across the full
fftSize/overlapFactor grid at 44.1/48/96 kHz: group delay shifts by
exactly 128 frames with output otherwise bit-identical, and
hop-aligned pitch ratios drop from 9-21 underruns per 20 s to zero at
every functional configuration. Arbitrary (unaligned) ratios are a
sustained stretch/transposer rate mismatch that outruns any fixed
cushion — the gate only delays onset there; matching the transposer to
the stage's realized tempo is the actual fix for that case. Configs
with hop <= 128 starve with or without the gate because
PhaseVocoder.process() handles at most one hop per call (unlike
Stretch, which loops until input is exhausted) — a pre-existing defect
to fix separately. The default of 0 leaves the other worklets' timing
unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@tzzo

tzzo commented Sep 23, 2026

Copy link
Copy Markdown

A data point from the WSOLA worklet (@soundtouchjs/audio-worklet), since this PR leaves it at startupHoldBlocks = 0 and the description reasons that the built-in Stretch, which loops until its input is exhausted, shouldn't need the gate.

It does. Feeding a constant stereo signal (0.5 / 0.25) through the processor built from main (2.1.1) in 128-frame quanta and counting output frames that deviate from the input after the first second, over 20 s of signal per cell:

semitones 48 kHz — frames / blocks / longest run 44.1 kHz — frames / blocks / longest run
−4 0 6 / 1 / 6
−3 0 6 / 1 / 6
−2 0 2 / 1 / 2
−1 0 4 / 2 / 2
0 0 0
+1 8 / 1 / 8 0
+2 11 / 5 / 3 3 / 1 / 3
+3 13 / 13 / 1 2 / 1 / 2
+4 5 / 4 / 2 21 / 15 / 2

Same mechanism as the description: the transposer emits a fractional frame count per block (128 / rate) and the stretch stage emits in seekWindowLength − overlapLength bursts, so the output buffer's trough sits at exactly one render block and any block where the carry comes up short is zero-filled. Which settings are affected depends on the side of rate = 1 and on the sample rate, which is why 48 kHz is clean below unity and 44.1 kHz isn't.

With one held block (a local patch that holds extraction for the first block that has output, i.e. startupHoldBlocks = 1 in this PR's terms), every cell above is 0 at both sample rates, for 128 frames of added latency.

So I'd suggest either overriding startupHoldBlocks to 1 in SoundTouchProcessor as well, or exposing it as a processorOptions field so hosts can choose. Happy to open a follow-up PR for whichever you prefer once this lands.

Repro script (Node; point it at a built soundtouch-processor.js)
// node holes.mjs packages/audio-worklet/.dist/soundtouch-processor.js 48000
import { readFileSync } from 'node:fs';
import { compileFunction } from 'node:vm';

const file = process.argv[2];
const SR = Number(process.argv[3] || 48000);
const Q = 128;
const SECONDS = 20;
const SKIP_S = 1;

const install = compileFunction(readFileSync(file, 'utf8'), [
  'AudioWorkletProcessor',
  'registerProcessor',
  'sampleRate',
]);
class Host {
  port = { onmessage: null, postMessage() {}, close() {} };
}

function run(semitones) {
  let Processor;
  install(Host, (_name, ctor) => (Processor = ctor), SR);
  const proc = new Processor();
  const params = {
    pitch: new Float32Array([1]),
    pitchSemitones: new Float32Array([semitones]),
    playbackRate: new Float32Array([1]),
  };
  const input = [[new Float32Array(Q).fill(0.5), new Float32Array(Q).fill(0.25)]];
  const out = [[new Float32Array(Q), new Float32Array(Q)]];
  let frames = 0, blocks = 0, longest = 0;
  for (let b = 0; b < (SECONDS * SR) / Q; b++) {
    proc.process(input, out, params);
    if (b * Q < SKIP_S * SR) continue;
    let run = 0, inBlock = 0;
    for (let i = 0; i < Q; i++) {
      const bad = Math.abs(out[0][0][i] - 0.5) > 1e-6 || Math.abs(out[0][1][i] - 0.25) > 1e-6;
      if (bad) { inBlock++; run++; longest = Math.max(longest, run); } else run = 0;
    }
    frames += inBlock;
    if (inBlock) blocks++;
  }
  return { semitones, deviatingFrames: frames, blocksAffected: blocks, longestRun: longest };
}

for (const s of [-4, -3, -2, -1, 0, 1, 2, 3, 4]) console.log(JSON.stringify(run(s)));

Measurements and draft written with Anthropic Fable 5.1.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants