feat(streaming): add speakerLabelsRevisionIntervalMs param - #188
Merged
Merged
Conversation
Add `speakerLabelsRevisionIntervalMs` to `StreamingTranscriberParams`, sent as the underscore-prefixed `_speaker_labels_revision_interval_ms` query param (the server's not-yet-GA wire name). It sets the cadence, in ms of audio time, at which the server emits mid-stream `speakerRevision` events when `speakerLabels` is enabled; unset or 0 sends only the end-of-stream revision, and values above the server default (300 000 ms) are clamped. `0` is sent explicitly rather than dropped. Bump version to 4.41.5 and add the changelog entry. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Verified live on prod: with speakerLabelsRevisionIntervalMs=60000 over 240 s of audio the first mid-stream speakerRevision arrived at 126 s (the server's 120 s min-content floor); the unset control got only the end-of-stream one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The server made the un-prefixed name canonical in DeepLearning#20528 (the `_`-prefixed spelling is now a legacy alias), and non-zero values are clamped server-side to 120 000–300 000 ms. Send the GA name, drop the "not officially supported" wording, document the clamp, and use 120 000 in the tests. Verified live on prod: interval=120000 got a mid-stream speakerRevision at 122.7 s; the unset control got only the final one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ccampbell-aai
approved these changes
Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
speakerLabelsRevisionIntervalMstoStreamingTranscriberParams, sent as thespeaker_labels_revision_interval_msquery param. That is the canonical GA name since DeepLearning#20528; the server still accepts the_-prefixed spelling as a legacy alias, but the SDK sends the GA form.How the server treats it (from
ConnectionParametersin the realtime api_v2 gateway):speakerRevisionevents. Unset or0gives only the end-of-stream revision;0is sent explicitly rather than dropped.maxSpeakers.speakerLabelsis enabled.Version bumped to 4.41.5 with a changelog entry, and the CLAUDE.md speaker-revisions section is updated.
Test plan
speaker_labels_revision_interval_ms=120000, and=0when explicitly set to 0speakerLabelsRevisionIntervalMs: 120_000(An earlier run using the legacy
_-prefixed name and a 60 000 value showed the same pattern, with its mid-stream revision at 126.3 s. The server clamps 60 000 up to 120 000.)Note:
tests/unit/utils.test.tshas 3 failures on Node 24 that also occur on cleanmain. They expect aruntime_envsegment in the user-agent string, and are unrelated to this change.🤖 Generated with Claude Code