Skip to content

Playback Trick Modes (Speed, Seek, Pause)

Playback speed is not a feature you add to the client. It is decided by who owns the RTSP session, and by default that is not the adapter — which is why a speed button wired straight to a CCTV adapter does nothing at all.

This page explains the constraint, the design that resolves it, and the traps that cost real debugging time on live hardware.


Why speed is special

Playback rate travels in the RTSP Scale header on a PLAY request, inside the session that carries the media (RFC 7826 §18.46). Two consequences follow, and both are absolute:

  1. Only the session owner can change the rate. There is no out-of-band way to ask a running session to go faster.
  2. go2rtc never sends it. Its RTSP client issues PLAY with no headers at all, and its Patch() only rewrites a producer's URL for the next reconnect. Producers also connect lazily, so there is nothing to pre-warm either.

So while go2rtc dials the NVR directly, Scale is unreachable. Declaring pq.command.video.playback.speed in that state produces a button the operator can press and a command the adapter can accept — with no effect on the picture. Do not ship that.

Direct pull — no trick modes possible
  NVR ──RTSP──► go2rtc ──WebRTC──► browser
                  ▲
        adapter only registers the alias here;
        it never sees the RTSP session

Adapter relay — trick modes reachable
  NVR ──RTSP──► adapter ──RTSP──► go2rtc ──WebRTC──► browser
                  ▲
        adapter owns the session: Scale, Range, PAUSE

The design

The adapter holds the RTSP session against the recorder and serves the media on for go2rtc to pull as an ordinary source. Every trick mode is then a mid-session request, which is what makes it seamless: the go2rtc alias never changes, so the viewer's WebRTC session is never renegotiated.

PQ command RTSP the adapter sends Effect
video.playback DESCRIBE + SETUP (+ PLAY once a viewer attaches) session up; capabilities read
video.playback.speed PLAY with Scale: N rate changes, picture keeps flowing
video.playback.seek PLAY with Range: clock=… position changes on the same session
video.playback.pause PAUSE recorder stops sending; nothing is discarded client-side
video.playback.resume PLAY resumes, and the response reports where
sequenceDiagram
    participant Op as Operator
    participant Cl as Client
    participant Ad as Adapter
    participant NVR
    participant G2 as go2rtc

    Op->>Cl: picks 2x
    Cl->>Ad: pq.command.video.playback.speed (rate=2)
    Note over Ad: clamp against the rates<br/>the device advertised
    Ad->>NVR: PLAY Scale: 2.000 (same session)
    NVR-->>Ad: 200 OK, Scale: 2.000, Range: clock=…
    NVR-->>Ad: RTP at 2x the data rate
    Ad-->>G2: same alias, same RTP stream
    Ad->>Cl: stream.ready (playback_from, playback_scales, playback_rate)
    Note over Cl: no WebRTC teardown,<br/>position re-anchors on the reported time

Pull, not push. go2rtc's ANNOUNCE handler resolves the alias with streams.Get(name), which returns nil for a name nobody registered — so pushing would need the stream created up front with some other source. Being pulled means the alias is registered exactly as before; only the URL changes, from the NVR to the adapter. go2rtc's lazy connect then still holds, so the recorder's session is opened only when a viewer actually attaches, and released when the last one leaves.


Rates are a device capability, never a constant

A recorder declares what it accepts in the SETUP response. A live iDS-7716NXI answers:

Media-Properties: Random-Access=1.0s, Unlimited, Immutable,Scales="-1, 0.5, 0.25, 0.125,:1, 2, 4"
Accept-Ranges: UTC

Read that list, do not invent one. On this device:

  • 8x does not exist. The CCTV wall used to offer ±8x; every press was going to be refused.
  • Slow motion does — 0.5, 0.25, 0.125 — and was never on the menu.
  • Reverse is a single rate, -1.

The list flows to the client as playback_scales on pq.event.video.stream.ready, and the speed menu is built from it. When a device reports nothing, the client falls back to a conservative 1, 2, 4 — never reverse, never beyond 4x.

The list is not always well formed

That device writes 0.125,:1 — a lo:hi range whose lower bound landed on the previous list item. PlaybackScales reads a colon token as "everything between the neighbouring values is allowed", which is the permissive-but-bounded reading. Parse defensively; a malformed capability must degrade the menu, never take playback down.

Clamping, and why the applied rate is published back

A requested rate is snapped onto something the device accepts (same direction, nearest on a log scale) before it is ever sent. If the direction is unsupported at all — reverse on a forward-only recorder — the command fails with a reason rather than silently doing nothing.

A clamped rate must be told to the client

The client advances the playback position as anchor + media-clock-elapsed × rate. If the adapter silently clamps 8x to 4x, the client keeps computing at 8x while the picture moves at 4x and the timeline drifts away from the image. The applied rate is therefore published as playback_rate and the client adopts it. Clamping without reporting is worse than refusing.


Verifying a firmware before trusting it

Scale being accepted does not mean it is usable. Measure media-seconds per wall-clock second: play at 1x for N seconds, then at Nx, and compare how far the RTP timestamps advanced.

Observation What the firmware does Consequence
~Nx media time, ~Nx bitrate rescales the media properly browser renders smooth fast-forward
~1x media time ignores Scale trick modes are out; fall back to jump-seek
~Nx media time, ~1x bitrate thins to key frames renderable, but a fast slideshow

Measured on the iDS-7716NXI (firmware V5.04.051): Scale: 2.000 gave 1.99x observed with 2.00x the bitrate and frames; 4.000 gave 3.92x with 3.99x. That is the good case.


Time is the device's, not ours

Every time sent to a recorder — replay URL starttime/endtime, the content-search window, a Range: clock= seek — is read by it as its own local wall clock. PQ times are UTC. Convert at the boundary or playback lands a whole timezone away from what the operator asked for.

client (UTC) ──► + device offset ──► recorder      (replay URL, search window, Range)
recorder ──────► − device offset ──► PQ (UTC)      (playback_from on stream.ready)

The offset comes from what the device itself reports (GET /ISAPI/System/time → localTime, e.g. 2026-08-20T21:40:00+01:00), and is cached per device.

Time sync must own the device's timezone, not just its clock

A recorder stores what you send as its local time, so the clock is only meaningful next to the zone it is interpreted in. Writing the adapter process's clock — UTC in a container — leaves the device skewed by the difference between the two zones, and that skew lands in the OSD burned into the picture, in every recording timestamp, and therefore in the search windows playback is built from.

The framework already models the intent: the time-sync scheduler hands SynchronizeTime a TimeProvider built from the timezone configured on that device, so LocalTimeZone and GetLocalNow() are what PQ means. Push both — the zone and the wall clock, in one write, because the clock is expressed in the zone. Leaving the device on its own zone means every conversion afterwards is guesswork.

A recorder can also just be wrong

One test device sat 2.5 months behind with timeMode: manual. A search window built from the host clock returns NO MATCHES, and a playback UI that draws a live edge from the operator's clock puts the slider months away from the picture. Anchor the timeline on what the adapter reports and advance it with the video element's media clock × applied rate — never wall-clock time, which keeps running through stalls, buffering and pauses and knows nothing about the rate.

A device left on the wrong zone reports an offset PQ did not intend: instants stay right if you convert by what it reports, but its OSD reads an hour off and so does every recording timestamp. That is why the zone is PQ's to write rather than the installer's to get right.

One caveat if you write it: a daylight clause can only be built from TimeZoneInfo.GetAdjustmentRules(), and that returns nothing on a chiseled runtime even where GetUtcOffset is correct — which silently produced a zone with no daylight rule on live hardware. Writing the effective offset and letting the periodic sync rewrite it after each changeover is the portable choice; the device is then an hour out for at most one sync interval, twice a year.

A clock-skew fallback must not disable the control it guards

A playback UI is right to stop a scrub past the live edge — there is no footage after now, so seeking there can only fail. But the live edge has to be expressed in the POSITION's basis, which is the device's, and when the two clocks disagree by more than the drawn window the operator's "now" is not a point on that timeline at all.

Falling back to the playhead in that case looks safe and is not: the entire future of the timeline becomes unreachable, so every forward scrub is clamped to exactly where playback already is while backward scrubs keep working perfectly. It reaches the operator as "seek forward does not work", with every command reporting success. Reported on live hardware 2026-09-09; the adapter and the recorder were both innocent, a forward seek driven directly moved the burned-in clock by exactly the hour requested.

Worse, the fallback fires in the ORDINARY case rather than an exotic one: reviewing footage from hours ago is the normal use of a recorder and needs no clock disagreement at all. A two-hour trust window and 09:00 footage on a 14:35 wall clock was enough.

Fall back to the end of the drawn window instead — it is derived from the device-reported anchor, so it is already in the right basis, it is what the slider can express anyway, and a target with no footage behind it fails with the recorder's own reason. Also keep the ceiling from ever dropping below the playhead: a long unattended playback can carry the position past the window. And name the parameter for what it is (a scrub ceiling, not "now") — as Now the clamp read as "never seek into the future", which is exactly what stopped anyone questioning it.


Relaying is a capability decision, not a setting

There is no flag for this, and there should not be one: the adapter registers the go2rtc source itself, so it already decides whether to hand over its own relay address or the recorder's. The right input for that decision is what the recorder can do, and it says so on SETUP:

  • advertises only 1x → nothing for an owned session to control, so the relay is released and go2rtc pulls the recorder directly. No extra hop, no media through the adapter.
  • advertises more → the session is kept and the alias points at the relay.
  • session will not open at all → direct pull, logged. Playback works, speed does not.

The rates are read on SETUP anyway to build the client's speed menu, so learning this costs nothing extra. One property remains, and it is infrastructure rather than policy:

Property Purpose
relay_host hostname go2rtc dials to reach the adapter (the compose service name). The listen port is assigned by the OS, so nothing needs mapping — but the name must resolve on the container network

Never serve the relay at a path equal to the alias

go2rtc's streams.Patch begins by checking whether the source is an RTSP URL whose path names an existing stream — a self-reference guard — and when it matches it returns that stream *untouched, while answering 200. The relay binds an ephemeral port, so every re-registration of the same transaction carries a new one: detaching the tile into its own window, or a seek that has to rebuild the session. Register rtsp://host:PORT/{alias} and those updates are silently dropped — go2rtc goes on dialing the port of the relay it was just told to forget, and signaling fails with 500 while the adapter's log cheerfully reports the new relay is up.

Serve under a prefix instead — rtsp://host:PORT/relay/{alias} — so the path can never collide with a stream name. Cheap to confirm on any deployment: PATCH a stream twice with different ports and read GET /api/streams back.


Driving the session with SharpRTSP

SharpRTSP (NuGet package SharpRTSP, branch dotnetcore) is enough to own the session. The whole trick-mode surface is four requests; what follows is the shape a shipping NVR adapter uses, trimmed to the parts that matter.

Opening it. DESCRIBE first — its SDP is what you answer a puller with — then SETUP over interleaved TCP. Note that no PLAY is sent yet: media should only start flowing when someone is actually watching.

RtspUtils.RegisterUri();

var transport = RtspUtils.CreateRtspTransportFromUrl(replayUri, credentials);
var listener = new RtspListener(transport, loggerFactory.CreateLogger<RtspListener>())
{
    AutoReconnect = false,
};
listener.MessageReceived += OnMessageReceived;   // completes the pending request, see below
listener.Start();

var describe = Send(
    new RtspRequestDescribe { RtspUri = replayUri, Headers = { { "Accept", "application/sdp" } } },
    "DESCRIBE");

var sdp = Encoding.UTF8.GetString(describe.Data.ToArray());

var setup = new RtspRequestSetup { RtspUri = ControlUri(sdp, replayUri) };
setup.Headers["Transport"] = "RTP/AVP/TCP;unicast;interleaved=0-1";

var setupResponse = Send(setup, "SETUP");
var sessionId = setupResponse.Session;

The SETUP response is where the capabilities are. This is the only place the device says which rates it will accept, so read it here and keep it:

// Media-Properties: Random-Access=1.0s, Unlimited, Immutable,Scales="-1, 0.5, 0.25, 0.125,:1, 2, 4"
var scales = PlaybackScales.Parse(Header(setupResponse, "Media-Properties"));

// Bind to the channel the device assigned, not the one you asked for.
var dataChannel = ReadDataChannel(setupResponse);   // parses interleaved=N-M from Transport
listener.DataReceived += OnInterleaved;

Changing rate or position is one request on the session you already hold — that is the whole point:

public bool Play(double scale = 1.0d, DateTime? seekTo = null)
{
    // Never ask for a rate the device did not advertise; it would answer 4xx and leave the session
    // in whatever state it was, with the client believing the new rate took.
    var applied = scales.Clamp(scale);
    if (applied is null)
        return false;

    var play = new RtspRequestPlay { RtspUri = _uri, Session = sessionId };

    if (Math.Abs(applied.Value - 1.0d) > 0.001d)
        play.Headers["Scale"] = applied.Value.ToString("0.000", CultureInfo.InvariantCulture);

    if (seekTo.HasValue)
        play.Headers["Range"] = $"clock={seekTo.Value:yyyyMMdd}T{seekTo.Value:HHmmss}Z-";

    var response = Send(play, "PLAY");
    if (response?.ReturnCode != 200)
        return false;

    // The device reports where it actually resumed - that is the authoritative position.
    Position = ReadRangeStart(Header(response, "Range")) ?? Position;
    Scale = applied.Value;
    return true;
}

PAUSE is Send(new RtspRequestPause { RtspUri = _uri, Session = sessionId }, "PAUSE"), and resume is just Play(Scale) again. SharpRTSP also ships RtspRequestPlay.AddPlayback(time, scale), which sets Scale and Range: clock= together — worth using if you go the ONVIF-replay route rather than this vendor dialect.

Reading the media. Interleaved RTP arrives on the listener, filtered by channel:

private void OnInterleaved(object? sender, RtspChunkEventArgs e)
{
    if (e.Message is not RtspData data || data.Channel != dataChannel)
        return;

    VideoRtpReceived?.Invoke(this, data.Data.ToArray());
}

Request/response pairing. RtspListener is event-driven, so a synchronous Send needs a completion the message handler fulfils, plus the one 401 retry digest auth requires:

private RtspResponse? Send(RtspRequest request, string method)
{
    var response = SendOnce(request, method);

    if (response?.ReturnCode != 401)
        return response;

    // First contact always 401s: build the digest from the challenge and replay the request.
    if (response.Headers.TryGetValue("WWW-Authenticate", out var challenge) && challenge is not null)
    {
        _authentication = Authentication.Create(_credentials, challenge);
        return SendOnce(request, method);
    }

    return response;
}

private RtspResponse? SendOnce(RtspRequest request, string method)
{
    if (_authentication is not null)
        request.Headers["Authorization"] = _authentication.GetResponse(++_nonceCounter, _uri.AbsoluteUri, method, []);

    var completion = new TaskCompletionSource<RtspResponse>(TaskCreationOptions.RunContinuationsAsynchronously);
    lock (_gate) { _pending = completion; }

    return listener.SendMessage(request) && completion.Task.Wait(RequestTimeout)
        ? completion.Task.Result
        : null;
}

Keeping it alive. The session dies on the timeout the device advertises (Session: 1835784359;timeout=60), and a dropped session takes the trick-mode state with it, so poll inside that window:

_keepAlive = new Timer(
    _ => Send(new RtspRequestOptions { RtspUri = _uri, Session = sessionId }, "OPTIONS"),
    null,
    TimeSpan.FromSeconds(25),
    TimeSpan.FromSeconds(25));

Keep the whole session — open, seek, rate change, teardown — in one playback-session class per transaction, so the keepalive and the recorder slot have a single owner.


Traps found on live hardware

One RTSP session per playback transaction

The relay holds a recorder session for every active playback, so release it when the last puller detaches and reopen it on the next DESCRIBE — which is what keeps an idle playback tile from occupying a recorder slot.

The design target is well under a hundred concurrent playbacks, and at that scale the relay is not the binding limit: the recorder's own cap on concurrent replay sessions is, and it is typically a fraction of that per device. Size the deployment against the recorder, not against the adapter.

Bandwidth scales with the rate

4x playback is literally 4x the data per second, and with a relay all of it traverses the adapter container. Fine for a few review sessions; never a design for live walls, which is why live stays a direct NVR → go2rtc pull.

A replay SDP is a list of candidates, not a description of what will be sent

A recorder can offer several video tracks and serve exactly one of them, and answer SETUP and PLAY with 200 OK on all of them. Verified on live recorder hardware: every replay SDP offers H265 (track0), JPEG (track3) and H264 (track1) on every channel regardless of what that channel recorded, and only the one matching the recorded codec carries RTP. On footage recorded as H.264: 0 bytes on track0 against 348 KB in five seconds on track1, same channel, same Range: clock=.

So take the first a=control: and you get a session that is healthy in every log line and never delivers a frame. Treat the SDP's video tracks as an ordered candidate list: play the first, wait a couple of seconds for the first RTP packet, and move on if it stays silent. Gate the probing on "no media has EVER arrived on this session" — once you know which track is real, a silence is a gap in the footage or a pause, and must never move the track underneath the viewer.

Two details decide whether this works:

  • Exclude non-video sections. An ONVIF metadata track (m=application) answers SETUP like any other and would burn a whole probe window on something that is never a picture.
  • Describe the track you actually serve. A relay forwards the RTP verbatim, payload type included, so the SDP it answers its own puller with must be the served track's media section. Announce the first track while forwarding the third's packets and the puller is told to expect H265 on PT 105 while H264 arrives on PT 98: every packet is discarded, and nothing reports an error — an ffmpeg puller sat on a healthy relay for 150 s and never printed a stream.

A device whose replay SDP has a single video track yields one candidate, so it behaves exactly as before and pays nothing: there is nothing to fall back to. The cost is only paid where the first track is wrong.

Hunt once per relay, and never with a puller waiting. Which track carries media is a fact about what the channel recorded, so it survives the session being closed and reopened — keep it outside the per-session state. Re-hunting on a reopen means seconds of not forwarding RTP on the very thread the relay answers its puller on, and go2rtc reads that as a dead producer and drops it: read tcp …: i/o timeout, then the relay detaches its puller, closes the device session, and the next thing the operator does fails on a session that no longer exists.

A seek must reopen the session, not rebuild the relay

Closing the device session when the last puller leaves is right — an idle playback tile should not hold one of a recorder's few replay slots. But it means an operator's seek routinely arrives on a session that is not open, and if PLAY simply fails there the adapter's fallback is to tear the relay down and build a new one.

That fallback is far worse than it looks, because the new relay listens on a new port while go2rtc goes on dialling the old one: connection refused, and the viewer sits on a tile that says CONNECTING for ever, even though every command returned success and the recorder was serving all along. HARDWARE-VERIFIED, 17 s per rebuild, and re-registering the same alias does not rescue the browser's existing WebRTC session.

So make the session reopen itself on PLAY. The relay's own listening socket never moves, the alias keeps its producer, and the seek costs one DESCRIBE/SETUP instead of a rebuild. Keep the adapter's rebuild path for a position the recorder genuinely refuses — just stop reaching it for the ordinary case.

SharpRTSP: interleaved RTP arrives on the listener

With SharpRTSP 1.11.1, read interleaved RTP from RtspListener.DataReceived and filter by RtspData.Channel. The RtpTcpTransport.DataReceived handler never fired against the verified firmware even though the same packets were landing on the listener — an hour of "0 packets" for nothing. Bind to the channel the device assigned in its SETUP response, not the one you asked for.

Read the right branch of the library

SharpRTSP's default branch is dotnetcore; master is a legacy .NET Framework 4.0 tree with a different API. The package ships RtspRequestPlay.AddPlayback(time, scale) (sets Scale + Range: clock=), ONVIF helpers, and an 0xABAC timestamp parser.

H.265 still costs a transcode

Trick modes are orthogonal to codecs. A recorder replaying H.265 still forces go2rtc to transcode for WebRTC — the relay changes who owns the session, not what the browser can decode.


What the device tells you about position

The PLAY response carries Range: clock=… at the instant it actually resumed, which is the authoritative playback position — no guessing needed once the adapter owns the session. That is also why real PAUSE is worth having: the recorder holds the position, so resume is a PLAY, not a re-seek.

The ONVIF replay RTP extension (0xABAC, an NTP timestamp on the first packet of every access unit) is the other way to learn per-frame time, but some vendor dialects do not send it — they use extension id 0x0001, vendor-private, instead. Do not build a position feature on it without checking first.


See also