Skip to content

Recording/streaming: no more game slowdown, correct A/V timing, Twitch-ready presets - #19752

Merged
LibretroAdmin merged 16 commits into
libretro:masterfrom
GavinDarkglider:recording-fixes
Oct 8, 2026
Merged

LibretroAdmin merged 16 commits into
libretro:masterfrom
GavinDarkglider:recording-fixes

Conversation

@GavinDarkglider

Copy link
Copy Markdown
Contributor

Summary

Recording or streaming with the FFmpeg driver slows the game down whenever the encoder falls behind: the frontend blocks in push_video/push_audio until the encoder catches up. On slower hardware (Switch/L4T, Raspberry Pi, older PCs) that makes GPU recording close to unusable. Several timing and rotation bugs also put broken video into the file even when the game ran fine.

This PR:

  • keeps the game at full speed by dropping frames, never stalling the frontend;
  • makes the timeline right for cores that don't present every tick and for cores that go silent;
  • makes the RTMP presets meet Twitch/YouTube ingest requirements;
  • adds an optional hardware H.264 encoder;
  • adds a cheaper 32-bit GPU readback;
  • fixes rotation, small-frame scaling and the SDL2 readback;
  • adds an opt-in "Record Game Only" mode that leaves the menu, notifications and on-screen messages out of the recording.

Default behaviour stays compatible, except for two changes. Drop-on-full is on by default; it can be switched off to block as before. The RTMP presets now use CBR with a fixed 2 s GOP.

Commits and the issues they address

# Commit Issues
1 recording: add video_record_allow_frame_drop and video_record_fifo_frames settings for #2
2 record_ffmpeg: drop instead of stalling when the encoder falls behind #16372, #14013, #7841 (slowdown while recording)
3 record_ffmpeg: make RTMP streaming presets ingest-compliant Twitch/YouTube: CBR, 2 s keyframes even with drops / half-rate cores, AAC 48 kHz
4 record_ffmpeg: optional hardware H.264 encoder with libx264 fallback new video_record_hw_encoder (nvv4l2, v4l2m2m, rkmpp, nvenc, amf, qsv, videotoolbox, mediacodec)
5 record_ffmpeg: run encoder threads at lower priority in drop mode #16372 (encoder competing with the core)
6 record_ffmpeg: route libav* logging into RetroArch's log codec errors visible in the log
7 video: BGRX readback for GPU recording (gl, vulkan) #16372, #7841, #14013 (GPU recording cost)
8 video: BGRX readback for glcore, gl1, sdl2 and sdl3; fix sdl2 readback SDL2 GPU recordings and screenshots were garbage
9 recording: performance counters for the recording path measuring the above (perfcnt_enable)
10 record_ffmpeg: timestamp video from the audio clock half-length video for 30 fps games on 60 Hz cores (e.g. Dolphin)
11 recording: optional "Record Game Only" (no menu, OSD or menu time) menu, volume/OSD messages, widgets in recordings
12 recording: apply the content rotation to raw (non-GPU) recordings #4819
13 record_ffmpeg: scale frames smaller than the output size to fill it #18871; also makes video_record_scale_factor > 1 work
14 record_ffmpeg: fill audio gaps when the core stops producing audio #3886
15 record_ffmpeg: drop the global ctx, close the output on failure exported global named ctx; file left open on failed start

Fixes #3886, fixes #4819, fixes #18871.
Should also address #16372 (GPU recording slowdown), #14013 and #7841 (Windows/NVENC, which I couldn't test); I'll leave those for the reporters to confirm.

Measurements

Setup

  • 2-vCPU Xeon @ 2.8 GHz with software rendering (Mesa llvmpipe / lavapipe under Xvfb), so GPU work also costs CPU, which is roughly what a small ARM board feels like.
  • Recording runs paced at 60 fps, 900 frames, best of 2.
  • The baseline is master plus the same performance counters and nothing else.
  • Main-thread cost is record_readback + record_push_video per frame.
  • The reproduction tools are attached (bench.py, avsync.sh, the pattern core).
Scenario master fps PR, drop off PR (default) frames dropped main thread per frame, master / PR
GPU rec, gl, 1440x1080, medium 49.4 51.1 57.6 24% 2.83 / 2.40 ms
GPU rec, vulkan, 1440x1080, medium 36.9 37.4 44.7 25% 4.11 / 3.13 ms
GPU rec, glcore, 1440x1080, medium 50.8 51.3 58.5 23% 2.94 / 2.44 ms
Twitch high (FLV), gl, 1440x1080 50.5 55.3 57.9 15% 2.89 / 2.61 ms
Raw 1280x720 core, medium 53.7 52.8 58.7 0% 1.45 / 0.96 ms
1 CPU only, GPU rec gl 1440x1080 27.1 25.8 52.6 71% 5.97 / 1.69 ms
Mr. Boom (real core), GPU rec gl 47.3 51.0 54.7 22% 3.18 / 2.84 ms

What this shows:

  • Game speed. With drop-on-full (default), the game stays at 54.7 to 58.7 fps where master runs at 27 to 54 fps.
  • Constrained CPU. In the 1-CPU case the game runs at 52.6 fps instead of 27.1 fps, about twice the speed, and the recording keeps every frame it can encode at the right timestamp.
  • Drop off. With drop-on-full off, the PR matches master speed, because the frontend still blocks on the encoder there.
  • BGRX readback. On x86 the readback itself costs about the same either way, because the BGR24 conversion is SIMD there. The BGRX path matters on ARM, where that conversion is plain C on the main thread. On a Switch (Lakka, L4T), Quake 3 (vulkan and gl) went from "laggy as all hell" to playable while recording. Switch numbers using recbench.sh will follow in a comment.

Correctness on real runs (RetroArch + pattern core)

Case master PR
30 fps core on 60 Hz (video every other tick), 40 s video ends at 20.0 s, audio at 40.0 s video 39.97 s, audio 39.98 s
Core silent for 5 s (#3886), 20 s video 20.0 s, audio 15.0 s video 20.0 s, audio 20.0 s
video_rotation = 1, raw recording (#4819) 320x240, unrotated 240x320, rotated
Twitch high preset, video every other tick, 20 s keyframes at every scene cut (~0.45 s apart), VBR, AAC 128k; video only 10 s long keyframes every 2.00 s, 6000 kbps CBR, AAC 160k at 48 kHz; video 20 s
SDL2 GPU recording every frame garbage correct
OSD message on screen, gl/glcore/vulkan in every frame (no option) Record Game Only off: in every frame, as before. On: only the first 1 to 4 frames, before the async readback starts

A pattern core paints each frame with its frame number in colour. The recordings are decoded back and checked for order, timestamps, orientation and pixel accuracy on gl, glcore, gl1, vulkan and sdl2, both with and without the threaded video wrapper. The threaded wrapper keeps master's own readback on the video thread.

Testing

  • samples/record/ffmpeg_pipeline (ASan + UBSan, TSan): passes. It gains a drop-mode pass that checks surviving frames come out in order, pixel-exact and at their own timestamps.
  • tools/lock_acquire_budget.sh: ffmpeg_push_video, ffmpeg_push_audio and ffmpeg_thread still take 0 locks.
  • Settings checks pass:
    • settings_check.py table --check
    • settings_check.py unity
    • settings_def_guard_check.py (with --selftest)
    • settings_capacity_check.py
    • settings_migrate_selftest.py
  • Each changed file compiles cleanly with C89_BUILD=1.
  • Every commit builds on its own.
  • Hardware tested on a Switch (Lakka, L4T FFmpeg) on the old base: h264_nvv4l2 encoder, gl and vulkan recording, GB/NES raw recording.
    • The nvv4l2 encoder needs the companion L4T FFmpeg fixes. Without them it can crash on start: it munmaps uninitialised plane pointers, which unmaps parts of the core.
    • Those fixes belong in the L4T FFmpeg tree, so they are a separate patch set for it, not worked around here.

Notes

  • video_record_allow_frame_drop defaults to on. Turning it off restores blocking, which keeps every frame at the cost of game speed.
  • video_record_hw_encoder defaults to off. It only replaces libx264 in the built-in presets; custom configs are untouched.
  • video_record_game_only defaults to off, so recordings still match the screen unless it's enabled.

GavinDarkglider and others added 16 commits October 8, 2026 17:28
…ames

Wires two record_ffmpeg queue knobs through the settings_def
single-source tables, record_params and the Recording settings list
(shown for the ffmpeg driver only).

  video_record_allow_frame_drop  bool, default true
  video_record_fifo_frames       uint, default 32, 8..128 step 4
                                 (Lakka: advanced only)

Both are read when a recording starts; the driver side is the next
patch. msg_hash_lbl_str.h regenerated with tools/gen_lbl_str.py, keeping
only the two new rows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
With the encoder behind, a full queue blocked the frontend in
push_video/push_audio, so a slow encoder slowed the game down.

With video_record_allow_frame_drop (default on):
- a frame that doesn't fit is dropped; the count rides with the next
  queued frame and the encoder leaves a pts gap for it, so video stays
  aligned with audio
- an audio chunk that doesn't fit is owed as silence and written ahead
  of the next chunk that does, and anything still owed is encoded at
  finalize, so the audio track keeps its length
- drops are reported every 64 frames and summarised at teardown

The video queue depth comes from video_record_fifo_frames (8..128).
The audio queue holds 2 s at the core's rate instead of ~0.36 s at
48 kHz. With drop off, both paths block as before.

ffmpeg_pipeline_test gains a drop-mode pass that checks the surviving
frames come out in order, pixel-exact and at their own timestamps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
The Twitch/YouTube/Facebook/Kick modes reused the recording presets:
CRF (unbounded VBR), the encoder's default GOP (~4 s for x264 at
60 fps), audio at whatever rate was closest to the core's and no
audio bitrate. Ingest servers want CBR H.264, a 2 s keyframe interval
and AAC at 44.1/48 kHz.

For the FLV/RTMP modes:
- CBR video at 2500/4500/6000 kbps (low/med/high), 1 s VBV buffer,
  x264 nal-hrd=cbr, no B-frames
- a fixed 2 s keyframe interval: x264 scenecut=0, and keyframes are
  forced on time rather than counted in frames, since pts gaps from
  dropped frames or cores presenting every other tick stretch a
  frame-counted GOP
- AAC 160 kbps at 48 kHz

Rate control goes through the generic AVCodecContext fields so it
carries over to other encoders. Local/custom streaming and the
recording presets are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
New setting video_record_hw_encoder (default on for HAVE_LAKKA_SWITCH,
off elsewhere). When set and the built-in preset would use libx264,
these encoders are tried in order and the first that opens is used:

  h264_nvv4l2, h264_nvmpi, h264_v4l2m2m, h264_rkmpp, h264_nvenc,
  h264_amf, h264_qsv, h264_videotoolbox, h264_mediacodec

All take system-memory frames; VAAPI is skipped because it needs a
hw_frames_ctx upload path. Input is YUV420P if the encoder takes it,
else NV12 (via swscale). x264 private options are not passed to them;
rate control/GOP come from the generic AVCodecContext fields, plus
rc=cbr for the RTMP presets. CRF presets get a bitrate derived from
0.05/0.1/0.2 bits per pixel per frame (low/med/high), floor 500 kbps.
If none open, libx264 is used as before. Custom configs are untouched.

ffmpeg_init_video is split so one encoder attempt is a self-contained
alloc/configure/open that frees its context on failure. Also sets
AVCodecContext.framerate, which some hardware rate controllers need.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
With a CPU-bound software core, libx264's worker threads, the
swscale/AAC work on the encoder thread and the core's main thread all
compete at the same priority, so recording slows the game even when
the queues never fill.

In drop-on-full mode, run everything encoder-side at nice +10:
- the driver's encoder thread renices itself on start
- threads spawned during the video/audio avcodec_open2() (x264's pool
  and lookahead) are found by diffing /proc/self/task and reniced
  afterwards; lowering the main thread around the open instead isn't
  reversible without CAP_SYS_NICE

The encoder then gets what the core leaves, and any shortfall shows up
as dropped frames rather than game slowdown. Linux only; skipped when
drop-on-full is off, since starving the encoder there means stalls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
FFmpeg logged to stderr through its default callback, so codec
messages never reached RetroArch's log or log file, which is what a
user can attach after a crash.

While recording, install an av_log callback that assembles lines
FFmpeg builds from several calls and forwards them as
RARCH_ERR/WARN/LOG/DBG by level, tagged [FFmpeg/lav]. A partial line
is flushed when the next piece comes from another context, so two
threads' messages are never spliced. FFmpeg's level follows the
frontend log level. The default callback is restored when the
recording is freed.

Adds verbosity_get_log_level().

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
GPU (post-shaded) recording, which hardware-rendered cores are forced
into, read every frame with read_viewport(), which converts the
32-bit readback to BGR24 on the main thread inside the frame budget.
The encoder thread then converts BGR24 to YUV again, and the frame is
queued at the viewport rounded up to powers of two.

- Add an optional video_driver_t::read_viewport_bgrx() that returns
  the viewport as 32-bit B,G,R,X plus its row order. It is the last
  slot, so positional driver tables leave it NULL.
- gl: desktop GL already reads back GL_BGRA, so rows are copied as-is
  (RGBA is swizzled on GLES), on both the PBO and synchronous paths.
- vulkan: B8G8R8A8 swapchains copy rows as-is, RGBA ones swizzle.
- Recording uses it when the driver has it, feeding the encoder
  ARGB8888: the main thread's work per frame becomes a plain copy and
  the encoder thread converts once. Little-endian hosts only, since
  FFEMU_PIX_ARGB8888 is a native-endian word. The threaded wrapper
  keeps its own readback on the video thread.
- The encoder queue is sized to the exact viewport (a resize already
  ends the recording): 1280x720 rather than 2048x1024 per frame.

The BGR24 path is unchanged for screenshots and other drivers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
Extend read_viewport_bgrx() to the other Linux-capable drivers.

- glcore: the async PBO readback reads GL_BGRA on desktop GL (GLES
  keeps RGBA), with the BGR24 scaler's input format to match, so
  recording gets a plain row copy as on gl.
- gl1: synchronous RGBA readback, swizzled to BGRX.
- sdl3: sdl3_capture_viewport() can emit BGRX, top row first.
- sdl2: read_viewport() converted the whole window surface
  (SDL_GetWindowSurface() on a window that has a renderer), copied
  surface->h * pitch into a buffer sized for the viewport, top row
  first where callers expect bottom row first, and leaked the
  converted surface on every call, so GPU recording and screenshots
  came out as garbage. Do what sdl3 does: redraw the last core frame
  into the game viewport and SDL_RenderReadPixels() exactly that, as
  BGRX or as flipped BGR24.

The drivers that pass OpenXR slots positionally get NULLs for them
so the new slot lands in place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
Register frontend performance counters for each stage of a recording,
reported at exit with perfcnt_enable, so its cost can be placed:

  record_readback      main thread: GPU readback
  record_push_video    main thread: queueing a frame for the encoder
  record_push_audio    audio path: queueing audio
  record_scale         encoder thread: colour conversion and scaling
  record_encode_video  encoder thread: video encode and mux
  record_encode_audio  encoder thread: audio resample, encode and mux

Counters are only touched with perfcnt_enable set, and each has a
single writer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
Video pts was one tick per frame handed to the encoder, i.e. it
assumed the core delivers a frame on every retro_run(). Many don't: a
30 fps game on a 60 Hz core (Dolphin's Wind Waker, for one) only
calls the video callback for new frames, and frame skipping and
run-ahead do the same. Those recordings played their video at the
wrong speed: a 26-minute Wind Waker session produced 13 minutes of
video under 26 minutes of audio, and RetroArch's own player, which is
clocked by audio, hung once the video track ran out.

Count every audio frame the core produced, which is the core's real
elapsed time, and place each video frame at that point on the
timeline through the same pts skip dropped frames use. Frames that
arrive every other tick get pts 0, 2, 4, ... Without audio, the
one-tick-per-frame accounting is kept.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
GPU recording reads the backbuffer at the end of the frame, so the
volume and other OSD messages, notification widgets, statistics,
input overlays and, with content running behind it, the menu end up
in the recording, and the menu's filler silence puts the time spent
in the menu into it too. That is what's on screen, which some users
expect, so this adds a switch rather than changing the default.

New setting video_record_game_only ("Record Game Only", default off).
It travels with the frame, like gpu_recording. When on:

- gl, glcore: the async PBO readback is issued right after the core's
  image (shader chain, PQ->SDR conversion) is in the backbuffer,
  before any UI is drawn. Under scRGB output the backbuffer only gets
  the image at the final composite, so that path reads back at the end.
- vulkan: after the final pass, end the render pass, copy the
  backbuffer to the readback staging buffer and resume with
  keep_render_pass (LOAD), so the UI draws over the unchanged image.
  HDR output keeps the end-of-frame readback, which needs its tonemap.
- audio_driver_menu_sample() doesn't push the menu's filler silence,
  so menu time is left out, as its frames already are.

When off, behaviour is as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
Raw recordings encoded the core's frames as handed over, ignoring the
rotation the display applies (core-requested or the user's video
rotation), so vertical games came out sideways. GPU recordings were
unaffected since they read back the already-rotated screen.

Pass the rotation to the recorder, describe the output with width and
height swapped for quarter turns (unless --size gave one), and rotate
each frame on the encoder thread before scaling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
ffmpeg_scale_input() clamped the scaling destination to the source
size whenever a frame was smaller than the output. That was meant for
the one pixel of padding added to reach even dimensions, but applied
to any smaller frame: frames below the core's base geometry landed in
the top-left corner of a black frame (Mupen64Plus-Next with Angrylion
and GPU recording off), and video_record_scale_factor > 1 had no
effect, since every frame is smaller than the scaled output.

Scale every frame to the output size; only a frame exactly one pixel
short of it in a dimension (the even-size padding) is copied 1:1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
A core that produces no audio for a while, rather than producing
silence, sent nothing to the recorder, so that time was missing from
the audio track and everything after it was out of sync (e.g.
Mupen64Plus' 64DD loading). With video timestamps taken from the
audio clock it would also stall the video timeline.

Once 4 video frames in a row arrive without any audio, owe one frame
period (sample_rate / fps, with the fraction carried) of silence per
video frame, backfilling the 4 already seen, through the existing
owed-silence mechanism. The silence goes ahead of later audio in both
drop and blocking mode.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
The muxer context also lived in a non-static global named 'ctx', an
exported symbol in every binary that links this driver, used by
finalize to close the output. Use handle->muxer.ctx throughout.

A recording that fails after the output was opened (an encoder that
won't open, say) never reached finalize, so its file stayed open.
ffmpeg_free() closes it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tpe3orhELpViSQpZcsTTjr
…nheritance, frame format across a threading switch

- A raw recording of a turned picture takes the display's aspect as it
  is, which is already the turned one; only an aspect worked out from
  the output size follows the turn, and --size keeps its own shape
  (record_raw_geometry).
- Drop mode says nothing while it records: dropped frames are counted
  and the total is logged when the recording ends.
- In drop mode the encoders are opened on a thread that has lowered its
  own priority, so the threads they start run lowered and no other
  thread is touched.
- Only the frontend-thread perf counters are registered, as the menu
  reads and resets them live.
- A GPU recording opened for 32-bit frames gets them widened from BGR24
  when the readback in charge hands over BGR24 - the threaded wrapper's,
  or a driver without the 32-bit one - into its own buffer.
- At the end of a recording the real audio left in the queue is encoded
  ahead of the silence still owed; the gap fill counts every frame the
  core presents, frame_drop_ratio or not; the output is closed whenever
  it is open; an encoder that leaves the CBR request unused is named.

ffmpeg_pipeline_test checks the geometry, that drop mode drops frames
without a warning or an on-screen message, and that its encoder
threads run lowered. threaded_video's gpu record format lane records
on Vulkan unthreaded and threaded with a 32-bit recording open, and
reads every row of what it is handed.
@LibretroAdmin
LibretroAdmin merged commit 885dc8d into libretro:master Oct 8, 2026
111 of 118 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants