Video Playback: Difference between revisions
| Line 501: | Line 501: | ||
Here’s a command that encodes video and audio while maintaining high time accuracy: | Here’s a command that encodes video and audio while maintaining high time accuracy: | ||
<pre> | <pre> | ||
ffmpeg -i input.mp4 -c:v libx264 -preset slow -crf 18 -vsync cfr -g 30 -c:a pcm_s16le -ar | ffmpeg -i input.mp4 -c:v libx264 -preset slow -crf 18 -vsync cfr -g 30 -c:a pcm_s16le -ar 48000 -fflags +genpts -async 1 output.mp4 | ||
-c:v libx264: Encode video using H.264. | -c:v libx264: Encode video using H.264. | ||
-preset slow: Optimize for quality and compression efficiency. | -preset slow: Optimize for quality and compression efficiency. | ||
Revision as of 08:00, 29 April 2025
When using video in your experiment, especially when presenting time-critical stimuli, special care should be taken to optimize the video and audio settings on multiple levels (hardware, OS, script), as many things can go wrong along the way.
This page outlines some best practices; however, we advise to always consult a TSG member if you plan to run a video experiment in the labs.
Video playback
Note that the Lab Computer displays are typically set to 1920×1080 at 120Hz. We found that this is sufficient for most applications. There are possibilities to go higher. Later in this wiki we will explain how to build audio and video. We will start with playing video, both with and without audio.
Python
Example demonstrating how to play a video with audio:
from psychopy import logging, prefs
prefs.hardware['audioLib'] = ['PTB']
prefs.hardware['audioLatencyMode'] = 2
from psychopy import visual, core, event
from psychopy.hardware import keyboard
# File paths for video and audio
video_file = "tick_rhythm_combined_30min.mp4"
win = visual.Window(size=(1024, 768), fullscr=False, color=(0, 0, 0))
video = visual.VlcMovieStim(
win, filename=video_file,
autoStart= False
)
kb = keyboard.Keyboard()
# Play the video
win.flip()
core.wait(3.0)
video.play()
video_start_time = core.getTime()
# Main loop for video playback
while video.status != visual.FINISHED:
# Draw the current video frame
video.draw()
win.flip()
keys = kb.getKeys(['q'], waitRelease=True)
if 'q' in keys:
break
win.close()
core.quit()
Example demonstrating how to play a video with audio disconnected:
from psychopy import logging, prefs
from psychopy import visual, core, sound, event
import time
prefs.hardware['audioLib'] = ['PTB']
prefs.hardware['audioLatencyMode'] = 2
# File paths for video and audio
video_file = "tick_rhythm_30min.mp4"
audio_file = "tick_rhythm_30min.wav"
win = visual.Window(size=(1280, 720), fullscr=False, color=(0, 0, 0), units="pix")
video = visual.VlcMovieStim(
win, filename=video_file,
size=None, # Use the native video size
pos=[0, 0],
flipVert=False,
flipHoriz=False,
loop=False,
autoStart=False,
noAudio=True,
volume=100,
name='myMovie'
)
# Load the audio
audio = sound.Sound(audio_file, -1)
# Synchronize audio and video playback
win.flip()
time.sleep(5)
audio.play()
time.sleep(0.04)
video.play()
video_start_time = core.getTime()
while video.status != visual.FINISHED:
# Draw the current video frame
video.draw()
win.flip()
# Check for keypress to quit
if "q" in event.getKeys():
audio.stop()
break
# Close the PsychoPy window
win.close()
core.quit()
Example demonstrating if video and audio encoding are correct:
import subprocess
import json
file_path = "C_dyad1_video2_241123.mp4"
def check_video_file(file_path):
try:
# Run ffprobe to get file metadata in JSON format
result = subprocess.run(
[
"ffprobe",
"-v", "error",
"-show_streams",
"-show_format",
"-print_format", "json",
file_path
],
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
text=True
)
metadata = json.loads(result.stdout)
except Exception as e:
print(f"Error running ffprobe: {e}")
return
# Check for video stream
video_stream = next((stream for stream in metadata['streams'] if stream['codec_type'] == 'video'), None)
if video_stream:
# Check video codec
video_codec = video_stream.get('codec_name')
if video_codec == 'h264':
print("Video codec: H.264")
else:
print(f"ERROR: Video codec is NOT H.264 (Found: {video_codec})")
# Extract and report frame rate
if 'r_frame_rate' in video_stream:
raw_frame_rate = video_stream['r_frame_rate']
calculated_frame_rate = eval(raw_frame_rate) # Convert string like "30/1" to float
print(f"Frame rate: {calculated_frame_rate:.2f} FPS (raw: {raw_frame_rate})")
else:
print("ERROR: Could not determine raw frame rate from metadata.")
# Check for constant frame rate
if video_stream.get('avg_frame_rate'):
avg_frame_rate = eval(video_stream['avg_frame_rate'])
if abs(avg_frame_rate - calculated_frame_rate) < 0.01:
print("Frame rate: Constant")
else:
print(f"ERROR: Frame rate is NOT constant (avg_frame_rate: {avg_frame_rate:.2f} FPS)")
else:
print("ERROR: Could not determine average frame rate consistency.")
# Check for frame drops
try:
frame_info_result = subprocess.run(
[
"ffprobe",
"-v", "error",
"-select_streams", "v:0",
"-show_entries", "frame=pkt_pts_time",
"-of", "csv=p=0",
file_path
],
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
text=True
)
# Filter out empty or invalid lines
frame_times = [
float(line.strip()) for line in frame_info_result.stdout.splitlines()
if line.strip() # Exclude empty lines
]
expected_interval = 1.0 / calculated_frame_rate # Expected time between frames
frame_drops = [
i for i, (t1, t2) in enumerate(zip(frame_times, frame_times[1:]))
if abs(t2 - t1 - expected_interval) > 0.01 # Tolerance for irregularity
]
if frame_drops:
print(f"ERROR: Detected frame drops at frames: {frame_drops}")
else:
print("No frame drops detected.")
except Exception as e:
print(f"Error analyzing frames for drops: {e}")
else:
print("ERROR: No video stream found")
# Check for audio stream
audio_stream = next((stream for stream in metadata['streams'] if stream['codec_type'] == 'audio'), None)
if audio_stream:
# Check audio codec
audio_codec = audio_stream.get('codec_name')
if audio_codec == 'pcm_s16le':
print("Audio codec: WAV (PCM)")
else:
print(f"ERROR: Audio codec is NOT WAV (PCM) (Found: {audio_codec})")
# Check sample rate
sample_rate = audio_stream.get('sample_rate')
if sample_rate == "44100":
print("Audio sample rate: 44.1 kHz")
else:
print(f"ERROR: Audio sample rate is NOT 44.1 kHz (Found: {sample_rate} Hz)")
else:
print("ERROR: No audio stream found")
# Check synchronization
if video_stream and audio_stream:
video_start_pts = float(video_stream.get('start_time', 0))
audio_start_pts = float(audio_stream.get('start_time', 0))
if abs(video_start_pts - audio_start_pts) < 0.01: # Tolerance for synchronization
print("Video and audio are synchronized.")
else:
print(f"ERROR: Video and audio are NOT synchronized. Start difference: {abs(video_start_pts - audio_start_pts):.3f} seconds")
else:
print("ERROR: Could not determine synchronization (missing video or audio streams).")
# Example usage
if __name__ == "__main__":
check_video_file(file_path)
Example demonstrating how to disconnect audio from video:
import os
import subprocess
input_file = 'tick_rhythm_combined_1min.mp4'
directory = os.path.dirname(input_file)
base_name = os.path.splitext(os.path.basename(input_file))[0]
output_video = os.path.join(directory, f"{base_name}_video_only.mp4")
output_audio = os.path.join(directory, f"{base_name}_audio_only.wav")
subprocess.run(['ffmpeg', '-i', input_file, '-an', output_video])
subprocess.run(['ffmpeg', '-i', input_file, '-vn', '-acodec', 'pcm_s16le', '-ar', '44100', output_audio])
print(f"Video saved to: {output_video}")
print(f"Audio saved to: {output_audio}")
Example demonstrating how to combine audio and video:
import os
import subprocess
# --- Inputs
video_file = 'tick_rhythm_combined_1min_video_only.mp4' # Your video-only file
audio_file = 'mic_segment.wav' # Your trimmed audio
output_file = 'final_synced_output.mp4' # Output file name
# --- FFmpeg command to combine
subprocess.run([
'ffmpeg',
'-i', video_file,
'-i', audio_file,
'-c:v', 'copy', # Copy video stream as-is
'-c:a', 'aac', # Encode audio with AAC (widely compatible)
'-shortest', # Trim to the shortest stream (prevents overhang)
output_file
])
print(f"Synchronized video saved to: {output_file}")
Video encoding
When recording video for stimulus material or as input for your experiment, please: Use a high-quality camera, with settings appropriate for your application (e.g., frame rate, resolution). Use a high-quality recorder or capture device, capable of recording at 1080p (1920×1080) and 60fps or higher. Stabilize the camera and avoid automatic exposure, white balance, or focus during recording to prevent inconsistencies. Record in a controlled environment with consistent lighting and minimal background distractions. You can use the facecam for high quality video recording.
Video Settings
We recommend using the following settings:
| File format | .mp4 (H.264 codec(libx264)) ik wil hier een link naar de dll? |
|---|---|
| Frame rate | 60 fps (frames per second) |
| Resolution | 1920×1080 (Full HD) or match your experiment's display settings |
| Bitrate | 10-20 Mbps for Full HD video |
| Constant Frame Rate (CFR) | enforce a constant frame rate |
Windows Settings
Windows 10 has a habit of automatically enabling video enhancements or unnecessary processing features, which can interfere with smooth playback. Therefore, please make sure these are disabled:
right click background → Display settings → Graphics Settings. If available, disable "Hardware-accelerated GPU scheduling" for critical timing experiments.
For specific applications (e.g., PsychoPy), under "Graphics Performance Preference," set them to "High Performance" to ensure they use the dedicated GPU.
Python
Example demonstrating how to record a video with a facecam:
#!/usr/bin/env python3.10
# -*- coding: utf-8 -*-
import datetime
import cv2
import ctypes
import ffmpegcv
#set sleep to 1ms accuracy
winmm = ctypes.WinDLL('winmm')
winmm.timeBeginPeriod(1)
def configure_webcam(cam_id, width=1920, height=1080, fps=60):
cap = cv2.VideoCapture(cam_id, cv2.CAP_DSHOW)
if not cap.isOpened():
print(f"Error: Couldn't open webcam {cam_id}.")
return None
# Try to set each property
cap.set(cv2.CAP_PROP_FRAME_WIDTH, width)
cap.set(cv2.CAP_PROP_FRAME_HEIGHT, height)
cap.set(cv2.CAP_PROP_FPS, fps)
# Read back the values
actual_width = cap.get(cv2.CAP_PROP_FRAME_WIDTH)
actual_height = cap.get(cv2.CAP_PROP_FRAME_HEIGHT)
actual_fps = cap.get(cv2.CAP_PROP_FPS)
print(f"Resolution set to: {actual_width}x{actual_height}")
print(f"FPS set to: {actual_fps}")
return cap
def getWebcamData():
global frame_width
global frame_height
print("opening webcam...")
camera = configure_webcam(1, frame_width, frame_height)
time_stamp = datetime.datetime.now().strftime('%Y-%m-%d %H-%M-%S')
file_name = time_stamp +'_output.avi'
video_writer = ffmpegcv.VideoWriter(file_name, 'h264', fps=freq)
while True:
grabbed = camera.grab()
if grabbed:
grabbed, frame = camera.retrieve()
video_writer.write(frame) # Write the video to the file system
frame = cv2.resize(frame, (int(frame_width/4),int(frame_height/4)))
cv2.imshow("Frame", frame) # show the frame to our screen
if cv2.waitKey(1) & 0xFF == ord('q'):
break
freq = 60
frame_width = 1920
frame_height = 1080
getWebcamData()
cv2.destroyAllWindows()
Audio encoding
Audio Settings
We recommend using the following settings for audio:
| Codec | lossless or high-quality codecs |
|---|---|
| PCM (WAV) | uncompressed |
| Sample Rate | 48 kHz |
Set your audio for low-latency, high-accuracy playback with ffmpeg:
ffmpeg -i input.wav -ar 48000 -ac 2 -sample_fmt s16 output_fixed.wav Explanation: -ar 48000 → Set sample rate to 48000 Hz (standard for ASIO/Windows audio, matches most soundcards) -ac 2 → Set 2 channels (stereo) -sample_fmt s16 → Use 16-bit signed integer samples
Windows Settings
Windows 10 Settings to check
sound → Playback → right-click → Properties → Advanced Tab: - Set Default Format to 48000 Hz, 16 bit, Studio Quality. - Disable sound enhancements. - In the same properties window, go to Enhancements tab → Disable all enhancements. - Exclusive Mode: - In the same Advanced tab. - Allow applications to take exclusive control of this device → CHECKED - Give exclusive mode applications priority → CHECKED
Python
Example demonstrating how to check and play your audio:
#!/usr/bin/env python3.10
import psychopy
print(psychopy.__version__)
import sys
print(sys.version)
import keyboard
from psychopy import prefs
from psychopy import visual, core, event
from psychopy.sound import backend_ptb
# 0: No special settings (default, not optimized)
# 1: Try low-latency but allow some delay
# 2: Aggressive low-latency
# 3: Exclusive mode, lowest latency but may not work on all systems
backend_ptb.SoundPTB.latencyMode = 2
prefs.hardware['audioLib'] = ['PTB']
prefs.hardware['audioDriver'] = ['ASIO']
prefs.hardware['audioDevice'] = ['ASIO4ALL v2']
from psychopy import sound
# --- OS-level audio device sample rate ---
default_output = sd.query_devices(kind='output')
print("\nDefault output device info (OS level):")
print(f" Name: {default_output['name']}")
print(f" Default Sample Rate: {default_output['default_samplerate']} Hz")
print(f" Max Output Channels: {default_output['max_output_channels']}")
# Confirm the audio library and output settings
print(f"Using {sound.audioLib} for sound playback.")
print(f"Audio library options: {prefs.hardware['audioLib']}")
print(f"Audio driver: {prefs.hardware.get('audioDriver', 'Default')}")
print(f"Audio device: {prefs.hardware.get('audioDevice', 'Default')}")
audio_file = 'tick_rhythm_5min.wav'
print("Creating sound...")
wave_file = sound.Sound(audio_file)
print("Playing sound...")
wave_file.play()
while not keyboard.is_pressed('q'):
pass
# Clean up
print("Exiting...")
win.close()
core.quit()
FFmpeg
Synchronization
Ensure the audio and video streams have consistent timestamps:
FFmpeg Options:
-fflags +genpts: Generates accurate presentation timestamps (PTS) for the video.
-async 1: Synchronizes audio and video when they drift.
-map 0:v:0 and -map 0:a:0: Explicitly map video and audio streams to avoid accidental mismatches.
Recommended FFmpeg Command
Here’s a command that encodes video and audio while maintaining high time accuracy:
ffmpeg -i input.mp4 -c:v libx264 -preset slow -crf 18 -vsync cfr -g 30 -c:a pcm_s16le -ar 48000 -fflags +genpts -async 1 output.mp4 -c:v libx264: Encode video using H.264. -preset slow: Optimize for quality and compression efficiency. -crf 18: Adjusts quality (lower = better; range: 0–51). -vsync cfr: Enforces constant frame rate. -c:a pcm_s16le: Encodes audio in uncompressed WAV format. -ar 48000: Sets audio sample rate to 48.0 kHz. -fflags +genpts: Ensures accurate timestamps. -async 1: Synchronizes audio and video streams.
Enumeration
- Ensure Low Latency: If you're processing video/audio in real time, use low-latency settings (e.g., -tune zerolatency for H.264).
- Avoid Resampling: If possible, use the original frame rate and sample rate to avoid timing mismatches.
- Testing: Always test playback on different devices or players to confirm synchronization.
Editing
Alternatively, you can use Shotcut, a simple open-source editor, available here: https://shotcut.org/
Another one is DaVinci Resolve for editing and converting video files. DaVinci Resolve is a free, professional-grade editing program, available here: https://www.blackmagicdesign.com/products/davinciresolve