Overview
The Scene Caption filter runs a Vision Language Model (VLM) on a sliding window of video frames and produces a natural-language scene description. The default prompt is generic; override FILTER_PROMPT to focus on specific use cases.
Changelog
Scene Caption filter release notes
Detection gating: throughput measurements (PLAT-1291)
Measured on ps-2x-a10 (2× NVIDIA A10), Cosmos-Reason2-2B on GPU 0, RT-DETR