Skip to main content

Tracking ReID Filter

The Tracking ReID filter enriches object tracks with appearance embeddings using deep learning-based Re-Identification (ReID) models. It is a secondary signal in multi-camera tracking pipelines, used to confirm object identity when geometry and temporal constraints are ambiguous.

📋 OverviewDirect link to 📋 Overview

This filter processes per-camera tracks and generates vector embeddings that capture the visual appearance of tracked objects. These embeddings are later used downstream for cross-camera association and global ID assignment.

Key Principle: ReID is a fallback mechanism. Geometry and time constraints are always the primary signals.


✨ FeaturesDirect link to ✨ Features

  • Dual Model Support

    • OSNet (lightweight, production-ready)
    • TransReID (transformer-based, more accurate, slower)
  • Appearance Embedding Generation

    • Extracts object crops from bounding boxes
    • Runs configurable ReID models
    • Generates L2-normalized feature vectors
    • Stores embeddings as Python lists for serialization
  • Flexible Input/Output Schema

    • Configurable input key (tracks, detections)
    • Configurable bounding box and embedding keys
    • Robustness to multiple image storage locations
  • Production-Ready Error Handling

    • Graceful handling of invalid bounding boxes
    • Minimum crop size validation (8px default)
    • Automatic fallback for missing/invalid data
    • Per-frame exception isolation
  • Smart Caching & Optimization

    • Skip re-computation if embedding already exists
    • L2 normalization for stable cosine similarity
    • Optional batch processing support
  • Model Metadata Tracking

    • Stores model type and name in output
    • Aids downstream debugging and audit trails

🛠️ Use CasesDirect link to 🛠️ Use Cases

  • Cross-Camera Appearance Matching

    • Confirm object identity when multiple candidates exist
    • Secondary validation after geometry filtering
  • Multi-Camera Linear Tracking

    • Objects moving sequentially across cameras (camera 1 → 2 → 3)
    • ReID disambiguates when spatial/temporal alone is insufficient
  • Production Monitoring

    • Audit embeddings for quality/distribution
    • Track model performance via metadata logs

⚙️ ConfigurationDirect link to ⚙️ Configuration

filter: FilterTrackingReid
config:
model_type: "osnet"
model_name_or_path: "osnet_x0_25"
device: "cuda"
input_key: "tracks"
bbox_key: "bbox"
output_embedding_key: "reid_embedding"
normalize_embeddings: true
min_crop_size: 8

Advanced Setup (TransReID)Direct link to Advanced Setup (TransReID)

filter: FilterTrackingReid
config:
model_type: "transreid"
model_name_or_path: "path/to/transreid_hf_export" # Local or HF repo
device: "cuda"
transreid_feature_token: "cls" # or "mean"
normalize_embeddings: true

Configuration ParametersDirect link to Configuration Parameters

ParameterTypeDefaultDescription
model_typestr"osnet"Model architecture: "osnet" or "transreid"
model_name_or_pathstr"osnet_x0_25"Model name/path: HuggingFace repo, local dir, or checkpoint
devicestr"cuda"PyTorch device: "cuda", "cuda:0", or "cpu"
input_keystr"tracks"Frame data key containing tracks/detections
bbox_keystr"bbox"Track field containing bounding box [x1, y1, x2, y2]
output_embedding_keystr"reid_embedding"Output field name for embedding vector
output_model_info_keystr"reid_model"Output field name for model metadata
normalize_embeddingsbooltrueL2-normalize embeddings (recommended)
skip_if_existsbooltrueSkip computation if embedding already present
min_crop_sizeint8Minimum valid crop width/height in pixels
transreid_feature_tokenstr"cls"TransReID feature extraction: "cls" or "mean"

📊 Input/Output FormatDirect link to 📊 Input/Output Format

Input (from upstream tracking filter)Direct link to Input (from upstream tracking filter)

{
"camera_id": "cam_1",
"timestamp": 1234567890.5,
"tracks": [
{
"local_track_id": 12,
"bbox": [100, 150, 200, 300],
"class_name": "person",
"score": 0.95
}
]
}

The frame must also contain an image accessible via:

  • frame.image (numpy array)
  • frame.frame (numpy array)
  • frame.data["image"] (numpy array)
  • frame.data["frame"] (numpy array)

Output (enriched tracks)Direct link to Output (enriched tracks)

{
"local_track_id": 12,
"bbox": [100, 150, 200, 300],
"class_name": "person",
"score": 0.95,
"reid_embedding": [0.12, -0.34, 0.56, ...], # L2-normalized
"reid_model": {
"model_type": "osnet",
"model_name_or_path": "osnet_x0_25"
}
}

🔗 Pipeline IntegrationDirect link to 🔗 Pipeline Integration

The ReID filter is positioned after geometry/time-based candidate filtering:

VideoIn
↓
RT-DETR (Detection + ByteTrack)
↓
Homography (World Position)
↓
Tracklet/Event Builder
↓
Candidate Gate (Geometry + Time Filter)
↓
[ReID Filter] ← Primary use here
↓
Global Association (Multi-Camera)
↓
Output (MQTT / DB / Logs)

Design Principle: ReID only runs when geometry/time cannot disambiguate. This minimizes inference overhead.


🚀 PerformanceDirect link to 🚀 Performance

  • Inference Speed (per object):

    • OSNet: ~10-20ms on V100 GPU
    • TransReID: ~50-100ms on V100 GPU
  • Memory:

    • OSNet: ~200MB VRAM
    • TransReID: ~1-2GB VRAM
  • Optimization:

    • Skip-if-exists reduces redundant computation
    • L2 normalization is done CPU-side (negligible cost)
    • Batch processing support for future enhancements

📝 Implementation DetailsDirect link to 📝 Implementation Details

OSNet BackendDirect link to OSNet Backend

  • Uses torchreid.utils.FeatureExtractor
  • Lightweight CNN-based architecture
  • Production-ready with minimal dependencies
  • Output dimension: typically 128-2048 features

TransReID BackendDirect link to TransReID Backend

  • Loads via HuggingFace transformers library
  • Transformer-based architecture for better accuracy
  • Supports both last_hidden_state and pooler_output
  • Flexible feature extraction: CLS token or mean pooling

Robustness FeaturesDirect link to Robustness Features

  • Crop validation: rejects boxes smaller than min_crop_size
  • Image format handling: PNG, JPEG, PIL Image, numpy arrays
  • Pixel coordinate clipping: prevents out-of-bounds access
  • Per-frame error isolation: exceptions don't crash the filter
  • Logging: detailed debug info for troubleshooting

⚠️ Important NotesDirect link to ⚠️ Important Notes

  1. Dependencies:

    • torchreid (for OSNet mode)
    • transformers + torch (for TransReID mode)
    • Both are optional; missing dependencies degrade gracefully
  2. Model Loading:

    • OSNet: Automatically downloaded from torchreid hub on first use
    • TransReID: Must be HuggingFace-compatible or local directory
  3. Normalization:

    • L2 normalization is always recommended for downstream similarity computation
    • Enables stable cosine similarity thresholding
  4. Compatibility:

    • Expects input tracks under configurable key (default: "tracks")
    • Field names must match configuration (e.g., bbox_key, output_embedding_key)
    • Compatible with OpenFilter Frame API