AI Training Datasets use case

YOLO labels from footage

The yolo labels from footage workflow turns uploaded video into a reviewable frame-level json dataset output. VidScanner analyzes the visual and spoken context, then ties every result back to the source recording with timestamps and evidence.

Security operations monitors showing camera feeds
Video to ML dataset

Workflow guide

How this use case fits into a repeatable video review process.

YOLO labels from footage is for teams that already have the video evidence but do not want to scrub through the entire file manually. Upload driving footage (urban / highway), let VidScanner process the frames and transcript, then review a structured result that points back to the exact moment in the source video.

This workflow fits inside VidScanner AI Training Datasets, so it uses the same search index, timestamps, screenshots, and export path as the rest of the product. That matters when a report, dataset, deck, or action list needs to be checked by another person before it becomes official.

For best results, capture the recording with stable framing, clear audio when narration helps, and enough time on the important visual evidence. The output is strongest when reviewers can see the subject clearly and hear the context that explains why the moment matters.

Sample input

Driving footage (urban / highway)

Sample output

Frame 12: person, forklift, pallet; Attribute: indoor warehouse lighting; Export: JSON or CSV

How it works

  1. Upload driving footage (urban / highway)
  2. Run VidScanner AI Training Datasets
  3. Review the generated output and evidence
  4. Export the result or continue searching the video library

Tips for this workflow

Up to 30 minutes per video (default cap)
Higher resolution = better label accuracy (Gemini multimodal benefits)
Mixed scene content produces a more balanced dataset than one continuous shot
If you need more frames, run multiple videos and concatenate the JSON exports

Review checklist

Confirm the important scene or statement is visible in the source recording.
Check timestamps before sharing the output with a customer, manager, or reviewer.
Export only after the structured output matches the evidence in the video.
Keep the original file available when the result will be used as an audit artifact.

FAQ

What video should I use for YOLO labels from footage?

Start with driving footage (urban / highway). The recording should show the important visual evidence clearly and include narration when spoken context changes the meaning of the scene.

What output does this workflow create?

VidScanner generates frame-level json dataset details such as Frame 12: person, forklift, pallet; Attribute: indoor warehouse lighting; Export: JSON or CSV. Each item stays connected to the source video for review.

Can the result be exported?

AI Training Datasets exports include JSON (full dataset — drop-in for most ML pipelines) and CSV (flat per-label rows — for spreadsheet / pandas analysis), depending on the app workflow and account plan.