Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

University of Central Florida

Video Content Understanding Using Text

University of Central Florida via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This advanced lecture summarizes research on using text to understand video, covering attention-based video representation and query testing, detection and correction of inaccurate content, and text-conditioned frame generation with GANs. It reports experiments on A2D and robotic data and discusses limitations and future work.

Syllabus

Intro
Motivation
Challenges
Algorithm
Training
Video Representation
Scoring function
Optimization - Updating Rules
Exemplar queries
Test on Unseen Queries
Qualitative results
Sentence Encoder
Spatial Attention Network • Which regions of the frames to look?
Temporal Attention Model
Inference Module
Experiments
Limitations
What is an Inaccuracy?
Formulation
Detection By Reconstruction
Visual Features
Inaccuracy Detection
Correction
Last two chapters
How about the opposite problem?
Problem Definition
Proposed Approach - Generator Block Diagram
Text Encoding
Start and End Distributions
Latent Path Construction
Conditional BatchNormalization (CBN)
Frame Generation
UpPooling Block Details
Proposed Approach - Discriminator
Loss Function - Generator
Hinge GAN-Loss on Discriminator
Evaluation Metrics
A2D Quantitative Results
A2D Results
Robotic Results
Dissertation Summary
Future Work

Taught by

UCF CRCV

Reviews

Start your review of Video Content Understanding Using Text

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.