Abstract
This paper describes the TUM approaches for violent scenes detection in movies, submitted for the MediaEval 2012 Affect Challenge. Score fusion is used to fuse Support-Vector Machine (SVM) confidence scores assigned to short fixed length windows within each movie shot. SVM predictors for acoustic and visual channels are trained. For the acoustic channel, a large set of acoustic features based on the set from the INTERSPEECH 2012 Speaker Trait Challenge is employed. A comprehensive set of common video low-level descriptors such as optical flow, gradients, and hue and saturation histograms is used for the visual channel.
| Original language | English |
|---|---|
| Journal | CEUR Workshop Proceedings |
| Volume | 927 |
| State | Published - 2012 |
| Event | Multimedia Benchmark Workshop, MediaEval 2012 - Pisa, Italy Duration: 4 Oct 2012 → 5 Oct 2012 |
Fingerprint
Dive into the research topics of 'Violent scenes detection with large, brute-forced acoustic and visual feature sets'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver