Video In Sentences Out

Andrei Barbu, Alexander Bridge, Zachary Burchill, Dan Coroian, Sven Dickinson, Sanja Fidler, Aaron Michaux, Sam Mussman, Siddharth Narayanaswamy, Dhaval Salvi, Lara Schmidt, Jiangnan Shangguan, Jeffrey Mark Siskind, Jarrell Waggoner, Song Wang, Jinlian Wei, Yifan Yin, Zhiqi Zhang
Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence, PMLR R10:100-110, 2012.

Abstract

We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it. Action class is rendered as a verb, participant objects as noun phrases, properties of those objects as adjectival modifiers in those noun phrases, spatial relations between those participants as prepositional phrases, and characteristics of the event as prepositional-phrase adjuncts and adverbial modifiers. Extracting the information needed to render these linguistic entities requires an approach to event recognition that recovers object tracks, the trackto-role assignments, and changing body posture.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR10-barbu12a, title = {Video In Sentences Out}, author = {Barbu, Andrei and Bridge, Alexander and Burchill, Zachary and Coroian, Dan and Dickinson, Sven and Fidler, Sanja and Michaux, Aaron and Mussman, Sam and Narayanaswamy, Siddharth and Salvi, Dhaval and Schmidt, Lara and Shangguan, Jiangnan and Siskind, Jeffrey Mark and Waggoner, Jarrell and Wang, Song and Wei, Jinlian and Yin, Yifan and Zhang, Zhiqi}, booktitle = {Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence}, pages = {100--110}, year = {2012}, editor = {de Freitas, Nando and Murphy, Kevin}, volume = {R10}, series = {Proceedings of Machine Learning Research}, month = {14--18 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r10/main/assets/barbu12a/barbu12a.pdf}, url = {https://proceedings.mlr.press/r10/barbu12a.html}, abstract = {We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it. Action class is rendered as a verb, participant objects as noun phrases, properties of those objects as adjectival modifiers in those noun phrases, spatial relations between those participants as prepositional phrases, and characteristics of the event as prepositional-phrase adjuncts and adverbial modifiers. Extracting the information needed to render these linguistic entities requires an approach to event recognition that recovers object tracks, the trackto-role assignments, and changing body posture.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Video In Sentences Out %A Andrei Barbu %A Alexander Bridge %A Zachary Burchill %A Dan Coroian %A Sven Dickinson %A Sanja Fidler %A Aaron Michaux %A Sam Mussman %A Siddharth Narayanaswamy %A Dhaval Salvi %A Lara Schmidt %A Jiangnan Shangguan %A Jeffrey Mark Siskind %A Jarrell Waggoner %A Song Wang %A Jinlian Wei %A Yifan Yin %A Zhiqi Zhang %B Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2012 %E Nando de Freitas %E Kevin Murphy %F pmlr-vR10-barbu12a %I PMLR %P 100--110 %U https://proceedings.mlr.press/r10/barbu12a.html %V R10 %X We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it. Action class is rendered as a verb, participant objects as noun phrases, properties of those objects as adjectival modifiers in those noun phrases, spatial relations between those participants as prepositional phrases, and characteristics of the event as prepositional-phrase adjuncts and adverbial modifiers. Extracting the information needed to render these linguistic entities requires an approach to event recognition that recovers object tracks, the trackto-role assignments, and changing body posture. %Z Reissued by PMLR on 04 October 2026.
APA
Barbu, A., Bridge, A., Burchill, Z., Coroian, D., Dickinson, S., Fidler, S., Michaux, A., Mussman, S., Narayanaswamy, S., Salvi, D., Schmidt, L., Shangguan, J., Siskind, J.M., Waggoner, J., Wang, S., Wei, J., Yin, Y. & Zhang, Z.. (2012). Video In Sentences Out. Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R10:100-110 Available from https://proceedings.mlr.press/r10/barbu12a.html. Reissued by PMLR on 04 October 2026.

Related Material