论文标题
视频字幕:对我们的位置的比较评论,哪个可能是路线
Video Captioning: a comparative review of where we are and which could be the route
论文作者
论文摘要
视频字幕是描述捕获其语义关系和含义的一系列图像的内容的过程。用单个图像处理此任务非常艰巨,更不用说视频(或图像序列)有多困难。视频字幕的应用的数量和相关性很大,主要是处理视频监视中的大量视频录制,或者帮助视力障碍的人,以提及一些视频。为了分析我们社区为解决视频字幕任务的努力以及更好的遵循路线所在的何处,该手稿对2016年至2021年期间的105篇论文进行了广泛的审查。结果,确定了最常用的数据集和指标。此外,使用的主要方法和最好的方法。我们根据几个性能指标来计算一组排名,根据其性能,可以获得视频字幕任务上最佳结果的最佳方法。最后,得出一些见解,即哪些可能是改善处理这一复杂任务的下一步或机会领域。
Video captioning is the process of describing the content of a sequence of images capturing its semantic relationships and meanings. Dealing with this task with a single image is arduous, not to mention how difficult it is for a video (or images sequence). The amount and relevance of the applications of video captioning are vast, mainly to deal with a significant amount of video recordings in video surveillance, or assisting people visually impaired, to mention a few. To analyze where the efforts of our community to solve the video captioning task are, as well as what route could be better to follow, this manuscript presents an extensive review of more than 105 papers for the period of 2016 to 2021. As a result, the most-used datasets and metrics are identified. Also, the main approaches used and the best ones. We compute a set of rankings based on several performance metrics to obtain, according to its performance, the best method with the best result on the video captioning task. Finally, some insights are concluded about which could be the next steps or opportunity areas to improve dealing with this complex task.