このエントリーをはてなブックマークに追加
ID 70079
FullText URL
fulltext.pdf 2.84 MB
Author
Haz, Amma Liesvarastranta Department of Information and Communication Systems, Okayama University
Brata, Komang Candra Department of Information and Communication Systems, Okayama University
Funabiki, Nobuo Department of Information and Communication Systems, Okayama University Kaken ID publons researchmap
Kyaw, Htoo Htoo Sandi Department of Information and Communication Systems, Okayama University
Fajrianti, Evianita Dewi Human Centric Multimedia Research Laboratory, Department of Informatic and Computer Engineering, Politeknik Elektronika Negeri Surabaya
Sukaridhoto, Sritrusta Human Centric Multimedia Research Laboratory, Department of Informatic and Computer Engineering, Politeknik Elektronika Negeri Surabaya
Abstract
With the rapid growth of online presentations, there has been an increasing need for efficient review of recorded materials. In typical presentations, speakers verbally elaborate on each slide, providing details not captured in the slides themselves. Automatically extracting and embedding these verbal explanations at their corresponding slide locations can greatly enhance the review process for audiences. This paper presents a Slide Annotation System that employs a robust hybrid two-stage detector to identify slide boundaries, extracts slide text through Optical Character Recognition (OCR), transcribes narration, and employs a multimodal Large Language Model (LLM) to generate concise, context-aware annotations that are added to their corresponding slide locations. For evaluations, the technical performance was validated on five recorded presentations, while the user experience was assessed by 37 participants. The results showed that the system achieved a macro-average 𝐹1 score of 0.879 (𝑆𝐷=0.024, 95% 𝐶𝐼[0.849,0.909]) for slide segmentation and 90.0% accuracy (95% 𝐶𝐼[74.4%,96.5%]) for annotation alignment. Subjective evaluations revealed high annotation validity and usefulness as rated by presenters, and a high System Usability Scale (SUS) score of 80.5 (𝑆𝐷=6.7, 95% 𝐶𝐼[78.3,82.7]). Qualitative feedback further confirmed that the system effectively streamlined the review process, enabling users to locate key information more efficiently than standard video playback. These findings demonstrate the strong potential of the proposed system as an effective automated annotation system.
Keywords
slide annotation
multimodal analysis
speech-to-text
LLM
SUS
Published Date
2026-02-01
Publication Title
Algorithms
Volume
volume19
Issue
issue2
Publisher
MDPI AG
Start Page
110
ISSN
1999-4893
Content Type
Journal Article
language
English
OAI-PMH Set
岡山大学
Copyright Holders
© 2026 by the authors.
File Version
publisher
DOI
Related Url
isVersionOf https://doi.org/10.3390/a19020110
License
https://creativecommons.org/licenses/by/4.0/
Citation
Haz, A.L.; Brata, K.C.; Funabiki, N.; Kyaw, H.H.S.; Fajrianti, E.D.; Sukaridhoto, S. A Slide Annotation System with Multimodal Analysis for Video Presentation Review. Algorithms 2026, 19, 110. https://doi.org/10.3390/a19020110