| ID | 70079 |
| FullText URL | |
| Author |
Haz, Amma Liesvarastranta
Department of Information and Communication Systems, Okayama University
Brata, Komang Candra
Department of Information and Communication Systems, Okayama University
Funabiki, Nobuo
Department of Information and Communication Systems, Okayama University
Kaken ID
publons
researchmap
Kyaw, Htoo Htoo Sandi
Department of Information and Communication Systems, Okayama University
Fajrianti, Evianita Dewi
Human Centric Multimedia Research Laboratory, Department of Informatic and Computer Engineering, Politeknik Elektronika Negeri Surabaya
Sukaridhoto, Sritrusta
Human Centric Multimedia Research Laboratory, Department of Informatic and Computer Engineering, Politeknik Elektronika Negeri Surabaya
|
| Abstract | With the rapid growth of online presentations, there has been an increasing need for efficient review of recorded materials. In typical presentations, speakers verbally elaborate on each slide, providing details not captured in the slides themselves. Automatically extracting and embedding these verbal explanations at their corresponding slide locations can greatly enhance the review process for audiences. This paper presents a Slide Annotation System that employs a robust hybrid two-stage detector to identify slide boundaries, extracts slide text through Optical Character Recognition (OCR), transcribes narration, and employs a multimodal Large Language Model (LLM) to generate concise, context-aware annotations that are added to their corresponding slide locations. For evaluations, the technical performance was validated on five recorded presentations, while the user experience was assessed by 37 participants. The results showed that the system achieved a macro-average 𝐹1 score of 0.879 (𝑆𝐷=0.024, 95% 𝐶𝐼[0.849,0.909]) for slide segmentation and 90.0% accuracy (95% 𝐶𝐼[74.4%,96.5%]) for annotation alignment. Subjective evaluations revealed high annotation validity and usefulness as rated by presenters, and a high System Usability Scale (SUS) score of 80.5 (𝑆𝐷=6.7, 95% 𝐶𝐼[78.3,82.7]). Qualitative feedback further confirmed that the system effectively streamlined the review process, enabling users to locate key information more efficiently than standard video playback. These findings demonstrate the strong potential of the proposed system as an effective automated annotation system.
|
| Keywords | slide annotation
multimodal analysis
speech-to-text
LLM
SUS
|
| Published Date | 2026-02-01
|
| Publication Title |
Algorithms
|
| Volume | volume19
|
| Issue | issue2
|
| Publisher | MDPI AG
|
| Start Page | 110
|
| ISSN | 1999-4893
|
| Content Type |
Journal Article
|
| language |
English
|
| OAI-PMH Set |
岡山大学
|
| Copyright Holders | © 2026 by the authors.
|
| File Version | publisher
|
| DOI | |
| Related Url | isVersionOf https://doi.org/10.3390/a19020110
|
| License | https://creativecommons.org/licenses/by/4.0/
|
| Citation | Haz, A.L.; Brata, K.C.; Funabiki, N.; Kyaw, H.H.S.; Fajrianti, E.D.; Sukaridhoto, S. A Slide Annotation System with Multimodal Analysis for Video Presentation Review. Algorithms 2026, 19, 110. https://doi.org/10.3390/a19020110
|