LLM Enhanced Video to Text Recognition for Authentic EFL Video Making Tasks: Evaluation Across GPT Versions and Video Lengths

Authors

  • Yi-Fan Liu Digital Affairs Department, Hsinchu County Government, Taiwan
  • Muhammad Irfan Luthfi Graduate Institute of Network Learning Technology, National Central University, Taiwan
  • Wu-Yuin Hwang Department of Computer Science and Information Engineering, National Dong Hwa University, Taiwan

Keywords:

Video to Text Recognition, large language models, EFL learning, video summarization, authentic video making tasks

Abstract

Authentic video making tasks are increasingly used in
English as Foreign
Language (EFL) learning because they support meaningful
communication and
provide rich evidence from speech, on screen text, and
visual context, but
teachers often face heavy workload when reviewing videos
and giving timely
support. This study proposes Large Language Model
(LLM)-enhanced Video to Text
Recognition (VTR), which transforms videos into text
outputs for learning
support. VTR generates sentence suggestions to support
learner narration and a
complete video summary to represent video content in clear
text. Because output
quality may vary by video length and model version, this
study conducted a
focused evaluation. Six English teachers from the language
department of a
public university participated. Using a web-based system
over two weeks,
participants evaluated 100 videos sampled from a public
dataset and grouped
into four duration categories with diverse scenes. They
rated the quality and
overall usefulness of sentence suggestions and video
summaries, judged whether
outputs were acceptable without edits, selected error tags,
and wrote reference
summaries to compare with system summaries. Results showed
clear differences
across Generative Pre-trained Transformer (GPT) versions
and video lengths.
Stronger GPT versions produced higher quality sentence
suggestions and
summaries and were more often acceptable without edits,
while longer videos
were more difficult and more likely to produce missing
information and overly
vague outputs. These findings suggest that VTR can support
EFL learning in
authentic video making tasks and highlight the need to
improve content coverage
and specificity, especially for longer videos, to increase
classroom
usefulness.

Downloads

Published

2026-06-03