Human-Centered Interfaces for Describing Complex Video Edits with Natural Language and Example Images

Authors

  • Nguyen Thai Duong Department of Information Systems, Thu Dau Mot University, Đường 30 Tháng 4, Phường Phú Thọ, Thủ Dầu Một, Vietnam Author
  • Le Thi Bao Ngan Department of Software Engineering, Quang Nam University, Đường Trần Hưng Đạo 102, Phường Tân Thạnh, Tam Kỳ, Vietnam Author

Abstract

Digital video editing is increasingly shaped by machine learning models that can transform content through high-level specifications, yet most interfaces remain grounded in low-level timelines, keyframes, and parameter panels. There is growing interest in describing edits using natural language, sometimes complemented by example images that express desired style, composition, or local modifications. However, the design of human-centered interfaces that allow users to specify complex, multi-step video edits through such multimodal descriptions remains only partially understood. This paper examines how natural language and example images can be combined into interfaces that allow users to describe detailed spatiotemporal changes while maintaining control, predictability, and the ability to iteratively refine results. The paper introduces a problem formulation in which user intent is modeled as a structured edit plan over a video, develops a multimodal representation that connects language, images, and video content, and sketches algorithmic mappings from interface-level descriptions to concrete edit operations. The discussion emphasizes ambiguity management, edit transparency, and support for exploratory workflows. Potential evaluation strategies grounded in usability, expressivity, and edit quality are outlined, with a focus on how different interaction patterns affect user understanding of the underlying automated system. The paper aims to provide a technically oriented perspective on the joint design of models and interfaces for human-centered, language-driven video editing supported by example imagery, with attention to both algorithmic structure and human factors.

Downloads

Published

2025-09-04

How to Cite

(1)
Duong, N. T.; Ngan, L. T. B. Human-Centered Interfaces for Describing Complex Video Edits With Natural Language and Example Images. PSDMVE 2025, 15 (9), 1-19.