CapSmith: Contextual Refinement of Additive Text Captions for Video Storytelling

Saehui Hwang, Jean-Peïc Chou, Anh Truong

2026ACM Creativity & Cognition (C&C '26)DOI 10.1145/3803784.3807559

Abstract

Additive Text Captions (ATC) are powerful storytelling devices that extend video narratives by progressing story (“meanwhile...”), focusing attention (“HUGE!”), and providing unspoken context (“crashing”). ATC authoring is difficult because it requires balancing creative decisions (word choice, tone, etc.) against practical constraints (timing, readability, visual harmony). Current video-editing interfaces treat text as timeline-bound assets, prioritizing temporal manipulation over narrative function. While LLMs can support creative writing, our formative study (n=8) showed creators resisted LLMs that write ATCs from scratch, citing control and authenticity concerns, but wanted help refining drafts. We designed CapSmith, an ATC- authoring system that features (1) a script panel that externalizes narrative structure and (2) a variation-from-example workflow that proposes variants grounded in inferred intent and context. A comparative study (n=10) showed that the script panel contextualized ATCs against the narrative and that variation-from-example workflow supported refinement without requiring creators to articulate tacit intent while preserving authorship.

Full text

This browser can’t show PDFs inline.

Open the PDF
CapSmith: Contextual Refinement of Additive Text Captions for Video Storytelling Open in a new tab