大規模言語モデル (LLM) を組み込んだマルチモーダルな創造性支援ツール(CST)の開発と効果の検証

2026. 03. 25

本研究は、大規模言語モデル(LLM)を組み込んだスケッチベースの創造性支援ツール(CST)において、言語的思考と視覚的思考を同一キャンバス上で統合したときに、創造的活動にどのような影響が生じるかを検討することを目的とする。従来の LLM 活用 CST では、チャット UIと言語的思考、キャンバス UI と視覚的思考が分離されることが多く、モダリティ間の断絶や思考の流れの中断が課題となる。そこで本研究では、LLM との音声によるリアルタイム会話と、その会話ログをキャンバス上の要素としてドラッグ&ドロップ・編集できる機構を中心に、同一キャンバス上で言語と視覚が並行して展開できる CST「Composer」を設計・実装した。評価として、提案システムと対象条件の 2 条件による比較実験を行い、収集したデータから検討した。その結果、SUS の一部項目において有意差が確認され、UI 上の統合感の向上が示唆され、LLM が発散や実例提示を担い、人間が収束や最終判断を担う役割分担が自然に形成された。

This study aims to investigate the effects on creative activities when linguistic thinking and visual thinking are integrated on a single canvas in a sketch-based Creativity Support Tool (CST) incorporating a Large Language Model (LLM). In conventional LLM-based CSTs, chat UIs for linguistic thinking and canvas UIs for visual thinking are often separated, leading to challenges such as discontinuity between modalities and interruption of the flow of thought. To address these issues, this study designed and implemented “Composer,” a CST that enables parallel development of language and visuals on a single canvas, featuring real-time voice conversation with an LLM and a mechanism that allows conversation logs to be dragged, dropped, and edited as elements on the canvas. For evaluation, a comparative experiment with two conditions (proposed system vs. control condition) was conducted, and the collected data were analyzed. The results showed significant differences in some SUS items, suggesting improved sense of integration in the UI, and a natural division of labor emerged where the LLM handled divergence and example presentation while humans handled convergence and final decision-making.