Multi Modal Input

Feed in text, voice, or audio clips to guide both visuals and sound in your video, combining ComfyUI nodes into one unified result.

Rich Guidance

Use text or voice or both to describe your creative vision. Provide nuanced direction for both visuals and audio.
Try Promptus for free ➜
promptus ai local app
promptus ai local app

Hands‑Free Option

Quick answer: Multimodal input means a generative AI system can accept more than just typed text — for example, spoken voice, an uploaded image, or a reference video — as the basis for generation. In Promptus, Multi Modal Input lets you create videos by speaking your idea instead of typing it out.

Create videos by speaking, without typing a word.

Multi Modal
Input FAQs

Text, voice narration, or short audio clips — used together to guide both the visuals and sound of a video generation.
No — you can create videos by speaking alone, without typing a word.
Yes — using text and voice together provides more nuanced direction for both visuals and audio.
Multi Modal AI Generation

Multi Modal Input adapts to your workflow, letting you speak or write as you prefer. It delivers richer, more precise video outputs by combining inputs.

Start using Promptus
Start running your first workflow
Go from idea to production-ready output in minutes.
Try Promptus for free ➜