Generate the source video first, then work on the subtitles.
Use local ComfyUI and MiniMax H3 to create a short spoken segment, check the audio and the person, then hand it to the spoken segment workstation to arrange subtitles, adjust the background, and export the final video. Here we provide the method and files, not a button that generates when you wait.
First check if it can run, then decide what to download.
This is not a lightweight browser model. The 35 minutes is reading and practice planning time, not including model download, environment setup, or generation queue time.
Already have a local environment
You need ComfyUI that supports H3 nodes, matching model components, a usable GPU, and sufficient disk space. The package lists the file names and parameters used in this test; check components first, it won't automatically modify your current environment.
Read model license first
The original tutorials and validation tools in this package are MIT, but that does not mean the H3 model or outputs receive the same license. The model has regional and commercial terms; check the current license before use; this page does not host the model and does not complete licensing for you.
Use an existing MP4 you have the right to use, import it into the workstation to practice subtitles, positioning, card backgrounds, and final export. You don't need to download a large model just to practice editing.
Generating the source video and post-production layout are handled separately. Subtitles are not drawn into the video by the model; when text needs changes, you don't have to regenerate the entire person segment.
01 · Generate source video with H3
First write a character, a scene, an action, and a spoken line. Use a short clip to verify identity, composition, action, and original audio; the output does not have additional burned-in subtitles.
Load the source video and SRT, drag or one-click center, customize card background color, transparency, corner radius, and text color. Save the project for later edits.
Keep or close the original audio, choose whether to add a watermark, export MP4 or WebM supported by the browser. Play back the start and end, check subtitles against audio, then decide whether to publish.
Start with a short clip, don't go straight to a long video.
The download package includes Skill, workflow API JSON, offline validation methods, prompts, shot plan, SRT, and workstation project sample. If nodes are missing or validation errors occur, stop and handle them, don't resubmit repeatedly.
Describe the image and audio clearly
Example: fictional adult lecturer, bust shot, full head, fixed camera, soft office lighting, one explanatory gesture, one line in Mandarin. Who the person is, how the shot is taken, and what is said should be described separately; don't pack multiple locations and actions into one short clip.
Validate offline, then you submit
First read the README, check model components, nodes, and output location. The validator in the package only reads selected files and generates request files; it does not connect to the internet, queue, or download models. Actual generation is initiated by you in your authorized local ComfyUI environment.
Watch the video, also listen to the full audio
Confirm the person is stable, the head is not cropped, gestures are natural, and the ending is not cut off. Generating an audio track does not prove the speech is accurate; listen sentence by sentence and correct the SRT, don't treat pre-written subtitles as calibrated transcription.
Make editable post-production in the workbench
First select the original footage, then import the SRT. Use horizontal centering to restore card horizontal position, choose background and font color; preview and export share layout. After saving the project, export the finished video, then reopen the downloaded file to check.
We hand over the parameters we ran, along with the boundaries.
Local case dated September 5, 2026: ComfyUI + H3 FL2VA, 8-step Turbo, 576 × 1024, 124 frames, 24 fps, single generation about 161 seconds. This is a machine measurement, not a speed promise for all devices.
AI-generated · actual workbench export
Local H3 original footage × custom subtitle background
This is an actual ~5-second finished video exported from the workbench, not a concept animation. The fictional instructor and voice are AI-generated; the yellow subtitle background, centered text, and no extra watermark come from workbench settings. Original footage is 576 × 1024, export canvas is 1080 × 1920; upscaling does not add native detail. Subtitles are an editing exercise; verify against the actual audio before using in formal projects.
Project file does not include video. You can open the project first, then select your own original footage; the tutorial learning pack does not include model weights.
Not H5, nor a paid cloud API. The workflow retains the audio decoding branch, and generation output passes full audio/video decoding checks. Whether the actual audio matches word-for-word and lip sync looks natural still requires human listening and viewing.
Don't confuse resolution
The generated original footage this time is 576 × 1024. The workbench can export the canvas as 1080 × 1920, but upscaling does not add native detail. Higher generation specs require re-validating VRAM, runtime, and results.
Download package contains no video or models
Provided are original methods, parameters, and handover samples; no weights, third-party programs, or generated assets are distributed. There is also no auto-install, hidden upload, or remote execution entry; fixed-package review does not guarantee absolute security in all environments.
If you run into issues, check these places first.
Prioritize narrowing down the problem scope; don't blindly add steps, switch models, or reinstall the entire environment.
Why can validation pass but generation still fail?
Offline validation only checks file structure and known parameters; it does not prove the model is installed, VRAM is sufficient, or all node versions are consistent. First verify components on this machine, then run a short clip; save error logs before handling.
Why is the spoken content not exactly the same as the pre-written subtitles?
Generative speech may change words, drop words, or vary in pacing. The sample SRT is an editing script, not a forced alignment result. First listen to the original audio, then fix subtitles; if dialogue requirements are strict, produce and calibrate voiceover and lip sync separately.
Why does the exported file look different?
First confirm that you are loading the same video and project, and check the watermark, background, frame size, and export range. If the completion area prompt differs from the current edit, you need to re-export. Real-time browser encoding may cause slight duration or frame rate differences and cannot replace frame-by-frame offline editing.
DOWNLOAD & REVIEW
Take away process files, keep the habit of checking.
Read the license and boundaries before starting. Review corresponds to the fixed version and hash below; no model hosting or commercial licensing agency is provided.
Not sure where to start? Copy a learning request.
You can download this reviewed version's resource package. If it includes SKILL.md, provide it and the related documentation to your AI tool; first read and check the materials and usage boundaries, and do not run anything automatically.
Prepare assets: start with fictional or de-identified examples.
Read the boundaries: confirm dependencies, inputs, outputs, and human checkpoints.
Review outputs: keep source and failure records before deciding to pilot.
You can also select the text to copy directly; read it first before executing.
This version has been reviewed.1.0.0
Based on one real local H3 generation and voiceover workbench export, this is an original teaching pack. Checked fixed file list, sanitized workflow, offline verification, and subtitle handover; contains no model weights, generated videos, or auto-installers. The sample SRT is an editing script pending audio calibration, not accurate transcription.
Maintainer
FORMWEFT
License
MIT (original tutorials and tools); models are separately licensed
Source type
FORMWEFT original
Network permission
On-demand online
Local programs
Offline verification script in the package; generation requires separate ComfyUI, Python, GPU, and licensed models
Review date
2026/09/05
Review scope
Fixed version
1.0.0
File count
13
Implementation method
Original teaching pack and offline verification
Network interfaces
Verifier does not connect to the internet; generation is submitted separately by the user