AI Image & Video Model News Roundup – October 10, 2026

The open-weights world kept moving this week. Below are the most notable open-source image and video model releases, fine-tunes, and ComfyUI tools from October 6 to October 10, 2026, each with links so you can try them yourself.

New image models

Qwen-Image-2.1-Turbo: official 8-step checkpoint

Alibaba’s Qwen team released Qwen-Image-2.1-Turbo, an accelerated checkpoint of Qwen-Image-2.1 that generates and edits images in just 8 denoising steps at CFG 1. It keeps the same 7B architecture and supports 2K output, transparent RGBA, and natural-language editing. The key change versus the community LoRA turbo builds: the 8-step schedule ships inside the official checkpoint, so no extra LoRA is needed. Note the license changed to the Qwen Research License, which requires a separate agreement for commercial use. Grab the weights and ComfyUI-ready repacks from the official Qwen-Image-2.1 repository; background reading at AI Weekly.

Kroma v0.3.1: distilled turbo for Krea 2

Lodestones shipped Kroma v0.3.1, an on-policy distilled (OPD) turbo checkpoint for Krea 2 that stays inside the base model’s distribution, with LoRA training support carried over. A practical upgrade path if you are already fine-tuning in the Krea 2 ecosystem.

New video models

Kandinsky 6.0 Video: MIT-licensed joint video and audio

Kandinsky Lab open-sourced Kandinsky 6.0 Video under the MIT license: a 29B Pro and a 3B Lite model that generate 5-second clips with synchronized 44 kHz audio, including lip-sync, in one pass. Text-to-audio-video and image-to-audio-video modes are supported, plus a separate super-resolution plugin that raises output to Full HD. Code, checkpoints, and diffusers integration are public, with day-one support in ComfyUI, FastVideo, SGLang, and vLLM-Omni. Start at the GitHub repository or read the hands-on breakdown.

Prism: 2K joint video-audio with 2.5x faster training

From Tencent Hunyuan with Fudan and Zhejiang University comes Prism, an MIT-licensed framework for natively training joint video-audio generation models at 2K. Its dynamic sparse attention scheme organizes tokens into spatiotemporal macro-zones, cutting training cost 2.5x versus full attention while matching or beating quality. Preview checkpoints render native 720p, 1080p, and 2K with structured audio tags such as music, sfx, and speech. Caveat: this is a research preview with steep hardware demands (720p needs one 80 GB GPU; 2K needs four or more). Explore the code on GitHub, the technical paper, and this summary.

ComfyUI workflows and tools

FastH3 Trim: MiniMax H3 video on 8 GB GPUs

The FastVideo team released FastH3 Trim, a pruned MiniMax H3 (42 transformer blocks instead of 50, rank-16 timestep conditioning) sampled in 8 steps. It is 4.2x smaller than base H3, and the int8 build runs synchronized video-with-audio on as little as 8 GB of VRAM. Repacked files for ComfyUI are ready to drop in: FastVideo/FastVideo-FastH3-Trim-Comfy on Hugging Face.

New LoRAs: person remover, storyboards, camera angles

Three community LoRAs worth a look this week, all covered on the ComfyUI Wiki news page: the H3 Person Remover (tracks a person with SAM 3.1, fills the mask, and rebuilds the background in overlapping windows), the Qwen-Image-2.1 Next-Scene LoRA (renders the next directed shot from a single frame and chains storyboards), and the Qwen-Image-2.1 Multiple-Angles LoRA (reshoots a subject from 72 azimuth/elevation framings).

Veda sparse attention: up to 3x faster H3

The Veda team shipped an official ComfyUI node that replaces MiniMax H3’s attention with a distilled sparse predictor, skipping 90% of attention tiles for up to 3x faster video generation. Pair it with FastH3 Trim if you are running on consumer hardware.

Bottom line

This week belonged to joint audio-video: Kandinsky 6.0 and Prism both push sound and picture in a single generative pass, one MIT-licensed and ready for ComfyUI, the other a research preview for heavy GPUs. On the efficiency side, FastH3 Trim and the Veda sparse-attention node keep pushing real video generation down to consumer cards. No new ComfyUI core release this week (still v0.38.0).

Further Reading

How to Use Qwen-Image 2.1 GGUF in ComfyUI
10Eros-Max: The New MiniMax H3 Model
https://www.kombitz.com/2026/06/26/how-to-use-krea-2-turbo-gguf-workflow-in-comfyui/

Be the first to comment

Leave a Reply