Civitai中国镜像,为方便没有魔法的炼丹师
Sign In

RedCraft | 红潮 | Hybrid H3 & Krea2DUAL MASO ver 双权重

Download

5 variants available

Type
Checkpoint Merge
Stats

38,576

1

Published

Dec 2, 2025

Base Model

ZImageTurbo

Hash
AutoV2
7DE772460E
default creator card background decoration
Reactions - 197859

197.9K

Followers - 18054

18.1K

Downloads - 894000

894K

".red" Badge

License:

Apache 2.0
REDZ15_00122_.png

RedCraft-红潮 RedFused

做好工具人 服务艺术家
Forever in memory of METAFILM Studio founder Mr. Yuan Bo

RedCraft Krea2 DUAL · MASO Edition

RedCraft-LowNoise-MASO-MangaDrama-Krea2Turbo-RedMix-v4
RedCraft-HighNoise-MASO-MangaDrama-Krea2Turbo-RedMix-v4

这是为 RedCraft / Dark Beast / REDGraft H3 配套制作的 Krea2 DUAL 双模型混合降噪版本
主要用于生成日本漫画风格的漫剧素材图。

DUAL 模式分别使用 HighNoise + LowNoise 模型接力采样,在高噪声阶段重新释放 Krea2 TURBO 被压制的随机多样性
因此它不是一个追求“每次都一样听话”的 fine-tune ——抽卡、重抽、再抽,才是正确打开方式。


WF of RedCraft Krea2 DUAL
https://civitai.red/api/download/models/3311198?fileId=3196644

严格指令遵循不是它的强项,随机性才是。
This is not a strict prompt-following model.

A Krea2 DUAL hybrid denoising model designed for RedCraft / Dark Beast / REDGraft H3, specifically tuned for Japanese manga-style visual assets and manga-drama production.

DUAL uses separate HighNoise + LowNoise models throughout the sampling process, bringing back more of Krea2's high-noise stochastic diversity.

It works best with IL / Pony / SD15-style Danbooru Tags, where randomness and prompt interpretation are part of the workflow.

The trade-off is simple:

More randomness → More cherry picking → More broken anatomy.

随机性 ↑ → 抽卡难度 ↑ → 肢体崩坏 ↑

Yes, sometimes it feels like A1111 from three years ago.

But when the pull hits, the model can produce high-dynamic poses, unusual compositions and motion combinations that conventional models rarely reach.

它尤其适合使用 IL / Pony / SD15 风格的 Danbooru Tags 进行提示。

严格指令遵循不是它的强项,随机性才是。

某些时候,你会感觉自己又回到了三年前的 A1111

但当它抽中时,能够得到一些常规模型很难产生的高动态、非常规构图与动作组合
——这也是这个版本存在的意义。

So this is, officially:

MASO Edition — 受虐狂最爱 A Masochist's Favorite

不是最听话的模型,但可能是最值得多抽几次的那个。


该模型 Paid Access 包含以下所有模型(及后续更新服务

RedCraft | 红潮 | REDMIX Hybrid A2A beta2 + LTX25 2k 16-bit HDR- 08/25/2026
beta2 版本融合了 LTX2.5 distilled 作为2k高清化方案,并且消除了音频bug。
The beta2 version integrates LTX2.5 distilled as a 2K HD solution and has eliminated audio bugs.
Any questions Contact me via Telegram @makemoviexxx


Dark Beast | 黑兽 H3 Director Edition 自动短剧生产线 08/28/2026

是原生H3单次采样直出2k分辨率的高清方案( 768P/VSR/RIFE TensorRT

所有示例短片均为 6-10 步 黑兽 H3 modified 版本单次采样生成
https://civitai.red/models/2242173/dark-beast-or-h3-director-edition

REDGraft | LTX 2.5 | 老同学 Fast 2K 16bit-HD| sulphur2 ported 移植版

https://civitai.red/models/1295569/redgraft-ltx-25-fast-2k-or-sulphur2-ported

LTX 16 bit HDR & Nvidia RTX VSR

The LogC3 transform and inverse transform are used prior and after generation (accordingly) in order to achieve the dynamic range of 16 bit within normal number range.

RIFE TensorRT Frame interpolation

工作流预览: https://civitai.red/models/579280


付费用户提供一对一配置服务

One-on-one configuration services are provided to paid users.

本地 A5000 24G 显存设备,输出2k 8秒视频≈240秒

商用 5090D 24G 显存设备,输出2k 8秒视频≈150秒

红潮/黑兽/LTX老同学 系列模型即将登录 Fal.ai 平台

届时可以体验 B200 飞速生产的效果(预计40-60秒)

红潮 / 黑兽 模型在线体验预览 https://asop.uk | https://app.asop.uk

The "RedCraft" "Black Beast" and "REDGraft LTX" model series are coming soon to the Fal.ai .
Will be the lightning-fast generation speeds of the B200 (estimated at 40–60 seconds).
RedCraft / Dark Beast – Online Model Preview https://asop.uk | https://app.asop.uk


Online Buzz Payment & Model Access

  • Online Buzz payment is supported, granting direct access to the models.

  • Payments in fiat currency or cryptocurrency are also accepted.

  • Regardless of the payment method chosen, you will gain access to all of my models.

  • If the service configuration requires the use of third-party models, I will pay the corresponding fees to the contributors.

Contact me via Telegram @makemoviexxx

let me know your requirements and configuration needs.


支持线上Buzz支付,同时支持等价法币或加密货币支付

Motion-context 多参导演台商用版本请拍下后私信 tg @makemoviexxx

Thank you! 🙏

to the Hong Kong team, and thank you to our friends from US

新的支持者!感谢来自杭州的团队


这套方案的核心不是「一个模型一次出 2K」,而是 H3 负责低分辨率语义草稿,LTX-2.5 负责 IC 引导的画幅放大与16bit HDR输出。先把内容做对,再把分辨率做高,所以能在 16GB 级显存上相对稳定高速地出 2K 视频。

为什么要双模型

MiniMax H3 擅长「生成」
低分辨率图生视频,同时出画面和音频。提示词、镜头、人物和声音都在这一段定下来。步数少(约 6 步),分辨率低,所以草稿快、语义稳。

LTX-2.5 擅长「增强」
不重新编故事,而是把 H3 成片当 IC-LoRA 的视频引导,做 V2V。蒸馏 8 步 + 潜空间 ×2 放大 3 步,再 RIFE 插帧到 48fps、VSR 拉到约 1440×2160。画质、边缘和 2K 细节主要靠这一段。

两者互补:H3 解决「像什么、怎么动、配什么声」;LTX 解决「够不够清、够不够 2K」。全程按原视频约束,比单模型直接冲 2K 更稳,也不容易跑偏。

效率从哪来

  • 先低后高:贵的 22B 级计算主要花在中分辨率 V2V,而不是从噪声直接出 2K。

  • 蒸馏 + 手动 sigma:LTX 侧大约 8+3 步,比全量采样短很多;自定义 sigma 用来压住细节,而不是靠加步数硬堆。

  • 模型体积可控:H3 用 int8 剪枝包,LTX 用 22B 裁切版,再加 IC spatial upscaler,而不是上完整大包。

  • 音频复用:声音在 H3 阶段生成,LTX 只做参考和同步,不必再跑一遍文生音频。

结果是:草稿快、放大可控、最终 2K 输出相对稳定

显存怎么看

显存预期 16GB 及以上

按当前参数可跑:H3 草稿 + LTX IC V2V + 2K 输出。这是这套方案的设计档位。

低于 16GB

能跑,但要接受更慢(更多卸载、tile、CPU 换算),或下调参数:分辨率、时长、IC 放大、tile、插帧档位。

低于 16GB 时,优先砍的是 LTX 分辨率和潜空间 ×2,而不是先砍 H3 的语义草稿——草稿错了,后面放大只会把错误放大。

一句话

H3 用低成本把「内容 + 声音」做稳,LTX-2.5 用 IC V2V 把「清晰度 + 2K」做上去。分工明确,所以比单模型直出 2K 更省、更稳;16GB 以上是舒适区,再往下就要用时间和画质换显存。


RedCraft | 红潮 | REDMIX Hybrid A2A beta1 众神归位 Lightning 8 - 08/13/2026

众神归位 Fusion of the Gods

Release Statement: MiniMax H3 RED-A2A Beta1

MiniMax H3 RED-A2A Beta1 is not an SFT-trained model. This release is designed to enable multi-format input and multimodal output generation within a single model and a single workflow.

This build integrates the outstanding work and research of many open-source masters in the community, including:

Turbo 1.0 FL2V accelerator by LightX2V ( lightx2v/Minimax-h3-Turbo ),
Ref2V weight lora separation by Kijai ( Kijai/MiniMax-H3-experimental ),
FL2V NSFW tuning by sexgod1979 ( SexGod's NaughtyTimes MM H3 Lora ),
FL2V aesthetic tuning by TenStrip ( TenStrip/10Eros-Max )

and the ComfyUI-MiniMax-ContextIR A2A nodes pack by Astral 星芒, as well as the selfless technical knowledge shared by numerous creators. (Special thanks to Aiwood, AIEveryThing, wuwukasi, and the many H3 promotion upper.)

The H3 A2A Beta1 release achieves "three-way connectivity" across first-frame, last-frame, and context reference inputs, along with "one-pass" multimodal generation. We will continue refining this direction in future versions, building MCP infrastructure for AGI content production — in conjunction with the Comfy MCP release — to deliver better automated content production services and accumulate more real-world experience along the way.

HHH / Hidden Spark Advisory Group / Metafilm / RedCraft

发布声明:MiniMax H3 RED-A2A Beta1

MiniMax H3 RED-A2A Beta1 并非 SFT(监督微调)训练模型。本版本的发布目的,是在单一模型与单一工作流中实现多参数格式的输入与多模态输出物的生成

本版本整合了开源社区众多"大神"们的劳动与研究成果,包括:来自 Kijai(Kijai/MiniMax-H3-experimental)的 Ref2V 权重分离、来自 LightX2V(lightx2v/Minimax-h3-Turbo)的 Turbo 加速器 1.0 FL2V 版本、来自 sexgod1979( SexGod's NaughtyTimes MM H3 Lora )的 NSFW 增强Lora(推荐尝试 PinkCherry MM H3 版本)、来自 TenStrip(推荐尝试 TenStrip/10Eros-Max )的 FL2V 版本美学调优,以及来自 Astral 星芒 的 ComfyUI-MiniMax-ContextIR A2A 节点包,以及众多博主无私的技术与知识分享。(特别感谢:AiwoodAIEveryThingwuwukasi 以及众多 H3 宣发账号和平台。)

H3 A2A Beta1 版本基本实现了首帧、尾帧、上下文参考的"三通",以及多模态生成的"一达"。我们将在后续版本中持续优化这一方向,为 AGI 内容生产提供 MCP 基础设施——结合 ComfyMCP 的发布——实现更好的内容生产自动化服务,并在过程中积累更多实战应用经验。

—— HHH / 隐世 Spark 顾问团 / 元影·Metafilm / RedCraft·红潮制作组


"Unlocked" versions of text-encoding models have not been shown to offer improved NSFW capabilities, and their impact on DiT models remains unclear due to the lack of alignment training. However, they can serve as upstream models for prompt enhancement and are highly beneficial for prompt optimization.

解锁版本的文本编码模型未被证明有NSFW能力的提升,并且由于缺乏对齐训练,对DIT模型的作用尚不明确。但是可以作为提示词增强前置模型,对提示词优化有很大帮助:

NVFP4 (4-bit) Qwen3-VL-32B text encoder Abliterated
HF Repo ( 6block/MiniMax-H3-Qwen3-VL-Abliterated-NVFP4 )

Qwen3-VL-32B INT8 TensorWise + ConvRot
HF Repo ( linjian257/qwen3vl_32b_minimax_h3_int8_convrot_uncensored-by-linjian257 )

Qwen3-VL-32B Heretic (MiniMax-H3 text encoder) — NVFP4
HF Repo ( sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4 )


Anything2Anything Application

A universal content generation solution powered by the multimodal generative model and driven by the Comfy MCP protocol, delivering a creation paradigm of "anything in, anything out."

The core positioning is a general-purpose multimodal generation engine. It accepts diverse input formats — whether plain text descriptions, structured JSON, Markdown documents, PPT presentations, DOCX reports, reference images, video clips, or audio assets — and after unified understanding, generates high-quality images, videos, audio, or digital human content. The entire pipeline is exposed as a standardized interface through Comfy MCP (Model Context Protocol)​, enabling any AI Agent platform to drive it without directly operating the underlying model or workflow editor.

Once connected via an MCP client, any AI Agent can invoke these tools like ordinary function calls. For example, in Claude or a custom Agent, the system automatically completes the full chain: document parsing → content extraction → Context-IR instruction generation → H3-Base inference → 2K upscaling → output delivery.

Typical Use Cases

Scenario 1: Document-to-Video Briefing The user uploads a PPT or DOCX quarterly report. The system automatically extracts key data and charts, generating a 15-second 2K video briefing with narration and background music, suitable for meeting presentations or social media distribution.

Scenario 2: Reference Video + Audio → Digital Human The user provides a 10-second real-person speech video and a new audio script. The system leverages the H3-Base Ref2VA checkpoint to generate a lip-synced, motion-consistent digital human video, applicable to multilingual content localization, virtual anchors, and similar use cases.

Scenario 3: Text + JSON Data → Brand Advertisement The user inputs ad copy and structured product parameters (color, dimensions, scene). The system generates a 2K brand-aligned advertising video with precise rendering of text and brand information.

Scenario 4: AI Agent Self-Service Orchestration An enterprise Agent receives a user query, automatically calls analyze_media to process the uploaded assets, and then invokes text_to_video to generate an explanatory video — all without any human intervention.

Core Competitive Advantage

The central differentiator lies in encapsulating powerful multimodal understanding and generation capabilities into standardized Agent tools via the Comfy MCP protocol. This means any AI application — whether a front-end conversational assistant, an enterprise automation pipeline, or a creative workstation — can integrate "multimodal content generation" as a fundamental on-demand capability. Input formats are unrestricted, output forms are not fixed, and orchestration is programmable — truly realizing the creative freedom of "Anything2Anything."

License

  • MiniMax H3 is released under the MiniMax H3 Community License Agreement.

  • MiniMax H3 依据 MiniMax H3 社区许可协议授权,版权所有 © 2026 MiniMax。

  • This model is a fine-tuned version of MiniMax H3, with modifications to the original weights."本模型依据 MiniMax H3 社区许可协议授权"


Anything2Anything 应用方案

基于全模态生成模型,通过 Comfy MCP 协议驱动的通用内容生成方案,实现"任意输入、任意输出"的创作范式。


方案定位

核心定位是一个通用多模态生成引擎。它接收用户提交的多种格式内容——无论是纯文本描述、结构化 JSON、Markdown 文档、PPT 演示文稿、DOCX 报告、参考图像、视频片段还是音频素材——经统一理解后,生成高质量的图像、视频、音频或数字人内容。整个链路通过 Comfy MCP(Model Context Protocol)​ 暴露为标准化接口,任何 AI Agent 平台均可驱动,无需直接操作底层模型或工作流编辑器。

任意 AI Agent 通过 MCP 客户端连接后,即可像调用普通函数一样调用这些工具。例如,在 Claude 或自定义 Agent 中,系统会自动完成文档解析 → 内容提取 → Context-IR 指令生成 → H3-Base 推理 → 2K 超分 → 输出回传的完整链路。


典型应用场景

场景一:文档转视频简报

用户上传一份 PPT 或 DOCX 季度报告,系统自动提取关键数据和图表,生成一段带旁白和背景音乐的 15 秒 2K 视频简报,适用于会议展示或社交媒体传播。

场景二:参考视频 + 音频 → 数字人

用户提供一段 10 秒的真人讲话视频和一段新的音频脚本,系统通过 H3-Base 的 Ref2VA 检查点,生成口型同步、动作一致的数字人视频,适用于多语言内容本地化、虚拟主播等场景。

场景三:文本 + JSON 数据 → 品牌广告

用户输入广告文案和结构化产品参数(颜色、尺寸、场景),系统生成符合品牌调性的 2K 广告视频,支持文字和品牌信息的精准呈现。

场景四:AI Agent 自助编排

企业内部 Agent 接收到用户查询后,自动调用 analyze_media 分析上传的素材,然后调用 text_to_video 生成解释性视频,全程无需人工介入。


核心竞争力

将强大的全模态理解与生成能力,通过 Comfy MCP 协议封装为标准化的 Agent 工具。这意味着任何 AI 应用——无论是最前端的对话式助手、企业级自动化流程,还是创意工作台——都可以将"多模态内容生成"作为一项基础能力按需集成。输入格式不受限,输出形态不固定,编排方式可编程,真正实现 ​"Anything2Anything"​ 的创作自由。

License

  • MiniMax H3 is released under the MiniMax H3 Community License Agreement.

  • MiniMax H3 依据 MiniMax H3 社区许可协议授权,版权所有 © 2026 MiniMax。

  • This model is a fine-tuned version of MiniMax H3, with modifications to the original weights."本模型依据 MiniMax H3 社区许可协议授权"


No Mosaics. 无码 2倍速
KREA 2 赤佬 Bastard3 Edition
REDZ 2 红潮 造相2 HDEdition

INT8/INT4 Convrot for ComfyUI 0.27 [Native 原生节点支持] Uploaded
File name: Krea2RedMix1.1-INT8-Convrot-ComfyUI (original native)
File name: Krea2RedMix2.1-INT8-Convrot-ComfyUI (no mosaics)
File name: Krea2RedMix3.1-INT8/INT4-Convrot-ComfyUI (visual effects)
File name: REDZimageTurbo2.0-INT8-Convrot-ComfyUI (ZIT-HD 2026)
File name: Krea2RedMix1.2-INT4-Convrot-ComfyUI (by Wikee Yang)

Hardcore version:

KREA 2 赤佬 黑兽3.0 Dark Beast 3 Krea2 Edition (Enhance Anatomy)

KREA2 GPT 逼真版 Grand PUSSY Truth | MIST2


REDZ-imageTurbo2 造相2.0 HD 高清版 05/07/2026 8 Steps

基于INT8 Convrot规格重制的商用高清图像处理引擎

作为ZimageTurbo模型家族的最新成员,这一版本红潮(byMetaFilm元影制作组)专为商业级高清图像处理需求而设计,选用海量商广正版4k/8K高清训练集,在性能、效率和输出质量方面均实现了突破性提升。

核心技术突破:完全基于INT8 Convrot量化重制

完全基于INT8 Convrot(卷积旋转)规格的架构重制。这一革命性的设计将传统的浮点运算转换为8位整数运算,在保持图像处理精度的同时,大幅降低了计算复杂度和内存占用。INT8 Convrot 技术通过优化卷积核的旋转机制,实现了更高效的张量运算路径。

性能飞跃:处理速度提升至2倍

得益于 INT8 Convrot 规格的深度优化,REDZ HD 2026 在处理速度上实现了质的飞跃。相比前代版本,新版本在相同硬件配置下能够将处理速度提升至2倍,这意味着商业用户现在可以在更短的时间内完成大量高清图像的处理任务,显著提升了工作流程效率。无论是批量图像增强、实时滤镜应用还是复杂特效渲染,都能提供前所未有的响应速度。

素材质量:精选商用高清素材库

其内置的精选商用高清素材库。这一素材库经过专业团队精心筛选和优化,包含了数十万种高质量的商业级图像资源,涵盖自然风光、城市景观、人物肖像、产品展示等多个类别。所有素材均经过严格的版权审查和质量控制,确保用户能够安全、合法地将其用于商业项目。素材库的智能匹配算法还能根据用户需求推荐最合适的图像资源,进一步提升创作效率。

应用场景广泛

红潮 ZimageTurbo HD 2026基于开放式商用协议发布,适用于多种商业应用场景:

  • 广告设计与营销材料制作

  • 电子商务产品图像优化

  • 社交媒体内容创作

  • 影视后期制作与特效处理

  • 游戏开发中的纹理生成与优化

  • 虚拟现实与增强现实内容创作

技术优势总结

红潮 2.0 HD 2026不仅是一次技术升级,更是图像处理工作流的重新定义。其INT8 Convrot架构带来的2倍速度提升,结合精选商用高清素材库,为专业用户提供了前所未有的创作自由度和效率保障。这一版本标志着ZimageTurbo模型在商业化应用道路上迈出了坚实的一步,为数字内容创作行业树立了新的技术标杆。


ZimageTurbo HD 2026: Commercial-Grade High-Definition Image Processing Engine with INT8 Convrot Architecture

ZimageTurbo HD 2026 represents a significant leap forward in image processing technology. As the latest member of the ZimageTurbo model family, this version is specifically designed for commercial-grade high-definition image processing needs, achieving breakthrough improvements in performance, efficiency, and output quality.

Core Technological Breakthrough: Complete INT8 Convrot Architecture Redesign

The core innovation of ZimageTurbo HD 2026 lies in its complete architectural redesign based on INT8 Convrot (Convolution Rotation) specifications. This revolutionary design transforms traditional floating-point operations into 8-bit integer operations, significantly reducing computational complexity and memory usage while maintaining image processing precision. The INT8 Convrot technology optimizes the rotation mechanism of convolution kernels, enabling more efficient tensor operation pathways and laying a solid foundation for real-time high-definition image processing.

Performance Leap: 2x Processing Speed Improvement

Thanks to the deep optimization of INT8 Convrot specifications, ZimageTurbo HD 2026 achieves a qualitative leap in processing speed. Compared to previous versions, the new version can increase processing speed by 2x under the same hardware configuration. This means commercial users can now complete large-scale high-definition image processing tasks in significantly less time, dramatically improving workflow efficiency. Whether for batch image enhancement, real-time filter application, or complex special effects rendering, ZimageTurbo HD 2026 delivers unprecedented response speeds.

Material Quality: Curated Commercial High-Definition Asset Library

Another standout feature of ZimageTurbo HD 2026 is its built-in curated commercial high-definition asset library. This library has been carefully selected and optimized by a professional team, containing thousands of high-quality commercial-grade image resources across multiple categories including natural landscapes, urban scenes, portrait photography, and product displays. All materials undergo rigorous copyright review and quality control, ensuring users can safely and legally incorporate them into commercial projects. The library's intelligent matching algorithm can also recommend the most suitable image resources based on user needs, further enhancing creative efficiency.

Wide Range of Application Scenarios

ZimageTurbo HD 2026 is suitable for various commercial application scenarios:

  • Advertising design and marketing material creation

  • E-commerce product image optimization

  • Social media content creation

  • Film and video post-production and special effects processing

  • Texture generation and optimization in game development

  • Virtual reality and augmented reality content creation

Technical Advantages Summary

ZimageTurbo REDZ2.0 HD 2026 is not just a technical upgrade but a redefinition of image processing workflows. The 2x speed improvement brought by its INT8 Convrot architecture, combined with the curated commercial high-definition asset library, provides professional users with unprecedented creative freedom and efficiency assurance. This version marks a solid step forward in the commercialization journey of the ZimageTurbo model, establishing a new technological benchmark for the digital content creation industry.


Krea2-RED-Mix2 赤佬无码版 01/07/2026 8 Steps

INT8 Convrot for ComfyUI 0.27 [Native 原生节点支持] Uploaded
File name: Krea2RedMix1.1-INT8-Convrot-ComfyUI (original native)
File name: Krea2RedMix2.1-INT8-Convrot-ComfyUI (no mosaics)

Usage :ER_SDE/Euler | Simple | CFG=1 | 8 Steps

链接: Dark Beast | 黑兽 🐱‍👤Krea2赤佬无码版 已发布 06/28/2026


ERNIE-Red-Mix 26/04/2026 10Steps Fine-Tuning

ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT) and paired with a lightweight Prompt Enhancer that expands brief user inputs into richer structured descriptions.

With only 8B DiT parameters, it achieves state-of-the-art performance among open-weight models. The model emphasizes both visual quality and controllability, making it highly effective for real-world generation tasks where precision matters.

ERNIE-Red-Mix adopts mixed-precision SFT (Supervised Fine-Tuning) with reference-based alignment, covering both AIO and DiT variants. It is trained on the mature RedCraft dataset, enabling broader generation flexibility, fewer constraints, and the ability to unlock novel visual styles beyond the base model.

Usage :EULER/DEIS | Simple | CFG=1 | 10Steps

🇨🇳 中文

ERNIE-Red-Mix 采用混合精度 SFT(参考式微调,对齐优化),覆盖 AIO / DiT 双版本,基于成熟的 RedCraft 训练集进行训练,在一定程度上解除生成限制,并显著提升风格扩展能力,可用于生成全新视觉风格

🇯🇵 日本語

ERNIE-Red-Mix は、混合精度による SFT(参照ベースのファインチューニング)を採用し、AIO / DiT の両バージョンに対応しています。成熟した RedCraft データセットで学習されており、生成制約の緩和とともに、新しいスタイル生成能力を拡張します。

🇹🇼 繁體中文

ERNIE-Red-Mix 採用混合精度 SFT(參考式微調),涵蓋 AIO / DiT 雙版本,基於成熟的 RedCraft 訓練集進行優化,在一定程度上解鎖生成限制,並提升風格拓展能力,可生成全新視覺風格

ERNIE-Red-Mix Highlights:​

  • Optimized speed ♥ Accelerated to 10 steps (CFG 1) while preserving BASE quality.

  • Compact but strong ♥ Performance on par with substantially larger models, with stable and accurate details and materials.

  • Text rendering preserved ♥ Fine-tuning does not compromise text rendering capability (posters, infographics, UI visuals).

  • Instruction following maintained ♥ Unaffected ability to reliably handle complex prompts with multiple objects.

  • Structured generation ♥ Continues to excel in posters, comics, and storyboards.

  • Broader style coverage ♥ More realistic photography and improved aesthetic outputs.

  • Lower VRAM footprint ♥ Mixed precision supports consumer GPUs with 8–12GB VRAM.

Limitations:​

  • Complex limbs prone to artifacts — An inherent ERNIE issue; generating more samples can help mitigate this.

  • Anatomical accuracy — Body anatomy and organ rendering still need further optimization and refinement.

  • Mixed precision impacts text rendering — Mixed precision has a noticeable negative effect on text quality; for text-heavy content, the BF16 BASE model is recommended.

Sample Images


ZImage DPO “AGILE” Now Uploading 08/03/2026 女神节祝福🌸

On this Day, may every woman feel the strength, grace, and boundless potential that lives within her. Thank you for your courage, your kindness, your resilience, and the countless ways you make the world brighter and better. Here's to equality, empowerment, joy, and endless possibilities—today and every day. Happy International Women's Day🌸

ZIDistilled FUN “AGILE” Now Released

Special thanks to the VideoXFUN team for releasing the groundbreaking Zimage Distilled Adapter 2603. By incorporating this latest update, AGILE achieves a refined balance between speed, diversity, and visual richnessunlocking more creative freedom while maintaining exceptional efficiency.

Agility in Motion, Diversity in Depth

The brand-new ZImage FUNAGILE” is built upon the cutting-edge ZIB acceleration framework. We deliberately reduced DPO & Distilled weight to preserve greater stochastic freedom, combined with the most recent training datasets, resulting in dramatically increased output variety, richer content details, and more imaginative compositions without sacrificing core stability.

For the first time, AGILE reaches a true quality parity challenge against the flagship “ZImage TURBO” in terms of overall image fidelity and sharpnesseven in complex scenes — while delivering faster iteration and superior responsiveness.


Key highlights:

True ZIT-level unlocked — photorealistic lighting, textures, and material rendering that now rivals or approaches ZImage TURBO quality, even at standard step counts.

Enhanced diversity & content richness DPO + Newest datasets = more varied poses, styles, atmospheres, intricate details, and unexpected creative sparks in every generation.

ZIB ecosystem ignition — exceptional native compatibility with ZIB-series LoRAs; your existing and future LoRAs now align faster, reproduce more faithfully, and shine brighter than ever before — officially kicking off the full ZIB LoRA era.

Agile workflows — seamless hybrid use with Klein 9B for refinement, ensemble boosting, or rapid prototyping; near-instant LoRA response with preserved high-entropy creativity.

Every generation is a step toward freer, bolder imagination.


欢迎体验 ZImage FUN “AGILE” —— 速度如洪创意如潮

Welcome to ZImage FUNAGILE” — where agility meets abundance, and your ideas finally run wild with unmatched fidelity and freedom.


ZImage DPO “Veris” Now Released 03/03/2026 元宵节快乐

I have uploaded more quantification and export VerisLoRA to HF repo. to avoid causing confusion for users in the community:
https://huggingface.co/GuangyuanSD/Z-Image-Distilled

版本太多网友容易迷糊,我导出了更多量化规格和 “Veris LoRA 版本,已发布抱脸仓库。

Special thanks to @Fok for providing the Flow-DPO technical adaptation. By skillfully integrating the training philosophy of Direct Preference Optimization (DPO) into the distillation weights, the Zimage distilled model achieves a major leap in lighting, color fidelity, and material authenticity — more natural light & shadow, more believable colors, and details that hold up under scrutiny.

特别感谢 @Fok 饼儿佬提供了Flow-DPO技术适配。通过巧妙地将直接偏好优化(DPO)的技术理念融入蒸馏权重,Zimage 蒸馏模型在光照色彩保真度材质真实性方面实现了重大飞跃——更自然的光影效果、更逼真的色彩,以及经得起仔细审查的细节。

The following example shows a comparison between ZIT and Flow DPO, intended to illustrate the effect of DPO, rather than a direct demonstration of ZIB Distilled


Speed of Truth, Fidelity of Flow

真实且极速,用忠诚在流动

The all-new ZIDPO “Veris” is powered by the latest-generation ZIB acceleration engine. Building on the RedZDX training data, we further distilled a more efficient, more refined Zimage-based model.

Now — solid, highly realistic generations in just 8 steps.(Better LoRAs alignment)

仅需8步即可生成更有层次感、高度逼真的图像。(LoRa对齐效果更佳

---

Key highlights:

  • Realism-first prototyping — near-zero latency for LoRAs, with lighting and color already very close to final training targets

  • High-entropy stochastic pre-sampling — delivers fast, high-quality realistic initial noise for ZImage pipelines

  • Hybrid realism workflows — seamless integration with Klein 9B for cascaded refinement or ensemble boosting, pushing visual fidelity and consistency even higher

  • Every step toward truth deserves full commitment.

---

欢迎体验ZIDPO“Veris”——您的LoRa训练结果不再只是“相似”,而是真正得到“复现”。

Welcome to experience ZImage DPO “Veris” — where your LoRAs generations are no longer just “similar”, but truly are.

同时,欢迎体验在 ZImage Turbo 模型上直接加载 DPO LoRA Adapter

抱脸(HF) https://huggingface.co/F16/z-image-turbo-flow-dpo

魔搭(境内) https://modelscope.cn/models/FFFFFFoo/z-image-turbo-flow-dpo


Redcraft DX3 ZIB🟥 Distilled models Zoo:

Full Model bf16 (19.11 GB)<- ComfyUI All-in-One Checkpoint BF16

Pruned Model BF16 (11.46 GB) <- ComfyUI Diffusion BF16 精度模型权重

Pruned Model fp8 (6.75 GB) <- ComfyUI Diffusion Scaled FP8 Mixed 混合精度

Pruned Model nf4 (6.73 GB) <- NVFP4 Mixed 混合精度BLACKWELL 50系加速

Training Data (3.75 KB) <- ComfyUI Simple Hybrid Workflow 简易混合采样工作流


Redcraft DX3 ZIB🟥 Distilled LoRA Adapter 02/19/2026

Additionally, I've exported Redcraft DX3 ZIB Distilled LoRA in Rank-256 format. The LoRA weight can be adjusted to adapt it to various ZIB fine-tune models, fully compatible with the Z-Image(non-turbo) base model.

Full Model fp16 (1.06 GB) <- 可以通过这里直接下载 LoRA 版本

[ZI Distilled HF repo.](https://huggingface.co/GuangyuanSD/Z-Image-Distilled)

上面是 Redcraft DX3 ZIB Distilled 导出为 Rank256 的LoRA版本,可以调整权重强度用于各种微调ZIT版本, 适配于 Z-Image(non-turbo) base 基底模型.


Redcraft DX3 ZIB🟥 Distilled LoRA adaptation models Zoo:

Z-Image-Base-GGUF <- Z-Image Base GGUF 量化模型

Z-Image Base <- Z-Image Base & TE (FP8/FP4) 模型

Z-Image Base FP8Mixed <- Z-Image Base FP8 混合精度模型

Text Encoder (ClipLoader use) <- Qwen3 4b FP16 文本编码器

Abliterated Huihui Qwen3 4B v2 (Q_8 GGUF) <- Z-Image Uncensored TE 文本编码器

VAE (Flux.1 16C VAE) <- 标准的 Flux.1 16 通道 VAE

Or download from the "Files" list below the "Details" on the right side of this page>>


Also available in NVFP4 quantized format, optimized for acceleration on Blackwell architecture GPUs.

Double speed, Half resources.

like RTX50XX, PRO6000, B200, and others

Verify environment is my ComfyUI 0.11

Also supports non-50 series GPUs (automatic 16-bit operation)


DF11 Lossless Compression RedZDX V3 came out! 2/15/2026

learn more: Dynamic-length Float (DFloat11)

[HF] mingyi456/Z-Image-Distilled-DF11-ComfyUI


Z-Image-Distilled v3 (RedZ DX3) 2/11/2026

Thanks to @Bubbliiiing VideoX-Fun&Alibaba-PAI Provided us with a more efficient distillation solution

Speed of Light, Power of Flow: The new ZID v3 "Lucis" is powered by the latest ZIB acceleration. Building on ZID v2 trainning sets, we've distilled a more efficient Zimage-based RedDX3. Now, in just 5 steps, you get solid results.

Rapid Prototyping: Test LoRA training hypotheses instantly with 'near-zero' latency.

Stochastic Pre-sampling: Serve as a high-speed, high-entropy source for ZiTurbo pipelines.

Hybrid Workflows: Pair seamlessly with Klein 9B for cascaded refinement or ensemble generation.

  • inference cfg: 1.0-1.5(建议1.0)

  • inference steps: 5(5-15步)

  • sampler / scheduler: Euler / simple

Welcome to the era of instant creativity. Welcome to 'Lucis'.

Preview images generated by Z-Image Hybrid Workflow of Distilled V3+Moody MIX V7(ZIT finetune) ,Just for showing the style difference between ZID(RedZDX3) and ZIT(fine-tunning) , no ranking intended =)

[ L = 'ZID v3', R = 'ZIT ft' ]

演示例图使用 ZIDistilled V3+Moody MIX V7 混合工作流程,不用做排名对比:

中国境内 [ modelscope ]AiMETATRON/Z-Image-Distilled | [ HF ] GuangyuanSD/Z-Image-Distilled


Z-Image-Distilled v2 (RedZ DX2) 2026/2/5

To a certain extent, the problem of ZIB color deviation has been reduced, but it is recommended to adjust the color appropriately according to the art style

  • inference cfg: 1.0(建议1.0)

  • inference steps: 10(10-15步)

  • sampler / scheduler: Euler / simple

感谢🙏这位作者完成了ZIB的FP8mixed混合量化方案:

https://huggingface.co/pachiiahri

已上传FP8版本,请给这位作者点赞👍

以上是FP8 scale&mixed 直出工作流(请不要再说我造假,我的所有例图工作流都是开放的)

精度混合方案来自 https://civitai.com/models/2172944/z-image-fp8


Comparison of RedCraft Zimages(bf16):

The art style leans towards realism

Retains ZIB's creative ability and reduces the collapse of Human anatomy.


REDZiBDX1·Demo accelerated Base-Model CFG1

Distilled form ZImage(non-turbo)base-model bf16

Now we have the LoRA version, thank to @anyMODE for Extract

in-site link https://civitai.com/models/2359857/z-image-base-distilled-lora-or-extracted

Pruned Model bf16 (11.46 GB) = Z-Image-Distilled / RedZDX-ZIB-Distilled-nocfg-10steps-BF16-Diffusion-models.safetensors 单独的扩散模型文件bf16剪裁精度

Full Model fp8 (16.87 GB) = Z-Image-Distilled / RedZDX-ZIB-Distilled-nocfg-10steps-FP8mixed-AIO-Checkpoints.safetensors 完整的Checkpoints(含TE/VAE)

Pruned Model fp8 (5.73 GB) = Z-Image-Distilled / RedZDX-ZIB-Distilled-nocfg-10steps-fp8-e4m3fn-Diffusion-models.safetensors 单独的扩散模型文件fp8剪裁精度及e4m3规格

Training Data (5.6 KB) = Z-Image-Distilled ComfyUI workflows 我自己使用的简易测试工作流

VAE (319.77 MB) = Flux.1 VAE ae.sft 常规的Flux.1 VAE

例图包含完整出图参数,点击右下角❕就可以看到,同时点击 COMFY Nodes 就可以复制源图工作流,并可以在ComfyUI界面中 Ctrl+V 粘贴至新的工作区。

All sample images contain full generation metadata.

Click the ❕ button (bottom-right) to view the complete parameters. Click COMFY Nodes to copy the original workflow JSON, then paste it (Ctrl+V) into a new ComfyUI workspace.


Z-Image-Distilled

本模型为基于 Z-Image 源版本(非Turbo)的直接蒸馏加速版,旨在测试Z-Image(non-turbo)版本上训练的LoRA效果,并显著提高推理/测试速度。模型完全没有融入Z-Image-Turbo的任何权重与风格,属于基于Z-Image的纯血版本,较好地保持了原版Z-Image的适配性、出图随机多样性以及整体图像风格。

相比官方Z-Image,推理速度更快(推荐10–20步即可获得较好效果);相比官方Z-Image-Turbo,本模型保留了更强的多样性、更好的LoRA兼容性与可微调潜力,但速度略慢于Turbo(仍远快于原始Z-Image的28~50步)。

模型主要适用于:

  • 希望在Z-Image非Turbo基底上训练/测试LoRA的用户

  • 需要比原版更快、但又不想牺牲太多多样性与风格自由度的场景

  • 艺术、插画、概念设计等对随机性与风格多样性有一定要求的生成任务

  • 适配 ComfyUI 的模型格式及层命名前缀

使用方法:

推荐推理参数:

  • inference cfg: 1.0–2.5(建议1.0~2.5区间,较高值可增强提示贴合度)

  • inference steps: 10–20(10步快速预览,15–20步品质更稳定)

  • sampler / scheduler: Euler / simple,或 res_m,或其他兼容sampler

LoRA兼容性良好,权重建议0.6~1.0,根据需求微调。

This model is a direct distillation-accelerated version based on the original Z-Image (non-Turbo) source. Its purpose is to test LoRA training effects on the Z-Image (non-turbo) version while significantly improving inference/test speed. The model does not incorporate any weights or style from Z-Image-Turbo at all — it is a pure-blood version based purely on Z-Image, effectively retaining the original Z-Image's adaptability, random diversity in outputs, and overall image style.

Compared to the official Z-Image, inference is much faster (good results achievable in just 10–20 steps); compared to the official Z-Image-Turbo, this model preserves stronger diversity, better LoRA compatibility, and greater fine-tuning potential, though it is slightly slower than Turbo (still far faster than the original Z-Image's 28–50 steps).

The model is mainly suitable for:

  • Users who want to train/test LoRAs on the Z-Image non-Turbo base

  • Scenarios needing faster generation than the original without sacrificing too much diversity and stylistic freedom

  • Artistic, illustration, concept design, and other generation tasks that require a certain level of randomness and style variety

  • Compatible with ComfyUI inference (layer prefix == model.diffusion_model)

Usage Instructions:

Basic workflow: please refer to the Z-Image-Turbo official workflow (fully compatible with the official Z-Image-Turbo workflow)

Recommended inference parameters:

  • inference cfg: 1.0–2.5 (recommended range: 1.0~2.5; higher values enhance prompt adherence)

  • inference steps: 10–20 (10 steps for quick previews, 15–20 steps for more stable quality)

  • sampler / scheduler: Euler / simple, or res_m, or any other compatible sampler

LoRA compatibility is good; recommended weight: 0.6~1.0, adjust as needed.

Current Limitations & Future Directions

Current main limitations:

  • The distillation process causes some damage to text (especially very small-sized text), with rendering clarity and completeness inferior to the original Z-Image

  • Overall color tone remains consistent with the original ZI, but certain samplers can produce color cast issues (particularly noticeable excessive blue tint)

Next optimization directions:

  • Further stabilize generation quality under CFG=1 within 10 steps or fewer, striving to achieve more usable results that are closer to the original style even at very low step counts

  • Optimize negative prompt adherence when CFG > 1, improving control over negative descriptions and reducing interference from unwanted elements

  • Continue improving clarity and readability in small text areas while maintaining the speed advantages brought by distillation

We welcome feedback and generated examples from all users — let's collaborate to advance this pure-blood acceleration direction!

当前不足与更新方向

当前主要不足:

  • 蒸馏过程对文本(特别是极小尺寸文字)有一定程度的破坏,渲染清晰度与完整性不如原版Z-Image

  • 色调整体与原版ZI保持一致,但在个别采样器下会出现偏色现象(特别是偏蓝色调的表现较为明显)

接下来优化方向:

  • 进一步稳定 CFG=1 情况下 10步以内 的生成质量,争取在极低步数下获得更可用、更接近原版风格的结果

  • 优化 CFG>1 时的负向提示词遵循表现,提升对负面描述的控制力,减少不需要的元素干扰

  • 持续改善小文字区域的清晰度与可读性,同时尽量维持蒸馏带来的速度优势

欢迎各位使用者提供反馈与生成示例,一起推动这个纯血加速方向的迭代!

Model License:

Please follow the Apache-2.0 license of the Z-Image model.

Please follow the Apache-2.0 open source license for the Z-Image model.

Thanks 🙏 DirectDistill(end-to-end) technology contributed

by Modelscope & DiffSynth Studio ,感谢佬同志团队提供的技术支持!


Dark Beast Z-造相·黑兽-DBZiT⚡️Final

1/26/26

[直达链接] https://civitai.com/models/2242173/dark-beast-z-or-dbz

迎接 Zimage non-turbo 版本特此放出基于ZiT的“黑兽”最终版本,期待官方EDIT版本的发布。我们会继续在ZI的基础上进行强化学习,并继续以 Turbo 方式发布全新 sft models

修正黑兽I-III越来越过激越界的调试方案,回归本质。

DBZ IV 是基于RedZ1.5 重新调试的修正畸变版本(没有额外突破性尝试

构图能力对标 ZiT base 原版,所有概念元素齐全LoRA适配性强。

---

Introducing Zimage DBZ IV sft FINAL model

Today we proudly release the final version of "Dark Beast ZiT" (黑兽),

fully rebuilt and refined on the ZiT foundation.

This iteration addresses and corrects the increasingly extreme and boundary-pushing tuning drifts seen in Dark Beast I–III, returning firmly to its core essence and original intent.

DBZ IV is a carefully re-tuned and distortion-corrected variant built directly on RedZ1.5.

It introduces no groundbreaking new capabilities beyond solid refinements

its strengths lie in stability and fidelity.

Composition quality matches or closely rivals the original ZiT base model.

Full coverage of conceptual elements and prompt adherence.

Excellent LoRA compatibility and strong adapter tolerance across a wide range of styles and concepts.

Enjoy the refined Dark Beast experience—clean, powerful, and true to its roots.

Feel free to tweak this further if you want a more hype-driven tone, shorter blurb,

or additional technical details!


这是「LSP」Lovin' Snuggly Positions