文章
Qwen-image2.1 提示词
图片编辑
上传图片,说清楚你想改什么就行。需要保留的地方可以特别说明,其他细节由 AI 帮你补充。
【通用写法】
把图里的【对象】改成【想要的样子】,保留【不想改变的部分】。
【多图参考】
用图1的【人物/产品】,参考图2的【服装/场景/画风】,做成【想要的画面】。
【常见用法】
1. 服装换色
把人物的白色上衣改成深红色,其他不变。
2. 整套换装
让图1的人穿上图2的衣服,脸和姿势不变。
3. 更换背景
把背景换成海边日落,保留人物。
4. 更换发色
把头发改成银白色,发型不变。
5. 更换表情
让人物笑得开心一点,其他不变。
6. 更换光线
改成温暖的夕阳光,人物和场景不变。
7. 白天变夜晚
把白天改成夜晚,让街边的灯亮起来。
8. 更换天气
改成下雪天,屋顶和地面有一点积雪。
9. 添加道具
给人物右手加一杯咖啡。
10. 移除道具
去掉桌上的花瓶,其他不变。
11. 移除路人
去掉人物后面的路人,保留主角。
12. 抠出主体
把这只猫抠出来,背景透明。
13. 添加文字
在图片上方加上“夏日特惠”,用醒目的白色大字。
14. 替换文字
把“夏日特惠”改成“限时优惠”,位置和样式不变。
15. 产品换色
把这辆车改成黑色,款式和标志不变。
16. 产品换背景
把这瓶香水放在大理石台面上,背景简洁一点。
17. 扩图
把图片左右扩宽成横图,保留原来的内容。
18. 局部修改
把圈出来的椅子换成木椅,其他不变。
19. 保持姿态换场景
让这个人保持现在的姿势,站到雪山前。
20. 多人物合影
让图1和图2的人站在一起,在公园里合影。
21. 转换画风
把图1画成图2的漫画风格,保留人物特征。
22. 制作多宫格
用图里的人物做九宫格漫画,讲他们打游戏输了后吵起来的故事。多图
图1女生穿着图2的衣服(只参考衣服),右手拎着图3的包,脖子上系着图4的丝巾,带着图5的帽子,摆成图6的姿势,但是只参考姿势,其他不要出现,保持图1人物的面部一致性。
图生图
---
name: qwen-image-21-image-to-image
description: 附带图像时编写 Qwen 图生图提示词,涵盖原图编辑、参考创作、人物与风格分工、多图合成和多宫格。
---
# Qwen 图生图提示词生成器
## 输入范围
当前任务附带一张或多张可查看的图像时使用本文件。明确沿用且仍可查看的历史图片也算当前输入;无关历史图片不自动纳入。
先看图再写提示词。要求修改某图而图像缺失或打不开时请求补图,不假装已经看过。只上传图却没有目标时询问希望如何修改或创作。完全没有图像、只要求文字创作时使用文生图文档。
## 使用与交付
本文件独立可用,不需要另一模式的文档。只编写提示词,不自动调用图像服务。默认按用户语言输出一版完成稿,以代码块放可复制正文,画幅参数放在正文外。用户指定英文时使用英文,画面内文字仍保留用户指定的准确字串。
用户需要 JSON 时再输出结构化字段。positive_prompt 是本适配版正文名,原包 PE 使用 rewritten_prompt,两套格式按用户接入需求选择,不混用。negative_prompt 默认空字符串;只在用户要求且渠道支持时填写。
长度由信息量决定,不硬凑字数;模板占位符在实际交付时填实。示例的题材不是默认题材,媒介与风格由用户要求决定。
## 图生图内部的两个分支
### 2.1 编辑现有画面
典型需求:换色、换装、换文字、去物体、局部重绘、换背景、改光、扩图、风格转换。
先确定被编辑的底图,再分清要改的属性和需要保留的属性。提示词写成可执行指令:
> 编辑动作与范围 → 目标效果 → 必要的融合关系 → 保留项。
对未涉及的内容用一句清楚的保留声明,不逐个重新描绘所有细节,以免让模型误以为要重画它们。
### 2.2 基于参考图创作新画面
典型需求:让参考人物进入新场景、按参考画风画新剧情、人物合影、海报、三视图、六宫格或九宫格。
仍属于图生图,因为输出依赖输入图片;但它不是对原构图的局部修改。
明确参考图提供的身份、服装、风格、场景或结构,主动设计用户要求的新姿态、新构图与剧情。只保留被指定继承的属性,不用“其他所有内容保持不变”锁死原图姿势、镜头和布局。
例如,用户说“图1人物、图2漫画风格,做两人打起来的九宫格”:图1负责人物辨识,图2负责画法;新动作、九格布局和剧情由需求决定。保留人物身份不等于保留照片的皮肤纹理和摄影光影。
## 图生图流程:看图后再写
1. 读取实际图像,区分可见事实与不确定细节。看不清的文字不猜作精确内容。
2. 判断是“编辑现有画面”还是“参考创作”。
3. 给每张参考图分配职责,明确底图是谁;新构图任务可以没有底图。
4. 列出目标属性与保留属性,确保二者不矛盾。
5. 将改动写成具体动词:替换、移除、增加、改为、延展、转换、安排。
6. 处理接触、遮挡、光影与透视。移除物体要补全露出的背景,添加道具要说明承托或握持关系。
7. 设置画幅,按用户需要交付纯提示词或 JSON。
### 引用规则
- 单张图:正文自然称“输入图像”“图中人物”;画幅字段仍可写 `ratio_follow: <image1>`。
- 多张图:正文使用 `<image1>`、`<image2>` 等,编号严格对应实际输入顺序。
- 分别说明每张实际使用的图提供什么,不让同一属性在多张图之间无原则混合。
- 不为排版方便擅自调换编号。线稿不必强行放最后,角色取决于任务而非上传位置。
- 多张图片提供同一人物的不同角度时,可以共同支持身份;区分补充视角与相互冲突的造型。
### 属性分工表
| 属性 | 确定方式 |
|---|---|
| 人物身份 | 指向身份参考;仅补充区分人物所需的简短特征 |
| 服装与配饰 | 用户未要求改变时继承人物图;指定换装时使用对应参考 |
| 风格与媒介 | 用户指定的风格图或文字要求决定 |
| 构图与姿态 | 编辑底图时按要求保留;新剧情或新镜头时重新设计 |
| 场景 | 由背景图、底图或明确场景要求决定 |
| 光线 | 局部编辑继承底图;换光任务或新场景按目标重建 |
| 道具 | 指定来源、数量、位置和握持者 |
风格转换时,画法是明确需要改变的属性。把照片变成漫画时保留身份和必要服装特征,允许将皮肤、衣料与立体光影转译为漫画线条、黑块和网点。
## JSON 输出
需要结构化接入时使用以下适配格式:
```json
{"positive_prompt":"完整编辑指令或参考创作提示词","negative_prompt":"","wh_ratio":"","ratio_follow":"<image1>"}
```
wh_ratio 与 ratio_follow 恰有一个非空。多图任务跟随实际底图,底图不一定是第一张;新场景和多宫格根据新构图设置画幅。
明确要求 PE-I2I 原始格式时输出 rewritten_prompt、wh_ratio、ratio_follow,不再混入 positive_prompt。单图正文自然称输入图像,多图按实际输入顺序使用标签;ratio_follow 始终用图像标签。
## 画幅与多格布局
按以下优先级选择:
1. 用户明确的宽高比或像素尺寸,像素尺寸先约为整数比。
2. 原图编辑且未改变画布:跟随底图。
3. 参考创作:依据新画面用途选比例。
4. 多格图:依据列数、行数与单格比例计算整张图比例。
| 用途 | 无明确要求时的参考值 |
|---|---|
| 横向普通场景 | 3:2 |
| 竖向人像 | 2:3 |
| 横屏分镜、演示画面 | 16:9 |
| 手机竖图 | 9:16 |
| 方形头像、图标 | 1:1 |
| 宽银幕画面 | 按需求使用 21:9 |
“4K、高清”不决定比例;场景里有“电影感”也不自动推翻已有比例。
等尺寸网格、不计窄边框时:
> 整图宽高比 =(单格宽 × 列数):(单格高 × 行数)。
- 九宫格,三列三行,每格 16:9 → 整图 16:9。
- 六宫格,两列三行,每格 16:9 → 整图 32:27。
- 六宫格,三列两行,每格 16:9 → 整图 8:3。
- 九宫格,每格正方形 → 整图 1:1。
明确区分整图画幅与单格形状。平台不支持计算出的比例时,说明适配方式,不能同时承诺互相冲突的整图与单格比例。接口严格禁止正文包含比例时,用“横向电影画幅”等表述描述单格,整图比值只写到画幅字段。
扩图明确向哪里延展,保留原图区域并补绘画外内容;依据延展方向设置新比例。四周等比例扩展可保持原比,但提示词仍应说明扩展画布。
## 多宫格与漫画分镜
先确定格数、行列、阅读顺序、整体媒介和人物来源,再逐格写关键瞬间。
每格包括:叙事作用、景别或机位、人物位置、一个核心动作、可见神态、必要道具状态、确有需要的文字。
### 镜头与叙事
- 远景建立环境和人物关系;中景说明互动与动作;近景看神态;特写揭示手部或道具;空镜承接停顿、悬念或笑点。
- 镜头变化服务剧情,避免每格都让人物站在中央面对观众。用户要求有呼吸感时,主动安排景别变化与适量空镜。
- 单格描绘一个可定格瞬间,不把“起身、跑过去、出拳、倒地”塞进同一格。
- 连续动作交代准备、接触或落空、后果。主体的动作方向和空间关系应可追踪。
- 记录关键道具状态:哪格拿着、哪格交给谁、何时飞出、下一格落在哪里。
- 对话适合当前情节,明确谁说;气泡位置避开脸部、手部和动作触点。用户不需要文字时不自动添加台词。
### 漫画风格具体化
“黑白”只规定颜色,不能单独确保漫画风格。根据风格参考写出:
1. 轮廓与结构线的粗细、疏密和概括程度。
2. 纯黑阴影、纸白亮部、网点与排线的分工。
3. 五官、头发、肌肉和衣褶如何漫画化。
4. 动作透视、夸张变形、速度线和冲击形状。
5. 背景在安静格保留多少细节,在动作格如何简化。
摄影的毛孔、逐根头发、真实布料纤维和柔滑灰阶塑形,只用于用户指定的相应画法。人物身份参考不应自动把目标漫画变成灰度照片。
### 可套用的多图漫画结构
```text
将 <image1> 中的人物重新绘制为 <image2> 的漫画画法。<image1> 提供人物身份、体型差异与服装辨识特征;<image2> 提供墨线、黑白分布、造型概括与动作透视。新剧情和分镜布局按以下描述设计。
整页共九格,三列三行,阅读顺序从左到右、从上到下。人物与场景连续,画法统一。
第一格:交代场景与事件起因的远景。
第二格:展示人物反应的近景。
第三格:揭示动作触发点的道具特写。
第四至第七格:按剧情分配准备、行动、应对和后果,交替使用可辨识全身动作与局部特写。
第八格:后果或情绪转折。
第九格:以远景、反应近景或空镜完成结尾。
```
此结构用于组织,不是每次强制照搬的剧情。实际交付要填成完整内容,不留下“按剧情”等占位语句。
## 文字、海报与信息图
画内每处可读文字明确写出:准确字串、位置、字体类型与相对字号、颜色、呈现方式。用英文双引号标注准确内容,不把章节名、工作指令和风格说明误当作画内标题。
文字语言与提示词语言分别确定:
1. 用户指定准确内容或语言时,按用户要求。
2. 编辑原图且未指定新语言时,保持原图文字的主导语言。
3. 原图没有文字且用户要求新增文字时,默认使用用户的语言。
中文提示词可以明确要求英文对白。需要双语时按用户要求排版,不自动把每句话翻译成双语塞进画面。
海报先确定标题、主体、说明与落款层级;信息图先确定区块和阅读顺序,再填文字与图形。图表的轴名、刻度、图例与数据也属于文字内容。
长文可能降低逐字准确度。保留用户要求的文字,不擅自删改;需要分块生成或后期排字时给出具体方案,不保证仅凭提示词完全还原。
## 常用任务配方
下表是编辑现有画面的写法要点,参考创作按“基于参考图创作新画面”分支设计。单图自然指代,多图使用实际编号。
| 任务 | 目标与保留范围 |
|---|---|
| 服装换色 | 指定哪件衣服及目标色;保留款式、褶皱、人物与原有光照。平涂仅用于目标画法需要时 |
| 整套换装 | 指定服装来源、覆盖范围和配饰处理;保留身份与未指定改变的姿态 |
| 更换背景 | 指定新环境及透视;保留主体身份和位置,按目标需要说明融合光影 |
| 更换发色 | 改颜色,保留发型轮廓与适合该媒介的明暗 |
| 更换表情 | 写眉眼、嘴角与视线的变化,保持身份;如嘴角上扬、眉间放松 |
| 更换光线 | 指定光源方向、软硬和色调;允许相应高光、阴影与反射变化 |
| 白天变夜晚 | 改环境亮度与相关光源关系,保留未指定改变的主体和场景结构 |
| 更换天气 | 描述雨雪雾和必要环境反应,如积雪或湿地;限定影响范围 |
| 添加道具 | 指定来源、位置、尺度、承托或握持关系和遮挡 |
| 移除道具 | 指定对象,并依据周边结构补全被遮挡区域 |
| 移除杂物或路人 | 明确对象或区域,保持其余人物、建筑和透视 |
| 抠主体 | 去除背景,保留主体与必要半透明边缘;要求实际 alpha 通道 |
| 添加文字 | 给出准确字串、位置、字体、颜色、大小与表面材质 |
| 替换文字 | 指明旧内容与新内容,保持用户未要求变化的排版属性 |
| 产品换色 | 改产品指定部件颜色,保留设计、数量、标志与材质响应 |
| 产品换背景 | 保留产品身份,按新承托面建立合理接触阴影或倒影 |
| 扩图 | 指定延展方向、画布目标、原区域保留及新增场景的透视延续 |
| 局部重绘 | 指定可编辑区域、替换目标与外部保留项 |
| 保持姿态换场景 | 明确身份与姿态继承,背景重新设计,光影是否调整由需求决定 |
| 多人物合影 | 逐一绑定人物来源,设计站位、尺度、视线与统一环境 |
### 局部编辑输入
圈选、涂鸦、独立掩码都属于图生图。看清哪张是底图、哪张是编辑范围辅助图;按实际工具约定确认掩码颜色语义,不假设所有渠道都用白色表示可改区域。用于定位的圈线或涂鸦,完成编辑后按任务要求去除。
### 透明背景
透明背景不同于白底或画出来的棋盘格。需要透明素材时明确 RGBA / alpha 通道、主体轮廓与发丝、薄纱、玻璃等半透明边缘。是否保留投影由素材用途决定。
真正输出透明文件取决于生成渠道和保存格式,单靠提示词不保证实际文件带 alpha。本文只编写要求,不声称已生成或验证透明图。
## 共用资料:词典、模板与排版
本节已内嵌到文生图、图生图两个文件,两份均可独立使用,无需另外加载共用资料。
按目标媒介选词:摄影材质和布光词用于相应画法,漫画以线条、黑块、网点与留白表达形体。静态图里的运镜词只表示关键帧机位感,不要求单幅图完成连续动作。原包的色温数字作为风格参考,不等于所有真实光源的固定物理参数。色号、整块平涂和强调色按需求使用,不强制每张图采用同一种配色方法。
图生图局部编辑只借用与目标属性有关的词,参考创作才按新场景补充整体描述。模板中的文字、品牌和道具属于可选内容;没有相应需求时不擅自添加。实际交付填实占位符,画幅参数置于正文外。
### 词典:景别 / 视角 / 镜头
#### 景别(决定"看到多少")
大远景 / 远景 / 全景 / 中全景 / 中景 / 中近景 / 近景 / 特写 / 大特写(微距)
写法:`中近景镜头:…`、`大特写,只拍手与刀柄`。
#### 视角(决定"从哪看")
平视 / 俯视(高机位)/ 仰视(低机位)/ 航拍(俯瞰)/ 过肩 / 主观视角(POV)/ 正上方俯拍(top-down)
#### 镜头类型(决定"什么镜头味")
微距 / 超广角 / 广角 / 标准(35–50mm)/ 中长焦(85mm 人像)/ 长焦 / 鱼眼 / 移轴
#### 运镜(静止图里也可写"这一帧的机位感")
推近 / 拉远 / 横移 / 摇 / 升降 / 环绕 / 手持轻晃 / 固定机位
#### 组合范例
- 写实人像:`中近景,85mm 人像镜头,浅景深,平视`
- 压迫感:`低机位仰视,广角,主体占画左 1/3`
- 场面:`高机位俯视全景,超广角,前中后景分层清楚`
- 细节:`微距特写,只拍掌心与刀柄,焦点锐利`
### 词典:光位与氛围
#### 光位(写"光从哪来")
顺光 / 侧光(侧 45°)/ 侧逆光 / 逆光(轮廓光)/ 顶光 / 底光 / 窗光 / 灯笼烛火等**实用光源**(practical light)
#### 光质(写"硬还是软")
硬光(边缘锐利、明暗对比强)/ 柔光(过渡柔和)/ 漫射(阴天、雾面柔光罩)/ 高对比(灯火夜景)/ 低调(暗部为主)与高调(亮部为主)
#### 色温与色调
暖调(2700–3500K,烛火、钨丝)/ 中性(5000–5600K)/ 冷调(6500K+,月光、阴天)/ 冷暖对比(暖主体 + 冷背景)
#### 氛围词(放在收尾一句)
安静 / 压抑 / 一触即发 / 奢靡 / 苍凉 / 孤寂 / 温柔 / 肃杀 / 神圣 / 市井烟火 / 潮湿 / 尘雾弥漫 / 逆光的尘埃颗粒
#### 写法示例
`光源来自画面左侧的两盏宫灯,暖色偏软,主光从画左打在人物肩上,右侧处于低照度,地板上拖出偏右下的影子,空气里有被光照亮的浮尘。`
### 词典:材质与质感
写材质是让画面"看起来真"的最短路径——每个主体至少给一个材质词。
#### 织物
丝绸(柔滑、有高光)/ 棉麻(哑光、有纤维感)/ 锦缎(有暗纹与金线)/ 纱(半透、边缘透光)/ 皮革(有油亮与磨损)/ 毛呢(哑光、有绒面)
#### 木与竹
原木(可见年轮)/ 上漆木(硬高光)/ 老木(包浆、裂纹)/ 竹(青皮、竹节)
#### 金属
黄铜(暖调高光)/ 铁(冷调、氧化)/ 钢(镜面反光、刃口亮线)/ 金(柔和高光 + 高饱和)
#### 石与土
青石(细颗粒)/ 汉白玉(半透、温润)/ 夯土(粗糙、掉渣)/ 沙土(有脚印与滑痕)
#### 液体与地面
酒液(表面张力、挂壁)/ 水面(反射 + 涟漪)/ 湿地板(镜面反射、脚印)/ 灰尘(浮尘颗粒、落地积灰)
#### 皮肤与毛发
皮肤(毛孔、油光、薄汗、擦伤)/ 头发(发丝分层、有高光带)/ 胡须(硬毛、根部有皮肤过渡)
#### 组合范例
`上装为哑光棉麻质地的实色深红,褶皱处有明显压痕;腰带是氧化发暗的黄铜扣;地面老木地板有包浆与细裂纹,反射着烛火的暖光。`
### 词典:色彩与配色
#### 基础色(别只说"红色")
按明度与饱和度给限定:
深红 / 正红 / 朱红 / 绯红 / 暗绛;藏青 / 靛蓝 / 湖蓝 / 雾蓝;墨黑 / 炭灰 / 银灰;米白 / 牙白 / 铅灰
#### 给色号(最稳的强控)
- 直接写 hex:`纯色 #C0392B`、`#8C6A44`
- 配一句约束:`整块大面积平涂,不灰化、不洗淡、不换色`
#### 配色关系
| 关系 | 写法 |
|---|---|
| 冷暖对比 | 暖主体 + 冷背景(`暖色烛火照亮人物,背景是冷调夜色`) |
| 同色系 | 主色 + 邻近色(`整体暖棕色系,深棕与米白分层`) |
| 互补 | `红与青绿互为补色,红只用在主体上装` |
| 低饱和高级感 | `整体低饱和,只有主体上装是唯一高饱和色块` |
#### 色彩必须交代的三件事
1. **主色**:画面里面积最大的颜色
2. **强调色**:主体上唯一的视觉焦点色
3. **环境色**:背景与光的色调
#### 不要
- 一句话堆五个颜色(模型会平均化 → 变灰)
- "五颜六色""色彩丰富"(不可控)
- 主体与背景同明度同饱和(分不出主体)
### 词典:构图与画面平衡
#### 主体位置
居中(对称、庄重)/ 三分法(左 1/3、右 1/3)/ 画面右半 / 左边缘留白 / 前景压角
写法:`主体位于画面左侧三分之一处,面朝画右`。
#### 层次
前景(遮挡物、虚化)/ 中景(主体)/ 背景(环境、远景)。写清"前景是什么、中景是谁、背景是什么"。
#### 视线与动线
- 主体视线朝画左 / 画右 → 前方要留**更多空间**(不然像撞墙)
- 运动方向朝画右 → 右侧留空间
- 两人对话 → 面对面、视线交叉、中间留出对白空间
#### 空间感(让画面"有纵深")
引导线(走廊、柱子、桌沿)/ 大小对比(人物 vs 建筑)/ 空气透视(远近明度与饱和度递减)/ 地面反射
#### 平衡
- 一边重(大主体)→ 另一边用小元素或亮部配平
- 上方留白过多 → 下方加地面反光或影子压住
- 不要四角都塞满
#### 可直接用的句式
```
主体位于画面右三分之一处,面朝画左,左侧留出空场与三点透视的走廊;
前景是虚化的桌角,中景是两人对峙,背景是二层木楼与挂灯,远处因雾气而降低对比。
```
### 中英对照速查(写英文提示词时用)
#### 镜头与景别
大远景 extreme wide / 远景 wide / 全景 full shot / 中全景 medium wide / 中景 medium shot / 中近景 medium close-up / 近景 close-up / 特写 extreme close-up / 微距 macro
#### 视角与机位
平视 eye level / 俯视 high angle / 仰视 low angle / 航拍 aerial / 过肩 over-the-shoulder / 主观 POV / 正上方俯拍 top-down
#### 运镜
推近 push in / 拉远 pull back / 横移 tracking / 摇 pan / 升降 crane / 环绕 arc / 手持 handheld / 固定机位 static
#### 光
顺光 front light / 侧光 side light / 侧逆光 three-quarter backlight / 逆光 backlight / 轮廓光 rim light / 顶光 top light / 底光 under light / 实用光源 practical light / 柔光 soft light / 硬光 hard light / 暖调 warm / 冷调 cool / 高对比 high contrast / 低调 low-key
#### 材质
棉麻 matte cotton-linen / 丝绸 silk / 锦缎 brocade / 纱 gauze / 皮革 leather / 原木 raw timber / 上漆木 lacquered wood / 老木 weathered timber / 黄铜 brass / 铁 iron / 钢 steel / 青石 bluestone / 汉白玉 white marble / 夯土 rammed earth
#### 质感与效果
浅景深 shallow depth of field / 胶片颗粒 film grain / 湿面反光 wet reflection / 空气浮尘 dust motes / 体积光 volumetric light / 半透明 translucent / 哑光 matte
#### 常用指令句
保持…不变 keep … unchanged / 整块大面积平涂 fill as one large flat area / 不灰化 do not desaturate / 不要把色块打散 do not break it into patches / 不要把描边留在画面里 do not keep the black outlines / 输出为超写实电影剧照 the output is a photorealistic film still / 画面中未出现其他文字 the image contains no readable text
### 通用画面模板
用于文生图或参考图新画面创作。局部改图按编辑流程限定改动范围。
#### 通用骨架
```
首句:一幅<场景类型>:<场景>里,<主体>正在<做什么>。
清点:画面里有 <N> 个<主体类别>——<逐个交代身份/位置/朝向>。
走查:<从画左到画右或前景到背景,逐个写:服装材质 + 颜色 + 姿态 + 与谁互动>。
文字:画面内没有出现任何可识别文字。 ← 或 → 左上角写着"<逐字内容>",<字体/颜色/呈现方式>
光:光源来自<位置>的<灯具/自然光>,光线<软硬>、色温偏<暖/冷>,主光从<方向>打来,<受光面>有高光,<地面>上拖出<方向>的影子。
收尾:整体<氛围词>,像<一句类比>。
```
#### 人像(中近景)
```
一幅室内人像:<身份>坐在<位置>,<动作>。画面里只有一个人——
他/她<年龄段>,<发型>,穿<服装材质与颜色>,<姿态与手势>,视线朝<方向>,表情<情绪>。
画面内没有文字。光源是<位置>的<灯>,柔光,主光在<左/右>,脸颊<哪侧>受光,<眼/发>有高光。
背景<虚化程度>,整体<氛围>。
```
#### 产品(电商主图)
```
一幅产品图:<产品>置于<台面材质>上,背景<纯色/渐变>。
产品为<材质>、<颜色>,<高光位置>,<配件>摆在<位置>。
画面右上角有文字"<品牌名>",无衬线粗体,白色,字号小。
布光为<左上/右上>双灯柔光箱,产品左缘一条高光带,<台面>上有轻微倒影。整体干净、商业,像棚拍。
```
#### 场景 / 大场面
```
一幅<时间>的<地点>大景:<主体>处在<空间关系>,画面里有 <N> 个<元素>——<逐个交代>。
前景是<…>,中景是<…>,背景是<…>,<天气/雾气>让远景逐层退去。
光源来自<方向>,<暖/冷>调,<介质>被光照亮。整体<氛围>,像<类比>。
```
#### 海报 / 信息图
```
一幅竖版海报:上方大标题"<逐字>",<字体/颜色/字号>;中部<主体画面描述>;
下方一行小字"<逐字>",底部居中"<逐字>"。背景为<描述>,整体<配色>。画面中未出现其他文字。
```
### 文字与信息排版的具体写法
#### 写文字的四要素(每处都要给)
1. **内容**:引号内**逐字**写出(含标点、大小写、空格)
2. **位置**:左上 / 右上 / 居中 / 底部居中 / 中部偏下…
3. **字体**:无衬线粗体 / 衬线 / 手写 / 书法 / 像素 / 圆体;字号(大 / 中 / 小)
4. **呈现方式**:印刷 / 刺绣 / 霓虹灯 / LED 屏 / 木牌阴刻 / 投影 / 贴纸;颜色与描边
最后加一句总控:`画面中未出现其他文字。`
#### 信息图 / PPT / 漫画分镜
- 写法:**先给版式骨架**(几个区块、怎么排),再逐区块给内容;
- 区块顺序写清(左上 → 右上 → 左下 → 右下),否则会乱排;
- 漫画分镜:写清"几个分格、每格谁在做什么、对话框里的文字"。
#### 例子(竖版信息图)
```
一幅竖版信息图,白底,四个横向区块自上而下排列。
第一区块标题"<逐字>",黑色无衬线粗体,字号大;
第二区块是<图形/图示描述>,配一行说明文字"<逐字>",灰色小字;
第三区块是<…>;第四区块底部居中一行小字"<逐字>"。
画面中未出现其他文字。整体配色为<主色 + 辅助色>。
```
### 英文描述句式
仅在需要英文时选用,画面内文字仍按用户要求保留语言。
#### 骨架(与中文一一对应)
```
An <scene type> of <scene>: <subject> is <action> at <location>.
There are <N> <subject class> in frame — <each: identity, position, facing>.
<Walk the frame left to right / foreground to background — garment material + colour + pose + interaction.>
There is no readable text in the image. ← or → The top-left corner reads "<exact text>", <font/colour/placement>.
The light comes from <source position>, <soft/hard>, warm/cool; the key falls from the <side>, catching <surface> with a highlight and casting a shadow toward <direction>.
The overall mood is <mood> — <one-line simile>.
```
#### 动词(描述"画面里正在发生什么")
stands / sits / reclines / leans / turns / faces / looks toward / holds / presses / rests / reaches / raises / drops / walks past / blocks / points at
#### 材质与质感
matte cotton-linen / silk with soft sheen / brocade with woven gold thread / worn leather / lacquered wood / weathered timber / oxidised brass / polished steel / damp floorboards with mirror reflection / dust motes in the air
#### 光位
key light from camera-left / soft window light / practical lantern light / rim light from behind / low-key with a single warm source / cool ambient with warm practicals
### 常见问题的共用处理
| 问题 | 对应处理 |
|---|---|
| 人数或道具数量不对 | 明确总数,逐个安排位置,检查前后数量是否一致 |
| 视线无目标 | 指明看向哪位人物、哪个道具或哪一侧空间 |
| 手部结构不清 | 写清持物者、左右手、接触点与遮挡;保留用户要求的动作,不为避错擅自取消动作 |
| 摄影皮肤塑料感 | 根据画法加入适度纹理与光照变化,避免“完美无瑕”等空泛要求 |
| 漫画像灰度照片 | 明确墨线、黑块、纸白、网点和造型概括,减少摄影微观材质词 |
| 构图杂乱或太平 | 建立主次、前中后景、主体位置与必要留白 |
| 光影不一致 | 写清光源方向、受光面、投影落点和接触阴影 |
| 反射不合理 | 指明反射表面、对应主体与透视压缩,不随意增加镜面 |
| 颜色偏移 | 指定明确色名或用户色号,区分固有色和环境光;平涂只用于合适媒介 |
| 文字错漏 | 核对准确字串、字号和空间;长文必要时分块或后期排字,不擅自删改用户内容 |
| 多出无关文字 | 明确哪些位置有字,其余表面为空白或文字不可辨认 |
| 多格镜头重复 | 给每格分配叙事作用,结合远中近景、特写与空镜 |
先用具体正向描述解决问题。负向提示词只在渠道支持且用户需要时提供,按当前问题选择,不把“无文字、无描边”等约束套在需要文字或漫画描边的画面上。
## 完整案例
### 示例 B:附一张人物图,局部编辑
用户:上衣换成深红,其他别动。
模式:图生图|原图编辑。
```text
将输入图像中人物上衣的原有颜色改为深红色,保留原款式、纹样、材质、褶皱及光照形成的明暗层次。保持人物身份、面部表情、发型、姿态、配饰、下装、背景和构图不变,仅让上衣颜色产生明确变化。
```
画幅:`ratio_follow = <image1>`。
### 示例 C:附人物图和漫画图,参考创作
用户:图1两个人,图2画风,做四格游戏输了吵架的漫画。
模式:图生图|参考创作。
```text
以 <image1> 中的两个人物作为身份、体型和服装参考,以 <image2> 作为线条、造型概括和黑白明暗的风格参考,重新绘制一页四格漫画。画面采用两列两行的等大方形分格,按从左到右、从上到下阅读。两人的服装与体型差异贯穿四格,面部和材质按漫画画法重新表达。第一格为客厅远景:两人并排坐在沙发上握着手柄,面对显示游戏失败的老式电视,左侧男子前倾僵住,右侧男子转头瞪他;屏幕中央写着白色像素字 "FAILED"。第二格为左侧男子面部近景:他扬眉侧目,嘴角绷紧,右侧上方的白色气泡内写着黑色手写体 "不是你让我加速的?"。第三格为坐垫特写:两只手柄落在两人之间的空坐垫上,连接线弯曲,撞击点以短促墨线表现,背景只保留两人起身的局部身影。第四格为双人中景:两人面对面揪着对方衣领,眉头紧皱,右侧男子上方气泡写着 "我是让你开车,不是拆墙!"。全页用明确墨线、黑色阴影块和纸白亮部建立形体,以有限网点表现中间调,表情夸张而轮廓清楚;紧张动作与幼稚争吵形成喜剧反差。
```
画幅:`wh_ratio = 1:1`。此例假设两张实际图片确实提供上述人物和画风;处理其他图片时先看图再替换内容。
### 示例 D:单图编辑,JSON 输出
用户:把图里的杯子抠出来,要透明底,给 JSON。
```json
{
"positive_prompt": "从输入图像中提取杯子,移除背景,输出具有真实 alpha 通道的透明背景素材。保留杯子的完整轮廓、原有设计、颜色、材质和光照;若杯体包含半透明部分,保留其透明度过渡。主体位置与画布范围沿用输入图像。",
"negative_prompt": "",
"wh_ratio": "",
"ratio_follow": "<image1>"
}
```
### 示例 E:多图合成,底图由任务决定
用户:把图1的人放进图2房间里,站在窗边。
模式:图生图|多图合成。
```text
以 <image2> 为底图,保留房间的构图、家具位置与空间透视,将 <image1> 中的人物放置在房间窗边。人物身份、发型和服装取自 <image1>,身体尺度与房间门窗比例协调,双脚落在地面上,人物与窗边家具形成自然遮挡。人物明暗依据 <image2> 的窗光方向调整,脚下增加相符的接触阴影,边缘与室内环境融合。
```
画幅:`ratio_follow = <image2>`。
## 交付前检查与迭代
在内部核对,不把长篇检查过程附在每次交付后:
1. 模式是否由当前任务的有效图像输入确定?是否把无关历史图误当成参考?
2. 有图时是否实际看过?参考分工、编号和底图是否清楚?
3. 原图编辑与参考创作是否区分?需要改变的姿态或风格是否被保留语句误锁?
4. 人物、道具、位置、文字和用户指定颜色是否完整且一致?
5. 媒介描述是否统一?漫画是否混入不适用的写实材质要求?
6. 多格是否数量正确、动作衔接、道具状态连贯?整图与单格比例是否匹配?
7. 画内文字是否准确?叙述说明是否被误当成需要绘制的字?
8. 画幅字段是否互斥?JSON 是否可解析?正文是否留有占位符?
用户反馈结果不对时,先对应具体症状调整:
| 症状 | 优先调整 |
|---|---|
| 人数、道具数量不对 | 清点主体并逐个安排位置 |
| 人脸或配饰漂移 | 核对人物来源,明确身份和关键配饰的继承 |
| 风格像灰度照片 | 重新分配人物参考与风格参考职责,具体写线条、黑块与概括方式 |
| 镜头重复、缺乏呼吸 | 调整每格叙事作用和景别,必要时插入道具或环境空镜 |
| 改了不该改的部分 | 缩小编辑范围,整理保留项,减少无关重绘描述 |
| 风格或动作改变太弱 | 检查是否被“保持原图不变”的笼统要求抵消 |
| 光影与环境不融合 | 对齐实际光源、接触阴影、透视与材质响应 |
| 文字错漏 | 核对准确字串与文字空间,必要时分块或后期排字 |
| 请求报错或超时 | 查看实际渠道报错和限制,不仅凭现象断言提示词太长 |
保持已经有效的内容,只修改导致当前问题的部分。无需把每次失败都变成新的全局禁令。
## 整理说明
本文件包含本模式完整流程、对应 PE 提示词原文、案例,以及内嵌的镜头、光线、材质、配色、构图、中英术语、通用模板和文字排版资料。共用资料在两个模式文件中各保留一份,不依赖第三个文件。
此次仅交付文生图、图生图两个 MD,不附带自动入口、部署历史、技术参数手册或脚本源码。末尾保留的是本模式的提示词规则原文,不是部署技术附录。日常按前文适配规则及用户要求使用;明确要求原始 PE 契约时采用其语言、长度和字段约定。
原包对模型版本、能力及“官方来源”的陈述未在本次整理中联网核实。已有完整版继续留作原资料备份,但不是使用本文件的前置条件。
## 附录:PE-I2I 原文
````markdown
# 官方原文:Qwen-Image-2.1 提示词增强(图像编辑 I2I)
> 来源:QwenLM/Qwen-Image-2.1 官方仓库 `prompt_rewrite/` 与 `Qwen/Qwen-Image-2.1-PE-I2I` 的 `system_prompt.txt`。**逐字保留,不要改写。**
---
# Edit Prompt Enhancer — General (v2, 精简版)
**FIRST — there are TWO separate language decisions. Do NOT conflate them.**
**(A) Language of the rewritten prompt's DESCRIPTIVE prose — every word OUTSIDE double quotes (the description you write for the diffusion model, NOT the text painted into the image). This decision is final and non-negotiable:**
- User instruction is in Chinese → write the description in Chinese.
- User instruction is in English → write the description in English.
- User instruction is in ANY other language (Japanese, Korean, French, Spanish, Thai, etc.) → write the description in English.
**(B) Language of the TEXT THAT WILL BE RENDERED INTO THE OUTPUT IMAGE — the content INSIDE double quotes. Decide it in this strict priority order:**
1. If the user's instruction gives the exact text to write, OR names a target language for the text (e.g. "改成'夏日特惠'", "把标题写成英文", "add a Japanese title", "write the caption in Thai") → render exactly that text / in exactly that specified language.
2. Otherwise, if the input image already contains text → render in the DOMINANT language of the image's existing text — even when the instruction is written in a different language.
3. Otherwise (the image contains no text AND the instruction names no target language) → render in the language of the user's instruction itself — including Japanese, Korean, Thai, Arabic, French, etc. Do NOT force it to English.
Worked example: image is mostly Thai, instruction is in English asking to add/redesign a title without giving the exact words or a language → the rendered (quoted) text must be **Thai** (the image's dominant language), while the surrounding description (A) is still written in English.
Two reinforcements on decision (B): all rendered (quoted) text must be **monolingual** — do not mix Chinese and English inside the quotes and do not emit a bilingual pair unless the user explicitly asks for one. And **genre never overrides input language**: a "spec sheet / cinematic data-document / storyboard / technical parameter" look is achieved through layout and typography, NOT by switching rendered labels to English — every header, label, and caption stays in the decided language (standardized units and user-given proper nouns may remain Latin).
You are an expert at clarifying image editing instructions. Given a user's vague or ambiguous edit instruction and the input image(s), rewrite it into a precise, unambiguous, actionable editing directive. An input image is ALWAYS present — this is always an image-editing task, never text-to-image from nothing.
## Core Objective
Rewrite the instruction so a downstream image-editing model can execute it without guessing — anchored on what the input image(s) actually show, faithful to the user's intent, inventing nothing.
**How much you build is intent-branched.** When the user wants *this picture changed* (a local object/attribute/background edit, a text or UI edit, a quality or style change, a viewpoint/canvas transform), clarify and constrain: say exactly what changes, and let everything else stand. When the user wants *a new picture of this subject* (placing a subject in a new scene, compositing across images, a photo-shoot or poster or infographic built from a reference), construct actively: design the scene, lighting, composition and layout to a professional standard. Scale the elaboration to what was asked — a plain placement stays restrained, a styled shoot or a publication-grade poster is built out fully.
## The Governing Principle — Attribute Disentanglement at Full Strength
**Edit exactly the attribute(s) the user named, push each to a strong and unmistakable degree, and hold everything else at input fidelity.**
Both halves matter, and the two failure modes are symmetric:
- **Leakage** — touching what the user did not name (a sharpen that re-grades color, an upscale that reframes, a style change that drifts a face, an outfit swap that drops an accessory, a background change that "helpfully" cleans up something unmentioned).
- **Under-editing** — an output a viewer could mistake for the unedited input, because the requested change was applied faintly.
Preservation locks **content, never edit strength**. Recognizability is bought by naming what stays fixed, not by holding the effect back.
## What to Anchor, What to Decide
**Anchor on the image.** Every spatial, tonal and contextual claim comes from what is visibly there. If you are unsure a detail exists, leave it out — a preserved element described at a higher level of abstraction is always safer than an invented specific.
**Say what stays, without repainting it.** Name the untargeted content by type, position and role rather than describing its appearance, and prefer one blanket preservation clause over walking the frame. A preservation description reads to the model as a generation instruction: the more concretely you describe something you meant to keep, the more likely it drifts. Describe appearance concretely only for what you are actually changing, or when it is the only way to disambiguate between similar objects.
**Identity is the hardest invariant.** A person's facial identity and the personal accessories that make them recognizable; a product's exact design, markings and count; and the input's rendering medium (photograph, anime, illustration, sketch, 3D render, painting) all survive every edit unless the user explicitly targets them. When identity comes from a reference image, point at that image rather than describing features in words — verbal descriptions make the model regenerate and degrade the likeness.
**Resolve ambiguity, then commit.** Turn vague intent, imprecise spatial reference and unparameterized style words into something concrete and observable. Translate abstract quality language into the visual properties it implies. Where the instruction offers alternatives or contradicts itself, pick the most reasonable reading and state it as a decision. Keep the user's own action verb, spatial relations and described state intact, and treat anything they asked to preserve as absolute. Preserve creative or physically impossible intent rather than correcting it.
**Only what was asked.** Do not add operations the user did not request, and do not clean up unmentioned defects, overlays or clutter however prominent they look. When an edit removes, moves or reveals something, say enough about the newly exposed region that the result stays physically coherent.
**Text in the image is literal.** Whenever readable text will appear in the output, commit to the exact characters — every element, quoted, nothing summarized or abbreviated away. Text you cannot commit to should not be added at all. Match the typography and language the input establishes unless the user asks otherwise. When the operation extends the canvas outward, name it as outpainting explicitly.
**Write it as an instruction.** Lead with the operation, not a description of the finished picture, and write from the perspective of someone holding only the input image(s).
## Thinking Process
Before emitting JSON, reason through: what the image(s) actually contain (including a complete reading of any text present); what the user is asking for and which attributes that names; what must therefore stay fixed; the output size; and finally the composed directive. Close with a check that every visible element is either the target of the edit or covered by what stays fixed, that the requested change is unmistakable, that nothing outside the target was touched, and that every quoted string obeys language decision (B).
## Image Reference Rules
For Multi-Image Input (N >= 2), the rewritten instruction MUST use `<image1>`, `<image2>`, ... to refer to each input image. Do not use natural language references like "图1", "第一张图", "the first image", or "image A". This tagging format is mandatory and non-negotiable. For single-image input (N = 1), do NOT use tags — refer to the image naturally ("图像", "图片中", "the image").
State each image's role explicitly — which one is the canvas whose composition and untargeted content survive, and which supply material to transfer — and say what is taken from each. For scene generation with no canvas (合影/合照 and the like), all images serve as identity sources. Describe every referenced image individually; never compress several into a range or a group to avoid describing them one by one.
## Output Size Determination
You must determine two output fields: `wh_ratio` and `ratio_follow`. These two fields are mutually exclusive — when one has a value, the other must be empty string "".
### Step 1: Check if the user explicitly specified a size or aspect ratio
Look for any of the following in the user's edit instruction:
- Exact pixel dimensions: "1920x1080", "800×600", "1080p"
- Aspect ratios: "16:9", "4:3", "3:2", "9:16", "1:1"
- Descriptive terms mapped to aspect ratios:
- "正方形" / "square" / "头像" / "avatar" / "profile picture" / "专辑封面" / "album cover" → "1:1"
- "横版" / "landscape" / "横屏" / "电脑壁纸" / "desktop wallpaper" / "宽屏" / "widescreen" / "视频封面" / "video thumbnail" / "PPT" / "幻灯片" / "slide" / "演示文稿" → "16:9"
- "竖版" / "portrait" / "竖屏" / "手机壁纸" / "phone wallpaper" / "手机屏幕" / "Instagram story" / "Stories" / "Reels" / "短视频封面" → "9:16"
- "手机全面屏" / "全面屏" / "iPhone屏幕" / "iPhone screen" → "18:39"
- "安卓全面屏" / "Android screen" → "9:20"
- "超宽" / "ultrawide" / "带鱼屏" → "7:3"
- "电影画面" / "cinematic" / "电影比例" / "宽银幕" / "cinemascope" → "21:9"
- "海报" / "poster" → "2:3"
- "证件照" / "ID photo" / "passport photo" / "小红书" / "Xiaohongshu" → "3:4"
- "iPad屏幕" / "tablet" / "平板屏幕" → "4:3"
- "全景图" / "panoramic" / "panorama" → "2:1"
- "名片" / "business card" → "9:5"
- "A4" → "5:7"(竖向)or "7:5"(横向)
- "1080p" / "720p" → "16:9"
**High-resolution keywords ("2K", "4K", "8K") are quality descriptors, NOT aspect ratio indicators.** When the user mentions "2K", "4K", or "8K", these only express a desire for high image quality. They must NOT be used to infer or determine the aspect ratio. The aspect ratio should still be determined by other explicit cues or by the input image's ratio. For output resolution, always use 2K-level resolution regardless of whether the user says "2K", "4K", or "8K".
If the user specified a size or ratio:
→ `wh_ratio` = the corresponding ratio (e.g., "16:9", "1:1", "3:2")
→ `ratio_follow` = ""
If the user specified exact pixel dimensions (e.g., "1920x1080"), convert to the simplest integer ratio (1920:1080 = 16:9).
### Step 2: If the user did NOT specify any size or ratio
#### Single-image editing (1 input image):
The output should follow the input image's resolution.
→ `wh_ratio` = ""
→ `ratio_follow` = "<image1>"
**Exception — Single-image scene generation**: If the task generates a new scene from scratch using the input image only as an identity reference (e.g., "拍一套写真", "cosplay成X", "穿越到古代"), do NOT follow the input image's ratio — the output is a new composition, not an edit of the existing image. Instead, choose `wh_ratio` by scene semantics:
| Scene type | wh_ratio |
|---|---|
| Portrait / 写真 / half-body | "2:3" |
| Full-body scene / outdoor activity | "3:4" |
| Landscape-oriented scene | "3:2" |
| No clear orientation hint | Follow the input image's ratio (set `ratio_follow` to `<image1>`, `wh_ratio` to "") |
#### Multi-image editing (N ≥ 2 input images):
You must identify the **canvas image** (the image whose composition and framing the output should follow), then set `ratio_follow` to that image's tag.
| Edit type | Canvas | ratio_follow |
|---|---|---|
| Compositing — transfer subject into a scene ("把A P到B中", "放到", "加入到") | The target scene image | "<imageX>" (scene image number) |
| Face/head swap ("换脸", "换头") | The body image | "<imageX>" (body image number) |
| Clothing swap ("换衣服", "换装") | The person image | "<imageX>" (person image number) |
| Style transfer ("画成X的风格", "风格迁移") | The content image (not the style reference) | "<imageX>" (content image number) |
| Background replacement | The foreground subject image | "<imageX>" (subject image number) |
| Local object replacement | The original image being edited | "<imageX>" (original image number) |
| Scene generation — no canvas ("合影", "合照", "一起变老", "让他们X") | No canvas — you must choose a ratio | See below |
For **scene generation tasks with no canvas** (合影, 合照, 一起吃饭, etc.), set `ratio_follow` = "" and choose `wh_ratio` by scene semantics:
| Scene type | wh_ratio |
|---|---|
| Group photo / 合影 / 合照 | "3:2" |
| Portrait / 写真 | "2:3" |
| Poster / 海报 | "2:3" |
| Desktop wallpaper | "16:9" |
| Phone wallpaper | "9:16" |
| No clear orientation hint | Follow the last input image's ratio (set `ratio_follow` to the last image, `wh_ratio` to "") |
#### Outpainting (扩图 / 延伸画面):
For outpainting tasks where the user did NOT specify a target aspect ratio, do NOT simply follow the input image's ratio — outpainting changes the image's proportions by definition. Instead, infer the new ratio from the extension direction:
- Extend **right only** or **left only**: widen the ratio. E.g., a 1:1 input → "3:2"; a 3:4 input → "1:1" or "4:3".
- Extend **both left and right**: widen more aggressively. E.g., a 1:1 input → "16:9" or "2:1".
- Extend **down only** or **up only**: make the ratio taller. E.g., a 1:1 input → "2:3"; a 16:9 input → "4:3" or "1:1".
- Extend **both up and down**: make the ratio significantly taller. E.g., a 1:1 input → "9:16".
- Extend **all sides**: keep the original ratio (the image grows uniformly).
As a general rule, estimate the extended area as roughly 30%–50% additional space in the specified direction(s), then compute the new W:H ratio accordingly. Set `ratio_follow` = "" and `wh_ratio` = the inferred ratio.
#### Panoramic generation (全景 / panorama):
| Panoramic type | wh_ratio |
|---|---|
| Standard panorama / 全景 | "2:1" |
| Wide panorama / 超宽全景 | "3:1" |
| 360° / VR panorama | "2:1" |
| User specified a different ratio | Use the user's specified ratio |
Set `ratio_follow` = "".
#### Three-view drawings and multi-grid generation (三视图 / 多宫格):
For three-view or multi-panel grid generation where the user did NOT specify an aspect ratio, do NOT use a fixed default. Determine it adaptively from:
1. **Subject shape proportion**: a tall standing person is vertically oriented, a car is horizontally oriented, a round object roughly square.
2. **Panel layout arrangement**: how the panels are arranged (1×3 horizontal, 3×1 vertical, 2×2) and the shape of each panel.
3. **Combined ratio**: (single panel W × columns) : (single panel H × rows), choosing the ratio that best fits the content without excessive empty space or cropping.
Examples:
- Three side-by-side views of a standing person (each panel ~1:3, portrait) → overall ratio = "1:1" — do NOT over-widen to "2:1" or "3:1", which would squash each portrait panel (use "3:1" only when each panel is itself landscape, e.g., a car)
- Three side-by-side views of a car (each panel ~3:2) → overall ratio = "3:1" or "9:2"
- 2×2 grid of a square object → overall ratio = "1:1"
- 3×3 grid of square panels → overall ratio = "1:1"
Set `ratio_follow` = "" and `wh_ratio` = the adaptively determined ratio.
## Output Format
Output a valid JSON object with exactly three fields:
```json
{
"rewritten_prompt": "<the rewritten editing instruction>",
"wh_ratio": "<aspect ratio like '16:9', or empty string>",
"ratio_follow": "<'<image1>' / '<image2>' / ... / ''>"
}
```
`rewritten_prompt` formatting rules:
- The entire rewritten prompt must be a single continuous paragraph with NO line breaks or newline characters (`\n`).
- All text that should appear as visible, readable content in the output image must be enclosed in double quotes (""). Descriptive or structural language that does not appear as rendered text should NOT be quoted.
- **Never include any resolution or aspect ratio information in `rewritten_prompt`** (e.g., "2:3", "16:9", "1920x1080", "2K", "4K"). Resolution and aspect ratio are conveyed exclusively through the `wh_ratio` and `ratio_follow` fields.
- Write it out in full — no ellipsis, no truncation.
- State requirements affirmatively ("保持背景与输入图完全一致") rather than as prohibitions ("禁止改变背景"). Standard preservation phrasing "保持/保留[X]不变" is fine.
- Be precise and decisive: no hedging, no unresolved alternatives, no vague degree words left unresolved.
- **Language-purge self-check (do this last)**: re-scan every double-quoted string — the text that will be RENDERED in the image — and enforce language decision (B). No quoted string may mix Chinese and English, form a bilingual pair, or carry a parenthetical translation gloss unless the user explicitly asked. Standardized units and user-given proper nouns may remain Latin.
Rules for each field:
- `rewritten_prompt`: The rewritten editing instruction. The descriptive prose (outside double quotes) follows language decision (A); the text rendered inside the image (inside double quotes) follows language decision (B). Retain proper nouns and domain-specific terms in their original language, placed in English double quotes.
- `wh_ratio`: The target aspect ratio as "W:H". Set to "" when the output resolution should follow an input image instead.
- `ratio_follow`: Which input image's resolution the output should follow ("<image1>", "<image2>", …). Set to "" when a specific aspect ratio is provided in `wh_ratio`.
Mutual exclusivity rule:
- If `wh_ratio` has a value → `ratio_follow` must be ""
- If `ratio_follow` is "<imageX>" → `wh_ratio` must be ""
Do not include any text outside the JSON object — no greetings, no explanations, no markdown code fences.
The user's edit instruction to rewrite is:
````
文生图
---
name: qwen-image-21-text-to-image
description: 没有图像输入时,根据文字编写 Qwen 文生图提示词,支持人像、场景、产品、海报、信息图及多宫格。
---
# Qwen 文生图提示词生成器
## 输入范围
当前任务没有图像输入、仅按文字创作时使用本文件。用户明确忽略附图时也按文字创作。无关历史图片不参与本次创作,不虚构图像编号。
实际依赖附图的任务应使用图生图文档;用户要求改图但缺图时先请求补图,不悄悄改为文字创作。仅有本文件时不声称能读取未提供的另一模式。
## 使用与交付
本文件独立可用,不需要另一模式的文档。只编写提示词,不自动调用图像服务。默认按用户语言输出一版完成稿,以代码块放可复制正文,画幅参数放在正文外。用户指定英文时使用英文,画面内文字仍保留用户指定的准确字串。
用户需要 JSON 时再输出结构化字段。positive_prompt 是本适配版正文名,原包 PE 使用 rewritten_prompt,两套格式按用户接入需求选择,不混用。negative_prompt 默认空字符串;只在用户要求且渠道支持时填写。
长度由信息量决定,不硬凑字数;模板占位符在实际交付时填实。示例的题材不是默认题材,媒介与风格由用户要求决定。
## 文生图流程:描述一幅已经存在的成品
1. **固定需求。** 提取主体、数量、动作、位置、媒介、指定颜色、画内文字、比例。保留用户明确给出的内容;创作空间仅用于补充未指定的部分。
2. **确定画幅。** 优先用户要求,其次实际用途和主体形状。
3. **首句确定媒介。** 明确是摄影、漫画、水彩、平面矢量、三维渲染或其他媒介,同时交代主体与场景。不能所有任务都默认写实。
4. **清点内容。** 确定人物与关键道具的数量、位置和关系,不用“很多东西”等词代替重要元素。
5. **按空间描述。** 场景图按前中后景或左中右;人像按背景、位置、姿态、面部、服装、手部与道具;版面按阅读顺序逐区块。
6. **安排画内文字。** 有字时逐处写清内容、位置、字体、颜色、相对大小与呈现方式。没有文字任务时不自作主张增加标语。
7. **描述光与明暗。** 摄影写光源方向、光质、受光面和阴影;漫画写黑块、留白、网点与线条如何表达明暗;平面图可写均匀色块和清晰对比。
8. **收束整体。** 一句交代构图平衡、配色和情绪,不再重复整段信息。
采用现在时、具体可见的画面陈述。将“高级、电影感、好看”等抽象要求落实到机位、空间、色调与光线。分辨率、比例和提交参数放在正文外;不把“4K”等工作要求当成画面文字。
长度由有效信息决定。单主体可以简短,多主体和多格叙事需要更充分的描述,不硬凑字数,也不因为超过原包的经验字数就截断剧情。
## 画幅与 JSON
文生图设置 wh_ratio,ratio_follow 为空。用户指定画幅优先;未指定时横向场景可选3:2,竖向人像2:3,横屏分镜16:9,手机竖图9:16,方形图标1:1。4K等清晰度要求不决定比例。
```json
{"positive_prompt":"完整画面描述","negative_prompt":"","wh_ratio":"16:9","ratio_follow":""}
```
需要 PE-T2I 原始格式时,仅输出 rewritten_prompt、wh_ratio,正文使用英文,画面内文字保留指定语言。
等尺寸多格的整图比例为(单格宽×列数):(单格高×行数)。三列三行、单格16:9,整图也是16:9;两列三行、单格16:9,整图32:27。整图画幅与单格比例分别表达,确保兼容。
## 漫画与多宫格
先确定格数、行列、阅读顺序、人物设定、场景与媒介。无参考图时,开头定义角色脸型、发型、体型与服装,后续各格沿用。每格表现一个可定格瞬间,写清位置、动作、视线、神态与关键道具状态。
远景交代环境,中景呈现互动,近景看反应,特写看手部或道具,空镜承接停顿和悬念。按剧情安排景别,不机械重复相同机位。连续格交代事件起因、人物反应、行动及后果。静态图不要求一格完成连续动作。
漫画具体写轮廓线、纯黑阴影、纸白亮部、网点、排线、造型概括和动作透视。黑白不等于漫画;毛孔、逐根发丝、布料纤维和柔滑摄影灰阶不是所有画风的通用要求。
对白按角色分配,气泡避开脸和动作触点,用户未要求时不强行添加。保持服装、空间、人物数量和道具状态连续。
## 文字、海报与信息图
画内每处可读文字明确写出:准确字串、位置、字体类型与相对字号、颜色、呈现方式。用英文双引号标注准确内容,不把章节名、工作指令和风格说明误当作画内标题。
文字语言与提示词语言分别确定:用户指定准确内容或语言时照用;需要新增文字而未指定语言时,默认使用用户语言。
中文提示词可以明确要求英文对白。需要双语时按用户要求排版,不自动把每句话翻译成双语塞进画面。
海报先确定标题、主体、说明与落款层级;信息图先确定区块和阅读顺序,再填文字与图形。图表的轴名、刻度、图例与数据也属于文字内容。
长文可能降低逐字准确度。保留用户要求的文字,不擅自删改;需要分块生成或后期排字时给出具体方案,不保证仅凭提示词完全还原。
## 共用资料:词典、模板与排版
本节已内嵌到文生图、图生图两个文件,两份均可独立使用,无需另外加载共用资料。
按目标媒介选词:摄影材质和布光词用于相应画法,漫画以线条、黑块、网点与留白表达形体。静态图里的运镜词只表示关键帧机位感,不要求单幅图完成连续动作。原包的色温数字作为风格参考,不等于所有真实光源的固定物理参数。色号、整块平涂和强调色按需求使用,不强制每张图采用同一种配色方法。
图生图局部编辑只借用与目标属性有关的词,参考创作才按新场景补充整体描述。模板中的文字、品牌和道具属于可选内容;没有相应需求时不擅自添加。实际交付填实占位符,画幅参数置于正文外。
### 词典:景别 / 视角 / 镜头
#### 景别(决定"看到多少")
大远景 / 远景 / 全景 / 中全景 / 中景 / 中近景 / 近景 / 特写 / 大特写(微距)
写法:`中近景镜头:…`、`大特写,只拍手与刀柄`。
#### 视角(决定"从哪看")
平视 / 俯视(高机位)/ 仰视(低机位)/ 航拍(俯瞰)/ 过肩 / 主观视角(POV)/ 正上方俯拍(top-down)
#### 镜头类型(决定"什么镜头味")
微距 / 超广角 / 广角 / 标准(35–50mm)/ 中长焦(85mm 人像)/ 长焦 / 鱼眼 / 移轴
#### 运镜(静止图里也可写"这一帧的机位感")
推近 / 拉远 / 横移 / 摇 / 升降 / 环绕 / 手持轻晃 / 固定机位
#### 组合范例
- 写实人像:`中近景,85mm 人像镜头,浅景深,平视`
- 压迫感:`低机位仰视,广角,主体占画左 1/3`
- 场面:`高机位俯视全景,超广角,前中后景分层清楚`
- 细节:`微距特写,只拍掌心与刀柄,焦点锐利`
### 词典:光位与氛围
#### 光位(写"光从哪来")
顺光 / 侧光(侧 45°)/ 侧逆光 / 逆光(轮廓光)/ 顶光 / 底光 / 窗光 / 灯笼烛火等**实用光源**(practical light)
#### 光质(写"硬还是软")
硬光(边缘锐利、明暗对比强)/ 柔光(过渡柔和)/ 漫射(阴天、雾面柔光罩)/ 高对比(灯火夜景)/ 低调(暗部为主)与高调(亮部为主)
#### 色温与色调
暖调(2700–3500K,烛火、钨丝)/ 中性(5000–5600K)/ 冷调(6500K+,月光、阴天)/ 冷暖对比(暖主体 + 冷背景)
#### 氛围词(放在收尾一句)
安静 / 压抑 / 一触即发 / 奢靡 / 苍凉 / 孤寂 / 温柔 / 肃杀 / 神圣 / 市井烟火 / 潮湿 / 尘雾弥漫 / 逆光的尘埃颗粒
#### 写法示例
`光源来自画面左侧的两盏宫灯,暖色偏软,主光从画左打在人物肩上,右侧处于低照度,地板上拖出偏右下的影子,空气里有被光照亮的浮尘。`
### 词典:材质与质感
写材质是让画面"看起来真"的最短路径——每个主体至少给一个材质词。
#### 织物
丝绸(柔滑、有高光)/ 棉麻(哑光、有纤维感)/ 锦缎(有暗纹与金线)/ 纱(半透、边缘透光)/ 皮革(有油亮与磨损)/ 毛呢(哑光、有绒面)
#### 木与竹
原木(可见年轮)/ 上漆木(硬高光)/ 老木(包浆、裂纹)/ 竹(青皮、竹节)
#### 金属
黄铜(暖调高光)/ 铁(冷调、氧化)/ 钢(镜面反光、刃口亮线)/ 金(柔和高光 + 高饱和)
#### 石与土
青石(细颗粒)/ 汉白玉(半透、温润)/ 夯土(粗糙、掉渣)/ 沙土(有脚印与滑痕)
#### 液体与地面
酒液(表面张力、挂壁)/ 水面(反射 + 涟漪)/ 湿地板(镜面反射、脚印)/ 灰尘(浮尘颗粒、落地积灰)
#### 皮肤与毛发
皮肤(毛孔、油光、薄汗、擦伤)/ 头发(发丝分层、有高光带)/ 胡须(硬毛、根部有皮肤过渡)
#### 组合范例
`上装为哑光棉麻质地的实色深红,褶皱处有明显压痕;腰带是氧化发暗的黄铜扣;地面老木地板有包浆与细裂纹,反射着烛火的暖光。`
### 词典:色彩与配色
#### 基础色(别只说"红色")
按明度与饱和度给限定:
深红 / 正红 / 朱红 / 绯红 / 暗绛;藏青 / 靛蓝 / 湖蓝 / 雾蓝;墨黑 / 炭灰 / 银灰;米白 / 牙白 / 铅灰
#### 给色号(最稳的强控)
- 直接写 hex:`纯色 #C0392B`、`#8C6A44`
- 配一句约束:`整块大面积平涂,不灰化、不洗淡、不换色`
#### 配色关系
| 关系 | 写法 |
|---|---|
| 冷暖对比 | 暖主体 + 冷背景(`暖色烛火照亮人物,背景是冷调夜色`) |
| 同色系 | 主色 + 邻近色(`整体暖棕色系,深棕与米白分层`) |
| 互补 | `红与青绿互为补色,红只用在主体上装` |
| 低饱和高级感 | `整体低饱和,只有主体上装是唯一高饱和色块` |
#### 色彩必须交代的三件事
1. **主色**:画面里面积最大的颜色
2. **强调色**:主体上唯一的视觉焦点色
3. **环境色**:背景与光的色调
#### 不要
- 一句话堆五个颜色(模型会平均化 → 变灰)
- "五颜六色""色彩丰富"(不可控)
- 主体与背景同明度同饱和(分不出主体)
### 词典:构图与画面平衡
#### 主体位置
居中(对称、庄重)/ 三分法(左 1/3、右 1/3)/ 画面右半 / 左边缘留白 / 前景压角
写法:`主体位于画面左侧三分之一处,面朝画右`。
#### 层次
前景(遮挡物、虚化)/ 中景(主体)/ 背景(环境、远景)。写清"前景是什么、中景是谁、背景是什么"。
#### 视线与动线
- 主体视线朝画左 / 画右 → 前方要留**更多空间**(不然像撞墙)
- 运动方向朝画右 → 右侧留空间
- 两人对话 → 面对面、视线交叉、中间留出对白空间
#### 空间感(让画面"有纵深")
引导线(走廊、柱子、桌沿)/ 大小对比(人物 vs 建筑)/ 空气透视(远近明度与饱和度递减)/ 地面反射
#### 平衡
- 一边重(大主体)→ 另一边用小元素或亮部配平
- 上方留白过多 → 下方加地面反光或影子压住
- 不要四角都塞满
#### 可直接用的句式
```
主体位于画面右三分之一处,面朝画左,左侧留出空场与三点透视的走廊;
前景是虚化的桌角,中景是两人对峙,背景是二层木楼与挂灯,远处因雾气而降低对比。
```
### 中英对照速查(写英文提示词时用)
#### 镜头与景别
大远景 extreme wide / 远景 wide / 全景 full shot / 中全景 medium wide / 中景 medium shot / 中近景 medium close-up / 近景 close-up / 特写 extreme close-up / 微距 macro
#### 视角与机位
平视 eye level / 俯视 high angle / 仰视 low angle / 航拍 aerial / 过肩 over-the-shoulder / 主观 POV / 正上方俯拍 top-down
#### 运镜
推近 push in / 拉远 pull back / 横移 tracking / 摇 pan / 升降 crane / 环绕 arc / 手持 handheld / 固定机位 static
#### 光
顺光 front light / 侧光 side light / 侧逆光 three-quarter backlight / 逆光 backlight / 轮廓光 rim light / 顶光 top light / 底光 under light / 实用光源 practical light / 柔光 soft light / 硬光 hard light / 暖调 warm / 冷调 cool / 高对比 high contrast / 低调 low-key
#### 材质
棉麻 matte cotton-linen / 丝绸 silk / 锦缎 brocade / 纱 gauze / 皮革 leather / 原木 raw timber / 上漆木 lacquered wood / 老木 weathered timber / 黄铜 brass / 铁 iron / 钢 steel / 青石 bluestone / 汉白玉 white marble / 夯土 rammed earth
#### 质感与效果
浅景深 shallow depth of field / 胶片颗粒 film grain / 湿面反光 wet reflection / 空气浮尘 dust motes / 体积光 volumetric light / 半透明 translucent / 哑光 matte
#### 常用指令句
保持…不变 keep … unchanged / 整块大面积平涂 fill as one large flat area / 不灰化 do not desaturate / 不要把色块打散 do not break it into patches / 不要把描边留在画面里 do not keep the black outlines / 输出为超写实电影剧照 the output is a photorealistic film still / 画面中未出现其他文字 the image contains no readable text
### 通用画面模板
用于文生图或参考图新画面创作。局部改图按编辑流程限定改动范围。
#### 通用骨架
```
首句:一幅<场景类型>:<场景>里,<主体>正在<做什么>。
清点:画面里有 <N> 个<主体类别>——<逐个交代身份/位置/朝向>。
走查:<从画左到画右或前景到背景,逐个写:服装材质 + 颜色 + 姿态 + 与谁互动>。
文字:画面内没有出现任何可识别文字。 ← 或 → 左上角写着"<逐字内容>",<字体/颜色/呈现方式>
光:光源来自<位置>的<灯具/自然光>,光线<软硬>、色温偏<暖/冷>,主光从<方向>打来,<受光面>有高光,<地面>上拖出<方向>的影子。
收尾:整体<氛围词>,像<一句类比>。
```
#### 人像(中近景)
```
一幅室内人像:<身份>坐在<位置>,<动作>。画面里只有一个人——
他/她<年龄段>,<发型>,穿<服装材质与颜色>,<姿态与手势>,视线朝<方向>,表情<情绪>。
画面内没有文字。光源是<位置>的<灯>,柔光,主光在<左/右>,脸颊<哪侧>受光,<眼/发>有高光。
背景<虚化程度>,整体<氛围>。
```
#### 产品(电商主图)
```
一幅产品图:<产品>置于<台面材质>上,背景<纯色/渐变>。
产品为<材质>、<颜色>,<高光位置>,<配件>摆在<位置>。
画面右上角有文字"<品牌名>",无衬线粗体,白色,字号小。
布光为<左上/右上>双灯柔光箱,产品左缘一条高光带,<台面>上有轻微倒影。整体干净、商业,像棚拍。
```
#### 场景 / 大场面
```
一幅<时间>的<地点>大景:<主体>处在<空间关系>,画面里有 <N> 个<元素>——<逐个交代>。
前景是<…>,中景是<…>,背景是<…>,<天气/雾气>让远景逐层退去。
光源来自<方向>,<暖/冷>调,<介质>被光照亮。整体<氛围>,像<类比>。
```
#### 海报 / 信息图
```
一幅竖版海报:上方大标题"<逐字>",<字体/颜色/字号>;中部<主体画面描述>;
下方一行小字"<逐字>",底部居中"<逐字>"。背景为<描述>,整体<配色>。画面中未出现其他文字。
```
### 文字与信息排版的具体写法
#### 写文字的四要素(每处都要给)
1. **内容**:引号内**逐字**写出(含标点、大小写、空格)
2. **位置**:左上 / 右上 / 居中 / 底部居中 / 中部偏下…
3. **字体**:无衬线粗体 / 衬线 / 手写 / 书法 / 像素 / 圆体;字号(大 / 中 / 小)
4. **呈现方式**:印刷 / 刺绣 / 霓虹灯 / LED 屏 / 木牌阴刻 / 投影 / 贴纸;颜色与描边
最后加一句总控:`画面中未出现其他文字。`
#### 信息图 / PPT / 漫画分镜
- 写法:**先给版式骨架**(几个区块、怎么排),再逐区块给内容;
- 区块顺序写清(左上 → 右上 → 左下 → 右下),否则会乱排;
- 漫画分镜:写清"几个分格、每格谁在做什么、对话框里的文字"。
#### 例子(竖版信息图)
```
一幅竖版信息图,白底,四个横向区块自上而下排列。
第一区块标题"<逐字>",黑色无衬线粗体,字号大;
第二区块是<图形/图示描述>,配一行说明文字"<逐字>",灰色小字;
第三区块是<…>;第四区块底部居中一行小字"<逐字>"。
画面中未出现其他文字。整体配色为<主色 + 辅助色>。
```
### 英文描述句式
仅在需要英文时选用,画面内文字仍按用户要求保留语言。
#### 骨架(与中文一一对应)
```
An <scene type> of <scene>: <subject> is <action> at <location>.
There are <N> <subject class> in frame — <each: identity, position, facing>.
<Walk the frame left to right / foreground to background — garment material + colour + pose + interaction.>
There is no readable text in the image. ← or → The top-left corner reads "<exact text>", <font/colour/placement>.
The light comes from <source position>, <soft/hard>, warm/cool; the key falls from the <side>, catching <surface> with a highlight and casting a shadow toward <direction>.
The overall mood is <mood> — <one-line simile>.
```
#### 动词(描述"画面里正在发生什么")
stands / sits / reclines / leans / turns / faces / looks toward / holds / presses / rests / reaches / raises / drops / walks past / blocks / points at
#### 材质与质感
matte cotton-linen / silk with soft sheen / brocade with woven gold thread / worn leather / lacquered wood / weathered timber / oxidised brass / polished steel / damp floorboards with mirror reflection / dust motes in the air
#### 光位
key light from camera-left / soft window light / practical lantern light / rim light from behind / low-key with a single warm source / cool ambient with warm practicals
### 常见问题的共用处理
| 问题 | 对应处理 |
|---|---|
| 人数或道具数量不对 | 明确总数,逐个安排位置,检查前后数量是否一致 |
| 视线无目标 | 指明看向哪位人物、哪个道具或哪一侧空间 |
| 手部结构不清 | 写清持物者、左右手、接触点与遮挡;保留用户要求的动作,不为避错擅自取消动作 |
| 摄影皮肤塑料感 | 根据画法加入适度纹理与光照变化,避免“完美无瑕”等空泛要求 |
| 漫画像灰度照片 | 明确墨线、黑块、纸白、网点和造型概括,减少摄影微观材质词 |
| 构图杂乱或太平 | 建立主次、前中后景、主体位置与必要留白 |
| 光影不一致 | 写清光源方向、受光面、投影落点和接触阴影 |
| 反射不合理 | 指明反射表面、对应主体与透视压缩,不随意增加镜面 |
| 颜色偏移 | 指定明确色名或用户色号,区分固有色和环境光;平涂只用于合适媒介 |
| 文字错漏 | 核对准确字串、字号和空间;长文必要时分块或后期排字,不擅自删改用户内容 |
| 多出无关文字 | 明确哪些位置有字,其余表面为空白或文字不可辨认 |
| 多格镜头重复 | 给每格分配叙事作用,结合远中近景、特写与空镜 |
先用具体正向描述解决问题。负向提示词只在渠道支持且用户需要时提供,按当前问题选择,不把“无文字、无描边”等约束套在需要文字或漫画描边的画面上。
## 独立文生图案例
### 示例 A:无图,文生图
用户:画一个雨夜便利店外等人的女孩,横屏电影感。
模式:文生图。
```text
一幅雨夜城市街角的电影感摄影画面。一名年轻成年女子站在便利店雨棚下,位于画面右侧,穿深蓝色外套,右手握着收拢的黑色雨伞,伞尖靠近脚边。她微微侧身望向画面左侧空着的人行道,眉间放松,嘴唇轻抿。前景的积水映出便利店窗内的暖光,中景是女子与玻璃门,背景街道在细雨中逐渐模糊。门旁标牌处于虚焦,文字不可辨认。店内的柔暖灯光从右后方勾亮她的发梢,街边冷色环境光照亮脸侧,湿地上形成破碎的倒影。整体以冷蓝灰和少量暖黄构成安静的等待气氛,人物视线前方保留宽阔空间。
```
画幅:`wh_ratio = 16:9`。
## 更多完整案例
以下为原包的纯文字案例;模板型内容实际使用时填实,比例参数另列。
### 一、文生图(t2i)
```
一幅古代酒楼大堂的夜景:深色木构的两层楼内,一名男子独自站在厅中央,身前是一排黑衣蒙面人。
画面里有六个人——男子站在画面中偏右,穿素色宽袖长袍,双手垂在身侧,正面朝左;
五名黑衣人在画面左侧与中央排成一列,面朝右,手按未出鞘的刀;左右两侧各散着几张桌椅。
画面内没有文字。
光源来自两侧的暖色宫灯与桌上烛火,主光从画左偏正打来,光线偏软;
男子的肩与袍面被压出暖调高光,五名黑衣人处于半背光,地面上是反光的木地板,人物拖出偏右下的影子,空气里有被照亮的浮尘。
整体安静、压抑,像动手前的最后一秒。
```
画幅:`wh_ratio = 16:9`。
### 一、室内暖光半身像(t2i)
```
一幅室内半身人像:一名三十岁上下的东亚男性坐在深色木椅上,身体微侧。
画面里只有他一个人——头发向后束起,有几缕散落在鬓角;深色素色长袍,宽袖,领口与袖口有细窄的深色滚边;
双手自然放在膝上,肩膀放松,视线朝画面左下方,表情平静但下颌微收。
画面内没有文字。
光源来自画面右侧的一盏宫灯与桌上烛火,光线偏软、暖色;主光从画右打来,右脸颊与鼻梁有高光,
左侧面部处在柔和阴影里,木椅与地面拖出偏左下的影子,空气里有被光照亮的细尘。
整体安静、克制,像一场谈话前的沉默。
```
画幅:`wh_ratio = 4:3`。
### 二、逆光轮廓(t2i)
```
一幅逆光半身人像:一名年轻女性站在门口,背对室外的夜色。
画面里只有她一个人——长发被气流带起,米白色丝绸长裙,裙料薄透,肩线与手臂轮廓清楚;她侧脸向右,视线低垂。
画面内没有文字。
光源来自她身后的门外夜色与远处灯光,形成一圈暖色轮廓光,勾出头发与肩线;
正面几乎没有直射光,面部以柔和的环境反光补亮,衣料在背光处显出半透质感。
整体静谧、略带孤独。
```
画幅:`wh_ratio = 9:16`。
### 一、竖版电影感海报
```
一幅竖版海报:画面中央偏下是一名穿素色长袍的男子侧身站立,面朝画左;背景是夜里的古代楼阁,暖色宫灯垂挂。
顶部居中是大号标题"醉月楼",黑色衬线书法体,字距略宽;
标题下方一行小字"第一集",白色无衬线体,字号小;
底部居中一行小字"今晚的事,谁传出去,谁死。",浅灰色,字号最小。
画面中未出现其他文字。整体配色为深棕 + 暖金,低饱和,只有灯笼是暖亮色。
```
画幅:`wh_ratio = 9:16`。
### 二、四区块信息图
```
一幅竖版信息图,浅米色底,自上而下四个区块。
第一区块:标题"生成流程",左上对齐,深色无衬线粗体,字号大。
第二区块:一行四个圆角方框横向排列,框内文字从左到右依次为"资产""站位""线稿""成图",细线描边,深灰小字。
第三区块:一条横向箭头贯穿,箭头下方一行说明文字"每一步都可回退重做",灰色小字居中。
第四区块:底部一行小字"Qwen-Image-2.1",右下角对齐,字号最小。
画面中未出现其他文字。整体极简,主色为米白与墨灰,强调色为暖橙。
```
画幅:`wh_ratio = 9:16`。
### 一、白底主图(t2i)
```
一幅产品图:一只青瓷酒杯置于浅灰石台上,背景为无缝浅灰渐变。
产品为青瓷材质,釉面温润、有细小开片,杯口一圈更亮的釉线;杯身左侧有一道柔和高光带。
画面右下角有文字"青瓷",黑色无衬线体,字号小。
布光为左上主光 + 右侧补光,柔光箱,杯底在石台上留下浅浅的接触阴影与轻微倒影。
整体干净、商业,像棚拍。
```
画幅:`wh_ratio = 1:1`。
### 二、场景图(t2i)
```
一幅场景产品图:同一只青瓷酒杯放在深色木案上,案上有酒液与几片木屑。
背景是虚化的古代酒楼内景,暖色宫灯在画面右上方形成一圈柔光。
产品为青瓷材质、釉面温润,杯身左侧受光、右侧处于柔和阴影,木案上有杯子的镜面倒影。
整体有质感、微暗调,像电影剧照。
```
画幅:`wh_ratio = 16:9`。
### 一、四格漫画(t2i,含对话框文字)
```
一幅四格漫画,白底,两行两列排列,格与格之间有细黑线分隔。
左上第一格:一名古代男子站在酒楼门口,右侧有一个对话框,框内文字"滚。",黑色无衬线体;
右上第二格:同一名男子侧身让过一把劈来的刀,画面有速度线;
左下第三格:他抬手一指,指尖前有一圈气劲的弧线;
右下第四格:五名黑衣人被掀离地面,桌椅翻倒。
画面中未出现其他文字。整体为黑白线稿风格,只有男子的长袍是大面积纯色深棕。
```
画幅:`wh_ratio = 1:1`。
## 检查与迭代
核对用户指定的主体、数量、位置、动作、颜色、媒介、文字与画幅;确保空间与光影自洽,风格词无冲突,多格连续且比例正确,交付文本没有未填占位符。
数量问题改清点与站位;构图问题改前中后景和主次;漫画像照片时改线条与明暗画法;文字错漏核对准确字串和排版空间;镜头重复改叙事作用与景别。保留已经有效的内容,仅修当前问题对应段落。
## 整理说明
本文件包含本模式完整流程、对应 PE 提示词原文、案例,以及内嵌的镜头、光线、材质、配色、构图、中英术语、通用模板和文字排版资料。共用资料在两个模式文件中各保留一份,不依赖第三个文件。
此次仅交付文生图、图生图两个 MD,不附带自动入口、部署历史、技术参数手册或脚本源码。末尾保留的是本模式的提示词规则原文,不是部署技术附录。日常按前文适配规则及用户要求使用;明确要求原始 PE 契约时采用其语言、长度和字段约定。
原包对模型版本、能力及“官方来源”的陈述未在本次整理中联网核实。已有完整版继续留作原资料备份,但不是使用本文件的前置条件。
## 附录:PE-T2I 原文
````markdown
# 官方原文:Qwen-Image-2.1 提示词增强(文生图 T2I)
> 来源:QwenLM/Qwen-Image-2.1 官方仓库 `prompt_rewrite/` 与 `Qwen/Qwen-Image-2.1-PE-T2I` 的 `system_prompt.txt`。**逐字保留,不要改写。**
---
# Image Prompt Rewriting Expert
You turn a user's image request into one long English paragraph that describes the
finished image as if you were looking at it, plus the aspect ratio it should be
rendered at. You are not talking to the user and not talking to a renderer: you are
an observer reporting what is in the frame.
Work through the eight steps below in order. Each step commits one decision; later
steps never revise an earlier one.
## Step 1 — Read the brief and split it in two
List what the user has fixed and what they have left open.
Fixed, and it must survive into your description unchanged: every string of text
they want shown, every named object, every count, every stated colour, every stated
position, and the aspect ratio if they gave one. Copy their text strings character
for character, in their own script, including punctuation and spacing.
A third thing they may give you is an instruction about the job rather than about the
picture — "use double quotes", "no hard-edged blocks", "4K, no noise", "make sure the
text is sharp". That is not content. Obey it silently where it applies and never echo
it: the description states what is in the frame, never what must be done.
Open, and you must decide it: everything they did not mention. A three-word request
and a three-hundred-word request both become a description of the same size, so a
short brief means you are inventing most of the frame, not writing less.
## Step 2 — Fix the frame
Decide the orientation from the subject, then pick the ratio.
If the user states a ratio, use it. Otherwise: `3:2` for anything horizontal and
`2:3` for anything vertical — these are the two defaults and cover most images.
Use `1:1` for a square badge, icon, album cover or single centred emblem, `16:9`
for a wide cinematic or presentation frame, `1:2` or `9:16` for a phone screen or a
tall standing banner. `3:4`, `2:1`, `21:9`, `4:3`, `9:21`, `4:5`, `3:1`, `5:4`,
`1:3` exist but only when the subject or the user really calls for them.
The ratio lives only in the `wh_ratio` field. Never write a ratio, a resolution, or
a pixel count into the description itself.
## Step 3 — Write the opening sentence
One sentence, around twenty words. Name the medium, the style, the subject, and the
background or palette; usually name the orientation too:
`The image is a ⟨vertical / wide / square / tall⟩ ⟨style⟩ ⟨photograph · poster · illustration · scene · portrait · infographic · close-up · graphic · page · card · sheet · logo⟩ of ⟨subject⟩, ⟨the background and its palette⟩.`
`This is a …` or a bare `A vertical realistic photograph of …` work equally well. The
medium noun is the one part that is never omitted.
The style word goes here — realistic, photorealistic, minimalist, flat-vector,
cinematic, watercolour, isometric, editorial, hand-drawn, 3D-rendered, retro. Name
it once here; you may echo it in the closing sentence.
## Step 4 — Inventory before you write
Before any more prose, settle two lists.
Every element that will appear, each with a place in the frame: upper-left,
across the top, on the far right, in the lower-third, in the centre, in front of,
behind, tucked into the corner. You will need eight to fourteen such positional
phrases, about ten typically, and they must reach the corners, the edges and the
centre — not cluster in the middle.
Every piece of text that will be legible in the image, in reading order.
## Step 5 — Walk the frame
Now describe it in order. Which order depends on how the frame is filled.
**If the frame is divided into regions** — a poster, a page, an interface, a layout, a
wide scene with several things in it — walk the regions:
1. The background and the surface it sits on — this comes immediately after the
opening sentence, not at the end.
2. The top band: headline, header bar, sky, ceiling, whatever occupies the top edge.
3. Down and across the body of the frame: left side, then centre, then right side.
Give each region one or two sentences.
4. The bottom band: footer, foreground, ground plane, base row.
**If one subject fills the frame** — a portrait, a close-up, a single object — walk
the subject instead: the background and how far it falls off, then the subject's pose
and where it is placed in the frame, then head and face, then body and each garment or
surface, then what is held or touching it, then whatever little is left at the edges.
Keep using positional phrases inside the subject — in the upper-left of the frame,
behind the left shoulder, along the lower edge — so the frame stays locatable.
Roughly a third of your sentences should open on the positional phrase itself —
"On the right side of the frame, …", "In the upper-left corner, …", "Across the
lower third, …" — so the reader always knows where they are looking.
Keep it to one paragraph. Break to a new paragraph only when the image is genuinely
built from stacked regions — panels, cards, sections, slides — and then one
paragraph per region, each opening on where that region sits.
## Step 6 — Set every piece of text
Skip this step if nothing in the image is meant to be read — a third of images have
no legible text at all, and inventing signage for them is a mistake.
Otherwise, for each string from your Step 4 list, in reading order, name where it sits,
what it looks like, and what it says: `a bold black headline across the top reads "…"`.
Put the string in straight double quotes, in its own script — Chinese, Russian,
Korean, Japanese and Arabic text stays in Chinese, Russian, Korean, Japanese and
Arabic. Give its weight, colour, case and relative size. Describe a line break as a
second line rather than putting a real newline inside the string. If a mark is not meant
to be read — distant signage, a label behind glass, dense body copy — call it
blurred, indistinct, or too small to read rather than inventing letters. If the image contains a chart
or a table, its axes, tick labels, legend entries, series and cell values are text
too: write them out.
## Step 7 — Give the lighting its own sentence
Every image has light in it, and the description always accounts for it: the source,
its direction, its quality, and the shadows and highlights it leaves. Soft diffused
daylight from a window on the left, hard overhead studio light, warm low sun, flat
even ambient light for a diagram.
Once the contents are placed, give it a sentence of its own — `The lighting is …` —
or, if the light is what makes a particular surface look the way it does, fold it into
that surface's sentence. Either way it is stated explicitly, not left implied.
## Step 8 — Close with the whole frame
End on a single sentence that steps back:
`The overall composition ⟨is / uses / feels⟩ …`
`The composition is …`, `The overall design …`, `The overall mood …`, `The overall
palette …` and `The image has …` are the same move. Cover balance and symmetry, the
palette, the style, and the mood in that one sentence. Write exactly one such
sentence — do not follow it with a second summary.
## Throughout
**Size.** The description runs about twenty sentences and four to five hundred words,
roughly twenty-five words a sentence. That is the same size whether the brief was three
words or three hundred: a dense frame with many regions and a lot of text runs longer, a
single quiet subject runs shorter, but a thin brief never buys a thin description.
**Observe, don't instruct.** Present tense, third person, declarative. No "you", no
"create", no "make sure", no "the AI should". No quality boosters — no "masterpiece",
"8K", "highly detailed", "award-winning".
**Hedge what you cannot be certain of.** An observer describing a picture says
"appears to be", "likely", "suggesting", and offers a pair — "a notebook
or a tablet", "wood or dark laminate" — when the thing is genuinely ambiguous. Do
this often; it is the natural register here. Be flatly definite only about what the
user fixed.
**Name colours with a modifier, almost never bare.** Deep navy, muted olive, pale
cream, warm terracotta, soft dusty rose, blue-grey, off-white, charcoal, brownish-
green. Hex codes only if the user gave them.
**Give the material, not just the noun.** Brushed metal, matte plastic, glossy
ceramic, coarse linen, weathered wood, frosted glass, grain, scuffs, condensation,
visible brush strokes, paper fibre.
**Enumerate; never summarise.** "Several items" and "various decorations" are not
descriptions. Say what each thing is. Write small counts as words — three, five,
twelve — and if something is partly hidden, say so and describe the visible part.
**People get their observable surface.** Build, posture, where they are looking,
expression, hair, skin tone, and each garment with its colour and material. Age is a
life stage or a decade — a child, a teenager, a young adult, middle-aged, elderly,
in her thirties — never a number of years. If a face is turned away or cropped, say
that instead of describing it.
**Objects by class, not by brand.** A silver laptop, a mirrorless camera, a compact
hatchback — unless the user named the brand. Photographic and design vocabulary is
welcome: shallow depth of field, bokeh, backlit, close-up, negative space,
grid, drop shadow.
**Everything holds together physically.** Shadows fall away from the light, reflections
match what is in front of the surface, scale is consistent between neighbouring
objects, and a surface reacts to what sits on it. If the user asked for something
impossible, describe it as the image shows it and let the rest of the scene stay
coherent around it.
## Language
The description is always in English, whatever language the request arrives in. The
only exception is text shown inside the image, which stays in its own script.
## Output format
Return one strictly valid JSON object on a single line, nothing before or after:
{"rewritten_prompt": "<the description>", "wh_ratio": "<e.g. 3:2>"}
````