JoyAI-Image空间编辑完全指南:物体移动、旋转与镜头控制3类提示词模板详解
【免费下载链接】JoyAI-ImageJoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.项目地址: https://gitcode.com/gh_mirrors/jo/JoyAI-Image
JoyAI-Image 是一个开源的统一多模态基础模型,集图像理解、文生图生成与指令式图像编辑于一体。它的空间编辑(Spatial Editing)能力是最大亮点:只需一句话提示词,就能让图中物体精确移动、按指定方位旋转,甚至像操作真实相机一样旋转、俯仰、缩放镜头——全程无需手动修图。本文将带你完整掌握 JoyAI-Image 空间编辑的3 类提示词模板:物体移动(Object Move)、物体旋转(Object Rotation)与镜头控制(Camera Control)。
30秒看懂 JoyAI-Image 空间编辑
JoyAI-Image 由一个 8B 多模态大语言模型(MLLM)和一个 16B 多模态扩散 Transformer(MMDiT)组成,核心理念是理解、生成、编辑的闭环协作:更强的空间理解让编辑指令执行得更精准,而视角变换等生成能力又反过来辅助空间推理。
与常见的"改颜色、换背景"式编辑不同,空间编辑操作的是三维空间关系:
- 📦物体移动:把目标物体移动到图中指定的红框区域
- 🔄物体旋转:让物体转到指定方位(正面、左后侧等 8 种视角)
- 📷镜头控制:保持 3D 场景不变,仅改变相机的偏航角、俯仰角和缩放
快速开始:3步搭好空间编辑环境
环境要求:Python ≥ 3.10、CUDA 显卡。
第 1 步:克隆仓库
git clone https://gitcode.com/gh_mirrors/jo/JoyAI-Image cd JoyAI-Image第 2 步:创建环境并安装(详见 README.md 的 Quick Start 章节)
conda create -n joyai python=3.10 -y conda activate joyai pip install -e .第 3 步:下载模型权重(JoyAI-Image-Edit,可从 Hugging Face / ModelScope 获取),然后运行 inference.py 执行空间编辑:
python inference.py \ --ckpt-root /path/to/ckpts \ --image test_images/test_1.jpg \ --prompt "Move the apple into the red box and finally remove the red box." \ --output outputs/result.png \ --steps 30 --guidance-scale 4.0💡 如果你习惯图形化工作流,仓库内置的 joyai_image_comfyui/ 目录提供了 ComfyUI 自定义节点,拖拽即可跑空间编辑。
类型一:Object Move —— 物体移动提示词模板
适用场景:想图中某个物体移动到指定的红框区域(红框是你预先画在输入图上的目标位置标注)。
模板:
Move the <object> into the red box and finally remove the red box.三条使用规则:
<object>替换为目标物体的清晰描述,如apple、the red car;- 红框代表目标位置——编辑前需在输入图上画出红框;
- 句尾的
finally remove the red box表示最终成图里不要出现引导框,务必保留这句。
官方示例:
Move the apple into the red box and finally remove the red box.下面是仓库assets/application/moving/中的真实效果:输入图中散落的橙子被整体移动到了人物手部位置,其余画面保持不变。
编辑前输入图:
物体移动后输出图:
动态过程(注意背景与人物姿态的高保真保持):
类型二:Object Rotation —— 物体旋转提示词模板
适用场景:让物体旋转到某个标准视角(canonical view),常用于产品展示、3D 数据增广。
模板:
Rotate the <object> to show the <view> side view.支持的 8 种<view>取值(只能从里面选):
front, right, left, rear, front right, front left, rear right, rear left使用规则:
<object>替换为要旋转的物体描述;<view>必须是上面 8 个方向词之一;- 该指令只改变物体朝向,物体身份和周围场景尽量保持一致。
官方示例:
Rotate the chair to show the front side view. Rotate the car to show the rear left side view.官方演示中,红色跑车从正面视角旋转为左前斜侧视角,车漆反光、轮胎细节和背景玻璃幕墙都保持了高度一致:
旋转前输入图:
旋转后输出图:
类型三:Camera Control —— 镜头控制提示词模板
适用场景:不改动场景本身,只改变相机视角——偏航(Yaw)、俯仰(Pitch)和缩放(Zoom),相当于给模型下"运镜"指令。
模板(注意是四行结构,请尽量原样套用):
Move the camera. - Camera rotation: Yaw {y_rotation}°, Pitch {p_rotation}°. - Camera zoom: in/out/unchanged. - Keep the 3D scene static; only change the viewpoint.填写规则:
| 占位符 | 说明 | 示例 |
|---|---|---|
{y_rotation} | 偏航角,单位为度(正负均可) | 45、-90 |
{p_rotation} | 俯仰角,单位为度 | 20、-30 |
Camera zoom | 只能填in、out、unchanged三选一 | in |
⚠️ 最后一行
Keep the 3D scene static; only change the viewpoint.是关键约束:它明确告诉模型保持 3D 场景内容不变,只调整视角。删掉这句,模型可能会"顺便"改动画面内容。
官方示例:
Move the camera. - Camera rotation: Yaw 45°, Pitch 0°. - Camera zoom: in. - Keep the 3D scene static; only change the viewpoint.Move the camera. - Camera rotation: Yaw -90°, Pitch 20°. - Camera zoom: unchanged. - Keep the 3D scene static; only change the viewpoint.官方演示中,镜头从平视旋转俯转为近 45° 俯视,人物表情、滑板文字和人群布局都未失真:
运镜前输入图:
运镜后输出图:
进阶玩法:空间编辑还能做什么?
掌握 3 类模板后,你还可以把空间编辑当作"空间数据合成器":
- 3D 重建增广:对单视角点云/场景,用镜头控制合成更丰富的观察视角,补全稀疏输入(参考 assets/application/3dpoint/ 中的点云演示);
- 视频首尾帧生成:用物体旋转模板生成视频尾帧,再交给视频生成模型补间,实现平滑旋转转场(见 assets/application/rotation/ 的演示);
- 辅助空间推理:通过高保真的新视角消歧"谁在谁上方"这类空间关系问题。
空间编辑提示词速查表
| 类型 | 模板 | 关键约束 |
|---|---|---|
| 物体移动 | Move the <object> into the red box and finally remove the red box. | 图中需有红框;保留 remove 句 |
| 物体旋转 | Rotate the <object> to show the <view> side view. | <view>限 8 个方向词 |
| 镜头控制 | Move the camera.+ Yaw/Pitch/zoom 三行 | 保留 scene static 末行 |
效果优化小贴士:
- 提示词尽量贴近模板原文,官方说明模板遵循度越高、行为越稳定;
- 推理参数建议
--steps 30、--guidance-scale 4.0,输出分辨率用--basesize 1024; - 简单指令不满意时,可加
--rewrite-prompt让 LLM 先自动改写提示词(inference.py 支持)。
从"移动一颗苹果"到"给场景运镜",JoyAI-Image 用三套简洁的提示词模板,把三维空间操作变成了人人可用的能力。现在就挑一张图,试试你的第一条空间编辑提示词吧 🚀
【免费下载链接】JoyAI-ImageJoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.项目地址: https://gitcode.com/gh_mirrors/jo/JoyAI-Image
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考