SMF·FLAIR-AI
Can a prompt replace Flair AI?
AI image generation — AI image and video generation
Exhibit tracking slip
Verdict
Flair's difference from every other product-photo tool is spatial: you arrange the cutout and the props where you want them, and the model fills in the rest respecting that layout. Locally that is inpainting with a composition mask, and ComfyUI plus a good SDXL checkpoint gets you there in a sitting. The gap is control fidelity — the paid product's models were tuned for this task, and a general checkpoint produces a scene that ignores your arrangement more often than it respects it.
Exhibit A — The prompt
Received on31.07.2026Build a local product-scene composer: arrange elements on a canvas, then generate the scene around them with inpainting.
Canvas: a fixed-aspect working area. The user places a product cutout (PNG with alpha, or an image put through a local segmentation model), plus simple prop placeholders — rectangles and ellipses they position and label with a word like "stone plinth", "linen cloth", "leaves". Objects can be moved, scaled and rotated; the arrangement is the input, not decoration.
Mask construction, which is the technical core. From the arrangement, build:
- A keep mask covering the product cutout's alpha, dilated by a couple of pixels. Nothing inside it may be regenerated.
- A guide image where each labelled placeholder is filled with a flat colour, so the diffusion pass has structure to follow.
- A generate mask covering everything else.
Generation: submit to a local ComfyUI instance configured in the environment. Run an inpainting workflow that takes the composed canvas, the generate mask, and a prompt assembled from the scene description plus each placeholder's label positioned by its location in the frame ("stone plinth in the lower left"). Use a depth or edge control input derived from the guide image so the placeholders influence composition rather than being ignored.
Be honest about adherence: after each generation, show the result with the original arrangement overlaid as an outline so the user can see how far the model drifted from their layout. Do not hide the drift behind a nice crop.
Iteration: seed control, a variant grid of four at once, and a rerun-with-adjustment action that keeps the arrangement and changes only the prompt or seed. Store the full workflow JSON and every parameter beside each output, so a good result can be reproduced exactly.
Lighting: a simple relight pass — the user picks a light direction and strength, and the pipeline applies a gradient to the composited product so it matches the generated scene. Without this the product reads as pasted on, which is the giveaway that separates a usable product shot from an obvious composite.
Export: full resolution PNG and JPEG, plus a marketplace-square crop.
Out of scope: any hosted API, model training or fine-tuning, an account system, and a template marketplace.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 models tuned specifically for product scenes, which is most of the quality gap
- B.2 GPU speed — an iteration is seconds hosted and a minute locally
- B.3 the template library of prepared scene compositions
- B.4 reliable adherence to the layout you arranged
Prior art
Exhibit C — Why people still pay: frontier models, compute, and data
Because product photography is iterative, and a tool that returns four options in ten seconds gets used until something works. One that takes a minute per attempt gets abandoned on the third try.
Questions
Can I bring anything over from Flair?
Only the generated images. Flair's scene templates and brand kits are platform assets with no export, so compositions get rebuilt here. Product cutouts are yours and transfer directly.
How closely does the model follow my arrangement?
Less closely than the paid product, and that is the honest headline. A general SDXL checkpoint with a depth control input respects rough placement and ignores fine detail, so a prop you positioned precisely may end up elsewhere. The overlay comparison in the prompt exists so you can see this rather than discover it after publishing.
What does it cost to run?
Nothing per image once the models are downloaded, but you need a GPU with at least 8 GB of VRAM for SDXL inpainting at a usable speed. On CPU it works and takes several minutes per generation, which kills the iteration loop this tool depends on.
What is the one thing that does not survive the rebuild?
Iteration speed. Product photography is a numbers game — you generate twenty and keep one. Ten seconds per attempt makes that pleasant; a minute per attempt makes you settle for the third result instead of the twentieth.
Related tools
Receipt