Black Forest Labs' FLUX 3 Image Lets Developers Place Objects With Exact Coordinates
Black Forest Labs ships FLUX 3 Image with bounding box layout control, multi-reference composition, native 4K output, and pixel-perfect multi-turn editing.
- Black Forest Labs released FLUX 3 Image, their multimodal model's image generation and editing component
- Bounding boxes let you place every element on a 0-1000 canvas grid with ids and descriptions
- Multi-turn edits preserve untouched pixels exactly, enabling iterative refinement without drift
- Compose from up to 10 reference images cited inline as ref_image_0 through ref_image_9
- Native 2K and 4K output up to roughly 5456x3072 pixels preserves text and texture detail
- 50% off via API until October 8; commercial weights license available, open weights coming
Black Forest Labs has released FLUX 3 Image, the image component of its multimodal FLUX 3 family. Developers can define a scene with bounding boxes, element descriptions, and a caption, giving the model explicit instructions about where each object belongs.
Text-to-image systems usually infer composition from prose, which makes precise layouts difficult to reproduce. FLUX 3 Image turns layout into a structured input. That approach targets magazine covers, product composites, collages, and editing workflows where position and spacing matter as much as visual style.
Coordinates replace prompt wrangling
FLUX 3 Image maps every canvas to a normalized 0-to-1000 grid, independent of aspect ratio or pixel dimensions. Each element receives an ID, a description, and a box formatted as [y_min, x_min, y_max, x_max]. The first and third values define its vertical span; the second and fourth define its horizontal span.
A layout request combines the element table with a scene-level caption:
[
{
"id": "dome_1",
"bbox": [250, 150, 650, 850],
"desc": "a massive, smooth parabolic dome of pale concrete"
},
{
"id": "swimmers_1",
"bbox": [580, 200, 720, 800],
"desc": "dozens of small, silhouetted figures wading in dark water"
},
{
"id": "crowd_1",
"bbox": [740, 0, 1000, 1000],
"desc": "a large crowd seated on the beach in light summer attire"
}
]Developers can supply these boxes directly or use an LLM to generate the caption and element table from a short instruction and aspect ratio. IDs and coordinates remain editable between turns, so an application can store the layout as project state and let an agent revise individual elements.
Local edits without global drift
FLUX 3 Image supports multi-turn edits within selected boxes. Black Forest Labs says pixels outside those regions remain bit-identical, meaning their numeric color values do not change. Its product demo recolors a surfer’s wetsuit and board while preserving the wave, sky, and monochrome treatment elsewhere in the frame.
BFL attributes this behavior partly to a prompt upsampler, which expands a short request into the detailed caption format used during training. Box IDs and coordinates bypass that rewrite and reach the model unchanged. According to the company, generated details affect the result only when the expanded caption refers to them.
Ten references, one composition
A single generation can include up to 10 reference images. The API assigns tokens in upload order, beginning with ref_image_0, and prompts cite those tokens inline:
A fashion streetwear portrait in Times Square with ref_image_0,
ref_image_1, ref_image_2, ref_image_3, ref_image_4 and ref_image_5.The model determines each reference image’s placement and scale before rendering the complete frame. Stable upload ordering therefore matters when applications construct prompts programmatically. The feature supports product scenes, fashion lookbooks, moodboards, and composites assembled from existing assets.
Full-resolution output reaches 16.8 megapixels
Maximum output is approximately 16.8 megapixels, including dimensions around 5,456 × 3,072. BFL says the model renders at the requested resolution, avoiding a separate enlargement pass that can soften small text and fine textures. Its demonstration includes a Japanese soba shop sign with legible hand-painted characters about 225 pixels tall.
One model in a four-part family
BFL describes FLUX 3 as a multimodal model trained jointly across image, video, and audio data. Shared training is intended to give the family a common representation of scenes across media. The company’s launch report provides additional background.
The Freiburg-based lab was founded in August 2024 by researchers who had worked on Stable Diffusion at Stability AI. It organizes FLUX 3 into four product lines:
| Product | Scope | Announced status |
|---|---|---|
| FLUX 3 Video | Video generation | Early access |
| FLUX 3 Image | Image generation and editing | Playground and API access |
| FLUX 3 Action | Purpose not detailed in the Image release | Early access |
| FLUX 3 Dev | Developer-focused open release | Upcoming |
From playground to self-hosting
FLUX 3 Image is available through the BFL Playground and API. At launch, BFL advertised a 50% API discount through October 8. The API documentation describes the prompt format, reference tokens, and layout schema.
Companies can also license commercial weights for fine-tuning and deployment on their own infrastructure. BFL announced a separate open-weights edition of FLUX 3 Image for release within weeks, without providing a calendar date. FLUX 3 Dev, which the company describes as open source, remains a separate product line.
Where bounding boxes pay off
- Magazine covers and editorial layouts that combine typography with imagery
- Panel grids, collages, and lookbooks with fixed spatial relationships
- Product and e-commerce composites built from reusable reference assets
- Multi-turn edits that must preserve approved regions exactly
- Agent-driven workflows in which an LLM plans the composition
Release materials do not report benchmarked layout accuracy, API latency, or preservation rates across long editing sessions. Production evaluations will need to measure those variables alongside reference fidelity, text rendering, throughput, and current API costs.