Most image model comparisons follow the same pattern: put four vendor samples next to each other, pick a winner, and call it a day.
That does not help much when you are choosing an API. The annoying questions come later. Can I change one region without touching the rest? How many references can I send? What happens when the text is wrong three times in a row? How much did the image I could actually ship cost?
One caveat before getting into it: this is not a blind image-quality test. I did not run dozens of matched prompts across all four models, so I am not going to invent scores. I compared the published capabilities, the parameters currently exposed on reAPI, and public prices checked on August 22, 2026.
My short version:
- For a general image feature, I would start with Gemini 3.1 Flash Image.
- For masks, transparent backgrounds, and controlled edits, I would use GPT Image 2.
- For multilingual posters and dense layouts, Qwen Image 3.0 deserves its own test set.
- For products and multi-reference composition, I would put FLUX.2 in the first round.
Price and API surface
The reAPI figures below are prices for one delivered image. If a request returns several images, each output is billed. I have put the direct vendor price in its own column so the two are not easy to confuse.
| Model | reAPI 1K | Direct vendor | Best first test |
|---|---|---|---|
| GPT Image 2 | USD 0.030 | USD 30 / 1M output image tokens | References, ordinary generation |
| Gemini 3.1 Flash Image | USD 0.028 | USD 0.067 | General images, search grounding |
| Qwen Image 3.0 Standard | USD 0.024 | USD 0.030 | Multilingual text, infographics |
| FLUX.2 Pro | USD 0.028 | USD 0.030 (from 1 MP) | Products, multi-image composition |
For the three rows that line up cleanly, reAPI is about 58% lower for Gemini 1K, 20% lower for Qwen Standard, and about 7% lower for a 1K FLUX.2 Pro text-to-image request. Google's direct 2K and 4K prices are USD 0.101 and USD 0.151, versus USD 0.043 and USD 0.064 on reAPI, so the gap stays around 57% at those sizes.
For higher resolutions on reAPI, GPT Image 2 is USD 0.050 at 2K and USD 0.080 at 4K; Qwen Standard is USD 0.024 at both 1K and 2K; and FLUX.2 Pro is USD 0.039 at 2K.
Pricing note: OpenAI bills GPT Image 2 from a combination of input tokens, output image tokens, size, and quality. There is no single direct “1K image” price to put beside reAPI's flat basic price, so there is no savings percentage in that row. BFL rounds resolution up by megapixel, and USD 0.030 is its starting price for a 1 MP text-to-image request. Extra vendor charges for input images or search are not included here.
GPT Image 2 also has a separate rate card for the tier with masks, backgrounds, quality, and format controls. Price comparisons only make sense after choosing the tier the product will actually call.
Eight differences that matter more than vendor samples
| Selection question | Published capability and limitation |
|---|---|
| Do I need inpainting, transparency, or a specific file format? | Test the advanced GPT Image 2 tier. It exposes masks, transparent backgrounds, input fidelity, quality tiers, PNG/JPEG/WebP, and compression. The USD 0.030 basic tier in the price table does not include them. |
| How many references can one request accept? | Current reAPI limits are 16 for GPT Image 2, 14 for Gemini, 8 for FLUX.2, and 3 for Qwen. A higher limit says nothing about how well conflicting references will be combined. |
| How many candidates can I get in one request? | Qwen returns up to 6; Gemini and advanced GPT Image 2 return up to 4; basic GPT Image 2 and FLUX.2 return 1. Every delivered image is billed. |
| Does the image depend on recent events, places, or products? | Gemini can use Google text and image search. It reduces stale context, but dates, prices, and figures in the final image still need checking. |
| Do I need a long banner or unusual aspect ratio? | Gemini covers 0.5K through 4K and adds 1:4, 4:1, 1:8, and 8:1. Qwen accepts custom dimensions from 512 to 2048 pixels per edge, also within a 1:8 to 8:1 range. |
| Is this a menu, infographic, or dense multilingual page? | Test Qwen Pro with the hardest real copy. It accepts prompts of roughly 4.5K tokens, plus a negative prompt and optional prompt expansion. Those controls do not guarantee every character will be correct. |
| Must brand colors stay close to specified values? | FLUX.2 accepts HEX colors and structured prompts. Flex also exposes sampling steps and guidance; Pro is cheaper but does not expose those two controls. |
| Does the bill need to be predictable before launch? | Qwen Standard costs the same at 1K and 2K. Basic GPT Image 2 and FLUX.2 use flat per-image prices by resolution. Retry count remains the larger unknown. |
GPT Image 2 is for requirements that leave little room for interpretation
GPT Image 2 is easy to misread from the rate card. Basic 1K generation is USD 0.030 on reAPI, but the more interesting part is the advanced control surface: masks, transparent backgrounds, output format, compression, quality, and input fidelity.
Consider a product tool replacing the background behind a coffee mug. The mug, logo, and shadow are approved and must not move. The result has to be a transparent PNG. At that point, “make a similar image” is not the job. Being able to send a mask and an explicit background setting is more useful than repeating “do not change anything else” in the prompt.
For avatars, covers, and ordinary text-to-image work, I would not choose GPT Image 2 on reputation alone. Its price changes a lot with size, quality, and request type. Any cost estimate that leaves those details out is suspect.
Gemini is the baseline when I do not know the answer yet
If I had room for only one model in the first evaluation, it would probably be Gemini 3.1 Flash Image, also known as Nano Banana 2.
The reason is mundane: it covers 0.5K through 4K, accepts up to 14 references, can return four images, and supports Google text and image search. It also handles odd formats such as 1:4 and 1:8.
It may not lead every category, but it is less likely to reveal a missing parameter immediately after integration. Social posts, product-in-context images, and visuals that depend on recent information all fit within its published surface.
Search grounding still needs supervision. Dates, prices, maps, and chart labels in a generated image must be checked. Google also adds SynthID to generated images. That is not a problem for most marketing assets, but it matters if provenance is part of the product.
Qwen Image 3.0 is unusually focused on text and layout
Qwen Image 3.0 is marketed around long prompts, small text, multilingual rendering, and complex pages. On reAPI, Standard costs USD 0.024 at both 1K and 2K. Pro costs USD 0.032 at 1K and USD 0.064 at 2K. It can return as many as six images in one request, which is handy when exploring layouts.
I would test it with material that exposes those claims: a Chinese product poster with an exact price, a three-column menu, and a 3×3 infographic. A landscape photograph will not tell you much about why this model exists.
I would not call it “the best model for text” yet. Alibaba's demos make text and dense layouts the headline, but independent comparisons are still thin. The honest test is to feed it the hardest copy your product needs and check every character.
FLUX.2 makes the most sense to me for products and brand work
Black Forest Labs talks about FLUX.2 in terms of photorealism, multi-reference editing, text, and color control. The current reAPI surface accepts up to eight reference images. Pro costs USD 0.028 at 1K and USD 0.039 at 2K; Flex is USD 0.077 and USD 0.132.
For a product campaign, I would test FLUX.2 early: put the same shoe in several environments without letting its shape or colors drift, or use separate references for the product, person, location, and lighting.
More references can make the result worse. Conflicting angles and lighting give the model several incompatible answers. Assigning each reference a job is safer than filling every input slot: “Image 1 defines the product only. Images 2 and 3 define the location. Image 4 defines the lighting.”
The useful cost is cost per accepted image
Suppose Model A costs USD 0.024 but needs five attempts to produce one usable asset. Model B costs USD 0.050 and passes after two:
Model A: USD 0.024 × 5 = USD 0.120
Model B: USD 0.050 × 2 = USD 0.100
The cheaper model costs 20% more. That still ignores the time spent fixing text, cutting out products, or rebuilding a layout.
I would not judge these APIs from one prompt and one output. Use a small fixed task set, run each task three times, and record whether required text is correct, protected objects changed, and how many retries were needed. Then calculate:
Cost per accepted image = total generation spend / accepted images
That number is much harder to market and much more useful than a leaderboard position.
My current test order would be Gemini for the general baseline, GPT Image 2 for controlled editing, Qwen for multilingual and dense layouts, and FLUX.2 for product work. A real product can mix them. There is no prize for forcing every image through the same model.
The model names in the opening table link to the live reAPI pages. Prices move, so the figures here are an August 22, 2026 snapshot.