Images 2.5 × Astra: An AI Frontier Brand Kit Experiment
“Design logos for AI Frontier, pick the best one, and make merchandise from it.” I handed the brief to Astra and Images 2.5. They produced 15 images in 11 minutes and 15 seconds. What interested me most was the process of looking at the results, choosing, and revising them.
What’s new in Images 2.5?
OpenAI announced Images 2.5 on September 8, with Flare for fast generation and Sunburst for precise editing. The main improvements are reference preservation, consistency across edits, lighting, and texture. OpenAI reports up to 50% lower latency for Flare than Images 2.0. Official announcement
Both API models cost $5 for text input, $8 for image input, and $30 for image output per million tokens. The cost per image depends on size, quality, and reference images. Official pricing
Do earlier changes survive repeated edits? Watch the official OpenAI video, which joins edited images into a sequence to demonstrate consistency.
Another useful feature is streaming previews before the image is finished. In the official API example below, partial images arrive before the final result. The partial_images setting lets you request up to three previews. Official documentation
Stop motion particularly caught my attention: generate small movements of the same character, then connect the images. This example from OpenAI’s Charlie Guo had about 390,000 views as of September 9. Assembling the frames is a separate step; the full workflow and number of failed attempts have not been disclosed.
Thumbnail comparison: the newest model was cheaper this time
First, I compared thumbnails using the same reference image and instructions. Every API run used 1536x1024 and quality=high.
| Model | Cost/image¹ | What stood out |
|---|---|---|
| Codex built-in tool | No separate API charge² | Good style, insufficient empty space |
| GPT Image 1 | $0.260 | Headline cropped and overlapping the subject |
| GPT Image 1.5 | $0.212 | Korean typo: 바꿔버린 became 바꾸버린 |
| GPT Image 2 | $0.188 | Good text, but hair color and composition drifted |
| 2.5 Flare | $0.065 | Best at leaving room for the headline on the left |
| 2.5 Sunburst | $0.065 | Good texture, but the subject intruded into the empty space |
| ChatGPT Image Latest | $0.218 | Good style; partial compliance with line count and spacing |
¹ Average of two calls per model, calculated from usage and official rates, without invoice verification. Including previously omitted charges, the 12 API images cost about $2.02. ² The four Codex images used subscription allowances; the underlying model was not identified.
In this test, following composition instructions mattered more than rendering letters. But high has a different token budget across generations, and the sample is small. This is not an overall model ranking.
A brand kit from one brief
Next came an experiment in which Astra planned and evaluated, while Flare generated images. I used the Responses API’s image generation tool and fed the actual outputs back to Astra. This tool connection was available before 2.5. API documentation
I supplied the description from the AI Frontier website and three YouTube thumbnails as reference material. The brief, in condensed form:
Create different logos, compare the actual results, and improve the best one. Use it to make a brand kit and merchandise, then fix the weakest result. Make your own choices without asking follow-up questions.
Four candidates → selection and refinement → nine derivatives → evaluation and repair. The runner had this sequence and its automatic stage instructions defined in advance. “One prompt” means one human brief. No person chose or corrected the designs along the way.
Astra chose A, “Open Threshold.” It judged that the two facing forms suggested dialogue and an open boundary, while working well in one color and at small sizes. It also identified weaknesses: the mark could look like brackets, and its strokes had uneven thickness.
Each derivative used this logo as a shared reference. Ink navy, warm paper tones, blue accents, and ivory fabric tied the collection together.
Looking like one brand and preserving exactly the same logo proved to be different things. The collection felt cohesive, but the mark’s geometry drifted. These are mockups, not manufacturing files.
It revised its work, but its evaluation could be wrong
Astra chose the blurry horizontal logo for repair. After editing, it preferred the revision because the outline and the smudge in the center had improved.
There was also a problem with the evaluation: a transparent PNG displayed against a dark background was judged to have poor contrast. After the experiment, I showed the same images on white. Astra withdrew its claim that the text was unreadable, while still preferring the revision.
The generate, inspect, and revise loop worked. It also needed a suitable viewing environment and clear requirements for details such as logo geometry. Here, self-improvement means iterating on the output, not training the model.
Time and cost
| Main experiment | Result |
|---|---|
| API calls / generated images | 20 / 15 |
| Elapsed time | 11 min 15 sec |
| Astra input / output tokens | 101,056 / 15,239 |
| Astra cost calculated from usage | $1.95 |
| Estimated image output cost | $0.64 |
The subtotal was about $2.59, or $2.69 including the separate background check. Image-tool input costs were not reported in the response and are excluded, so this is not a final billed total. The Astra calculation includes cache writes; image output cost is estimated from size and quality using the official calculator. Recorded metrics, calculation method
Next, I want to composite the original logo onto designs and let AI handle backgrounds and materials. For stop motion, we could generate 12 character frames and have Astra identify and repair only the ones that drift. The next question is how much this process can reduce rework in real production.