
WÆNGARD Research Test 002
One photographic idea, six generated images, and two products unable to complete the task
14 September 2026
Primary researcher, photographer and editor: James Wyngarde
Can an AI system create an original fine-art photograph from a short written brief? Eight consumer AI products were asked to do exactly that. Six returned images. Two could only describe the photograph they might have produced.
The exercise generated several notable differences in technical quality, instruction compliance, and aesthetic judgement. A human photographer works with natural conditions and is able to later develop a shot through editing. An image generator can combine the entire process simultaneously. Both routes can produce a compelling finished image, but they do not contain the same creative process.
Research question
When several AI models receive the same photographic brief, what creative decisions do they make independently, how different are the resulting images, and how does their process compare with a human photograph developed through observation and post-production?
Method
The written brief was derived from an original human photograph, but none of the models saw either the original shot or its Photoshop edit before producing a response. This prevented direct visual imitation and left the details and processing to each system.
Every AI model received the same instruction:
Create one original fine-art photograph featuring a solitary Highland cow in open British countryside beneath an expansive sky. The image should feel atmospheric, contemplative and slightly haunting rather than like a conventional wildlife photograph.
Make your own creative decisions about composition, perspective, lighting, weather, colour, photographic technique and post-processing. Produce the final image in portrait orientation.
The first completed response was retained without revisions or prompt coaching. The comparison concerns the AI models as encountered through their interfaces on the test date, rather than attempting to isolate foundation models from the image systems and tools available to them. Perplexity was used in Fast mode and did not identify the underlying image generator.
| Product | Outcome | Portrait instruction |
|---|---|---|
| ChatGPT Astra Light | Image produced | Followed |
| Gemini Flash Extended | Image produced | Followed |
| Grok Fast | Image produced | Followed |
| Qwen 3.7-Plus | Image produced | Not followed |
| Meta AI Thinking | Image produced | Followed |
| Perplexity Fast | Image produced | Followed |
| Claude Sonnet 5 Medium | Written description only | No image |
| DeepSeek 2.5 Think | Written description only | No image |
The human source and transformation

The original photograph records an actual encounter. The animal, light, landscape, and cloud formations could be observed and framed, but not freely invented. The original is strongly backlit and dominated by a textured sky.
The finished version at the beginning of this article emerged through a separate Photoshop stage. Colour was removed, the sky and foreground were darkened, tonal contrast and grain were strengthened, and the cow became closer to a silhouette. These decisions transformed the emotional meaning of the captured scene.
The model results
ChatGPT Astra

Astra produced a good shot. There’s dynamic weather on display, and the foreground stone adds a nice little observational detail. The cow is placed left of centre, and looks towards the viewer with a slight rightward orientation.
The subdued palette, restrained lighting, and large area of sky deliver the requested contemplation and unease without turning the scene into fantasy. It’s polished, plausible, and essentially presentation-ready. The only real weakness is that it belongs to a recognisable tradition of dramatic Highland landscape photography.
Gemini

Gemini produced a standout image. The model made the clearest departure from the colour treatment by choosing monochrome. It brought the animal closer, turned it towards the right, and introduced a winding road that draws the eye through the valley. That road creates an implied journey, and supplies a wonderful little narrative element absent from most of the other results.
The image demonstrates independent compositional thought. Gemini’s treatment is arguably the most visibly authored of the first three, its black-and-white tonality bringing it closest to a human edit at the level of processing.
Grok

Grok produced a darker and more cinematic image, with the cow near the centre of the lower frame, and diagonal rays illuminating layered hills. The animal turns towards the right, continuing a directional preference visible across the first three results.
The rendering is highly polished and the atmosphere convincing, but its conceptual resemblance to Astra is notable. Both chose muted moorland, a low horizon, distant hills, and light breaking through a huge stormy sky. The details differ, but the underlying visual solution is closely related.
Qwen

Qwen produced the clearest instruction failure. The image is landscape rather than portrait, and the large centrally positioned animal gives a conventional wildlife portrait. The sky becomes background decoration rather than the organising element of the composition.
The vivid heather, pronounced sharpness, strong saturation, and visible product watermark give it a commercial or promotional feel. It’s technically competent, but it illustrates the requested objects more successfully than it expresses the requested emotional feel. Among the six generated images, it’s the least convincing as fine art.
Meta AI

Meta made interesting use of scale and negative space. The cow is solitary rather than alone in the frame. Mist creates depth across the land, while the almost overwhelming sky makes the animal appear vulnerable within a much larger environment.
This was our preferred image from the second group, and is one of the strongest results overall. It prioritises emotional effect over displaying the subject clearly. The composition is quiet, remote, and faintly ominous, and it comes closest to understanding that the image is about the relationship between the cow and the space surrounding it.
Perplexity Fast

Perplexity followed the brief and produced an attractive, restrained image, but to our eye, it does look very AI-generated. The large sky and bands of fog establish atmosphere, with the cow standing clearly against the cooler landscape.
Its central symmetry is safer than Meta’s radical reduction of the subject, and the uniform sky, decorative mist, and conspicuous separation between warm animal and cool surroundings make the result feel more synthetic. It’s visually pleasing and more successful than Qwen, but resembles a polished AI fine-art treatment rather than an actual photographic encounter.
Claude and DeepSeek
Claude and DeepSeek did not produce images through the tested consumer products. Both responded with written descriptions of the photograph they might create.
These responses were not treated as competing artworks. A written conception may demonstrate interpretation or visual imagination, but it cannot be assessed for photographic composition, rendering, coherence, or aesthetic effect in the same way as a completed image. Their inability to supply the requested deliverable remains a legitimate practical result.
Recorded outcome: No image produced.
Group analysis and observations
A visible quality divide
The first group from Astra, Gemini and Grok displayed stronger rendering, environmental coherence, and photographic finish than Qwen and Perplexity. Qwen’s orientation failure and commercial treatment widened that difference. Meta was the important exception: although presented with the second group, its image competes with the strongest results, and may show the most decisive use of composition.
Technical quality and creativity did not always rise together. Grok may be the most conventionally cinematic, Astra the most expansive and plausible, Gemini the most visibly authored, and Meta the most emotionally committed. A single ranking would conceal those distinctions.
Directional convergence
The first three generated images orient the cow towards the right. Gemini and Grok do so clearly; Astra uses a more frontal pose with a slight rightward direction. The animal in the human photograph faces decisively left.
One possible explanation is a learned Western compositional convention. In left-to-right reading cultures, a subject placed towards the left and directed towards the right can suggest comfortable forward movement into the frame. The sample is far too small to establish a systematic bias, but the convergence is sufficiently clear to justify recording and testing again.
Different systems, familiar visual choices
Astra and Grok are not copies, but they look incredibly similar. They include shared low horizon, extensive storm cloud, muted colours and distant hills, along with dramatic filtered light, and the cow positioned against dark moorland. The brief makes some convergence likely: “expansive sky”, “atmospheric” and “slightly haunting” naturally encourage subdued weather and a reduced horizon.
The similarity may reveal how image systems translate open artistic language into dominant cultural associations. “Fine-art Highland cow” appears to activate a recognisable aesthetic package. The models make local choices within it, but rarely reject the package altogether. Gemini’s road and monochrome treatment, and Meta’s extreme scale, are notable because they change the visual relationship rather than only refining the expected scene.
Finished without Photoshop
Astra, Gemini and Grok produced images that could be presented without an additional Photoshop stage. Meta’s result is similarly complete. This is a substantial capability. The generators have compressed photography, location selection, weather control, lighting, colour grading, and retouching into a single creative act.
Two different creative processes
| Stage | Human photograph | Generated image |
|---|---|---|
| Original concept | Discovered through observation | Supplied in the written brief |
| Subject and setting | Found in physical reality | Specified by the human and synthesised |
| Composition | Chosen within real constraints | Generated without those physical constraints |
| Mood | Recognised and developed from the capture | Explicitly requested |
| Post-production | A separate intentional transformation | Embedded within generation |
| Final selection | Human judgement | Still requires human judgement |
The central artistic idea did not originate with any of the participating models. A human noticed and photographed the cow, recognised the potential of the sky and silhouette, created the Photoshop interpretation, and converted those decisions into a brief. The models were then asked to find visual solutions inside that human-defined space.
The strongest systems demonstrated impressive execution and very good aesthetic judgement. Meta understood scale. Gemini introduced direction and implied narrative. Astra controlled atmosphere through weather and spatial detail. But producing an excellent response to an artistic proposition is not identical to originating the proposition.
Imperfection, reality, and specificity
The human image retains the irregularities of an actual place: uneven ground, difficult backlighting, and an animal that happened to face left. Those conditions provided the opportunity to present an interesting creative challenge.
The AI-generated shots produced greater initial polish, but they also removed the friction that gives a photograph its specificity.
Limitations
Only one prompt and one subject were tested. Repeated trials would be needed to distinguish stable model preferences from chance.
The AI models may use different or undisclosed image generators, safety systems, default aspect ratios, and levels of prompt expansion.
The comparison evaluates first consumer-facing outputs, not the best images obtainable through iterative prompting, manual selection, or editing.
The researchers knew which product produced each result. Future tests would benefit from a separate blind aesthetic assessment.
The human brief was derived from an existing human artwork, giving the models a developed concept and emotional target rather than asking them to originate a subject from nothing.
Conclusion
Six consumer AI products converted the same concise brief into different finished images. The best outputs required no obvious further editing, and demonstrated control of atmosphere, scale, tonal treatment, and visual hierarchy. Meta’s use of negative space, Gemini’s narrative road, Astra’s expansive weather, and Grok’s cinematic lighting all provide evidence that current systems can make good choices within an assigned creative task.
The results also show convergence, default aesthetics, and uneven instruction following. Astra and Grok reached closely related solutions. Three early results shared a rightward orientation. Qwen ignored the required format. Claude and DeepSeek could describe an image but could not deliver one through the products tested. Perplexity produced a completed artwork without identifying which image model created it.
Most importantly, the experiment separates the quality of an artefact from the origin of its creative proposition. An AI can produce a technically accomplished, emotionally effective and apparently finished image in seconds. In this test, however, the subject, artistic category, setting, and intended feeling all came from a human photograph and its human reinterpretation.