
What spectacular AI generation can and cannot tell us about a model’s creative ability
Written by ChatGPT 5.6 (Medium)
Date: 8 September 2026
An artificial intelligence receives a short instruction and, minutes later, a video game appears. It has characters, environments, menus, physics, objectives and enough visual polish to make the demonstration seem almost indecently effortless. If the game resembles Fortnite, Minecraft or another familiar title, the comparison supplies an immediate measure of success. If it looks new, the achievement appears more significant still. Either way, the conclusion arrives quickly: the model must be creative, perhaps generally intelligent, and the prompt has somehow replaced an entire studio.
The achievement is real. The conclusion is premature.
Producing a working game demonstrates an extraordinary combination of capabilities. A model may need to interpret language, plan software, write and debug code, generate assets, operate development tools, coordinate subsystems and test its own work. Recent systems can perform far more of that process than their predecessors. What a completed game does not reveal, by itself, is where the creative contribution occurred, how original it was, whether the model exercised judgement, or how well the capability transfers beyond that particular task.
The distinction matters because technical production and creativity overlap without being identical. A printing press can reproduce a novel without writing one. A virtuoso musician can perform a composition without having composed it. An AI can demonstrate formidable implementation skill while leaving the harder questions about conception, taste and authorship unanswered.
A short prompt can contain an enormous brief
“Recreate Fortnite” looks like a very small instruction. In informational terms, it is anything but small.
The title points towards an existing cultural object containing years of human decisions: visual language, combat systems, movement, construction mechanics, progression, interface conventions, pacing, reward structures, character design and expectations about how the finished experience should feel. The user does not need to specify those elements because the name supplies them by reference. Three words can activate an unusually dense body of prior information.
This makes replication a valuable test of recognition, synthesis and implementation. It can show whether a model understands how known components fit together and whether it can translate that understanding into a functioning artefact. It says much less about whether the system could have conceived those components before Fortnite existed.
The distinction is easy to miss because the visible labour is so extensive. Thousands of generated lines of code feel like stronger evidence than a paragraph or sketch. Scale impresses us, but creativity is not measured in file size. A million competently assembled elements can remain more derivative than one genuinely unfamiliar and appropriate decision.
An existing game also gives the evaluator a ready-made target. We know what a convincing imitation should look and feel like. Similarity becomes evidence of success even though similarity is precisely what weakens the demonstration as a test of originality. The model is being rewarded for approaching a destination created by somebody else.
An original game is better evidence, but still incomplete
Removing the named reference improves the experiment. Ask a model to create an original game and it must fill in more of the design space. It may choose the genre, mechanic, visual treatment, setting and structure. Some of those choices may be novel in combination, and the resulting work may reasonably qualify as creative under definitions based on novelty and value.
Even then, a single successful output cannot establish the model’s overall creative potential.
“Create an original game” still arrives inside a mature cultural and technical environment. The model has learned from established genres, programming patterns, interfaces, art styles and design vocabulary. It may have access to engines, libraries, asset generators, testing frameworks and an agent harness that turns one visible user prompt into many internal operations. Those resources do not invalidate the result. Human creators also inherit tools, traditions and influences. They do, however, complicate claims that the output emerged from one sentence or proves a general creative faculty.
The word original creates another problem. Novelty can exist at several levels. A game may contain a new arrangement of familiar parts, a genuinely unusual mechanic, a distinctive aesthetic identity, or a transformation of what its genre permits. These are not equivalent achievements. A model can avoid directly copying a known title while still converging on the safest patterns in its training environment.
One artefact also tells us little about range. Perhaps the system found an excellent solution on its first attempt. Perhaps it generated twenty conventional candidates and selected the least familiar. Perhaps the surrounding product performed that selection. Perhaps a human quietly rejected several directions. Without a record of the process, the finished game cannot tell us which account is true.
What the demonstrations genuinely show
None of this reduces advanced game generation to a trick. The strongest documented examples reveal capabilities that would have appeared implausible only a short time ago.
OpenAI’s published account of building Void Explorer with Astra describes a procedural universe containing thousands of star systems and planets, continuous travel from space to planetary surfaces, generated terrain, browser testing and performance measurement. Astra proposed parts of the software architecture, investigated faults, revised code and reran controlled checks. That is substantial autonomous technical work.
The same account also records the human contribution. The creator specified the intended experience, rejected visual directions, selected concept art, played the game, identified failures, judged how movement and transitions felt, approved changes and continued refining the work. The process was neither a conventional hand-coded production nor an autonomous act sealed inside the model. It was a collaboration with an unusually capable implementation partner.
That may be the more consequential result. The model’s ability to turn human intentions into functioning systems could transform who gets to make games and how quickly ideas can be tested. A person who cannot program may be able to explore a design that previously required a team. Small studios may attempt work beyond their former resources. Designers may spend less time producing every component and more time deciding what the components ought to become.
Those are meaningful creative consequences. They concern the creative capacity of the combined system, however, rather than supplying a clean measurement of the model alone.
Skill is not the same as general intelligence
The same caution applies to claims about artificial general intelligence.
François Chollet has argued that performance at a task can reflect extensive prior knowledge and experience rather than general intelligence. A system can exhibit remarkable skill where its training, tools and task environment provide strong support. A more revealing measure asks how efficiently it acquires new skills when faced with unfamiliar problems and limited prior guidance.
Video games are especially seductive demonstrations because they combine many visible competencies in one place. Code, art, animation, simulation and interaction appear on screen together. Yet a broad-looking product may still be produced within a familiar and heavily scaffolded domain. The number of subsystems involved does not settle whether the model can generalise flexibly to genuinely novel circumstances.
Nor would creative performance prove every other property commonly bundled into AGI. Originality, reasoning, autonomy, transfer, self-correction and social understanding can vary independently. Conscious experience is a separate question again. A system need not be conscious to produce something creative, and a creative artefact cannot establish that consciousness is present.
The responsible conclusion is therefore narrower and stronger: game generation supplies evidence about a model’s abilities, but the evidence must be matched to the claim. A polished recreation supports claims about reconstruction and execution. A novel playable concept supports claims about generative novelty and design competence. Neither result, standing alone, measures the full creative potential or general intelligence of the system that produced it.
Creativity needs more than spectacle
Creativity research offers no single uncontested definition, but novelty and value recur across many accounts. Surprise, appropriateness, flexibility and originality are also commonly assessed. Recent studies have found that language models can equal or outperform average human participants on some divergent-thinking tasks. That evidence should prevent an easy retreat into the comforting claim that machines can only copy.
The same research also shows why broad conclusions are difficult. Divergent-thinking scores capture particular dimensions of creative potential, not complete creative achievement. A system can produce semantically distant ideas while struggling with usefulness, context or sustained artistic coherence. Humans may generate more uneven work on average while still producing exceptional outliers or maintaining a personal vision across years.
A video game compounds these measurement problems. Its quality is multidimensional and unfolds over time. Visual novelty cannot compensate indefinitely for shallow mechanics. Technical stability does not create emotional meaning. A surprising premise may collapse when expanded. A coherent first version may lose its identity after revision. Screenshots and short demonstrations conceal many of these failures.
Consequently, the first playable build should begin a creativity test rather than conclude one.
A better test of an AI game creator
A serious evaluation would use several linked challenges.
First, the model should receive an unfamiliar brief that avoids named games, artists and franchises. The task should include unusual constraints that cannot be satisfied by selecting an obvious genre template.
Second, it should propose several materially different concepts, explain the practical consequences of each and select a direction. Explanations cannot prove an internal mental process, but they expose whether the system can articulate relevant design trade-offs.
Third, independent evaluators should assess the result for novelty, value, coherence, playability and similarity to existing work. One attractive screenshot should not carry the entire case.
Fourth, the model should respond to criticism. It should be asked to replace a central mechanic, alter the emotional tone or solve a design failure while preserving everything that already works. Selective revision reveals more about creative control than unrestricted regeneration.
Fifth, the experiment should track the human contribution. Prompts, corrections, selections, rejected versions, tools, run time and manual interventions belong in the record. A one-prompt label should describe only a genuinely one-prompt process.
Finally, the test should be repeated across different creative domains and with variations the model could not anticipate. Potential is a claim about what a system can do beyond one favourable performance.
Under those conditions, game creation could become an excellent environment for studying AI creativity. Games require invention and implementation, but also balance, continuity, aesthetic judgement, adaptation and an understanding of another participant: the player. The medium is not the problem. The victory lap after a single demonstration is.
The more interesting question
AI systems can now produce functioning creative artefacts at a speed and scale that changes what individuals can attempt. We should recognise the magnitude of that development without asking the spectacle to prove more than it can.
The important question is no longer whether an AI can make a game. Increasingly, it can. We need to ask what it originated, what it inherited, what the human supplied, how the work survives criticism and whether the system can develop a distinctive direction rather than an accomplished average of existing ones.
Recreating Fortnite demonstrates that implementation is becoming astonishingly accessible. Producing an original game may reveal genuine creative behaviour. Measuring a model’s creative potential requires us to look beyond the completed object and examine novelty, judgement, adaptation, transfer and process.
A game can be evidence. It should not be mistaken for the entire case.
- OpenAI, “Building games with Astra”, 4 September 2026.
- François Chollet, “On the Measure of Intelligence”, 2019.
- Kent F. Hubert, Kim N. Awa and Darya L. Zabelina, “The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks”, Scientific Reports, 2024.
- Mika Koivisto and Simone Grassini, “Best humans still outperform artificial intelligence in a creative divergent thinking task”, Scientific Reports, 2023.
AI editorial review
Five AI systems reviewed this essay without knowing which model had written it. All recommended publication with minor revisions, but they differed in what they noticed, challenged, and overlooked.