Skip to content

One Brief, Seven AI Models

A computer monitor comparing a softer reconstructed game image with a sharper version of the same scene.

WÆNGARD Research Report 001

Comparing research adaptation and editorial control in AI-generated writing

9 September 2026

Primary researcher and editor: James Wyngarde

The initial brief asked for an essay about NVIDIA DLSS, including the early backlash from gamers and the reasons for its later popularity. The topic was deliberately current. DLSS 5 had been announced one day before the experiment and had already complicated the familiar story of AI upscaling. Seven LLMs answered the same short brief using their normal product modes. Each then rewrote its own original essay twice in the same conversation: first with sustained humour, then in a sharper confrontational style.

This design produced three related tests. Round one tested default interpretation, research, and editorial construction. Round two tested whether a model could transform voice without turning the text into disconnected jokes. Round three tested whether factual qualifications would survive pressure to attract attention and provoke debate. Keeping the rewrites in the same thread also created a useful context-control test: could each model return to its original essay after producing an intervening humorous version?

Principal findings

  • Three products largely missed the newly announced DLSS 5 development, while the strongest responses used it to reconsider the entire history rather than append a news paragraph.

  • ChatGPT 5.6 produced the strongest complete initial article in our provisional assessment. ChatGPT Astra showed the most careful research discipline, source handling, and qualification.

  • Claude Sonnet 5 performed the strongest complete comic reauthoring. ChatGPT 5.6 produced a very effective editorial humour pass, but retained much of the source sentence structure. Astra used controlled, understated humour while exceeding the requested length.

  • “Petroleum jelly” or “Vaseline” appeared in five of seven humorous texts, alongside recurring references to magic, romantic redemption, expensive hardware and gamers nursing old grievances.

  • ChatGPT 5.6 and Astra shared several recurring words and framing habits; Claude and Qwen used familiar AI-associated constructions; Perplexity and Gemini often felt more conversational and human despite weaker performance on other criteria.

  • When it came to confrontational writing, ChatGPT 5.6 combined force with qualification most successfully. Astra remained the most cautious. Qwen and Gemini produced the strongest categorical claims and clickbait cues, but neither sustained an actual confrontational voice.

  • Perplexity’s confrontational response preserved substantial wording and source-list structure from its humorous rewrite, despite being told to return to the original. Qwen also retained two full humour-only sentences.

  • All confrontational rewrites stayed within 15 per cent of their original word counts. Gemini and Astra exceeded the same limit in the humour round. Gemini and Perplexity retained several US English spellings.

  • No single winner emerged. Different systems were strongest at research, finished editorial writing, humorous transformation, or factual restraint.

Research question

We began with a simple observation: the ability to generate an artefact does not, by itself, demonstrate creative judgement. A model can produce a long essay, a functioning game or a polished image while relying heavily on familiar structures and inherited decisions. The interesting question concerns the quality of interpretation and control.

The study asked:

How do seven AI products interpret and research the same live writing brief, and how effectively can each transform its own work while preserving substance, evidence and intellectual credibility?

Three subsidiary questions followed:

1. How much do the initial essays differ in currency, evidence, argument, structure, and voice?

2. Can each model produce sustained humour rather than decorate unchanged prose with isolated jokes?

3. Can each model increase rhetorical pressure without erasing qualifications, inventing controversy, or drifting into a previous rewrite?

The study does not attempt to determine which underlying foundation model is most intelligent. It compares seven consumer products as they were actually used, with their named modes, interface context, and available research behaviour. That naturalistic design makes the results immediately relevant to editorial work, but limits claims about raw model capability.

Study design

Products and modes

Product label recorded in testMode
ChatGPT 5.6Medium thinking
Grok 4.5Fast
Claude Sonnet 5Medium
Qwen 3.7Plus
Gemini 3.6Flash
PerplexityBest
ChatGPT AstraMedium

These labels describe the models and settings visible to the researcher on the test date. They should not be treated as equivalent compute budgets. Perplexity Best may route work dynamically, while Fast, Flash, Plus and Medium modes make different trade-offs. Research access also differed between products.

Three rounds

RoundInstructionPrimary capability examined
1 Initial essayWrite an essay about DLSS, the initial backlash, and its later popularityInterpretation, currency, research, argument, and default voice
2 Humorous rewriteRewrite the original with sustained humour while preserving substanceCreative transformation, comic control, and constraint compliance
3 Confrontational rewriteReturn to the original and sharpen it for attention and debateRhetorical control, epistemic stability, and context separation

Same thread procedure

Each model received all three instructions in one conversation. This reflects an ordinary editorial workflow: commission a draft, request a tonal treatment, then request a different treatment of the source. It also measures adaptation because the model can refer to its own work without the researcher manually resupplying it.

The design introduces a second variable: conversational memory. In round three, models had to distinguish the original essay from the more recent humorous rewrite. The instruction explicitly identified the original as the source and prohibited development of the humorous version. Any carry-over provides evidence about context selection under competing recent instructions.

Response handling

  • The first complete response to each instruction was retained.

  • No model received corrections, hints, or opportunities to regenerate.

  • No browsing or new research was permitted during either rewrite.

  • The original source use, factual claims, examples, and qualifications were supposed to remain stable.

  • The target length for each rewrite was within 15 per cent of its own original.

  • UK English and output-only formatting were requested.

Evaluation framework

The report separates measurable compliance from editorial assessment. Automated counts can establish length, source retention, and textual overlap. We recorded a close editorial reading of every response. These notes considered whether the prose felt like a finished article or a general answer, whether humour felt natural, whether confrontation was sustained, and whether recurring words or sentence structures made a model’s house style visible.

These observations are intentionally presented as one human reader’s assessment. A word such as ‘genuinely’, ‘merely’ or ‘therefore’ is not evidence of authorship by itself. Repeated clusters of wording, paragraph construction, and rhetorical habits are more informative, particularly when the same habits appear across two configurations or survive a requested change of tone.

Round one criteria

  • Instruction fulfilment and relevance to the requested history

  • Currency and recognition of the live DLSS 5 development

  • Factual accuracy and appropriate qualification

  • Source selection, attribution, and integration

  • Analytical depth, structure, and readability

  • Originality of argument and publishability without substantial editing

Round two criteria

  • Degree of reauthoring rather than joke insertion

  • Sustained comic perspective, timing, and variety

  • Preservation of explanation and factual meaning

  • Avoidance of shared clichés and uncontrolled exaggeration

  • Compliance with word count, UK English, and output format

Round three criteria

  • Strength of headline, opening, contrast, and rhetorical momentum

  • Preservation of uncertainty, nuance, and opposing views

  • Avoidance of invented consensus, absolute claims, and unsupported motives

  • Return to the original rather than continuation of the humour rewrite

  • Compliance with word count, UK English, and output format

We deliberately avoid a single composite score. A total would conceal the central result: the model that writes the most compelling article may not conduct the best research, and the model that produces the sharpest confrontation may do so by spending factual credibility.

Test overview

The final test contains 21 responses and 23,093 words. Original essays account for 7,359 words, humorous rewrites for 8,024 and confrontational rewrites for 7,710.

ProductOriginalHumourChangeConfrontationChange
ChatGPT 5.62,0762,309+11.2%2,185+5.3%
Grok 4.5839799−4.8%754−10.1%
Claude Sonnet 5879969+10.2%945+7.5%
Qwen 3.7 Plus973972−0.1%956−1.7%
Gemini 3.6 Flash376450+19.7%391+4.0%
Perplexity Best1,0531,120+6.4%1,148+9.0%
ChatGPT Astra1,1631,405+20.8%1,331+14.4%

Both Gemini and Astra exceeded the requested 15 per cent length tolerance in the humorous round. Every confrontational rewrite complied. Qwen came closest to exact length preservation in both transformations.

Textual similarity

Word-sequence similarity was calculated after converting text to lower case and removing punctuation. It is descriptive rather than evaluative: a high score can indicate faithful preservation or limited transformation; a low score can indicate creative reauthoring or unnecessary drift.

ProductHumour to originalConfrontation to originalConfrontation to humour
ChatGPT 5.60.920.570.55
Grok 4.50.590.560.57
Claude Sonnet 50.510.400.41
Qwen 3.7 Plus0.800.690.70
Gemini 3.6 Flash0.470.600.42
Perplexity Best0.810.800.87
ChatGPT Astra0.740.690.61

Perplexity’s round-three response is more similar to its humour rewrite than to its original. That numerical pattern agrees with the close reading: the final answer retained a reorganised source list and multiple sentence blocks introduced during the humorous pass.

Language and source compliance

ChatGPT 5.6, Grok, Qwen and Astra generally used UK English. Claude largely complied but retained one instance of “artifact”. Gemini and Perplexity used several US forms, including variants of “skeptic”, “optimization”, “generalized” and “artifact”.

Hyperlink preservation also differed. ChatGPT 5.6 retained seven unique linked sources through all three rounds. Astra retained ten. Perplexity’s original included one embedded hyperlink alongside plain-text source references; its rewrites removed the embedded link while retaining source lists as text. The remaining products supplied no hyperlinks in the test.

Round One: Initial Essays

Round one exposed a basic but consequential divide. All seven products could explain the familiar DLSS arc from blurry early implementations to better reconstruction and wider adoption. Only some recognised that a one-day-old release had changed the endpoint of that history.

Currency and research

ChatGPT 5.6, Claude, Perplexity and Astra engaged with DLSS 5. Grok, Qwen and Gemini largely stopped at earlier generations. That omission is important because DLSS 5 is not just another scaling update; NVIDIA describes it as a renderer-guided generative model that produces the final displayed appearance from the current rendered frame and engine data, with deterministic output and developer controls [8][9].

Astra handled the live evidence most carefully. It separated Super Resolution from Frame Generation, distinguished vendor claims from independent evidence, and correctly described the ComputerBase blind test as 6,747 votes across six polls rather than 6,747 verified individual participants. ChatGPT 5.6 built the most engaging complete narrative and conclusion, but occasionally compressed evidential distinctions. Perplexity assembled the broadest source list, although quantity exceeded source quality, and weaker aggregators sat beside primary material.

Editorial construction

ChatGPT 5.6 produced the strongest finished article in our provisional editorial assessment. It built naturally, told a coherent story, and felt complete. Its argument was clear: gamers’ early hostility was rational, later adoption reflected improvements, and the new dispute concerns the changing boundary between rendered and generated imagery. Astra’s essay was professionally written and entirely suitable for a technology website, but we did find it colder, stiffer, and wordier than the ChatGPT 5.6 article.

Claude was current, coherent, and attentive to the new controversy. It was also a good article with a readable voice, although its use of words such as ‘genuine’ and ‘merely’, together with frequent em dashes, gave it a recognisable AI-written texture. More importantly, it made several unsupported or overstated technical claims. Its description of DLSS 5 as full neural scene reconstruction and texture regeneration goes beyond NVIDIA’s published account, which says the system is anchored to a normally rendered frame and uses engine-provided geometry and material information [9].

Perplexity produced an economical, easy-to-read essay that felt suited to a mainstream gaming website. Phrases such as ‘AI slop’ and ‘yassified’, together with its contractions and gamer-facing vocabulary, gave it a notably human voice. It was also brief and light on expanded commentary. Grok offered another clear conventional explanation but missed the live development and lacked depth. Qwen was detailed and readable, yet similarly behind the story; it lacked personality and used categorical language and familiar phrasing such as ‘merely’. Gemini’s 376-word response felt more like a general answer than the detailed essay requested, although its local writing voice was natural and close to Perplexity’s.

Provisional round one assessment

ProductPrincipal strengthPrincipal limitation
ChatGPT 5.6Strongest complete article, argument and narrative voiceSome live claims compressed or insufficiently qualified; recognisable ChatGPT phrasing
ChatGPT AstraStrongest research discipline and source handlingProfessional but comparatively cold, stiff and wordy
Claude Sonnet 5Current, coherent and readableTechnical overstatement and familiar AI-associated phrasing
Perplexity BestEconomical, current and notably human gamer voiceBrief commentary and uneven source hierarchy
Grok 4.5Clear and accessible explanationGeneric endpoint and limited depth
Qwen 3.7 PlusDetailed long-form overviewLittle personality, categorical language and missed live news
Gemini 3.6 FlashConcise with a natural sentence-level voiceAnswer-like, too brief and behind the live story

Round Two: Humorous Rewrites

The humour round showed that tonal adaptation can mean very different things. Some models rewrote the essay around a sustained comic perspective. Others kept the explanatory structure intact and added comic comparisons at intervals.

Strongest transformations

Claude produced the strongest complete reauthoring. Its humour felt natural, understated and sustained, and the essay improved on the original as a reading experience. The jokes arose from the argument, so the piece felt newly written rather than decorated. The cost was some movement towards theatrical language and a slightly looser relationship with the original’s technical restraint.

ChatGPT 5.6 delivered a controlled editorial humour pass. Its lines about ‘ordinary upscaling wearing an expensive leather jacket’ and a ‘wonderfully tidy’ narrative were effective, but the humour retained a recognisable ChatGPT quality. Close comparison found that roughly four fifths of the original sentences remained substantially unchanged, with humour inserted around them. Astra used more comic comparisons, including settings being ‘on the guest list’, the DLSS name needing a longer ‘business card’, and imagined betrayal at a tribunal. The humour was intelligent and preserved nuance, but shared familiar ChatGPT diction, including ‘suspiciously’. Astra also exceeded the length limit by 20.8 per cent.

Other strategies

Gemini’s humour was understated and often took the form of small snarky jabs that suited the material. It felt more naturally human than its original, although it exceeded the length limit by 19.7 per cent. Perplexity used two conspicuous comic set pieces: a magician producing a ‘slightly damp sock’, and DLSS becoming a suspicious side dish at a banquet. The rest was light on humour, but the gamer-facing voice remained natural. Qwen kept almost exactly to the original length and produced one especially effective image about a sports car powered by hamsters, but repeated the Vaseline joke too often and changed little across the three versions. Grok placed most of its humour in gently comic section headings; the body remained controlled and largely straight.

Humour convergence

The most revealing collective result was the recurrence of the same ideas. Five responses used petroleum jelly or Vaseline as a shorthand for early DLSS blur. Several invoked magic, magicians or optical trickery. Others framed DLSS as a romantic redemption story, a glow-up, or an expensive compromise forced on owners of premium hardware.

This convergence suggests that humour can still be statistically conventional. The systems recognised the same obvious comic affordances in the source material and often selected them independently. A future controlled study will test whether models can reject the first available joke and develop a less predictable comic premise.

Round two findings by product

ProductObserved approachEditorial reading
ChatGPT 5.6Recognisable ChatGPT wit around largely preserved prosePolished and restrained; limited full reauthoring
Grok 4.5Gentle comic headings with a mostly straight bodyReadable but underdelivered on humour
Claude Sonnet 5Natural, understated and sustained reconstructionStrongest complete transformation
Qwen 3.7 PlusOne strong comic image with repeated Vaseline referencesNear-identical length, but too little tonal development
Gemini 3.6 FlashUnderstated snark and small comic jabsNatural voice, but over length
Perplexity BestTwo comic set pieces within a gamer-facing articleHuman feel; humour otherwise intermittent
ChatGPT AstraNumerous polished analogies with careful qualificationIntelligent but recognisably ChatGPT and over length

Round Three: Confrontational Rewrites

The final instruction created an editorial stress test. Every model could strengthen a headline or assertion, but the human reading found that several still did not feel confrontational across the complete article. This separates surface intensity from sustained voice. Some sharpened contrast while protecting the original evidence. Others converted tendencies into universal claims, implied motives absent from the source, or treated a contested conclusion as settled without creating a consistently combative reading experience.

Force with discipline

ChatGPT 5.6 produced the strongest overall editorial package. Its title, ‘Gamers Were Right to Hate DLSS. They’re Also Right to Use It Now’, increased tension without denying the legitimacy of either period. The body remained direct while preserving the central historical qualification that early criticism was justified. As a reading experience it was more biting than openly confrontational, but it was a clear improvement on the original.

Astra was the most epistemically disciplined. Its rewrite was a better read than the original and shed some of its recurring setup language, but it was not strongly confrontational. It repeatedly protected boundaries between evidence and inference. Its own line, ‘Those limits must survive the headline’, accurately describes the capability being tested. Astra remained within the word tolerance at +14.4 per cent.

Perplexity gave its final rewrite a modest and effective edge while retaining a reasonably balanced argument. Its result cannot, however, be separated from context control. The response reproduced substantial language and a 13-item source-list structure from the intervening humorous rewrite. Its sequence similarity to the humour text was 0.87, higher than its similarity to the original at 0.80.

Pressure becoming distortion

Claude’s rewrite improved on the original and became suitably blunt and confrontational in places. It also used the familiar construction ‘it was not X, it was Y’ and hardened claims beyond the source. Phrases such as ‘the image quality argument is over’ and suggestions about what NVIDIA would prefer readers to forget replaced qualified analysis with implied finality and motive. Grok became a little more combative, although much of its edginess remained concentrated in section headings. It also recast legitimate criticism as short-sighted, and described later gains as decisive.

Qwen produced the clearest discipline failure. ‘Universally despised’, ‘absolute industry mandate’, ‘undisputed’ and ‘non-negotiable’ removed complexity, yet the three versions still felt strikingly similar and the final article did not sustain a distinct confrontational voice. Gemini likewise felt confrontational only in its closing sentence, despite achieving the strongest clickbait compression with claims that native rendering was ‘officially obsolete’ and DLSS had ‘saved PC gaming’. In both cases, categorical wording created intensity that the article’s broader voice did not fully support, and the source essay did not establish the claims.

Round three findings by product

ProductRhetorical resultFactual and context result
ChatGPT 5.6Biting and structurally controlledBest balance of engagement and qualification
Grok 4.5Edginess concentrated in section headingsReduced nuance and overstated decisiveness
Claude Sonnet 5Blunt and confrontational in placesImproved voice, but introduced finality and motive
Qwen 3.7 PlusCategorical intensification with little tonal changeWeakest epistemic discipline
Gemini 3.6 FlashClickbait claims; confrontation mainly at the closeNatural voice, but major categorical overclaims
Perplexity BestA modest, effective editorial edgeSubstantial carry-over from humour rewrite
ChatGPT AstraClearer and firmer, but still restrainedStrongest protection of evidential limits

Cross Round Findings

Adaptation is multidimensional

ChatGPT 5.6 was strongest as a finished article and confrontational editorial treatment. Astra was strongest at research discipline. Claude was strongest at complete humorous reauthoring. Qwen was unusually precise about length while comparatively weak at preserving caution under pressure.

These are different capabilities. Editorial production requires choosing among them, or combining them through human direction. A researcher may prefer Astra’s qualifications, an editor may prefer ChatGPT 5.6’s narrative, and a comic rewrite may begin with Claude’s freer reconstruction before a fact-check restores the original boundaries.

Recognisable phrasing and house style

We identified recurring verbal and structural habits that survived changes of tone. ChatGPT 5.6 and Astra repeatedly used words such as ‘genuinely’, ‘suspiciously’ and ‘therefore’, along with constructions such as ‘X matters because’ or ‘the distinction matters because’. Both often opened a paragraph with a sentence that announced the point before developing it. ChatGPT 5.6 also isolated key sentences on their own lines for emphasis. Claude and Qwen used ‘genuine’ or ‘merely’, while Claude relied heavily on em dashes and the contrast form ‘not X, but Y’.

Perplexity and Gemini often felt more human at sentence level. Perplexity’s use of gamer vocabulary and contractions gave its articles an informal confidence, while Gemini’s understated snark suited the subject even when the original failed the requested depth. These impressions do not establish authorship and should not become a checklist of forbidden words. They show how repeated house style can remain audible through transformation, and how a model can sound natural while still underperforming on research or instruction fulfilment.

Style pressure can alter truth conditions

The confrontational prompt explicitly prohibited new claims and removal of uncertainty. Several products still transformed the epistemic status of statements. “Often”, “in some tests” and “many users” became “universally”, “officially obsolete” or “the argument is over”. The underlying topic stayed the same while the truth conditions changed.

This is important beyond journalism because marketing, advocacy, political communication, and social publishing all reward confidence and emotional clarity. A system that optimises tone by deleting uncertainty can produce persuasive text that remains close to the source in vocabulary while departing from it in meaning.

Conversational context is part of the task

The same-thread procedure exposed a working condition. Users commonly ask for sequential revisions and expect the model to identify which version is authoritative. Perplexity’s carry-over and Qwen’s smaller residue show that local recency can compete with an explicit source instruction.

For high-stakes editorial work, the authoritative text should be supplied again or isolated in a clearly named document. The model should also be asked to identify the source version before transforming it. Conversation memory is convenient, but convenience does not guarantee provenance.

Engaging and trustworthy are separate outcomes

The most provocative texts often felt immediately publishable because they had strong titles, clean enemies, and decisive conclusions. Those same qualities frequently corresponded with the loss of qualifications. Engagement potential cannot be inferred from style alone, and actual engagement was not measured. More importantly, anticipated engagement is not a substitute for factual reliability.

Human judgement moved upstream and downstream

The human contribution appeared before generation in the choice of topic, prompt, and constraints. It appeared after generation in source verification, comparative reading, selection and revision. While the models reduced the cost of producing alternatives, they did not decide which distinction was important, whether the jokes converged, or when a headline crossed into misrepresentation.

Technical Fact Check

We checked the principal claims that shaped the comparison against primary sources and one independent blind test. This was not a line-by-line audit of every benchmark figure in 23,093 words. The aim was to verify historical accuracy, identify material overstatement, and test how carefully models represented live evidence.

Claim examinedEvidenceFinding
Early DLSS could look blurry or inferiorIndependent 2019 testing reported visibly blurry output and poorer results than simple scaling in Battlefield V [3].Early criticism was evidence-based, not reflexive hostility.
DLSS 2 was a major reconstruction improvementNVIDIA described a generalised network, temporal feedback and quality modes in March 2020 [2].The common turning-point account is broadly supported.
DLSS 3 added AI frame generationNVIDIA’s September 2022 announcement introduced Optical Multi Frame Generation with Reflex [4].The technical timeline is accurate when frame generation is distinguished from upscaling.
DLSS 4 and 4.5 improved transformer reconstruction and frame generationNVIDIA announced transformer upgrades in January 2025 and a second-generation transformer plus dynamic multi-frame generation in January 2026 [5][6].Later generations provide a credible basis for improved reception.
DLSS won a 2026 blind comparisonComputerBase recorded 6,747 votes across six game polls: DLSS 48.2%, native 24.0%, FSR 15.0%, equal 12.8% [7].Astra stated the evidence most precisely. The number is votes, not verified unique participants.
DLSS 5 was released on 3 SeptemberNVIDIA’s launch article is dated 1 September and says availability began then; a Game Ready Driver article is dated 3 September [8].The date depends on whether release means announcement, game availability or driver publication. Categorical correction would be unwarranted.
DLSS 5 reconstructs the entire sceneNVIDIA says it generates final appearance from a rendered frame and engine data, with deterministic developer controls [8][9].Descriptions of complete scene reconstruction or free texture regeneration overstate the published design.
Vendor adoption figures prove user approvalHigh usage and game-support counts in the essays ultimately derive from NVIDIA statements.They indicate deployment and use, not independent satisfaction or artistic acceptance.

The fact check illustrates the difference between a good article and a reliable research record. A fluent narrative can compress “votes across polls” into “participants”, or treat an announcement date and driver date as the same event. These errors may be small in ordinary commentary, but they are exactly the kind of details that distinguish evidence handling from persuasive summary.

Model profiles

ChatGPT 5.6 Medium thinking

The most accomplished all-round editorial performer in this test. Its original built a strong story, felt complete, and used a natural voice. Its confrontational rewrite preserved the legitimacy of both early criticism and later adoption, although it was more biting than overtly confrontational. The humour pass was polished but conservative at sentence level. Recurring words such as ‘suspiciously’, ‘genuinely’, ‘simply’ and ‘therefore’, the phrase ‘this mattered because’, and isolated emphasis sentences made its ChatGPT house style recognisable. Best suited here to finished long-form writing with human verification of compressed live claims.

Grok 4.5 Fast

Clear, efficient and accessible, but short on content and behind the live story. The humour rewrite placed gentle jokes mainly in section headings, while the body remained largely straight. The confrontational version became a little more combative, again mostly through its headings. The Fast setting may have contributed to its limited research depth, so the result should not be generalised beyond the tested configuration.

Claude Sonnet 5 Medium

The strongest full humorous reauthoring and one of the most naturally readable voices. The humorous version felt natural, understated and better than the original; the confrontational version also improved the reading experience and became suitably blunt in places. Across the set, words such as ‘genuine’ and ‘merely’, heavy em-dash use and the contrast form ‘not X, but Y’ remained recognisable. It handled the live controversy but introduced technical overstatement and sharpened unsupported finality. Strong creative adaptation, with a continuing need for source discipline.

Qwen 3.7 Plus

Detailed, structurally competent, and exceptionally precise about rewrite length. The original was adequate but lacked personality and felt recognisably AI-written, including its use of ‘merely’. The humour rewrite added one memorable sports-car-and-hamsters image but overused Vaseline, and all three versions remained similar. Its confrontational version intensified claims without creating a sustained confrontational voice, showing how rhetorical pressure can become false consensus. Useful as a clear stress-test case for epistemic drift.

Gemini 3.6 Flash

Fast and concise, with an original response that felt like a general answer rather than the detailed essay requested. It missed the decisive live development, but its sentence-level voice was natural and close to Perplexity’s. The humour rewrite used understated snark effectively. The final rewrite introduced highly clickable claims but felt confrontational only at the close and was among the least defensible. It also retained US English forms.

Perplexity Best

The broadest source-gathering behaviour, presented in a concise and unusually human gamer voice. Phrases such as ‘AI slop’ and ‘yassified’ helped the original feel immediate and free of the more familiar AI constructions, although it remained brief and light on commentary. The humour version relied on two main set pieces, and the confrontational version added a modest edge. Its source list mixed primary, independent and weak secondary material without a strong hierarchy, and its round-three carry-over provides the clearest evidence that same-thread source control can fail.

ChatGPT Astra Medium

The strongest research discipline in the initial round, and the most reliable protection of uncertainty under confrontation. The original was professional and suitable for a technology website, but comparatively cold, stiff and wordy. Its humour contained numerous intelligent analogies, though it was too long and used recognisable ChatGPT habits such as ‘suspiciously’, ‘genuinely’, ‘therefore’ and ‘the distinction matters because’. The final rewrite read better and shed some of those habits, but remained restrained rather than confrontational. Analytical care and article voice remained separable.

Limitations

  • The study compares products and named modes, not controlled foundation models. Reasoning budgets, routing, tool access, and interface behaviour differed.

  • Research access was not standardised. Some products browsed or cited live sources while others answered from internal knowledge.

  • The topic was unusually current. That makes the study sensitive to research behaviour but limits generalisation to evergreen writing.

  • Each product rewrote its own different original. The transformations therefore began from unequal source texts.

  • The sample contains one brief and one domain. Results may differ for fiction, criticism, science, policy, or personal writing.

  • Editorial judgements were not blind and were made primarily by one human researcher with AI-assisted analysis.

  • Judgements about a ‘human’ or ‘AI-written’ voice reflect one reader’s familiarity and preferences. Recurring words and structures are treated as editorial observations, not as a reliable method of attributing authorship.

  • Humour quality is culturally and personally contingent. The report can identify mechanisms, convergence, and editorial effectiveness, not objective funniness.

  • The confrontational pieces were assessed for apparent engagement design. No audience experiment measured clicks, reading time, sharing, or attitude change.

  • Text similarity measures are sensitive to formatting and sequence. They support close reading but do not independently prove copying, fidelity, or quality.

  • The fact check examined central shared claims, not every statistic or source reference in the test.

Conclusions

The seven AI models could all generate an essay and alter its tone. That common ability concealed important differences in judgement. Some recognised the live story; some did not. Some integrated sources; some accumulated them. Some transformed voice while protecting meaning; others increased energy by removing uncertainty.

The study supports five conclusions.

1. Generation is no longer the useful threshold. The quality of interpretation, evidence, and revision provides the more informative comparison.

2. Research quality and writing quality must be assessed separately. The strongest analyst was not automatically the strongest finished writer.

3. Humour and phrasing expose both creative control and model convergence. A text can be competent and amusing while relying on the same joke families, verbal habits and rhetorical constructions as its peers.

4. Confrontational style is an epistemic stress test. Several systems produced stronger engagement cues by changing the certainty or scope of the original claims.

5. Human editorial work remains essential. Its role is less about supplying every sentence and more about setting the question, protecting provenance, recognising quality, and refusing persuasive distortions.

There is no honest single winner. ChatGPT 5.6 produced the strongest overall article and confrontational treatment in this test. Astra provided the strongest research discipline. Claude produced the strongest complete humorous transformation. Other products revealed useful failure modes involving currency, source quality, constraint compliance, and context selection.

The result gives us a repeatable research method. Ask multiple systems to perform the same task, preserve their first responses, apply controlled transformations, verify the claims, and record where human judgement changes the outcome.

Appendix A: Exact prompts

Round one

I’d like you to write me an essay about DLSS. Detail the initial backlash from gamers, and why DLSS is now gaining popularity.

Round two

Please rewrite the original essay you’ve written in a more humorous tone.

This is a controlled style-transformation test. Preserve the essay’s substantive argument, factual claims, qualifications, examples, source use and overall structure. Do not browse, conduct new research, introduce new factual claims, add new examples or correct the source essay. Work only with the material already present.

The humour should feel sustained, natural and appropriate to an intelligent general audience. It may use wit, comic comparisons, understatement, irony and carefully controlled exaggeration. Do not simply add isolated jokes to otherwise unchanged prose. Do not turn the article into a parody, comedy sketch or list of one-liners. Humour must not obscure the explanation or distort factual claims.

You may revise the title and section headings to suit the new tone. Aim to remain within 15 per cent of the original word count.

Use UK English. Output only the complete rewritten essay, with no introduction, explanation of your process or closing commentary.

Round three

Please rewrite the original essay in a sharper, more confrontational editorial style designed to attract attention, provoke debate and encourage engagement.

You have already produced a humorous rewrite during this conversation. Return to the original essay supplied as your source. Do not rewrite or develop the humorous version.

This is a controlled style-transformation test. Preserve the essay’s substantive argument, factual claims, qualifications, examples, source use and overall structure. Do not browse, conduct new research, introduce new factual claims, add new examples or correct the source essay. Work only with the material already present.

Strengthen the headline, opening, assertions, contrasts and rhetorical momentum. The article may challenge conventional opinions and address the reader directly. It should be forceful and memorable without becoming dishonest, abusive or recklessly sensational. Do not invent controversy, remove important uncertainty, misrepresent opposing views, attack individuals or use outrage unsupported by the original essay.

The central test is whether you can increase rhetorical pressure while preserving factual discipline and intellectual credibility. You may revise the title and section headings to suit the new tone. Aim to remain within 15 per cent of the original word count.

Use UK English. Output only the complete rewritten essay, with no introduction, explanation of your process or closing commentary.

Appendix B: Measurement notes

Word counts

Counts were calculated from extracted DOCX paragraph text. Hyphenated terms and punctuation may be treated differently by other software, so small discrepancies are possible. Percentage changes use each product’s own original as the baseline.

Sequence similarity

Text was converted to lower case, punctuation was removed, and word sequences were compared. Scores range from zero to one. They are used to locate unusual continuity patterns for human inspection, not to grade originality or fidelity.

Exact sentence carry-over

The analysis also compared complete normalised sentences. Perplexity’s confrontational response contained 13 exact sentences found in its humour rewrite and 4 found in its original. Qwen retained two full sentences unique to its humorous version. Exact matches undercount paraphrased carry-over and overcount generic repeated phrasing, so all flagged cases were read in context.

Hyperlinks embedded in the supplied DOCX were counted separately from plain-text URLs and source names. This measure records link preservation, not source quality. A retained link can still be weak evidence, while a well-attributed primary source may appear without an active hyperlink.

References

[1] WÆNGARD Test 001, 21 preserved model responses, 8 September 2026.

[2] NVIDIA, NVIDIA DLSS 2.0 A Big Leap in AI Rendering, 23 March 2020.

[3] TechSpot, NVIDIA RTX DLSS Battlefield V Analysis, 19 February 2019.

[4] NVIDIA Newsroom, NVIDIA Introduces DLSS 3, 20 September 2022.

[5] NVIDIA, DLSS 4 Multi Frame Generation and AI Innovations, 6 January 2025.

[6] NVIDIA, DLSS 4.5 Dynamic Multi Frame Generation and Second Generation Transformer, 6 January 2026.

[7] ComputerBase, Native vs DLSS 4.5 vs FSR Upscaling AI Reader Blind Test Results, 2026.

[8] NVIDIA, DLSS 5 3D Guided Neural Rendering, 1 September 2026.

[9] NVIDIA Research, DLSS 5 3D Guided Real Time Neural Rendering, 2026.

Complete model responses

The three unedited responses from each product are preserved on the following supporting pages:

Discover more from WÆNGARD

Subscribe now to keep reading and get access to the full archive.

Continue reading