GPT-6 Astra enters the Short-Story Creative Writing Benchmark at #3
GPT-6 Astra has achieved the #3 position in the Short-Story Creative Writing Benchmark, outperforming GPT-5.6 Sol (high) with a score increase from 2.5 to 3.5. Astra distinguishes itself by locating harm within the protagonist's excellence or intimate circle, converting prompt phrases into behavior rather than verbatim labels, and concluding stories with physical acts or resumed obligations instead of aphorisms. While GPT-5.6 Sol (high) unifies attributes, objects, settings, and themes, Astra's approach remains descriptive.
- Published
- Sep 5, 2026, 22:33
- Source type
- Dev community
- Tier
- Community
- Source status
- Sync delayed
Times shown in UTC
More details
The Creative Writing Benchmark tests how well models turn constrained briefs into complete 600-800-word stories. Each brief requires 10 elements, including a character, object, setting, motivation, and tone, that must meaningfully shape the story. Judges assess prose, originality, coherence, characterization, and how effectively those ingredients work together.
Models write to the same prompts. Three judges from different model families compare each story pair, shown in both orders to reduce position bias. The main leaderboard now covers 50 models and 79,507 evaluator judgments.
A separate self-judging experiment produced a striking result: with model names hidden, Astra chose its own story in 692 of 700 comparisons. It acknowledged just 1 of the 57 losses assigned by the regular judging panel.
Given the same short-fiction prompts, GPT-5.6 tends to write stories that put things back. A hidden design gets decoded, the wrongdoer is exposed or converted, the loss that started the story is returned, and a closing sentence names the lesson [C013]. GPT-6 Astra writes stories that spend something instead: its protagonists are usually implicated in the harm they are repairing, and the repair costs them something they do not get back [C034]. One prompt shows the split cleanly. Handed a dead wife's last surviving recording, GPT-5.6 returns the woman herself from the phonograph's hum; GPT-6 Astra breaks the wax disc into a lamp's reservoir to save a stranger, listens for her in the hiss, and finds only burning [C001]. The habits extend to the assignment itself: GPT-5.6 writes the required phrases into the prose where a reader can see them, while GPT-6 Astra converts them into behaviour and never names them [C007] — the source of its best effects and, occasionally, of a requested mood that never quite arrives. Across fifty prompts given to both, the difference is lopsided, and the fiction that results reads less like a fable than like a deposition.
The recurring difference between GPT-6 Astra (high) and GPT-5.6 Sol (high) is not skill level but what a story is *for*. Across all fourteen blinded packets, GPT-5.6 Sol (high) writes restorative parables: a concealed design is decoded, an antagonist is exposed or converted, the loss that motivated the story is returned, and a closing sentence names the lesson [C013] [C028] [C048]. Astra writes consequence fiction: the protagonist is implicated in the harm she is repairing, the fix costs something that is not refunded, and the ending leaves the cost running [C001] [C063] [C082]. Four packet critics reached this independently under different blindings, and the quantitative record is one-sided in the same direction: 43 focus wins, 6 ties, 1 comparison win across 50 pairs, mean margin +1.513 (CI half-width 0.263), median +1.667.
Three structural corollaries recur. First, the fault line: Astra locates the harm inside the protagonist's own professional excellence or her intimate circle, while GPT-5.6 Sol (high) locates it in villains, ministries, and inherited wrongs [C004] [C015] [C027] [C034] [C055] [C065]. Second, how requirements enter: GPT-5.6 Sol (high) inscribes prompt phrases verbatim as narrator labels; Astra converts them into behaviour, object, and procedure, usually without naming them [C007] [C030] [C043] [C052] [C062]. Third, how meaning is delivered: GPT-5.6 Sol (high) terminates on a portable aphorism, often a "not A, but B" construction; Astra terminates on a physical act or resumed obligation and declines to summarise [C003] [C011] [C023] [C074] [C084].
GPT-5.6 Sol (high)'s own portrait is not a deficit portrait. It is the stronger inventor of systems and worlds: nested adversarial architecture [C019], mandated objects converted into a story's central proposition [C025], sensory settings that carry exposition causally [C039], multi-function symbolic devices [C070], mechanisms whose rules are identical to the story's ethics [C091], and the larger, stranger conceits [C078] [C085]. It also reliably manufactures jeopardy where Astra sometimes omits it [C046].
Astra's reasoning shows most clearly where a rule must forbid something. Its mechanisms specify failure modes and then obey them, including past the climax [C036] [C076] [C042], and its supernatural channels emit images requiring interpretive labour rather than plain-English instructions [C088]. GPT-5.6 Sol (high)'s plots more often depend on a design laid in advance by a dead figure who predicted the protagonist's arrival, so the task is recognition rather than invention [C090]; at its weakest this produces devices revealed in the sentence that resolves the plot [C018] [C061] or a stated prohibition suspended for a quotable benediction [C042].
Aesthetic judgment diverges in ways that are mode, not rank, and several critics explicitly refused to score them: reveals that obligate versus reveals that absolve [C008]; machinery that witnesses versus machinery that retrieves [C047]; carried guilt versus concealed guilt [C051]; climax as conversation versus climax as broadcast [C060]; paperwork ethics versus ceremony ethics [C092]; two opposite doctrines of emotional control on the same prompt [C072]; maxim closers versus gesture closers [C069] [C033]. These are the honest core of the "different intelligences" minority reading.
Corroboration from story-level evaluation is mostly aligned but not total. One material disagreement: the packet critics twice praised Astra's hydraulic reasoning on 0240 as load-bearing where GPT-5.6 Sol (high)'s mycelial mechanism restated the theme [C002] [C036]; the story-level evaluator preferred GPT-5.6 Sol (high) there on conceptual originality and prose, and the pair scored a near-tie (+0.217). Causal auditability is a real behaviour, but it is not uniformly what gets rewarded.
Neither writer escapes a template. Both fill a fixed kinship slot every time — Astra a dead or silenced mother who bequeathed an object and a method, GPT-5.6 Sol (high) a lost sibling absorbed into a machine or song [C080]. Astra's withholding is itself a repeatable closing device: deferral, unverifiability, or an unfinished task, five of five in one packet [C093], and Astra aphorises when a prompt presses [C003] [C023] and occasionally inserts tone words verbatim [C006] [C052]. GPT-5.6 Sol (high)'s terminal thesis and one-sentence lineation are equally constitutional [C011] [C040] [C050].
The evidence favours a *raised floor* over a special ceiling. GPT-5.6 Sol (high)'s ceiling is genuinely high — the packet critics' own reversals name its most conceptually ambitious pieces [C025] [C078] [C070] [C091], and story-level evaluation independently preferred it on 0342 (fractal recursion as structural motif), 0174 (calligraphy as philosophical instrument) and 0276 (transliteration and constellation integration). Astra's ceiling stories are strong but its distribution is what wins: 43 of 50 pairs, and in the extension cohort not a single loss. Its floor failures are narrow and specific (below). GPT-5.6 Sol (high)'s floor failures are causal — soft joints under baroque climaxes [C018], rules broken at the emotional peak [C042], powers declared at need [C061].
GPT-5.6 Sol (high), on evidence spanning multiple packets and corroborated by story-level evaluation: conceptual invention and contraption architecture [C019] [C078] [C085]; mandated objects turned into theses [C025]; systems whose technical rules *are* the moral argument [C091]; multi-function devices with sequential payoff [C070]; setting rendered as causal atmosphere [C039]; motif braiding and image density [C032]; seeded clue engineering [C059]; installed jeopardy — deadlines, enforcers, stated penalties — where Astra sometimes proceeds without opposition [C046]; and appetite for collective scale [C053].
Astra: everything organised around consequence — retained costs that generate later plot [C022] [C087], complicity [C034], auditable mechanism [C002] [C036], error-correcting plots [C081], objections as engine [C021], institutions that bargain rather than forbid [C005] [C035] [C017], administrative and consent-shaped climaxes [C012] [C049], priced wonder [C058], relinquished gifts [C056], held ambivalence [C079], permitted comedy inside grief [C045], and mundane cost-accounting [C024] [C064].
This split is partly prompt-dependent. GPT-5.6 Sol (high) is at full strength on civic and forensic premises and weaker on lyrical ones [C026] (single-packet, low confidence), and its advantage concentrates where a prompt rewards spectacle, jeopardy or exotic tone vocabulary.
0343 (+2.883) — enactment as advantage. Astra dramatises the required tone in two concrete beats and never names it; GPT-5.6 Sol (high) names both required terms and then explains the label [C007]. The renunciation of glory is staged as cartography — passenger names written where the discovery's title belongs, spellings checked aloud [C012]. The evaluator independently records the comparison stating the tone rather than dramatising it.
0199 (+3.000) — occupation versus oracle. Astra's protagonist reorders seed baskets by pioneer species [C044] and jokes inside her own competence [C045]; GPT-5.6 Sol (high)'s prophecy issues a plain-English imperative containing the prompt's verb [C088], and the required infinitive survives ungrammatically inside the prose [C043].
0260 (+3.000) — refused mercy. Astra names and declines the tactically useful forgiveness and files the protagonist's own signature as evidence [C004]; the ending is a physical arrangement, "held securely without anything having been made whole" [C003].
0367 (−1.500) — the single comparison win. Astra's method is explained rather than dramatised, its "accidentally prophetic" attribute confined to backstory guilt, and its rescue compressed into one clause [C026]; GPT-5.6 Sol (high) supplies the escalating mechanism and live obstacle its jeopardy habit reliably produces [C046]. This is the clearest instance of Astra's floor and GPT-5.6 Sol (high)'s ceiling meeting.
0154 (−0.083) — the cost of never naming. Astra's non-inscription strategy leaves a required tone unfulfilled as both phrase and concept, the explicit risk attached to that strategy [C030] [C007].
0240 (+0.217) — critics against evaluator. Astra's culvert hydraulics predict their own payoff [C002] [C036]; the evaluator preferred GPT-5.6 Sol (high)'s mycelial concept and prose. Auditable causality is a genuine behaviour, not an automatic scoring advantage.
0342 (−0.250) — GPT-5.6 Sol (high)'s ceiling. The recursive coin unifies attribute, object, setting and theme [C078], while Astra's handling of the same attribute stays descriptive; here it is GPT-5.6 Sol (high) that refuses consolation.