Altworld's Hemmingway-1 Skips the Fluff and Delivers Paste-Ready Messages
Altworld's 27B fine-tune of Qwen3.8 tops GPT-6 Astra on everyday message writing and reads as human 26 points more often.
- Altworld released Hemmingway-1, a 27B Qwen3.8 fine-tune for everyday message writing, Apache-2.0 licensed.
- Beats GPT-6 Astra by 50 points on CommunicationBench across 80 blind matchups.
- Wins 72% of hard asks versus GPT-6 Astra's 9%; 26 points clear on human-likeness.
- Placed 3rd on public EQ-Bench 4, ahead of GPT-5.5, Opus 4.7 and Opus 4.8.
- Returns just the message text instead of options and preambles, unlike most competitors.
- 262K context, runs on vLLM/SGLang/Transformers; English-only, not for medical, legal, financial use.
Hemmingway-1 targets paste-ready messages
General chatbots often surround a requested message with introductions, alternatives, and tone notes. Altworld’s model card presents Hemmingway-1 as a simpler option: ask for a text, email, or reply, and receive the copy itself. The release applies a 27B open-weight model to short, tone-sensitive communication, a common workload that conventional reasoning and coding benchmarks rarely measure.
Built for the compose box
| Model size | 27 billion parameters |
|---|---|
| Base model | Qwen3.8-27B |
| Maximum context | 262,144 tokens |
| License | Apache-2.0 |
| Distribution | Hugging Face weights, quantizations, hosted playground, Mac app, and Android app |
Altworld released the weights alongside a hosted playground and a public code repository. The Apache-2.0 license permits commercial use, modification, and redistribution subject to its terms.
Eighty prompts, pairwise judging
Altworld created CommunicationBench from 80 real-world writing requests. For each matchup, a judge model received two shuffled answers to the same prompt without model labels. Altworld says it ran comparisons in both orders and used a judge model distinct from the systems being evaluated, reducing label and ordering bias.
On Altworld’s reported scoring scale, Hemmingway-1 ranks above Fable 5.1 and leads GPT-6 Astra by 50 points. The published chart also places Kimi K3, GLM-5.3, Grok 4.6, and DeepSeek V4 Pro behind it.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.