AI Twitter Highlights · 2026-09-28

AI Twitter/X Highlights Digest · 2026-09-28 (Mon)

Key Takeaways

  1. Opus 5.5 keeps dominating the timeline — from “magic” praise to burning through Max quotas, plus ~4–6× limit headroom vs Fable 5.1; some runners say OpenClaw + Opus completes tasks better than GPT-6 Astra.
  2. antirez argues Astra needs a new skillset; Codex is being asked to thin CI; UFO launches as a multiplayer agent harness OS.
  3. Cline ships Ember-1 (Kimi K3 post-train, ~40% fewer tokens); GitHub Copilot pushes parallel agents; plus DeepSeek Harness plugins, Jev-as-a-Judge, and a Chinese-timeline pivot to OpenCode + DeepSeek after a Claude ban.

1. Opus 5.5: from “magic” to maxing Max

Summary: @dhh wrote that Anthropic “did magic with Opus 5.5.” @theo said he burned through roughly 3.5 $200 Claude Code accounts in five days — Opus 5.5 is incredible and the Max sub is a steal. He also broke down the quota feel: Opus is gentler on cost, Fable is capped at 50% of the sub, so moving Fable 5.1 High → Opus 5.5 High is about a 4.3× limit bump (and ~6.6× from Fable xhigh). @AravSrinivas migrated 50–100 workflows from Fable 5.1 to Opus with minimal differences, but still feels FOMO about not using the “smarter” model.

Why it matters: The debate has shifted from raw IQ to how far a quota goes and whether orchestration stacks survive a model swap.

Opus vs Fable limits


2. Opus 5.5 × OpenClaw: finishing work better than Astra?

Summary: @garrytan said Opus 5.5 with OpenClaw feels “strangely smarter” and better at completing tasks than GPT-6 Astra — and that the gap was surprising. Model × harness combos are back on center stage.

Why it matters: The same day people praise Opus quota efficiency, others put it inside a concrete agent runtime and compare completion rates — default stacks may swap model and shell together.

Opus with OpenClaw


3. antirez: struggling with GPT-6 Astra may mean outdated skills

Summary: @antirez notes that “good programmers” say they use GPT-6 Astra without good results. His explanation is blunt: yesterday’s good programming and today’s required skillset overlap but are no longer the same.

Why it matters: Reframes “the model is bad” as a human–AI collaboration skill migration — complementary to quota and harness talk.

antirez on Astra skillset


4. Codex: let the model decide which CI tests to run

Summary: In a thread about CI becoming every team’s top bottleneck, @steipete praised sponsor @useblacksmith but said load still needs spreading; his plan is to let Codex decide which tests actually need to run, slash CI, and run tests hourly. Separately, Chinese timeline notes from Codex app lead @ajambrosino: whenever a new model drops, ask it to rewrite the Codex app in GPUI + Rust.

Why it matters: Coding agents are expanding from “write code” to “decide what to test” — CI spend is becoming part of the agent product surface.

Codex for CI


5. UFO: multiplayer agent orchestrator / memory system is live

Summary: GitHub Copilot co-creator @alexgraveley announced UFO — a scalable multiplayer agent orchestrator and memory system built for agent-first business pain points. The product account calls it a multiplayer agent harness OS you can run a company on, hosted or open source, good at code and work (ufo.ai). This is a product-go-live narrative, not a claim that it was “just open-sourced for the first time today.”

Why it matters: Another independent harness grown out of big-lab agent experience, pitched as a company OS rather than an IDE plugin.

UFO harness


6. Cline: Ember-1 (Kimi K3 post-train, ~40% fewer tokens)

Summary: @cline highlighted Fireworks Research’s Ember-1: built on Kimi K3, using roughly 40% fewer tokens at matched benchmark performance by post-training the model to think less repetitively. In a live coding A/B test it used ~71% fewer reasoning tokens and ~39% fewer total tokens at the same success rate. Available in Cline (including Desktop); @omarsar0 frames it as a meaningful push on the token Pareto frontier.

Why it matters: In agent loops, “thinking too much” is a bill — specialized post-training is now aimed directly at that cost.

Cline Ember-1


7. GitHub Copilot: run agents in parallel in the app

Summary: @github reminded users you can run agents in parallel in the GitHub Copilot app — each session gets its own Git worktree and context so you can build, review, and test at the same time.

Why it matters: The default big-vendor client is making multi-worktree parallelism a first-class feature, matching the multi-session story of indie harnesses.

Copilot parallel agents


8. DeepSeek Harness: dsh-better-sidebar as a base plugin

Summary: @tianyi continued the DSH plugin series with dsh-better-sidebar: sidebars, bottom bars, split panes, and floating panels for DeepSeek Harness, positioned as a base capability other plugins can build on. After official basic sidebar support landed, the plugin reuses those component APIs and extends them. It continues the earlier narrative that ~60% of users install at least one third-party plugin.

Why it matters: Domestic harnesses are treating plugin composability as product differentiation, not just a model shell.

DeepSeek Harness sidebar


9. Jev-as-a-Judge: catch alignment failures with a classifier

Summary: @omarsar0 covered using Jev as a cheap alignment/failure detector: ask one generic yes/no about a model response and use the probability as a score. With no extra training, median AUROC is about 0.886; across 19 benchmarks a Jev pass cost about $0.30 vs about $18.96 for the LLM judges those benches use. A follow-up argues harness builders should deliberately mix System One (classification) and System Two (generation) models.

Why it matters: After Jev Router, Jev is moving into the judge/gate layer — cheap calibrated decision models are becoming standard harness parts.

Jev as a Judge


10. After a Claude ban: OpenCode + DeepSeek as a fallback stack

Summary: @imwsl90 said getting banned on Claude clarified things — most of the real work already ran on the DeepSeek API. No appeal planned; the path forward is herdr + OpenCode + DeepSeek, much cheaper than Claude, with an Omarchy switch from Claude to OpenCode done in a few sentences. A live Chinese-timeline sample of “default harness is replaceable.”

Why it matters: When a subscription account is a single point of failure, portable open harnesses plus open/domestic models matter more than any one SOTA checkpoint.

OpenCode DeepSeek