AI Twitter Highlights · 2026-10-04

AI Twitter/X Highlights Digest · 2026-10-04 (Sun)

Key Takeaways

  1. Codex / OpenCode both crowd-source pain points in public: OpenAI’s Tibo asks what’s missing in Codex, while OpenCode 2 asks for biggest issues—product iteration happens on the timeline.
  2. On-device and distribution: OpenClaw’s Android app stays stuck in Google review for over a week while v2026.9.8 adds GPT-6.1 Sol; Pi Durable demos a phone-native, live-editable, multiplayer agent.
  3. Models into IDEs / harnesses: Antigravity ships Opus 5.5 / Sonnet 5.5 for paid users; Cline offers Ling 3.1 Flash free for a limited time; AutoCompact-style work trains compaction into the model; Claude Code 2.1.289 and Copilot’s side-by-side review target agent workflows.

1. Codex: official public wishlist

Takeaway: @thsottiaux (OpenAI) asks, “What’s one thing that’s missing in codex that you wish we had?”—and the replies flood in. Separately, community posts still push GPT-6.1 Sol at high effort as a stable daily driver with less model-switching.

Why it matters: A flagship coding agent runs requirements gathering in public, which surfaces real friction faster than a changelog.

Codex wishlist


2. OpenClaw: Android review limbo + v2026.9.8 with GPT-6.1 Sol

Takeaway: @steipete says OpenClaw’s Android app has been stuck in Google review for over a week and asks for help. The same day, @openclaw ships v2026.9.8: GPT-6.1 Sol support, agent replies routed back correctly, lower memory use, plus update and Windows startup fixes (43 PRs / 8 contributors).

Why it matters: Store-gate distribution friction sits next to a desktop/CLI release that already tracks frontier models—“can we ship the app” becomes part of the product story.

OpenClaw Android and release


3. Antigravity: Opus 5.5 / Sonnet 5.5 for paid users

Takeaway: @_mohansolo confirms Antigravity added Claude Opus 5.5 and Sonnet 5.5 for all paid users, framing it as access to the best frontier models, with a broader Gemini 4 Argon rollout teased.

Why it matters: An independent coding product packs the newest Claude generation into its paid tier, competing with single-lab IDE narratives.

Antigravity Opus Sonnet 5.5


4. Pi Durable: on-phone multiplayer agents

Takeaway: @badlogicgames spends the weekend building a personal Pi Durable project on Android: fully on-device (no cloud VM), live-editable, multiplayer, any provider/model, plus artifacts—claiming it works better than Claude for Android or ChatGPT for Android. Follow-ups show ngrok-exposed multiplayer, hot-reload self-modification (“building pim with pim”), and a January belief that agents should run on phones, not cloud sandboxes. Explicitly not an official Earendil product—just a Durable stress test.

Why it matters: Pushing a Durable runtime onto a pocket device challenges the default that agents must live in cloud sandboxes.

Pi Durable on phone


5. OpenCode 2: founder asks for the biggest issues

Takeaway: @thdxr posts, “what are your biggest issues with OpenCode 2?” and gets a high-volume reply thread. It mirrors the same-day Codex wishlist—both products converging requirements in public.

Why it matters: The next major open coding-agent release ties its feedback loop directly to the timeline.

OpenCode 2 issues


6. T3 Code: a usage view that shows where spend goes

Takeaway: @theo ships more improvements to the T3 Code usage view—clearer spend attribution and how models behave on your own data—and notes it measures all Claude Code and Codex usage on your machines, not only usage inside T3 Code.

Why it matters: Once multi-harness / multi-model is normal, spend and quota visualization becomes a core coding-UI feature.

T3 Code usage view


7. AutoCompact: training models to decide when to compact

Takeaway: @omarsar0 connects AutoHarness → AutoContext → AutoCompact: models increasingly absorb work that used to live only in the harness. AutoCompact trains agents to choose when to compact, what working state to keep, and how to resume; after judge-corrected trajectories plus SFT/RL, pass rates rise +9.2 on SWE-bench Verified and +5.0 on SWE-PolyBench Verified—even with a 256K window that never overflows. A same-day CMU harness-learning paper trains an RL proposer to edit harness code while freezing the solver weights.

Why it matters: The next coding-agent leap is model–harness co-design—“when to compact / which harness edit”—not only bigger context windows.

AutoCompact paper


8. Cline: Ling 3.1 Flash free until October 13

Takeaway: @cline announces Ling 3.1 Flash inside Cline, free until October 13. The MoE is 560B total / 25B active and is positioned alongside frontier open weights like Kimi K3 and DeepSeek V4 Pro.

Why it matters: Coding IDEs keep using time-boxed free frontier open weights to lower trial cost and drive model switching.

Cline Ling 3.1 Flash


9. Claude Code 2.1.289: teammates can spawn shared agents

Takeaway: @ClaudeCodeLog notes Claude Code 2.1.289 is live (~27 CLI changes). Highlights: teammates can spawn shared agents via agent.spawn; agent IDs are unified with clearer idle/waiting states; Read deny rules now cover @-mentioned files so mentions can’t bypass read restrictions; plus plugin/MCP sign-in and terminal freeze fixes.

Why it matters: Multi-agent collaboration moves from “open more sessions by hand” toward programmatic shared agents—with tighter permission rules alongside.

Claude Code 2.1.289


10. GitHub Copilot: side-by-side review for agent code

Takeaway: @github highlights the Copilot app’s side-by-side review: instead of tab-hopping while reviewing agent-written code, keep the diff, terminal, and browser together to check, run, and preview.

Why it matters: As agent output volume rises, the bottleneck shifts from writing code to verifying it in one surface—IDEs/review UIs are being rebuilt for agent workflows.

GitHub Copilot side-by-side