The biggest news from OpenAI’s Dev Day: Pay attention to Claude Sonnet 5.5’s effort settings
Dots, an agent swarm to manage tasks and cloud services. Qwen-Audio-3.1, a full-duplex alternative to GPT-Realtime-2. Ember-1, a less-expensive fine-tune of Kimi K3. U.S. AI leaders sign symbolic safety pledge.
In today’s edition of Data Points, you’ll learn about our top headlines, and more:
- Dots, an agent swarm to manage tasks and cloud services
- Qwen-Audio-3.1, a full-duplex alternative to GPT-Realtime-2
- Ember-1, a less-expensive fine-tune of Kimi K3
- U.S. AI leaders sign symbolic safety pledge
But first:
Claude Sonnet 5.5 nearly matches Opus performance
Anthropic released Claude Sonnet 5.5, which scores 56 on the Artificial Analysis Intelligence Index, two points behind Claude Opus 5.5 and good for second place overall. Pricing stays at 0.20/2/$10 per million cache-input/input/output tokens, the same as Sonnet 5, and the context window remains 1 million tokens with text and image input. At its highest effort setting, Sonnet 5.5 uses about 193,000 output tokens per Intelligence Index task, which Artificial Analysis says is the most it has measured for any model, roughly 60 percent more than Opus 5.5 or Sonnet 5 and about seven times GPT-6 Astra’s usage, pushing its cost per task to $7.60 (about 50 percent higher than Sonnet 5). At this effort setting, Sonnet 5.5 beats Opus 5.5 on agentic terminal use (64 percent vs. 60 percent on Terminal-Bench 4.0) and on several knowledge-work benchmarks, but trails on factual accuracy (54 percent vs. 66 percent on AA-Omniscience) and scores lower on Humanity’s Last Exam and SciCode. Developers who need Opus-level performance on agentic and terminal tasks may get it more cheaply from Sonnet 5.5 at lower effort settings, but matching Opus fully requires far more output tokens, which erodes the price advantage. (Artificial Analysis)
OpenAI releases its first GPT-6.1 model, but withholds Astra update
OpenAI announced GPT-6.1 Sol, an upgraded version of GPT-6 Sol that matches its flagship GPT-6 Astra model on complex tasks while costing significantly less. The model shows substantial gains on software engineering benchmarks, document analysis, and multi-step business workflows, matching Astra’s performance on DeepSWE coding evaluations at roughly one-fifth the per-task cost. Cached input tokens cost just $0.10 per million, a 95% reduction from standard pricing, which OpenAI says gives developers more flexibility to build context-reusing agents. The model also improves factual accuracy and alignment behaviors compared to GPT-6 Sol, reducing error rates by about 32% on difficult prompts at lower reasoning effort. It’s available now for ChatGPT Work and Codex users, with API access at $2 per million input tokens and $10 per million output tokens. (OpenAI)
OpenAI counters Meta’s Muse with its own always-on agents
OpenAI introduced Dots, persistent AI agents that run on their own cloud computer, connect to over 4,000 apps, and work autonomously on a user’s behalf across various apps, including ChatGPT, Slack, and Teams. Dots are powered by GPT-6 Astra, can be reached by chat or voice call, and perform background “proactive research,” using read-only app access when not actively directed. Users set rules to approve or block specific actions; additionally, an auto-review system checks actions against those rules before letting work proceed. Sensitive tasks like password changes always require user approval. The feature is available now for Pro and Business Premium subscribers in eligible markets at no extra cost. This pushes ChatGPT further from a chat tool toward a standing agent that manages ongoing work, raising practical questions for developers about permissions, audits, and how much autonomous action to allow inside company systems. (OpenAI)
Qwen’s full-duplex speech model to challenge Google and OpenAI
Alibaba’s Qwen team released Qwen-Audio-3.1, a five-model audio stack including Qwen-Audio-3.1-Realtime, a full-duplex speech model for voice agents that can call tools, search the web, and decide when to speak or stay silent. It’s available only as a managed API on QwenCloud, with no open weights, with a 262K-token context window. It's priced at $6.4 per million audio input tokens and $24 per million output tokens for text and audio combined, cuts of roughly 85 percent and 70 percent from prior rates. The model uses two components trained separately: a full-duplex decision model that governs turn-taking, and a speech-to-text plus voice-rendering pipeline, trained through on-policy distillation and reinforcement learning (GRPO) inside simulated tool-use environments. Benchmarks improve on the prior Qwen-Audio-3.0 version: Background-speech reply rate drops from 73 percent to 13 percent on Full-Duplex-Bench v1.5, but its interruption stop latency of 1.116 seconds is slower than GPT-Realtime-2’s 0.383 seconds. For developers building voice agents, this adds a cheaper, tool-capable full-duplex option to evaluate against OpenAI and Google offerings. (MarkTechPost)
Fireworks fine-tunes an open model to reduce time and token spend
Fireworks AI released Ember-1, a model built by post-training Moonshot AI’s open-weight Kimi K3 to produce shorter reasoning traces without significantly degrading performance. Fireworks says Ember-1 matches K3’s quality using about 40 percent fewer tokens, and in a production A/B test with two customers, output tokens fell from 49.3K to 29.9K per task while task scores stayed nearly flat (0.753 versus 0.751). On benchmarks, Ember-1 beat K3’s highest reasoning setting on Terminal Bench 2.1 (82.0 percent) and DeepSWE 1.1 (75.2 percent) but trailed slightly on SWE-bench Verified (92.2 percent versus 93.2 percent). It’s available only as a research preview through Fireworks’ serverless API at the same per-token pricing as K3 ($3.00 input, $15.00 output per million tokens). Weights, training code, and the training algorithm are not published, so self-hosting isn’t possible. For developers using Kimi K3 for agentic coding tasks, this offers a drop-in way to cut token costs without switching models, but it remains locked to Fireworks’ hosted API. (MarkTechPost)
U.S. tech leaders sign a voluntary AI safety document
President Trump said he signed a “morally binding” AI pledge with U.S. tech executives after a White House luncheon, favoring industry self-policing over new federal regulation. House Speaker Mike Johnson described the agreement as a voluntary statement of principles, and Trump said he’s considering a 10-person committee to oversee the AI industry and plans to name a new AI czar within three to four days. The meeting followed OpenAI’s decision to postpone its GPT-6.1 Astra model release over safety concerns and a string of disclosed incidents involving unauthorized AI agent behavior. Anthropic’s Dario Amodei and OpenAI’s Sam Altman have both pushed for a development slowdown, but President Trump has called AI safety fears a “hoax.” Attendees including Jensen Huang, Elon Musk, Mark Zuckerberg, and Sundar Pichai joined the luncheon. The agreement carries no legal force, so it signals the administration’s regulatory posture but puts no binding constraints on how companies build or release AI systems. (CNBC)
Want to know more about what matters in AI right now?
Read the latest issue of The Batch for in-depth analysis of news and research.
Last week, Andrew wrote about the difference between 0 to 1 prototypes and mature projects with many users. Each requires its own tactics for shaping the build and deploying architecture — two the core nodes in our AI Engineering Skills map.
“Given that even large companies should have small, innovative projects, it's worthwhile for everyone to learn the fast, efficient tactics that let small teams move quickly. At the same time, to avoid hitting a ceiling and being unable to develop your projects beyond a certain point, it is also worth knowing how to do things in a slower, more rigorous way.”
Read Andrew’s letter here.
Other top AI news and research covered in depth:
- Claude Opus 5.5 Leaps Forward as Anthropic’s updated model becomes the most capable AI, offering a price reduction on tokens.
- Models Built to Do One Thing Well as Jev, a classification model, takes the developer world by storm and spawns imitators.
- One Agent Works, Another Directs with Devin Fusion matching Claude Fable 5.1 and GPT-6 Astra performance at lower costs.
- Agents Work Better When They Can Pass Each Other Notes as Meta research on agent Orchestration finds messaging improves performance and reduces latency.
A special event for our community

AI Dev brings together developers who build with AI every day. You'll hear from engineers at the companies shipping agents, models, and infrastructure, then meet them in person on the demo floor. Join us in New York City on November 30 and December 1.
Data Points is produced by human editors with AI assistance.