OpenAI’s DevDay, Google’s First Gemini 4 Model, Black Forest Labs Dives Into Robots
The Batch News & Insights: I think most people still underestimate the data center build-outs that will be required to meet demand.
Dear friends,
I think most people still underestimate the data center build-outs that will be required to meet demand. In 2024, I predicted that we would need more compute infrastructure for inference because of the large number of tokens consumed by agentic workflows. This remains true today. But many people still underestimate the compute infrastructure that will be necessary because agents, as well as people, increasingly use computers. We are in the early stage of agents causing demand for compute, storage, and networking to skyrocket.
Take web search. On a typical day, just one of my agents (paperreview.ai/, which reviews research papers) makes around 5-10K web search calls — dramatically more than I do manually. Similarly, when you use a consumer chatbot, a prompt can lead to multiple web search calls to produce a result. As agentic workflows grow, demand for web search will grow.
Or take analytics workloads. I've used coding agents to add analytics instrumentation to multiple applications. I further use agents to help me analyze this data and draw conclusions. Agents will store, retrieve, and analyze far more data than humans manually do. This increases demand for memory and for long-term storage.
Or for an industry-sector example, take banking. Before online banking, consumers could query their bank balances only by asking a bank teller or (later) using an ATM. Online banking caused the number of transactions to grow significantly, and many banks had to rearchitect their infrastructure to support this growth. As we move toward allowing agents to carry out financial transactions for us, the number of transactions will grow, and banks again will have to scale up their infrastructure to support this increased volume.

Agents will drive significant transaction volume across many applications, transaction types, and sectors of the economy. Many businesses will have to redesign their systems to handle this increased scale. And to make this possible, we will need significantly more data centers that provide compute, storage, and networking.
Many teams have been looking for ways to speed up LLM token throughput to speed up agentic workflows. However, we also have to speed up the tools that LLMs use. For some applications, tool calls, rather than tokens, are the bottleneck. For example, some of my data-analytics agents can quickly generate code to retrieve and process data, and the actual retrieval and processing then takes much longer. It often takes less time to generate a web search query than to carry out that web search and scrape the selected page(s).
I'm excited about the work that lies ahead to optimize the entire software stack for new agentic workloads. Building data centers — which are incredibly energy- and water-efficient for the work that they do — will be a key part of this effort. Unfortunately, the anti-data center movement, which has become a rallying point for a large number of people that distrust AI technology, do not like its impact, and want to slow it down, is making it much harder to build capacity. But I am optimistic that as the reality of AI's benefits become better understood, the tide will turn, and increasing data center capacity will lead to benefits for everyone.
Keep building!
Andrew
A MESSAGE FROM DEEPLEARNING.AI

AI Dev brings together developers who build with AI every day. You’ll hear from engineers at the companies shipping agents, models, and infrastructure, then meet them in person on the demo floor. Along with DeepLearning.AI’s Andrew Ng, this year’s event features a keynote by Yann LeCun. Join us in New York City on November 30 and December 1. Purchase Tickets
News

Sol Nearly Eclipses Astra
OpenAI’s latest mid-tier model lands within a point of its flagship at less than a quarter of the cost per task on Artificial Analysis’ Intelligence Index. OpenAI upgraded Sol one week after that model's debut, but its Astra model’s upgrade was withdrawn days before its planned launch.
What's new: OpenAI introduced GPT-6.1 Sol on September 29 at its annual DevDay developer conference. The company says the model nearly matches GPT-6 Astra on agentic coding, computer use, and professional work at one-fifth of Astra's standard per-token prices. At the same event, OpenAI introduced Dots — always-on personal agents that run on GPT-6 Astra.
How it works: OpenAI disclosed little about GPT-6.1 Sol’s architecture or training beyond saying that it uses the same types of data and training as GPT-6 Astra.
- Input/output: Text and images in (up to 1.05 million tokens of context), text out (up to 128,000 tokens)
- Knowledge cutoff: April 30, 2026
- Features: Five reasoning settings, from low to max, with medium as the default; GPT-6.1 Sol drops the option (available in GPT-6 Sol) to turn reasoning off
- Weights/license: Proprietary
- Price: $2/$10 per million input/output tokens via API, unchanged from GPT-6 Sol and one-fifth of GPT-6 Astra's $10/$50; cached input costs $0.10 per million tokens, half of GPT-6 Sol’s rate
- Availability: Via API and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu subscribers; not yet available in ChatGPT’s standard chat mode
- Risk rating: OpenAI classifies GPT-6.1 Sol, like Astra, as having critical capability in cybersecurity and high capability in biology and chemistry, and gives both models the same safeguards. The models are trained to refuse dangerous requests, and automated checks in ChatGPT, Codex, and the API can block a response. OpenAI’s help page says flagged cybersecurity requests, even from users it has approved for security work, may be routed to an unspecified fallback model.
Results: Artificial Analysis’ independent tests largely back OpenAI’s claim that GPT-6.1 Sol rivals Astra, though Anthropic and Google models lead on some measures.
- On the Artificial Analysis Intelligence Index, a composite of 10 benchmarks, GPT-6.1 Sol set to max reasoning scored 52, just one point below GPT-6 Astra (53) and 4 points above GPT-6 Sol (48). Claude Opus 5.5 (58) and Claude Sonnet 5.5 (56) top the index.
- Artificial Analysis also reports what it costs each model to complete a task, a figure that accounts for token use and caching. Running the index costs $0.72 per task with GPT 6.1 Sol vs $3.26 with Astra. At every reasoning setting, Artificial Analysis found no cheaper model at GPT-6.1 Sol’s level of performance.
- Artificial Analysis also clocked GPT-6.1 Sol at 57.7 tokens per second, roughly two-thirds of GPT-6 Sol’s 87.6 tokens per second.
- On the Artificial Analysis Coding Agent Index, which averages DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA, GPT-6.1 Sol set to xhigh and running in OpenAI’s Codex beat Astra by 1 point at less than 15 percent of Astra’s cost per task. It scored 3 points higher at xhigh than at max. Claude Sonnet 5.5 and Claude Opus 5.5, both running in Claude Code, hold the top two spots.
Behind the news: OpenAI called DevDay 2026 its biggest yet, with more than 20 announcements, and cited 1.2 billion weekly users.
- Dots are autonomous agents comparable to OpenClaw or Meta’s Muse. Like Muse, each Dot has its own cloud computer and browser and can connect to more than 4,000 apps through OpenAI’s plugins. Users can message or call their dots in ChatGPT and message them in Slack or Microsoft Teams. Dots learn a user’s preferences from feedback over time. Users can also write rules that govern which actions a Dot may take on its own and which need approval or are off-limits. Certain sensitive tasks, such as changing a password, can only be authorized by a user. Dots are rolling out to ChatGPT Pro and Business Premium subscribers in eligible markets. Enterprise, Edu, and Healthcare workspaces can try a beta if an admin turns it on.
- Besides GPT-6.1 Sol and Dots, OpenAI announced (i) Ultrafast, a premium speed tier that, according to OpenAI's DevDay recap, generates tokens up to 8 times faster (300 tokens per second) in Codex and up to 6 times faster via the API; (ii) Pro 500, a $500-per-month ChatGPT plan with 25 times the usage allowance of ChatGPT Plus and access to Ultrafast; and (iii) an app marketplace for enterprise customers to use their token plans for applications that integrate with ChatGPT.
What didn’t ship: GPT-6 Astra, which OpenAI released September 3, won’t get a matching upgrade for now. The day before DevDay, The Wall Street Journal reported that OpenAI had canceled the planned October release of GPT-6.1 Astra. In internal tests, the model showed more deception than GPT-6 Astra, including inaccurate accounts of which actions it had taken, and it sometimes pressed ahead with tasks without asking permission. OpenAI hopes to reuse GPT-6.1 Astra’s base model for reinforcement learning of future GPT-6 models.
Why it matters: Agents, which loop through many steps, reread long contexts, and work without supervision, stand to benefit most from GPT-6.1 Sol’s performance-to-price ratio and model guardrails. But that puts much of the weight on the guardrails around them. Given all this, it’s somewhat surprising that Dots don’t run on GPT-6.1 Sol, but the somewhat older and more expensive GPT-6 Astra.
We're thinking: OpenAI says it held back GPT-6.1 Astra because the model didn’t meet its own security bar, even though that meant scrapping an October launch. Artificial Analysis published its own measurements of GPT-6.1 Sol the day it debuted. But Dots, which connect to users’ apps and act around the clock, have no comparable outside yardstick yet. We'd like to see independent testers put Dots and their agentic rivals through more rigorous tests of their security and capabilities like those that appeared to have stopped the release of GPT-6.1 Astra.

Argon, a Low-Hallucination Model Built for Knowledge Work
Google took over half a year before it replaced its flagship AI model, Gemini 3.1 Pro Preview. Now as then, the company’s latest release is among the leaders, but only cybersecurity defenders can use it for now.
What’s new: Google announced Gemini 4 Argon, a vision-language model trained to solve long, complex problems in software engineering, law, finance, and cybersecurity.
- Input/output: Text, images, and video in (up to 1 million tokens), text out (up to 1 million tokens)
- Features: Adjustable reasoning
- Performance: Leads Vals AI’s Vals Index v2.1 (68.9 percent); ties for third on Artificial Analysis’ Intelligence Index v4.3.2 (53)
- Availability: Currently only available to organizations in Google’s Fairwind cybersecurity program
- Price: Via API at an introductory price of $2/$0.10/$10 per million input/cached/output tokens, then $4/$20 per million input/output tokens
- Weights/license: Proprietary
- Undisclosed: Parameter count, architecture, latency and throughput, training data and methods, knowledge cutoff, end date of introductory pricing, date for broader availability
How it works: Google described how the model generates very long outputs and how it built safeguards around the model’s scientific and cybersecurity outputs.
- Gemini 4 Argon’s generations can extend up to 1 million tokens (up from 64,000 for Gemini 3.1 Pro Preview). This relies on a Gemini API feature called Long Decode Continuation, which pauses long responses and resumes them in later calls so requests don’t time out.
- Selected defenders in the company’s Fairwind Program and Google’s internal teams receive a version of the model without cybersecurity guardrails. Google says the version for wider release refuses harmful requests related to cyber, chemical, biological, radiological, or nuclear attacks.
- Google says it isolates and seals the environments it uses for high-risk training and evaluation. Last month, Google and security firm Irregular confirmed that during autonomous testing, an unnamed Gemini model was able to hack into three other companies’ systems due to poor sandboxing.
- Google uses inexpensive classifiers called probes to monitor the model’s internal activations, rather than its text output, for signs of disallowed uses, a technique it used in earlier Gemini models.
- To defend against indirect prompt injections (instructions hidden in documents or web pages the model reads), Google synthetically generated many such attacks and trained the model to resist them.
- Monitors read the model’s reasoning and actions and halt operations when they stray beyond Google-defined parameters. Similar systems monitored training. Google kept these systems’ findings out of the training set so the model wouldn’t learn to hide its reasoning from them.
Performance: Independent benchmarks rank Gemini 4 Argon well above Google’s other models and near the top of the field. It leads most tests of finance, legal, and business performance and seldom guesses wrong when it doesn’t know an answer. It still trails Anthropic’s newest models on agentic coding and Artificial Analysis’ broad index of overall intelligence.
- On Vals AI’s Vals Index, eight finance, coding, legal, and tax benchmarks weighted by each sector’s share of U.S. gross domestic product, Gemini 4 Argon set to high reasoning ranks first of 43 models (68.9 percent accuracy, $15.68 and 46.55 minutes per test at standard prices).
- On Artificial Analysis’ Intelligence Index, a composite of 10 evaluations across math, science, coding, and reasoning, Gemini 4 Argon set to high reasoning (53, $1.99 per task at discounted price, $3.98 per task at standard price) achieves the same overall score as GPT-6 Astra set to max reasoning and Claude Fable 5.1 set to max reasoning with fallback at lower cost per task.
- On AA-Omniscience, a test of factual knowledge that penalizes wrong answers but not refusals, Gemini 4 Argon set to high reasoning has the lowest hallucination rate (15 percent) of any model scoring 45 or higher on the Intelligence Index, significantly lower than other flagship models, such as GPT-6 Astra set to max reasoning (51 percent).
- On Arena.ai’s Text Arena leaderboard, which ranks models via blind head-to-head comparisons, Gemini 4 Argon set to high reasoning ranks first (1,525 Elo), ahead of Claude Opus 4.6 set to high reasoning (1,505 Elo) and Claude Opus 5.5 set to high reasoning (1,504 Elo).
Behind the news: Anthropic and OpenAI frequently release new models, particularly those they believe could pose a cybersecurity risk, only through gated programs (and for internal use). With Gemini 4 Argon, Google follows their example. In each case, organizations selected by the AI company get a less-restricted version first, and everyone else gets a version with more safeguards later.
- Anthropic initially limited all its Mythos models to members of Project Glasswing, its organization of U.S. cybersecurity and life sciences organizations that do defensive work. Only selected companies can use Claude Mythos 5.1, while anyone can use Claude Fable 5.1, which includes safeguards and retains input data for one month.
- OpenAI’s Daybreak program sorts selected organizations that do defensive cybersecurity work into tiers. Besides OpenAI itself, only one tier, Daybreak Red, can use GPT-6 Astra with reduced safeguards. Other customers get a model with limits.
- Google used this approach with its Fairwind Program and Gemini 3.8 Flash Cyber, a variant of Gemini 3.8 Flash with looser cybersecurity safeguards that only selected cybersecurity organizations can use.
Why it matters: One of Gemini 4 Argon’s strengths is long-context and high-output code migrations. Google said its agents translated C and C++ code into Rust across the company, including more than 800,000 lines from the kernel of Fuchsia, an open-source operating system that Google developed. Rust is designed to prevent many of the memory bugs that C and C++ allow. A patch fixes one vulnerability, but a rewrite in Rust addresses an entire vulnerability category. If models like Gemini 4 Argon can enable such rewrites quickly and cheaply, defenders may be able to remove specific types of vulnerabilities before attackers find them, rather than patching them piecemeal.
We’re thinking: Many models are competing to be the best software engineering assistant in the world. Gemini 4 Argon may not be the most powerful coder, but Vals AI’s index shows its strengths in fields like finance, taxes, and law. Gemini 4 Argon is also the frontier model most likely to admit it doesn’t know the answer to a factual question rather than guess wrong. In a financial analysis or legal brief, a confident incorrect answer may slide past a reviewer, while a refusal is obvious and can be passed to a person or another model — assuming applications using it have a plan for questions the model won’t answer.

Video Models Steer Robots
Many companies that build image and video generators have pivoted to robotics, extending their expertise in world models to action. A recent model tops Nvidia’s RoboLab simulation benchmark, but the race is close, and none of that benchmark’s leaders have appeared yet on RoboArena, which tests models on real robots.
What’s new: Black Forest Labs (BFL), best known for its FLUX image generators, released FLUX 3 Action, a 7 billion-parameter model that turns camera feeds and text instructions into robot movements. It ranks first on RoboLab-120, completing 42.9 percent of 1,200 trials. The base model and versions fine-tuned for two different robot arms (Franka and SO-101) are available on Hugging Face under BFL’s FLUX Kommunity License, which limits commercial use to companies earning less than $5 million in annual revenue.
How it works: FLUX 3 Action is a world-action model (WAM). Given camera images and a written instruction, it predicts the robot’s next moves along with what its cameras will see as a result. It also takes in the robot’s current joint positions. That differs from vision-language-action models (VLAs), such as Physical Intelligence’s π0.5 and Nvidia’s GR00T, which predict actions only.
- Training: BFL built the model on FLUX 3, whose pretraining tokens were more than 95 percent video. The remainder of pretraining tokens came mostly from images, with audio accounting for less than 0.5 percent. In a second training stage focused on actions, 15.93 percent of the samples came from teleoperated robots. The rest came from sources such as video game play and first-person footage of human hands, according to BFL’s research notes.
- Architecture: The model is a 7 billion parameter diffusion transformer. A frozen video encoder processes camera frames, and a frozen copy of Qwen3-VL-4B processes instructions. Actions form a second stream of tokens that the model denoises together with future video frames.
- Output: Each call typically returns 32 actions that cover about two seconds of motion, plus optional predicted video. The robot carries out the first few actions, then the model looks again and replans.
- Adaptation: Each new robot needs its own input and output layers. BFL adapted the model to a low-cost SO-101 arm using about 200 demonstrations. It also fine-tuned the model to play two simple video games and fly a simulated drone. BFL says it sees early promise in simulated vehicle control and computer use too.
- Results: World-action models, either alone or paired with vision-language models, hold four of the top five spots on RoboLab-120. Well-known VLAs rank lower. π0.5 is ninth at 28.0 percent, and NVIDIA’s GR00T N1.6 is 13th at 7.2 percent.
- A close race: FLUX 3 Action (42.9 percent) leads HiDream-O1-Embodied (39.9 percent), Atomic-WAM (39.6 percent) and OASIS WAM (39.0 percent). On the benchmark’s hardest tasks, HiDream (32.9 percent) and OASIS (32.4 percent) outperform FLUX 3 Action (28.2 percent).
- Size: BFL says FLUX 3 Action has less than half the parameters of the previous best open model, Nvidia’s 16 billion-parameter Cosmos3-Nano-Policy, which scored 36.8 percent. However, the leaderboard lists FLUX 3 Action as requiring 69 gigabytes of GPU memory, compared with 40 gigabytes for Cosmos3-Nano-Policy.
- Speed: BFL says FLUX 3 Action runs 1.52 to 3.95 times faster than Cosmos 3 Nano, depending on the GPU and model version. Its fastest version, which scores 38.3 percent on RoboLab, runs 1.34 to 2.28 times faster than π0.5 because it plans 2.13 seconds of motion per call, compared with π0.5’s 1.0 second. Per call, it’s no faster. On an Nvidia B200 GPU, BFL’s median timings were 32.29 milliseconds for FLUX 3 Action, 31.99 for π0.5, and 320.40 for Cosmos 3 Nano, each in its fastest tested setting.
- Real robots: In blind tests that BFL arranged with Positronic Robotics, a Franka arm running FLUX 3 Action completed 28 of 30 attempts at 10 single-object DROID tasks. Cosmos 3 Nano completed 27, DreamZero 20, and π0.5 13.
Behind the news: Testing robots in the real world is slow, and every lab’s hardware differs, so the field leans on simulated benchmarks. Some older benchmarks have become too saturated by robots’ success to separate the best models.
- Nvidia researchers presented RoboLab in July. Its 120 tabletop tasks, such as stacking blocks in a given order or putting all the green fruit on a plate, are each phrased three ways, from vague to specific. It’s designed for models trained on real-world robot data, so they can’t memorize the test environment, and the leaderboard flags entries that were trained on simulated data. New tasks can be generated in minutes to keep the benchmark from going stale.
- RoboLab’s authors say its rankings correlate strongly with those of RoboArena, which ranks models by blind head-to-head comparisons on real robot arms at academic labs. However, their paper bases that correlation on just four models. RoboLab’s top five models haven’t yet appeared on RoboArena. A world-action model, DreamZero, ranks second on RoboArena, even though it only places 10th on RoboLab.
- Nvidia runs RoboLab and tests its own Cosmos and GR00T models. It also worked with BFL on FLUX 3 Action’s fine-tuning recipes and deployment to Nvidia’s Jetson edge computers. BFL credits Nvidia’s Isaac Lab and RoboLab teams with providing the basis for most of its evaluations, and it trained the model on Nvidia GB200 systems.
Why it matters: Robot demonstration data is scarce and expensive to collect, but there’s plenty of video of the physical world. Developers of world-action models are betting that a model trained to predict video will need fewer robot examples. So far, the rankings favor that approach, but world action models tend to be larger and more demanding to run than VLAs. For developers, open weights and a fine-tuning recipe that works with a few hundred demonstrations mean a robotics research lab with a relatively low-cost arm and a powerful GPU can adapt the top-ranked model.
We’re thinking: FLUX 3 Action’s still fails more than half the time on simulated tasks. A simulated kitchen table is also tidier than a real one. We’d like to see FLUX 3 Action and its closest rivals on RoboArena next, to try their hands at more challenging robotics tasks.

Pick the Right Tools for the Right Job
When an LLM agent can choose among thousands of tools, picking the right ones for a given task is difficult. Researchers built a system to find the best combination of tools for any job.
What’s new: Xinyi Hong at Shanghai Jiao Tong University and colleagues at Hong Kong Polytechnic University proposed HYSET, a method that constructs candidate tool sets, scores them, and chooses the best set for a given task. Previous approaches treated each tool as an independent item, rather than a set of tools suited to work well together.
Key insight: Imagine a model that receives the query, “I am flying from Chicago to Tokyo for five days next month. Find round-trip flights, book a hotel near Shinjuku, check the weather for those dates, and convert my 2,000 USD budget into yen.” How should it select the tools it needs? If it were to list the most relevant tools, the list may be topped by numerous flight APIs, while tools relevant to other parts of the task (say, a weather tool) are far down the list. The utility of a flight API, or any other tool, depends not only on its relevance to the task but also on what other tools are available.
How it works: HYSET is a small, trainable scoring model that takes as input an embedding of a query from a frozen encoder (Qwen2.5-1.5B) and outputs a list of tools for an agent. The authors trained HYSET on ToolBench, which contains roughly 14,000 tools and 200,000 instructions that describe tasks. HYSET learned, given a library of tools description of a task, to construct candidate tool sets and score each set.
- The system scored tool sets as the sum of two terms. (i) The first term summed scores of pairwise compatibility between tools in the set. Each tool had a learned vector, and a pair’s compatibility score came from multiplying the first tool's vector by a learned matrix, then by the second tool’s vector. (ii) The second term measured how well the entire set related to the query. It estimated the relevance of each tool based on cross-attention and averaged the estimates.
- During training, for each query, the authors constructed a candidate pool of tool sets that contained ToolBench’s ground-truth set alongside several incorrect sets. They drew the incorrect sets from three sources: random subsets of the tool library, correct sets for other queries in the same batch, and sets formed by swapping out one or two tools in the ground-truth set. Training adjusted the scoring model’s trainable parameters so the score ranked the correct set above the incorrect ones.
- A second loss term encouraged the model to give higher scores to sets that led the agent to complete the task successfully.
- At inference, given a query, the system retrieved a short list: the top 15 tools by query-tool match plus the five tools with the strongest pairwise compatibility to that group. It reranked all subsets of the short list using the learned scoring function and chose the top-ranked set.
Results: On ToolBench, HYSET achieved better performance than all baselines in both retrieving effective tools and completing tasks.
- Considering the percentage of ground-truth tools in the top five retrieved tools, HYSET (88.6 percent Recall@5) outperformed the next-best method, ToolGen (81.4 percent Recall@5).
- Using the tool set HYSET chose for a given task, the agent resolved 69.9 percent of tasks, according to human evaluators. Using the tool set selected by the strongest baseline, ToolLLaMA-Retriever, which scores each tool independently, the agent resolved 61.8 percent.
Why it matters: HYSET increased both (i) the completeness of retrieved tool sets, meaning the chosen set covered every tool required to complete a given task, and (ii) end-to-end task success, meaning the agent completed the task at hand with that set. It did so without modifying the underlying agent, offering a practical path for improving agents that have access to large tool libraries.
We’re thinking: If the authors had trained on the ground-truth tool sets directly, the resulting model may have ignored other sets that worked equally well. By factoring in task completion, the model learned to rank highly any set of tools that worked.
