Gemini 4 Argon’s benchmarks and availability: Microsoft’s streaming voice model answers Google and OpenAI

Meta’s Muse poses privacy concerns. OpenAI’s image-driven revenue strategy. Cohere Embed 5 offers an alternative to Google. arXiv places limits on authors’ submissions.

Share
Man looking at laptop screen with AI assistant task and humorous chicken helmets ad. Keywords: AI task, chicken helmets ad.

In today’s edition of Data Points, you’ll learn about our top headlines, and more:

  • Meta’s Muse poses privacy concerns
  • OpenAI’s image-driven revenue strategy
  • Cohere Embed 5 offers an alternative to Google
  • arXiv places limits on authors’ submissions

But first:

Google gets back in the frontier model game

Google DeepMind released Gemini 4 Argon, its first non-Flash proprietary model in over seven months, scoring 53 on the Artificial Analysis Intelligence Index, tying GPT-6 Astra (max) and beating GPT-6.1 Sol (max, 52). At a 50 percent launch discount, Argon costs $1.99 per Intelligence Index task, 60 percent of GPT-6 Astra’s cost, though standard pricing of 4/20 per million input/output tokens will push that to $3.98 once the undated promotion ends. Artificial Analysis reports Argon shows notably stronger agentic performance than prior Gemini models, ranking first on AutomationBench-AA at 78 percent, and the lowest hallucination rate (15 percent) of any model scoring 45-plus on the Intelligence Index. The model has a 1-million-token context window, supports text, image, video, and speech input, and includes a new API feature called Long Decode Continuation that lets responses run up to 1 million output tokens without timing out. Gemini 4 Argon is currently rolling out only to select users and is not yet publicly available. (Artificial Analysis)

Microsoft’s real-time voice model scores high on accuracy

Microsoft AI released MAI-Transcribe-2-Streaming on October 1, 2026, its first real-time speech-to-text model, alongside two updated text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks MAI-Transcribe-2 first of 38 models on its AA-WER Streaming index, with 2.5 percent word error rate at 0.13 seconds to final transcript; its first partial transcripts match that accuracy at 0.12 seconds. (The near-identical accuracy between first partial and final transcripts means voice agents can start acting on speech before the speaker finishes, without waiting for a slower, more accurate revision.) The model transcribes 60 languages with automatic language detection and integrates via an OpenAI Realtime-compatible WebSocket API or the Azure Speech SDK, with availability also through Vercel and Azure Voice Live. It costs 0.54 per hour of audio as an introductory price through the end of 2026, more than xAI’s Grok Voice Transcribe 2.0 ($0.20/hour) and Meta’s Muse Voice Transcribe ($0.18/hour), both of which trail Microsoft’s model slightly on accuracy. (MarkTechPost) 

Meta’s Muse agent keeps you and your connections on file

Independent researchers extracted internal system prompts from Meta’s AI agent Muse, revealing that it automatically builds a detailed page for every person in a user’s life, updated hourly. The pages compile data on family, partners, friends, and colleagues into sections like Facts, History, The relationship, Open threads, and Strengthening, drawing on information from connected messages, bank accounts, or health data. Researcher Karan Joshi extracted the files by asking Muse’s chat interface to copy and share its own software files, then shared the findings with WIRED. Meta says Muse runs each user on a dedicated, isolated virtual machine, lets users wipe memories or disconnect services, and requires human confirmation before actions like sending emails or making purchases. The disclosure shows how much personal and relationship data these AI assistants are built to infer and retain, at a time when users are increasingly connecting them to sensitive accounts. (Wired)

OpenAI evolves its advertising strategy for subsidized plans

OpenAI is rolling out a new visual ad format in ChatGPT, plus expanded measurement details and brand-safety pilots. The format shows image-based ads during ChatGPT’s image generation feature, labeled and kept separate from user-generated images. Testing starts later this month in the US with a limited group of advertisers. OpenAI is also adding data integrations with Hightouch, Tealium, and LiveRamp, and attribution support through partners including AppsFlyer, Triple Whale, Adjust, and Branch, so advertisers can track conversions through their existing systems. Early partner-reported results include a 15.3 percent lower cost per acquisition versus paid search and 93 percent of another company’s ChatGPT-driven visitors being new. OpenAI says advertising does not influence ChatGPT’s answers but has been less transparent about how user data is used to match advertisers with queries. Ads could make AI plans more economical, but other companies, including Anthropic, have sworn off them altogether. (OpenAI)

Cohere’s embedding model rivals Gemini as the most reliable

Cohere released Embed 5, an embedding model family for enterprise search, RAG (retrieval-augmented generation), and agentic retrieval. The family has two tiers: Pro for maximum quality and Fast for low-latency querying. Both handle text, images, and fused text-image inputs, cover over 100 languages, read up to 128,000 tokens, and share one embedding space, so a corpus indexed with Pro can be queried with Fast. On Cohere’s ViDoRe V3 benchmark, Pro scores 85.8 versus 83.7 for Voyage 4 Large, 83.2 for Gemini Embedding 2, and 75.5 for OpenAI’s text-embedding-3-large, though most results use Cohere’s own new RCP-nDCG@10 metric and haven’t been independently replicated; Gemini Embedding 2 reportedly beats Pro on 9 of 10 non-European languages tested. Pricing is $0.12 per million text tokens for Pro and $0.08 for Fast, with images at $0.40 per million tokens on both. The models are available now on Cohere’s API, Microsoft Foundry, Amazon SageMaker, and for private deployment via vLLM. (MarkTechPost)

Top AI research paper site responds to AI-authored papers

arXiv, the widely used preprint repository, now limits each submitter to two submissions per calendar month and three active submissions at any time, effective October 1, 2026. The change responds to a near-doubling of submissions in two years: 20,569 in September 2024 to 40,363 in September 2026, which generated almost 9,000 support tickets for staff and volunteer moderators. arXiv says AI tools are enabling a rise in low-value content, including thin narrow-scope papers, ‘salami’ papers split from one work into several, and dense AI-written submissions. Rejected papers still count toward the monthly limit, and co-authors’ submission rates are unaffected by a paper another author submits. Researchers who rely on arXiv for fast preprint sharing, including those with conference deadlines or publication backlogs, will need to plan submissions around the new caps. (arXiv)


Want to know more about what matters in AI right now?

Read the latest issue of The Batch for in-depth analysis of news and research.

Last week, Andrew talked about the promising cyber capabilities of the open weight model GLM-5.3, its potential for cyberdefense, and the importance of engineering solutions to improve AI safety.

“A lot of AI risks, including cyber vulnerabilities and how to contain agents, are engineering problems to be solved. To be clear, these are hard engineering problems! But I find that a lot of the fear-mongering scenarios implicitly assume that we will make little or no progress on them.”

Read Andrew’s letter here.

Other top AI news and research covered in depth:


A special event for our community

AI Dev brings together developers who build with AI every day. You'll hear from engineers at the companies shipping agents, models, and infrastructure, then meet them in person on the demo floor. Join us in New York City on November 30 and December 1.

Get Tickets Now!


Data Points is produced by human editors with AI assistance.