What we know about preview model Ox Alpha: An experimental model takes DeepSeek-V4-Flash multimodal
GLM-5.3 jumps to top of open weight leaderboards. Harvey’s Tenet offers law firms a specialized model. OpenAI introduces safety program for data-conscious customers. New Google research gives the transformer a makeover.
Welcome back after our three-week summer hiatus! We know some of you missed us. We missed you too. Now we’re back!
In today’s edition of Data Points, you’ll learn about our top headlines, and more:
- GLM-5.3 jumps to top of open weight leaderboards
- Harvey’s Tenet offers law firms a specialized model
- OpenAI introduces safety program for data-conscious customers
- New Google research gives the transformer a makeover
But first:
Mystery vision-language model impresses, puzzles early testers
Ox Alpha appeared on OpenRouter and OpenCode on August 20. The model is free, multimodal, has a 1-million-token context window, and offers 100 trillion tokens per day for a week at no cost. Early evidence for its maker pointed to Chinese AI lab Zhipu: matching tokenizer math, video encoding pipelines, near-identical responses, and a pattern of live-testing unreleased models under ‘Alpha’ codenames. Ox Alpha scored 87.5 percent on Kingbench, just behind GLM-5.3’s 91.25 percent, but has vision and video capabilities GLM-5.3 lacks publicly, suggesting an upgraded, unreleased variant. DeepSeek was another candidate, given V4’s similar ‘Hunter Alpha’ debut, though its recent price hikes on V4-Flash and mismatched tokenization weakened the case. The latest wrinkle: Ox Alpha uses OpenAI’s cl100k_base tokenizer, which undercuts the Chinese-origin theories and has shifted speculation toward Microsoft’s upcoming MAI 2. (WCCF Tech)
DeepSeek adds vision capabilities to popular V4-Flash model
DeepSeek released V4-Flash-Vision-Exp, an experimental multimodal model, on its API platform. The company says it matches the text-only V4-Flash on agents, reasoning, and world knowledge while posting notably stronger multimodal agent benchmark scores, scores DeepSeek claims rival Anthropic’s Opus-4.8. (The Opus-4.8 comparison is self-reported and worth holding lightly until independent testing arrives.) Developers can access it via model=’deepseek-v4-flash-vision-exp’. DeepSeek also shipped Harness 0.1.1 with out-of-the-box support for the new model. (DeepSeek via X)
Zhipu’s update posts a 60 on Artificial Analysis’ Intelligence Index
Z.ai released GLM-5.3, which reuses GLM-5.2’s 743-billion-parameter base unchanged, gaining in performance via scaled post-training alone. The jump is most notable on long-horizon coding: Terminal-Bench 3.0 climbs from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9. Cybersecurity gains were a bigger surprise: Z.ai added vulnerability-discovery data expecting modest single-bug gains, but as training scaled the model began forming coherent multi-step exploitation plans, pushing CyberGym to 84.5%, ahead of Mythos 5 and GPT-5.6 Sol. It still trails Claude Fable 5 on tougher coding benchmarks and lags Mythos 5 on ExploitBench and ExploitGym. The model is live on Z.ai’s API, Coding Plan, and ZCode; weights follow in roughly two weeks, pending safety evaluation. (MarkTechPost)
Fine-tuned legal analysis models hit the market
Harvey released Tenet, a research preview of its first post-trained model, fine-tuned from the open-weight Kimi K3 base via asynchronous reinforcement learning. The model targets long-horizon legal tasks like M&A diligence, contract review, and redlining. Harvey claims Tenet nearly doubles task completion on its own Legal Agent Benchmark and improves 20 percent on LAB: Contracts over the base model. Tenet was trained on roughly 150 Nvidia B300 GPUs over two months. Harvey also describes three specialist components for diligence, contract review tables, and firm-knowledge retrieval, each built with partners Baseten, Applied Compute, and Engram and all posting large cost or accuracy gains. Harvey states no customer data was used in training. There’s no public API, model card, or weights. An independent audit found most headline numbers are either self-reported or run on Harvey’s own benchmarks rather than verified on public leaderboards. (MarkTechPost)
OpenAI unveils change to its data retention plans
OpenAI rolled out Private Safety Processing, a new safety layer for its Zero Data Retention (ZDR) program. Existing ZDR safety systems check each interaction in isolation, but real abuses, like coordinated attacks across accounts, or an agent that keeps acting after being told to stop, only become visible across multiple interactions. The new system extends pattern detection across related interactions without exposing data to OpenAI staff — stored data is encrypted with customer-controlled keys OpenAI can’t access. When something looks risky, OpenAI receives only a narrow signal about the activity type, leaving customers to investigate with their own logs. Full rollout and a technical white paper are expected in September. (OpenAI)
Information loops help transformer-based models hold on to context
Google published research on Recirculation, a technique that helps transformers retain context as processing continues. Normally, information flows through transformer layers once, but Recirculation feeds richer representations from deeper layers back into shallower ones, so understanding built early stays available for later inputs. This happens at inference, and weights are frozen: There’s no retraining, just a change in how information moves through the model. The results are notable: a 60 percent drop in contextualization errors, 23 percent lower perplexity, and a 21 percent gain on GSM8K. (arXiv)
Want to know more about what matters in AI right now?
Read the latest issue of The Batch for in-depth analysis of news and research.
Last week, Andrew talked about the essential skills for building and deploying AI applications, focusing on LLM foundations, grounding models with data, and evaluation-driven development.
“The key difference between AI applications and non-AI software is that the former’s output is less predictable. You don’t know in advance what an LLM will output, or what predictions a supervised learning algorithm will make. Because of this uncertainty, building AI systems is a much more iterative process than building traditional software — it is harder to plan the process in advance.”
Read Andrew’s letter here.
Other top AI news and research covered in depth:
- Grok’s Cursor Alliance Pays Off as Grok 4.6 rivals Claude Opus 5 and GPT-5.6 Sol at a lower price.
- Anthropic details how Claude’s watermarks work with a new invisible text marking system for future versions of Claude.
- Qwen3.8-Max lands with a bang as Alibaba unveils two new models, Qwen3.8-Max and Qwen3.8-27B.
- Agents come to speech recognition with AgenticASR, which incorporates user corrections to edit speech recognition on the fly.
A special offer for our community
DeepLearning.AI’s first-ever subscription plan for our entire course catalog includes foundational classics like the Machine Learning and Deep Learning specializations plus seminars on the latest tools and frameworks you need.
As a Pro Member, you’ll immediately enjoy access to:
- Nearly 200 short and long AI courses from Andrew Ng and industry experts
- Labs and quizzes to test your knowledge
- Projects to share with employers
- Certificates to testify to your new skills
- A community to help you advance at the speed of AI
Enroll now to lock in a year of full access for $25 per month paid upfront, or opt for month-to-month payments at just $30 per month. Both payment options begin with a one-week free trial.
Explore Pro’s benefits and start building today!