AI News Summary 2026-08-29
AI News Summary 2026-08-29
GAFAM and major AI companies
OpenAI and Google tighten the agent stack
OpenAI’s Codex is now generally available with a new Slack integration, an SDK, and workspace admin tools, while Google introduced flexible usage limits for Gemini Notebook that refresh every five hours and scale with compute-heavy behavior such as longer chats and more sources. Together, the two launches show the same trend from opposite sides of the market: more capable AI assistants, but with more explicit controls for teams and power users. OpenAI Codex GA • Gemini Notebook usage limits
Influencers and technical blogs
Technical digest signals worth following
DeepLearning.AI’s Aug 28 The Batch digest highlights fresh model-watch items such as DeepSeek-V4-Pro Gets Refreshed and GLM-5.3 Makes Cybersecurity Gains, which is a useful signal for readers tracking open and frontier-model movement. Simon Willison’s Aug 27 analysis of Claude Code auto mode is also a strong nearby reference for the current safety debate around coding agents. The Batch tag page • Simon Willison
Generative imagery
Canva and Google keep turning AI images into editable work
Canva’s Magic Layers now appears directly inside Gemini and ChatGPT, letting people turn generated images into fully editable Canva designs. Google Flow also added more creative control, including start and end frames, lower-cost draft renders, and higher-quality exports. The practical result is less “generate a file” and more “ship a usable creative asset.” Canva Magic Layers • Google Flow controls
Chatbots and agents
Claude Code on the web is today’s lead story
Anthropic introduced Claude Code on the web as a browser-based way to delegate coding tasks. It supports parallel sessions, isolated sandboxes, real-time progress tracking, automatic PR creation, and an iOS preview. That makes the product feel less like a terminal tool and more like a browser-native command center for coding work. Claude Code on the web • Claude status
Local AI and serving
Serving stacks keep getting more practical
vLLM’s recent AMD GPU post walks through speculative decoding and related tuning for latency- and throughput-sensitive inference. It’s a good reminder that the local-serving story is still about practical engineering wins: better decode efficiency, more predictable performance, and fewer compromises when you want to run models on your own hardware. vLLM speculative decoding on AMD GPUs