AI News Summary 2026-07-15
.. lang: en
AI News Summary July 15, 2026
Today’s overarching theme was clear: companies are treating AI as an operational layer with metrics, governance, and real-world deployment; at the same time, agents are moving beyond chat and approaching the operating system, and local serving continues to be refined for low-latency workloads.
GAFAM and Major AI Companies
How to Manage AI Investments in the Agentic Era
OpenAI published a guide for investing with greater confidence in the agentic era and states that the cost per million tokens fell by 97% from GPT-4 to GPT-5.4; furthermore, GPT-5.6 improves the Artificial Analysis Coding Agent Index with 54% fewer output tokens and 57% less time per task. OpenAI article · OpenAI news
How Gemini is speaking the language of Southeast Asia
Google says that active Gemini users in Southeast Asia have more than doubled in the last year and that Gemini Spark is being rolled out in local languages for Gemini Advanced subscribers. Google article · Innovation & AI
Influencers and Tech Blogs
TIL: Using uvx in GitHub Actions in a cache-friendly way
Simon Willison shares a practical pattern for caching uvx in GitHub Actions using UV_EXCLUDE_NEWER, with the goal of avoiding repeated downloads from PyPI on every run. post · Atom feed
Generative imaging
Introducing Muse Image and Muse Video
Meta introduced Muse Image for following instructions, editing, and composition with multiple references, and Muse Video with native audio support. Although it isn’t the freshest news of the day, it remains the most solid highlight to emerge from the generative image source review. Meta AI · blog
Chatbots and Agents
Personal Computer Is Here
Perplexity has begun rolling out Personal Computer, an expansion of Perplexity Computer that brings multimodal orchestration to local files, native apps, connectors, and the web, while keeping the user in control of sensitive actions. Perplexity blog · hub
On-Premises AI and Serving
vLLM x TileRT: Specialized Decode for Latency-Critical Serving
vLLM introduces a pluggable decode path with TileRT for low-latency workloads, while keeping the prefill and the rest of the vLLM stack intact. vLLM article · blog