PM Status Report - 3 August 2026
Two developments in the eight days to 3 August do the most for a delivery team’s actual workload. Anthropic rewrote the protocol most agents use to reach outside tools, building enterprise identity controls into its core rather than bolting them on afterwards. And a wave of ultra-cheap model releases out of China pushed the price of capable AI down again, hard enough that a cost case you built a month ago is probably wrong now.
Google DeepMind also gave its robots a planning brain, which matters less this week and more as a signal of where physical-world AI is heading. And the EU AI Act’s enforcement powers went live on 2 August.
The Protocol Behind Every Agent Just Grew Enterprise Controls
Anthropic published a substantial overhaul of the Model Context Protocol, the open standard most AI agents now use to reach outside tools and data, on 28 July. The core connection model changed from a persistent, stateful session to a stateless request/response pattern, which is the difference between something you can run on ordinary web infrastructure and something that needs dedicated always-on servers. That alone should widen who can afford to build and host a connector.
The more relevant change for anyone running a PMO is identity. The spec now builds in OAuth 2.0 and OpenID Connect, so an organisation can centrally control which agents get access to which tools rather than trusting each connector’s own login screen. It also adds “MCP tunnels”, still in research preview, that let a cloud-hosted agent reach an on-premise system, an EHR, a council asset register, a mainframe nobody wants exposed to the internet, without opening a public endpoint. OpenAI moved the same direction with a beta identity feature, Sign in with ChatGPT, letting users authenticate into Airtable, GitLab, HubSpot, Notion, Supabase and Vercel with their ChatGPT account.
Read together, both labs are solving the same problem from opposite ends: standardising how an agent proves who it is and what it’s allowed to touch. That’s the unglamorous, structural work that has to exist before an organisation can hand an agent real access without it turning into a governance incident.
The Cost Floor Dropped Again, This Time From China
A cluster of releases out of Chinese labs over 27-31 July pushed the price of a capable multimodal model down further than the waves earlier in the year. Alibaba’s Qwen3.7 Flash undercuts Google’s cheapest comparable model by roughly an order of magnitude on both input and output pricing, shipped with no published benchmarks to back the capability claims. Moonshot published full open weights for Kimi K3, by scale the largest open-weight language model released anywhere to date, under a bespoke licence that isn’t quite standard open source. DeepSeek moved its V4-Flash model out of preview into public beta with real gains on agentic and tool-calling tasks.
None of this comes with independent verification yet. Capability claims for the cheapest tier remain vendor-stated, and self-hosting something the size of Kimi K3 is its own infrastructure project, not a download. But the direction holds: the gap between frontier-adjacent capability and cheap, fast capability keeps narrowing, and it’s narrowing from below as much as from the top labs cutting their own prices.
Robots Got a Planning Brain, Not Just Better Hands
Google DeepMind released Gemini Robotics ER 2 on 30 July, a model built to act as the high-level “brain” for a robot rather than control its motors directly. It watches continuous video, tracks whether a task is actually progressing, corrects course when something goes wrong, plans multi-step work, and hands the low-level movement off to separate vision-language-action models. In demonstrations it coordinated an Apptronik Apollo 2 humanoid through kitchen and tidying tasks, and directed a Boston Dynamics Spot unit through navigation and fetching alongside a manipulator arm, splitting the work by which robot was actually suited to it.
This is a research and early-access release, not something you’re deploying on a construction site next quarter. What it signals is more useful than what it does today: labs are building the coordination layer for multiple physical agents working a shared space, which is the same problem a site supervisor or facilities manager already solves with people. Worth watching if your projects touch physical assets, not worth budgeting for yet.
What This Means for Your Projects
Use identity federation to unblock the systems everyone’s given up on. MCP tunnels and OAuth-based agent access mean a cloud agent can now reach an on-premise system, a hospital’s patient record system, a council’s asset register, a manufacturer’s decades-old ERP, without punching a hole in the firewall for it. That’s the barrier that’s stopped a lot of legacy-system automation from getting past a pilot. Worth revisiting any project that stalled because IT wouldn’t expose an internal system to an external agent.
Re-run the cost case again. Two weeks after the last price drop, capable models got cheaper again. If you built a budget for AI-assisted document review, procurement analysis or contact-centre triage more than a month ago, it’s probably overstated. Keep the model choice abstracted from your workflow so you can chase price without re-integrating, and treat the newest ultra-cheap Chinese-hosted models as unverified and out of scope for anything involving regulated or client data until you’ve run your own tests and checked data residency.
Watch physical AI, don’t plan around it yet. Multi-robot coordination is moving from lab demo to early-access product, which matters for construction, logistics, warehousing, healthcare facilities and anywhere else people currently coordinate expensive physical assets by radio and clipboard. It isn’t procurable at project scale yet. File it as a horizon item for your next capability review, not a line in this quarter’s budget.
Where Things Stand
The more durable development this week isn’t regulatory. Identity and access control for agents, the unglamorous infrastructure that has to exist before agentic delivery moves from pilot to production, is now shipping rather than being discussed. Combine that with a cost curve that keeps compressing and the arithmetic for what’s worth automating changes again, a few weeks after it last changed.
Frequently Asked Questions
Is the MCP identity overhaul something we need to act on immediately? Not urgently, but it’s worth flagging to whoever owns your agent tooling. If your organisation has stalled an automation project because a legacy system couldn’t safely be exposed to an external agent, the tunnel and OAuth features are the reason to revisit it.
Should we move workloads to the new ultra-cheap Chinese models? Not yet, and not for anything involving regulated or client data. The capability claims are vendor-stated with no independent benchmarks published, and hosted versions raise data-residency questions that self-hosted Western open-weight models don’t. Worth testing on low-stakes, high-volume tasks; not worth betting a production workflow on until someone else has verified the numbers.
Is Gemini Robotics ER 2 relevant if we don’t run physical operations? Not directly, but the coordination pattern it demonstrates, one model planning, several specialised systems executing, is the same pattern showing up in software agent platforms this week. Worth reading as a preview of how multi-agent orchestration gets structured generally, even if the robots themselves aren’t on your roadmap.
For project delivery this week the useful frame is portability: which model does the job, and which systems it’s allowed to reach, should both be decisions you can change quickly, not commitments you’re stuck with.
If the cheapest model capable of doing the job changed again next month, would your workflow notice, or would you?
Yes - AI helped me to write this :)
Unsubscribe