The operating system for an AI team

The advantage is in the system around the model

Most of us are running the same personal setup.

A capable model or two. A handful of tools. Usage climbing month over month. And no real view of what any of it costs. (These Max plans feel almost unlimited...)

That works fine until it doesn't. And then you hit the wall hard.

This week's news was four separate signals that the interesting work has moved somewhere else – to the operating system around the model. How a serious team measures agents. How it keeps sensitive work inside the right perimeter. What it decides to run locally instead of in someone else's cloud. And how it bounds an agent's responsibility.

None of this requires Uber's budget. But it does require answering four questions, and I'll come back to them at the end.


01 — Uber shows the unglamorous work that makes an AI software team run

What happened

Uber published a look inside its software delivery organization, and the numbers are large.

The company reports that more than 70% of its pull requests now involve local or cloud agents. Its engineers have built more than 3,600 skills, which run more than 30,000 times a day. From February to mid-August 2026, Uber reports weekly active users grew 7x and weekly agentic requests grew 9.4x. Same as last week, we're seeing usage grow over time as companies get used to working with AI agents as a daily part of their work.

The cost figures tell a slightly different story. Holding the model fixed for its own internal comparison, Uber reports cost per 1,000 requests fell almost 34% from its peak, and cost per session fell 52% from its June peak.

Those are Uber's first-party numbers about Uber's environment. This is interesting as a read on the state of the industry for silicon valley, but not necessarily a benchmark for yours (unless your company also has around 34k employees and a market cap of 160.99 billion). I did fine, however, some important ways we can apply what they shared.

What changes

What's changed is the public way they've shared this information and especially how they're managing AI-enhanced employees.

Uber's agents have outcome-denominated metrics – cost per merged pull request, per review, per alert, per cleanup. Its model routing comes from a benchmark built out of real work, evaluated behind a common harness. They also display session costs on screen while people work. So users can see in real time what this action costs. And there are human review and escalation paths to move particularly prolific users to the next tier of token spend if and when it makes sense for the project or the area of the company.

None of these are technical features. They're the change management – the operating model – that surrounds this new way of working.

What to do next

Six translations for a solo operator or a small team:

1. Measure an outcome, not activity. Cost per completed deliverable, accepted recommendation, merged change, or resolved request. Prompts and tokens are inputs; they tell you almost nothing.

2. Match the cache to the work horizon. Now forgive me because this one is a bit technical. When you're making calls to models to get answers to questions (inference) you can store those questions and those answers for a limited time so that you don't have to spend token to repeat the work if they're asked again in a certain window. This is called caching.

Uber lengthened caches for human sessions because people stop for meetings and come back (to 1 hour). It kept subagent caches short because those tasks finish relatively quickly. A longer cache is efficient when it prevents rework and pure waste when you're paying to preserve context for work that already ended a while ago.

3. Keep cost visible at the moment of choice. A live counter or a periodic check changes behavior in a way that a post-hoc surprise never does.

The subtle implication here is a simple question:

Is the work I'm doing worth it?

This is a uniquely human question, at least for the time being 😉. Only we can decide, with our judgment and experience, what's worth spending our time on. That's an even more important skill when your decisions have leverage that compels others (either artificial or human teams) to spend their time on that work as well.

4. Pool the view across tools. I love that in this instance Uber, a company that's not really known for putting people at the center of things, put the human at the center of their analytic data.

Centralizing this around the individual user rather than the tool allowed them to measure whether the whole working system – not just Claude Code or Codex in isolation – is returning value per person and per outcome.

5. Route bounded work to the cheapest model that clears your quality bar. This follows the Law of Comparative Advantage or, as it's know around my house, the "Don't do for the kids what the kids can do for themselves" principle. This allows you to be as efficient as possible with your usage and frees up the opportunity cost of the stronger models or the human owner to maintain judgment and evaluation without becoming overwhelmed.

6. Turn repeated friction into an ever improving workflow. Capture the judgment while the work is fresh, then update the skill or the process. Otherwise your team's learning evaporates when the session ends. (Note: You don't need to be working with AI to do this.)

One caveat on context graphs, since it comes up in the article: elaborate context graphs can be worth it, but for most personal or small business setups, flat files plus local search (like grep or rg) are cheaper to run and vastly easier to maintain.

Read Uber's writeup


02 — Google packages a governed room for legal work

What happened

Google Cloud launched Gemini Enterprise for Legal in preview.

Google says the product preserves a firm's existing document and matter permissions, supports ethical walls, provides a governed control plane, and keeps customer data and outputs private to the organization rather than using them to train or fine-tune Google foundation models.

What changes

Legal work shouldn't be pasted into a consumer AI chat – see the US v Heppner ruling earlier this year. To manage this, Google required the permissions, confidentiality boundaries, and audit requirements to travel with the work regardless of how capable the model is. Google is selling the control plane, the access model, and the data perimeter as part of the product, for a profession where those things are non-negotiable.

Expect that packaging to spread to every field with the same constraints.

None of this makes a workload automatically privileged, and Google is not the first vendor to offer secure legal AI (see CoCounsel and Harvey). Privilege and confidentiality still depend on the firm's configuration, engagement terms, data handling, retention, review process, and applicable law. Also note: Google's product (and extension of their enterprise AI tier) is in preview.

What to do next

The word "safe" doesn't tell you anything about an AI tool on its own.

Ask instead whether the system preserves the four things the work requires:

  1. Data security and privacy
  2. Privileged access and security
  3. Auditability and reporting
  4. Ethical oversight

If you can't answer all four for a given workflow (and you must), that's the workflow to fix first.

Read Google Cloud's announcement


03 — Apple's new desktops make “keep it local” a real option

What happened

Apple announced an M6 Mac mini and new M5 Max and M5 Ultra Mac Studio models.

The memory ceilings haven't improved much from before but that doesn't tell the whole story. 32GB unified memory on the M6 mini, up to 128GB on the M5 Max Studio, and up to 512GB on the M5 Ultra Studio.

Bandwidth splits the two Studio tiers – Apple lists up to 614GB/s on the M5 Max and 1.2TB/s on the M5 Ultra. Availability starts September 22, 2026, and Apple says the 512GB configuration arrives later this year.

What changes

Local AI is emerging from a hobbyist exercise to a viable deployment strategy for certain privileged or offline workflows.

Memory capacity determines which models can practically live on a machine at all (like how many items you can keep in your short term memory). Bandwidth largely determines whether using them feels responsive (like whether it feels like you're talking to a quick-thinking rabbit or well-meaning sloth). The M6 mini is a credible always-on box for smaller private workflows. The Studio tiers make very capable local models realistic for anyone who can justify the price and the operational responsibility that comes with owning and maintaining the hardware.

Two cautions.

  1. These are pre-announced and expected to ship on September 22, 2026, so no one has real inference benchmarks yet. We can expect that they'll be better than what we've seen in the past from Apple, which has been pretty good.
  2. And local is not automatically safer – endpoint security, access control, backups, logging, and physical control of the machine all still apply, and now they're your job.

What to do next

Treat "local" as a data-governance and economics decision, not a speed decision.

Start from the workflow whose content should never leave your controlled environment. If you have one, you have your answer. If you don't, the cloud is probably still the cheaper and faster operational choice — for now.

Read Apple's Mac mini announcement · Read Apple's Mac Studio announcement


04 — The deployments are starting in the high-risk work

What happened

Salesforce published The State of Agentic AI in the Enterprise, a vendor-sponsored survey gated behind a registration form.

Two findings stand out. Organizations are applying AI to high-risk areas – the places where the need is most acute – rather than starting safe.

And orchestration is becoming common: a central AI coordinating a set of specialist agents. The report's executive summary reports 30% of organizations deployed and 47% experimenting.

What changes

High-cost, high-friction work is exactly where an organization has enough motivation to build the controls, evaluation, and accountable operating model that AI deployment requires. Low-stakes work never generates enough pain to justify that effort, so it never gets the attention.

The orchestration finding is interesting for a different reason. A coordinating agent is a design choice, not an autonomous executive: one layer that delegates bounded work to specialists, with clear controls and one accountable human at the top. The delegation moves down the chain. Accountability stays with the human at the top. This is the model I've been personally using for months and that I've been deploying with great success for others. It's nice to see it finally getting some love.

What to do next

Look at your most expensive, most annoying, most error-prone process – the one everybody complains about.

That's probably where an agent is worth the governance overhead. I've worked in and around enterprise technology for long enough to know you automate for two primary reasons:

  1. speed
  2. accuracy

And to be clear AI is just automation with style.

Read the Salesforce report


Also worth knowing — Model standards are reaching into hardware

Anthropic published a research preview of a Model Hardware Standard.

The near-term relevance is narrow: this is for manufacturers and operators thinking about models reaching into physical systems.

The reason to watch it is that model standards are starting to have consequences outside software. Anthropic's previous standards work has traveled further than expected, which is a reason to pay attention.

Read the research preview


Four questions, then. Pick one workflow and answer them:

  1. What is one real outcome I would measure for an agent?
  2. What data, permissions, and review boundary does that agent need?
  3. Which work should stay local because privacy or cost matters?
  4. Where can a coordinating agent delegate to specialists without erasing accountability?

Which one is hardest for you right now?

Hit reply and tell me. I read every reply.

–KM

p.s. I finally got the sending email and domain set up correctly. Thanks to those of you who noticed and then dragged me for it 😂. Add kevin@kevinmcmurphy.email to your contact list to make sure these continue to get through. Until next week...

Worth Your Attention

A weekly briefing for business leaders who want the few AI developments that matter.