The AI labs are nervous. What should you do?

An overview of my video editing folder flow

The AI labs are nervous. What should you do?

Dario Amodei's call to slow frontier AI development was the story I kept coming back to this week. For those of us already using agents, there's a practical question: how much of their work can we see?

Microsoft's internal case study and the new premade business workflows give us a few places to look. A friend and fellow reader wrote me this week that I need to share more workflows. I agree. (Thanks, Mo). So, I'll close our this issue with a short video of how we (me and my agent team) have built human review into our own video editing process.

1. Slow the frontier. Watch the work.

What happened

Amodei proposes pacing improvements in AI capabilities so safety work can keep up. Anthropic commits to embedding outside evaluators with access to its work; broader industry and international coordination remain proposals. There is no enacted global pause in this announcement. (Just the hope that there might be). Read the proposal.

A related development: OpenAI introduced a misalignment reporting framework and released six reports from training and evaluation over the previous six months. They include concealed mistakes and unauthorized file uploads. These are individual cases, not a measure of how often agents misbehave, and they involve multiple models.

What changes

I want more of this evidence available. It gives us something concrete to inspect when a company says its agents are safe enough to take on more work. And it provides more context on what "dangerous" actually means in practice. I see two primary schools of thought on this: One who thinks the insiders know something we don't, and another who thinks this is a ploy to regulate away competition. Which do you think it is?

Regardless of your answer, We have to recognize that agents can no longer be "set-it-and-forget-it" (at least at the frontier level).

Pay attention to what your agents do. Read the output and check where they went to produce it.

What to do next

Keep a record of tool actions, files read or changed, outside destinations, approvals and errors. Redact sensitive values and limit access to those logs. A log gives you something to investigate; it doesn't prevent a bad action.

Here is what staying informed looks like with my personal agent team.A simple activity board on Moria shows recent edits, writes and task-list updates. A separate Grafana dashboard shows usage, cost and session activity. These are different views of the work so that I can understand if my agents are behaving as I expect them to and catch it early if there's something strange going on.

Activity from my last Claude sessions. File paths are redacted; the timestamps and edit/write labels are unchanged.

My Grafana telemetry dashboard: usage, cost and session activity over time.

Pick one task this week. Check which files the agent changed, review the result, and compare the recorded activity with what you asked it to do. Give you claude agent this page and ask it to build something to consume and show your analytics in a local Grafana dashboard

2. Microsoft's results deserve a look at the footnotes

What happened

Microsoft reports that selected cloud supply-chain planning cycles fell from roughly ten business days to under 2.5 across five cycles, April through August 2026. The effort involved more than 150 people, with more than 111 agents deployed by September; simpler processes and shared data came first. Case study and methodology.

What changes

That's encouraging, but the results aren't unbiased or separated from process changes. Microsoft also sells this technology and chose the examples. The before-and-after comparison bundles agents with process changes and a substantial team; it can't tell us how much improvement came from agents alone, but it can tell us that people with access to the technology and the appropriate change management can get substantial results.

The same article compares frequent and infrequent Copilot users in sales. That comparison is observational. Better starting performance or stronger support could help explain the difference.

Somewhat surprisingly, we don't have evidence here that high-agency people benefit most. In a different setting, Brynjolfsson, Li and Raymond's customer-support research found larger gains for novice and lower-skilled workers using an AI assistant (because it normalizes their output to a mean). Skill, initiative and improvement are different things.

What to do next

I think taking ownership of the work matters. Decide what a good result looks like, inspect what comes back, and change the instructions when it misses.

Pick one recurring task and track time to an accepted result, including corrections. Give yourself a baseline before you add an agent. That will tell you more about your own work, probably surprise you, and it will give you a baseline for efficiency improvements when working with an agent.

3. Start premade. Make it fit your work.

What happened

Anthropic announced 43 small-business workflows and 27 new integrations across paid Claude plans. Examples include a Monday business briefing, CRM follow-up and payroll preparation, with a person submitting payroll.

Claude is also bringing chat and Cowork together, alongside document and presentation tools. Rollout and beta limits apply.

What changes

I think premade is what a lot of people think they want because it's up and running immediately. But the real value of these tool comes from working with them and making them better little-by-little. Once you have a starting point, you can adapt the inputs and rules to your work, define a good result, and decide where someone needs to stop and review it. This incremental improvement will give you an output that will absolutely trounce a pre-made prompt.

Here's the first of many examples of how I'm doing this internally for my companies.

Project showcase: our video-editing workflow

We built this flow to turn a recording into clips, with a person choosing what gets made and reviewing the result. The 39-second animation below walks through the stages and human checkpoints.

  1. Recording → transcript. Put the source video in the intake folder. An agent starts the work; the tools produce a transcript with timestamps.
  2. Transcript → clip choices. The agent proposes three to five clips with reasons. Human gate: I choose which clip to make and can even customize what to emphasize.
  3. Chosen clip → edited versions. Tools trim, apply a static face-centered vertical crop, add captions and prepare platform files.
  4. Drafts → accepted files. Human gate: I check the words, crop, pacing and ending. Corrections go back through editing. The final video is loaded to a temporary webpage for me to view and/or download along with any supporting documents to offer as downloads for ManyChat.
​

For each stage, write down Input / Action / Done Well / What's next. This is an agent workflow built entirely on files and folders which is gaining some traction as Interpretable Context Methodology (Van Clief & McDermott, 2026). For me it's easier to think of them as Folder Flows.

Start with this prompt:

Help me turn this recording into short video clips. First, ask about the audience, platform, target length, and editing preferences.

Organize the work into separate folders for intake, clip proposals, editing, and review. Give each folder a context.md describing its inputs, instructions, what a good result looks like, and when to stop.

Read the transcript and suggest three clips with timestamps, a proposed opening, and why each would work. Flag anything that could lose its meaning out of context. Save the proposals and wait for me to choose.

Before editing, show me options for framing, captions, and pacing. After I approve, use the available tools to create drafts. Put them in the review folder with a summary of changes and anything I should check. Wait for my feedback before finalizing. Don’t publish anything.

Which recurring task would you like to see turned into a flow that fits your work? Hit reply and tell me where it gets stuck.

–KM

Worth Your Attention

A weekly briefing for business leaders who want the few AI developments that matter.